VLDB 2026 Research / reviewers in the wild / expert
Razib Iqbal
dblp:92/3835
· DBLP profile ↗
37ranked-venue papers
13as first author
15since 2021 · last 2025
0000-0002-9293-2993ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Verbal Abuse Detection from Short ConversationsabstractSmart assistants and smart microphones can contribute to detecting verbal abuse by analyzing speech-to-text contents. In this paper, we compare different large language models (LLMs) for detecting verbal abuse and propose a framework that first identifies emotion from short audio conversations by extracting Mel Frequency Cepstral Coefficient (MFCC) and Mel Spectrogram (MEL) features. It then utilizes transfer learning with a fully connected neural network incorporating an attention mechanism and SBERT encoding. To demonstrate the efficacy of the proposed framework, we prepared a custom dataset containing instances of verbal abuse. Evaluation results show that our framework is lightweight and achieves commendable accuracy compared to the existing LLM models. Fahim Ahmed Irfan, Christina Behl, Razib Iqbal |
CCNC | 3 |
| 2025 | DyCoDe: Dynamic Context Detection from Voice Commands and Conversations in Smart HomesabstractAs voice assistants become integral to smart home environments, the need for accurate context interpretation of voice commands is critical for ensuring seamless and autonomous device control. Traditional smart assistants like Siri, Amazon Alexa, and Google Home often depend on users remembering specific commands, which can limit the user experience. Hence, this paper proposes a novel dynamic context detection framework, DyCoDe, that aims to eliminate the need to remember pre-defined commands. Additionally, with the advancement of technology, newly invented appliances and sensors are being employed in IoT smart home environments, creating a diverse and unique nature for each smart home. Our framework incorporates descriptions and capabilities of these new appliances to ensure accurate context detection. DyCoDe continuously processes conversational data, employing techniques such as clustering, topic identification, and zero-shot learning to detect and adapt to new contexts automatically. This adaptive model evolves alongside its environment, identifying and enhancing autonomous actuation without requiring manual intervention. Our dynamic context detection framework builds on our prior context recognition model, which combines a transformer architecture, attention mechanism, and fully connected neural network. We have extended this model to offer an unsupervised and adaptive detection framework for diverse smart home environments. Our evaluation results show that DyCoDe significantly improves context detection in smart homes, allowing voice assistants to automate a broader range of daily tasks more effectively. Jeniya Sultana, Razib Iqbal |
CCNC | 2 |
| 2025 | CoRe: A Comparative Study of Transformer Models for Context Recognition in Smart ClassroomsabstractWith the increasing adoption of smart assistants, voice-enabled interactions have the potential to transform traditional classrooms into intelligent learning environments. A voice-enabled smart assistant can be designed to recognize context from student questions and brief conversations, enabling efficient query handling with minimal human intervention. This allows instructors to focus more on students who require personalized assistance, improving the overall learning experience. However, accurately identifying context remains a challenge due to the complexity and variability of student interactions. Therefore, in this paper, we present a comparative study of transformer models for context recognition (CoRe) in K-12 smart classrooms. We curated a custom dataset comprising short commands and conversations related to day-to-day classroom operations and trained multiple transformer models to identify the best-performing model for recognizing context. To ensure data quality and support data-driven decision-making, we performed topic modeling. Our evaluation results show that a transformer model enhanced with an attention mechanism outperforms existing transformer models while maintaining low computational costs, making it a viable solution for real-world smart classroom applications. Fahim Ahmed Irfan, Razib Iqbal, Dawn Eckstein, Molly Strickland |
COMPSAC | 2 |
| 2025 | SEAD: Sensor Event-Based Anomaly Detection for Smart Home AutomationabstractAs smart IoT devices become increasingly common in our homes, there is a growing demand for seamless automation and synchronization between these devices. A key challenge in achieving this automation is accurately identifying and grouping related sensors, which are essential for generating automated operational policies to control the actuators. However, this process is often hindered by anomalous data in sensor readings, which can obscure the sensor relationships identified during the grouping process. These anomalies disrupt the accuracy of the sensor groupings and, as a result, compromise the effectiveness and reliability of the automation policies generated for managing smart environments. In this paper, we introduce SEAD, a novel approach for detecting anomalies by first calculating the total number of sensor events within a specified time window, followed by the use of an unsupervised learning method for anomaly detection. We evaluate the effectiveness of this method by leveraging existing sensor inference techniques and testing it on three custom datasets and one public dataset. Our experimental results show that applying SEAD to remove anomalies improves the quality of sensor groupings for smart home automation. Md. Asif Tanvir, Fahim Ahmed Irfan, Razib Iqbal |
COMPSAC | 3 |
| 2024 | Word Embedding with Emotionally Relevant Keyword Search for Context Detection from Smart Home Voice CommandsabstractVoice-enabled virtual assistants have received widespread popularity in smart homes. Adding a context detection feature in voice conversations with virtual assistants can offer a more personalized experience in smart homes such that it maintains awareness of the ongoing conversation and responds appropriately. In this paper, we present a novel word embedding with emotionally relevant keyword search (WERKS) approach for context detection. This WERKS approach makes use of a combination of emotion detection, keyword search, and word embedding for context detection from voice commands and short conversations with virtual assistants. The TPOT classifier was applied over RAVDESS and a custom data set to obtain experimental results, which demonstrated a 15 and 12 percent increase in prediction accuracy of our defined contexts. Brent Anderson, Razib Iqbal |
CCNC | 2 |
| 2024 | SeReIn-M: Sensor Relationship Inference in Multi-Resident Smart HomesabstractModern smart homes comprise a large amount of sensors and actuators. Identifying sensor relationships can contribute to automating operational policies of the actuators. However, in a multi-resident smart home, it is difficult to identify sensor relationships due to a variety of simultaneous sensor events. In this paper, we propose a novel two-step approach, which initially extracts features from time series data generated by sensor events to cluster related sensors, and then identifies how each sensor is related to other sensors revealing their physical proximity and enabling a sensor's ability to participate in multiple sensor groups. Experimental results show that our approach performs well in multi-resident homes even if the sensor events are not equally distributed. Fahim Ahmed Irfan, Razib Iqbal, Ayesha Siddiqua |
CCNC | 2 |
| 2024 | TIM-MARL: Information Sharing for Multi-Agent Reinforcement Learning in Smart EnvironmentsabstractInformation sharing among agents to jointly solve problems is challenging for multi-agent reinforcement learning algorithms (MARL) in smart environments. In this paper, we present a novel information sharing approach for MARL, which introduces a Team Information Matrix (TIM) that integrates scenario-independent spatial and environmental information combined with the agent's local observations, augmenting both individual agent's performance and global awareness during the MARL learning. To evaluate this approach, we conducted experiments on three multi-agent scenarios of varying difficulty levels implemented in Unity ML-Agents Toolkit. Experimental results show that the agents utilizing our TIM-Shared variation outperformed those using decentralized MARL and achieved comparable performance to agents employing centralized MARL. Ayesha Siddiqua, Siming Liu 0001, Razib Iqbal, Fahim Ahmed Irfan, Logan Ross, Brian Zweerink |
CCNC | 3 |
| 2024 | A Framework for Context Recognition from Voice Commands and Conversations with Smart AssistantsabstractRecognizing contexts or meanings from voice commands and conversations with smart assistants can contribute to the autonomous control of smart home devices and appliances. To perform context recognition from short voice conversations, we can convert audio to text and apply natural language processing (NLP) techniques to analyze the textual content. However, the existing text classification systems that apply NLP techniques to extract meaningful information from texts require large training data. In this paper, we propose a novel framework to extract contexts from short-spoken texts requiring smaller training datasets. This framework exploits the power of transfer learning and uses a fully connected neural network aided with SBERT encoding, and an attention mechanism. Our proposed framework has been evaluated using two datasets containing short smart home commands. Evaluation results demonstrate that our model achieves higher accuracy in context recognition with low computational costs and less training time compared to other methods like BERT and deep neural networks. Jeniya Sultana, Razib Iqbal |
CCNC | 2 |
| 2024 | Six Weeks with ROSE: Teacher Perspectives on Computer Science Professional DevelopmentabstractThis innovative practice full paper describes a novel twofold approach to analyze the data collected during a six-week research-based summer professional development workshop for middle and high school STEM teachers in Southwest Missouri. In the fast-evolving field of Computer Science (CS), particularly within the Internet of Things (IoT) domain, the demand for a skilled workforce is increasing. To meet this demand, it is crucial to provide K-12 students with education in computing, computational thinking, and other broader Science, Technology, Engineering, and Mathematics (STEM) disciplines. The existing STEM curriculum could be enhanced to develop skills such as research-based problem-solving more effectively, aiming for a more comprehensive skill set among students. Therefore, empowering STEM teachers with a solid foundation of research-based problem-solving skills can significantly boost their readiness for the classroom, thus enriching their students' educational experiences. The Research Opportunity for Smart Environments (ROSE) program, a three-year initiative funded by the National Science Foundation (NSF), aims to prepare middle and high school STEM teachers to effectively introduce CS concepts through innovative IoT applications in Smart Environments, including Smart Homes and Smart Classrooms. To support STEM education in rural and other underrepresented areas, a cohort of in-service teachers from rural Southwest Missouri was selected for the inaugural ROSE summer workshop. Our data collection approach included open-ended questions and effective formative assessment techniques to gather the teachers' perspectives during the initial summer session. We adopted a twofold data analysis strategy, using the qualitative coding tool MAXQDA and sentiment analysis tools Valence Aware Dictionary and sEntiment Reasoner (VADER) and TextBlob, to analyze the teachers' responses. The research questions explored the impact of the ROSE program on educators and how their experiences and outlooks influenced their engagement and professional development within the program. The findings indicate that the ROSE program has positively influenced the teachers' personal and professional growth. The teachers experienced increased confidence and knowledge in research and teaching CS concepts, despite facing various challenges, which acted as motivation for learning. Their reflections also indicated that the mentorship and resources provided by the ROSE program have promoted the development of innovative teaching methods and a shift towards a student-centered approach, showcasing the program's success in fostering the personal and professional development of teachers for the advancement of future CS and STEM professionals. Zihan Zan, Razib Iqbal, Ajay K. Katangur, Siming Liu 0001, Diana Piccolo |
FIE | 2 |
| 2023 | An Unsupervised Learning Approach for Smart Home Operational Policy GenerationabstractWith the rise of the Internet of Things (IoT), smart homes can provide intelligent services to monitor household appliances remotely and automate user tasks. However, a significant amount of human intervention is expected in the deployment and operation of such services, making it inconvenient for human users who are less tech-savvy. Therefore, such systems should be trained to learn user behavior patterns to automatically configure and adapt their actions according to the preferences and daily routines of the occupants with minimal user involvement to enhance user experience. This paper uses an unsupervised learning approach that can be integrated with a generative policy framework. It enables automatic operational policy generation by analyzing the continuous data from sensors and smart devices. In order to generate policies according to user preferences, our process infers users' behavior patterns by looking for the patterns in their daily routine activities. We compared the performance of our proposed learning approach with existing approaches used in generative policy frameworks. Evaluation results show that our proposed approach positively contributes to the automatic policy generation to automate user tasks in smart homes. Santhi Priya Challa, Razib Iqbal, Siming Liu 0001 |
CCNC | 2 |
| 2022 | Multi-Scale Speaker Vectors for Zero-Shot Speech SynthesisabstractRecent advances in deep learning have allowed the task of synthesizing speech from text, otherwise known as Text-to-Speech (TTS), to be addressed with more powerful and effective techniques than previously possible. In this paper, we propose a novel speaker encoder model utilizing some of those techniques, such as self-attention and transformers. The proposed model results in a speaker embedding vector called Multi-Scale Speaker (MSS) vectors. MSS vectors aim to make improvements over the current state-of-the-art speaker embeddings for more natural and similar-sounding synthesized speech for unseen speakers in a zero-shot speech synthesis model. The proposed architecture relies on encoder layers which are coupled with a multi-scale approach for learning both local and global mel-spectra features of a reference speaker. We compared our proposed model against a modified generalized end-to-end (GE2E) speaker encoder using mel-cepstrum distortion and cosine similarity measures as well as mean opinion scores. Evaluation results confirm our encoder's advantages over a state-of-the-art speaker encoder model. Tristin Cory, Razib Iqbal |
COMPSAC | 2 |
| 2022 | DESCo: Detecting Emotions from Smart CommandsabstractWith the proliferation of smart home devices like Amazon Alexa and Google Home, automatic emotion detection from user commands and interactions with smart assistants, in other words smart commands, can enable personalized services for the inhabitants in a smart-home environment. In this paper, we compare different machine learning algorithms to identify a suitable classification technique for emotion detection. We also propose four new audio features named Chunk Gap Length, Mean Chunk Duration, Mean Word Duration Per Chunk and Per Chunk Word Count in addition to the existing Mel Frequency Cepstral Coefficient (MFCC) and Mel Spectrogram (MEL) features for emotion classification. We used the publicly available RAVDESS dataset for our initial experiment and then generated a custom dataset consisting of 5000 smart-home voice commands covering five emotional states - happy, normal, sad, fearful, and angry. Evaluation results show that combining our proposed features with MFCC and MEL provides better accuracy in classifying the correct emotions for individual users than MFCC and MEL only. Sunanda Guha, Razib Iqbal |
COMPSAC | 2 |
| 2022 | Comparison of Multi-Scale Speaker Vectors and S-Vectors for Zero-Shot Speech SynthesisabstractWe compare a novel speaker encoder model, called Multi-Scale Speaker (MSS) Vectors, with state-of-the-art s-vectors model for zero-shot speech synthesis. The s-vectors model relies on a modified transformer self-attention network for its architecture. The MSS vectors model introduces a multi-scale approach to the s-vectors model. Results demonstrate that our model produces more natural and similar-sounding synthesized speech for unseen speakers in a zero-shot speech synthesis system. Tristin Cory, Razib Iqbal |
ISM | 2 |
| 2021 | Towards a Template Matching Approach for Human Fall DetectionabstractThe technology and products related to fall detection has always been in high demand within the security and healthcare industries. A fall detection system must provide a reliable notification mechanism to significantly reduce the risk and medical care costs associated with falls. The rapid developments in Smart Environments and the Internet of Things paradigms, together with the increasing number of low-cost cameras, form a favorable setup for vision-based fall detection systems. In this paper, we propose a novel vision-based fall detection technique that uses fall templates. Fall templates represent human shape variations in case of a fall. We implemented different template matching techniques to evaluate the performance of our proposed template-based approach. Performance evaluation shows encouraging results for single camera-based live deployments. Snigdha Chaudhari, Razib Iqbal |
COMPSAC | 2 |
| 2021 | An Analysis of Lightweight Convolutional Neural Networks for Parking Space Occupancy DetectionabstractCommercial parking space occupancy detection systems used to be mostly sensor-based. Very recently, we have seen great success in computer vision techniques which allow us to utilize the CCTV camera feed in real-time. In this paper, we review multiple existing convolutional neural network models of various sizes to analyze the benefits of using each model for parking space occupancy detection. We measure the accuracy, required floating point operations, and parameter counts of each model. We then compare model performance over different conditions such as camera perspectives, weather types, different parking lots, and lighting conditions. Based on our observations and experience, we introduce three novel architectures, two based on the DenseNet architecture - Mini DenseNet and Simple DenseNet, and CoarseNet - a multi-layer perceptron. We compare our proposed models to other models of various sizes. Performance results show that these three models have advantages over existing models in parameter counts, accuracy, and resilience to new camera perspectives. Joshua D. Ellis, Anthony Harris, Naseem Saquer, Razib Iqbal |
ISM | 4 |
| 2020 | SoCeR: A New Source Code Recommendation Technique for Code ReuseabstractMotivated by the idea of reusing existing source code from previous projects within a software company, in this paper, we present a new source code recommendation technique called "SoCeR" to help programmers find relevant implementations or sample code based on software requirement specifications. SoCeR assists programmers to search existing code repositories using natural language query. Our proposed approach summarizes Python code into sentences or phrases to match them against user queries. SoCeR extracts and analyzes the content of the code (such as variables, functions, docstrings, and comments) to generate code summary for each function which is then mapped to the respective functions. For evaluation purposes, we developed a web-based tool for users to enter a textual search query and get the relevant code search results that were most relevant to the query. In SoCeR, users can also upload new code to enrich the code base with tested code. If adopted, then SoCeR will benefit a software company to build a trusted code base enabling large-scale software code reuse. Md. Mazharul Islam 0004, Razib Iqbal |
COMPSAC | 2 |
| 2018 | Joint Intra and Multiple Description Coding for Packet Loss Resilient Video TransmissionabstractMultiple description coding (MDC) is a technique for video transmission over error prone networks where the descriptions are routed over multiple paths. Intra coding such as MDC provides error resiliency but coding in this mode must be decided with care since it degrades the compression ratio. In this paper, we present our investigation results for a new intra coding approach in MDC. We have found that, in MDC streams, the best policy is to encode selective frames as I-frame instead of coding some macroblocks of frames in intra mode. In order to find the most suitable I-frame positions within a given video stream, we developed a cost function based on which intra/inter frame type is decided. The MDC scheme with the proposed intra coding criterion, with and without redundancy optimization, is implemented in the H.264/AVC reference software, JM16.0. Based on the experimental performance evaluation, we show that our method achieves higher average PSNR compared to the other optimized MDCs found in the literature. Mohammad Kazemi 0002, Razib Iqbal, Shervin Shirmohammadi |
IEEE Trans. Multim. | 2 |
| 2017 | A Cloud-Based Multi-threaded Implementation of View Synthesis SystemabstractIn multiview video applications, view synthesis is a computationally intensive task that needs to be done correctly and efficiently in order to deliver a seamless user experience. In order to provide fast and efficient view synthesis, in this paper, we present a cloud-based implementation that will be especially beneficial to mobile users whose devices may not be powerful enough for high quality view synthesis. Our proposed implementation balances the view synthesis algorithm's components across multiple threads and utilizes the computational capacity of modern CPUs for faster and higher quality view synthesis. For arbitrary view generation, we utilize the depth map of the scene from the cameras' viewpoint and estimate the depth information conceived from the virtual camera. The estimated depth is then used in a backward direction to warp the cameras' image onto the virtual view. Finally, we use a depth-aided inpainting strategy for the rendering step to reduce the effect of disocclusion regions (holes) and to paint the missing pixels. For our cloud implementation, we employed an automatic scaling feature to offer elasticity in order to adapt the service load according to the fluctuating user demands. Our performance results using 4 multiview videos over 2 different scenarios show that our proposed system achieves average improvement of around 1.7 times for speedup, 50% and 52% for efficiency and resource utilization, respectively. Parvaneh Pouladzadeh, Razib Iqbal, Shervin Shirmohammadi, Omid Fatemi |
ISM | 2 |
| 2017 | Redundancy Allocation Based on the Weighted Mismatch-Rate Slope for Multiple Description Video CodingabstractMultiple description coding (MDC) is a robust coding technique for video transmission over error prone networks, whereby the video is encoded into multiple descriptions with some redundancy between the descriptions. This redundancy leads to error resiliency in the case of packet loss during the network transport. However, the amount of this redundancy has a critical role in MDC performance. Therefore, a crucial problem in MDC is to find what the optimum amount of redundancy budget is, and then how this redundancy budget can be optimally allocated to the frames. To solve this problem, we propose a scheme in which the redundancy budget is allocated to the frames based on the weighted mismatch-rate slopes so that this additional bitrate can attain maximum distortion reduction. The redundancy is added gradually so that fine tuning of the utilized bitrate is achievable. We have verified our proposed scheme by implementing it in H.264/AVC reference software JM16.0, and running experiments against two representative reference methods. Our experiments show that our scheme not only minimizes the end-to-end distortion with a rate-distortion performance that is better than the reference methods, especially for high PLRs, but also entirely uses the available bandwidth, unlike the reference methods. Mohammad Kazemi 0002, Razib Iqbal, Shervin Shirmohammadi |
IEEE Trans. Multim. | 2 |
| 2016 | A Selective Intra-Coding Approach for Multiple Description Video CodingabstractMultiple Description Coding (MDC) is a technique for video transmission over error prone networks where the descriptions are routed over multiple paths. Intra coding technique and MDC both provides error resiliency for real-time video transport over unreliable networks. In this paper, we present our investigation results for a new intra coding approach for MDC where selective frames are fully encoded in intra mode instead of coding selective macroblocks in intra mode in many frames. We implemented our scheme in the H.264/AVC reference software, JM16.0. Based on the experimental performance evaluation, we show that our method achieves higher average PSNR compared to the other optimized MDC schemes available in the literature. Mohammad Kazemi 0002, Razib Iqbal, Shervin Shirmohammadi |
ISM | 2 |
| 2015 | A Dynamic Alpha Congestion Controller for WebRTCabstractVideo conferencing applications have significantly changed the way people communicate over the Internet. Web Real-Time Communication (WebRTC), drafted by the World Wide Web Consortium (W3C) and Internet Engineering Task Force (IETF) working groups, has added new functionality to the web browsers, allowing audio/video calls between browsers without the need to install any video telephony applications. The Google Congestion Control (GCC) algorithm has been proposed as WebRTC's congestion control mechanism, but its performance is limited due to using a fixed incoming rate decrease factor, known as alpha (a). In this paper, we propose a dynamic alpha model to reduce the available receiving bandwidth estimate during overuse as indicated by the over-use detector. Experiments using our specific testbed show that our proposed model achieves a 33% higher incoming rate and a 16% lower round-trip time, while keeping a similar packet loss rate and video quality, compared to a fixed alpha model. Rasha Atwah, Razib Iqbal, Shervin Shirmohammadi, Abbas Javadtalab |
ISM | 2 |
| 2013 | Modeling and Evaluation of a Metadata-Based Adaptive P2P Video-Streaming SystemabstractIn this paper, we present a multi-parent adaptive video-streaming system (MAVSS). MAVSS is a cooperative video-streaming system, based on the peer-to-peer (P2P) content distribution concept, to simultaneously adapt and stream video contents to heterogeneous users. In order to ensure video codec independence, we emphasize on structured, metadata-based adaptation. Therefore, MPEG-21 generic Bitstream syntax description is chosen to describe the parts of the video contents selected for adaptation operations. Additionally, without an optimal solution, it is hard to judge the efficiency of such systems. Therefore, in this paper, we first present the fundamental properties of MAVSS. We then present a mathematical model of the adaptive streaming system, where we define efficient streaming as an optimization of a cost function that can be solved as an integer linear programming problem. The model illustrates the relation among all the parameters that affect the resource contribution, resource utilization, load balancing and service fairness. We use this model to analyze the trade-offs that exist between service fairness and system efficiency. Razib Iqbal, Shervin Shirmohammadi, Behnoosh Hariri |
Comput. J. | 1 |
| 2011 | Data Visualization Using Shape Preserving C2 Rational SplineabstractA rational cubic spline is developed to provide smooth curves(positive, monotone and convex). To control the shape of the curve, two families of parameters are introduced in its representation. Three schemes using rational cubic spline are elaborated to obtain positive curves through positive data, monotone curves through monotone data and convex curves through convex data. As well as degree of smoothness attained is C2. Muhammad Sarfraz 0001, Malik Zawwar Hussain, Tahira Sumbal Shaikh, Razib Iqbal |
IV | 4 |
| 2009 | A light-weight federated video adaptation system for P2P overlaysabstractThis paper presents a lightweight but effective mastersender-driven multiple parent approach for online video adaptation and streaming to heterogeneous clients. The proposed design facilitates the cooperation of participating clients for improving the utilization of the spare resources in an overlay. Our design uses standard H.264/AVC video streams. It enables peers to contribute both CPU and bandwidth in a multi-parent fashion, such that no dedicated adaptation server and streaming server is required, although we need a reliable rendezvous point for the video stream originators and the peers in the overlay. We present the live video processing paradigm using metadata followed by the video distribution overview. A brief performance evaluation supporting the design choices is also presented. Razib Iqbal, Shervin Shirmohammadi |
ICME | 1 |
| 2009 | A Cooperative video adaptation and streaming scheme for mobile and heterogeneous devices in a community networkabstractIn this paper, we aim to present the fundamental properties of a community-driven adaptive P2P streaming scheme. We show that if the participants in a community network agree to not only share their bandwidth, but also their computing resources according to the design principles mentioned here, then mobile and heterogeneous devices can be accommodated in the P2P paradigm ensuring adequate resource utilization, with respect to resilience to peer dynamics. We present simple design principles for a multimodal P2P system considering available bandwidth, computing power, and delay to build the video overlays. Some evaluation results supporting our design principles are also presented. Razib Iqbal, Shervin Shirmohammadi |
ICME | 1 |
| 2009 | MPEG-21 based temporal video adaptation for heterogeneous devices and mobile environmentsabstractIn the era of universal multimedia access (UMA), applications like video conferencing, surveillance and streaming are challenged by the multiplicity of devices. Since adaptation is the newly applied practice for digital video content customization, in this demo, we present a simple approach to adapt H.264 videos, in the compressed domain, for heterogeneous devices and mobile environments. Our approach is in contrast to most existing approaches that do adaptation in a series of cascaded decode/re-encode operations. We show that the adaptation operations can be expedited, using our approach, when adaptation systems are designed to adapt contents according to the encoding structure but in an intermediary node following a codec-independent technique. We present our adaptation system as a utility box where adaptation is performed on-demand based on its generic bitstream syntax description (gBSD), and in compressed domain. Razib Iqbal, Shervin Shirmohammadi |
ICME | 1 |
| 2009 | Compressed-domain temporal adaptation-resilient watermarking for H.264 video authenticationabstractIn this paper, we present a DCT domain watermarking approach for H.264/AVC video coding standard. This scheme is resilient to compressed-domain temporal adaptation. A cryptographic hash function is used to generate a semi-fragile watermark to provide content-based authentication. The embedded watermark can withstand frame-dropping due to temporal adaptation, yet it is able to detect malicious attacks such as content modification, transcoding etc. Simulation results demonstrate that the watermarking scheme is computationally efficient and suitable for practical use. Sharmeen Shahabuddin, Razib Iqbal, Shervin Shirmohammadi, Jiying Zhao |
ICME | 2 |
| 2009 | A compressed-domain spatio-temporal adaptation system for video deliveryabstractIn this demo, we present a working system that we have built for metadata-based compressed-domain spatio-temporal video adaptation of H.264/AVC videos. Our approach is in contrast to most existing approaches that do adaptation in a series of cascaded decode/re-encode procedures. We show that by applying compressed-domain operations, adaptation can be expedited for real-time application scenarios like news/sports broadcasting. Moreover, the system can be applied as a tool box for distributed adaptation within P2P video distribution applications to support heterogeneous devices. Razib Iqbal, Sharmeen Shahabuddin, Shervin Shirmohammadi |
ACM Multimedia | 1 |
| 2009 | Compressed domain spatial adaptation for H.264 videoabstractIn this paper, we present a metadata-based compressed-domain spatial adaptation scheme for H.264/AVC video. We have enhanced the H.264/AVC encoder with our proposed adaptation strategies in order to reduce video size by cropping individual frames in an intermediary node prior to transmitting that video to heterogeneous devices. In this regard, we exploit the sliced architecture of the video frames within the first version of the H.264/AVC specification and devise different slicing strategies. The compressed-domain bitstream modification is performed at the intermediary nodes, avoiding the need for any cascaded operations. Here, we briefly present our adaptation scheme as well as evaluation results showing the effectiveness of the slicing strategies on the bitrate reduction and processing time. A comparison of our approach with an existing cropping scheme is also presented. Sharmeen Shahabuddin, Razib Iqbal, Ali A. Nazari Shirehjini, Shervin Shirmohammadi |
ACM Multimedia | 2 |
| 2009 | DAg-stream: Distributed video adaptation for overlay streaming to heterogeneous devices
Razib Iqbal, Shervin Shirmohammadi |
Peer-to-Peer Netw. Appl. | 1 |
| 2008 | Online adaptation for video sharing applicationsabstractThe main concept of Peer-to-Peer (P2P) streaming is that viewers will contribute their bandwidth to the overlay and act as a relay for the video streams. In this paper, we introduce how a peer may implement an adaptive streaming scheme to serve peers in a P2P application. The technical contribution of this paper is to present the effectiveness and feasibility of utilizing the available computing power of the participating peers to serve mobile and heterogeneous clients by adapting the video content on the fly. The benefit is that there is no need for a dedicated adaptation or streaming server deployed in the system for video streaming/sharing applications. We emphasize on structured metadata-based adaptation and streaming utilizing MPEG-21 gBSD. Here, we briefly illustrate our scheme and present some experimental evaluations supporting our design choices. Copyright 2008 ACM. Razib Iqbal, Shervin Shirmohammadi |
ACM Multimedia | 1 |
| 2008 | Modeling and evaluation of overlay generation problem for peer-assisted video adaptation and streamingabstractIn this paper, we consider the problem of overlay generation for video adaptation and streaming applications in a way to efficiently utilize the bandwidth and computing power of the participating peers. Therefore, the proposed architecture performs regular streaming functions as well as video adaptation functions, moving the video contents adaptation computation load away from dedicated media-streaming/adaptation servers to the participating peers. To verify the performance of our design, we followed an analytical approach based on 0-1 Integer Linear Programming method to model the system and to calculate the optimum overlay. The performance of our scheme is evaluated by simulations. Preliminary results demonstrate that our design performance nearly follows the optimal boundary in terms of resource utilization. Razib Iqbal, Behnoosh Hariri, Shervin Shirmohammadi |
NOSSDAV | 1 |
| 2008 | Distributed Video Adaptation and Streaming for Heterogeneous DevicesabstractIn this paper, we present a new concept of distributed adaptation and P2P streaming supporting the expansion of context-aware mobile P2P systems. Here, we propose to adapt video contents in a distributed manner to address the peer heterogeneity. It also forms an overlay network and organizes video streaming sessions among the participating peers. We have followed a master-sender-driven approach (one-to-many), with the help of cooperative peers. The essential incentive behind this work is that with the modern computing capabilities, a single peer can carry out sufficient adaptation operations and stream media contents to other peers. Overall design and adaptation approach along with the performance evaluation is presented in brief. Razib Iqbal, Dewan Tanvir Ahmed, Shervin Shirmohammadi |
PerCom | 1 |
| 2007 | Hard Authentication of H.264 Video Applying MPEG-21 Generic Bitstream Syntax Description (gBSD)abstractWhile trivial research has been conducted in watermarking and authentication of H.264 video in recent years, most techniques require cascaded operations within a video adaptation scenario. In this paper, we propose an authentication scheme for adapted H.264 video content to detect integrity at the receiver's side without the need for cascaded operations. The proposed scheme utilizes MPEG-21 gBSD for hard authentication of H.264 video in the compressed domain and does not necessitate any cascaded decompression and recompression. The design uses content-based authentication which is derived from a hash value. The authentication data is embedded as a fragile watermark, and the marking space is selected during the adaptation process of the H.264 video by parsing the gBSD. The authentication information is embedded in already encoded videos during adaptation. Proof of concept and performance evaluation is also presented. Razib Iqbal, Shervin Shirmohammadi, Jiying Zhao |
ICME | 1 |
| 2006 | Secured MPEG-21 Digital Item Adaptation for H.264 VideoabstractSeamless adaptation and transcoding techniques to adapt the digital content have achieved significant focus to serve the consumers with the desired content in a feasible way. With the succession of time we sense that secured adaptation should also be taken care of for not only serving sensitive digital contents but also to offer security as an embedded feature of the adaptation practice to ensure digital right management and confidentiality. In this paper, we propose an encryption framework for a transcoder while adapting H.264 video conforming to MPEG-21 DIA. Encryption mechanism is applied on the adapted video content thus reducing computational overhead compared to that on the original content Razib Iqbal, Shervin Shirmohammadi, Abdulmotaleb El Saddik |
ICME | 1 |
| 2006 | Compressed-Domain Encryption of Adapted H.264 VideoabstractCommercial service providers and secret services yearn to employ the available environment for conveyance of their data in a secured way. In order to encrypt or to ensure personalized security of the video contents in an intermediary node, it is necessary to have the content structure conforming to an international standard. Moreover, pressure to satisfy user preferences and device requirements seamlessly are raising the need for content to be customized providing the best possible experience. In this paper, we present perceptual encryption scheme for video encryption that is incorporated with a dynamic temporal adaptation technique of the H.264 video conforming ISO/IEC MPEG-21 Digital Item Adaptation. Encryption is performed on demand directly from the adapted bitstream and its generic Bitstream Syntax Description (gBSD). Razib Iqbal, Shervin Shirmohammadi, Abdulmotaleb El Saddik |
ISM | 1 |
| 2006 | MPEG-21 Based Temporal Adaptation of Live H.264 VideoabstractThe diversity of devices in both wired and wireless networks via which multimedia contents are desired to be accessed and interacted with has grown significantly. Applications like video conferencing, surveillance and chatting is challenged by this diversity which requires live adaptation to meet user requirements and device specifications. In this paper, we present an architecture for temporal adaptation of ITU-T H.264 video conforming to ISO/IEC MPEG-21 DIA for live video stream along with the adaptation module implementation detail. Adaptation is performed on demand directly from the live bitstream and its generic bitstream syntax description (gBSD) avoiding conventional approaches seen in traditional transcoders. As a result, any MPEG-21 compliant host can adapt the stream without requiring the video codec. A prototype, based on the proposed architecture, and experimental evaluations of the system and its performance supporting the architecture are also presented Razib Iqbal, Shervin Shirmohammadi, Chris Joslin |
ISM | 1 |