EDBT 2026 Demo / reviewers in the wild / expert
Shervin Shirmohammadi
dblp:10/2042
· DBLP profile ↗
136ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0002-3973-4445ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 99 · 8 first-author · 14 since 2021Computer networks · 24 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 first-authorArtificial intelligence and machine learning · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorSystems, architecture and hardware · 3Security and privacy · 3Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | iSR: Super-resolution for Immersive Cloud VR Gaming PlatformsabstractCloud-based Virtual Reality (VR) gaming enables immersive experiences without the need for costly high-end consumer hardware. However, it imposes substantial bandwidth requirements due to the need to stream high-resolution, high frame rate, and stereoscopic frames to maintain immersion and prevent motion sickness. Existing techniques like foveated rendering and encoding face challenges such as reliance on costly eye-tracking hardware and sensitivity to sudden gaze shifts. In addition, prior super-resolution methods can improve fidelity but are often too computationally heavy for practical deployment. To address these challenges, we propose iSR, a system that integrates stereo-aware colorization and super-resolution to reduce transmission cost while preserving visual quality. The key idea of iSR is that it first downsamples both stereo views to reduce the total number of transmitted pixels. Then, it transmits one view in full color and the other in monochrome. This removes redundant chrominance information and further reduces the required bandwidth. On the client side, iSR reconstructs full-color, high-resolution stereo frames by transferring chroma between views and enhancing spatial resolution. Extensive experiments across multiple VR games show that iSR achieves substantial bitrate reductions while maintaining high visual fidelity. These results highlight its potential for enabling high-quality VR streaming in bandwidth-limited environments. Ghazaleh Bakhtiariazad, Shervin Shirmohammadi, Ihab Amer, Mohamed Hefeeda |
MMSys | 3 |
| 2026 | JNTD-DS: A Benchmark Dataset for Just Noticeable frame rate-based Temporal Difference in Perceptual Video CodingabstractThe Just Noticeable Difference (JND) is defined as the maximum change in a visual stimulus (image or video) which the Human Visual System (HVS) can tolerate without perceiving visual distortion. Previous JND research has largely focused on spatial distortions, yielding several datasets and models that predict spatial thresholds based on parameters such as Quantization Parameter (QP) or Quality Factor (QF). However, temporal thresholds, specifically the maximum frame rate reductions which viewers cannot detect, remain largely unexplored, despite their critical importance for efficient video coding. To address this gap, we introduce JNTD-DS, which, to the best of our knowledge, is the first benchmark dataset specifically designed to measure the Just Noticeable frame rate-based Temporal Difference (JNTD). The dataset comprises 50 video scenes covering various content, and the JNTD level associated with them. T e video scenes are studied through extensive subjective tests, comparing the high frame rate videos with their temporally downsampled versions. This forms 1196 opinion scores from 78 subjects. Analyzing the collected data confirms that JNTD thresholds, which are fundamentally defined by the HVS, are inherently complex and vary across content. By providing critical insights into HVS sensitivity to frame rate changes, the dataset enables content-adaptive frame rate optimization for perceptual video coding, allowing more efficient compression in video streaming and bandwidth-limited applications without compromising visual quality. We further demonstrate the practical impact of these insights by developing a JNTD prediction model and integrating it into a video compression pipeline, achieving an average bitrate reduction of 13.62% with only a marginal quality loss. The JNTD-DS is publicly available at https://github.com/sanaznami/JNTD-DS. Sanaz Nami, Farhad Pakdaman, Sahab Taali, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
MMSys | 5 |
| 2026 | GameLab: AI-Enabled Cloud Gaming TestbedabstractWe present GameLab, an open-source, AI-enabled cloud gaming testbed built on WebRTC. Unlike existing open-source stacks and deployment-oriented pipelines (e.g., GamingAnywhere, Sunshine/-Moonlight, and Unity Render Streaming), GameLab is designed for AI-in-the-loop systems research: it provides programmable interfaces on both the server and client, with well-defined hook points to plug in machine learning modules on demand (e.g., super-resolution, denoising, object detection, QoE estimation, and learned rate control). GameLab also collects detailed transport and application traces for online and offline analyses of gaming sessions and network behavior. To enable reliable objective evaluation in interactive settings, GameLab embeds compact QR-based frame identifiers into the video stream, allowing accurate computation of full-reference quality metrics, such as PSNR, SSIM, and VMAF, even under frame loss, reordering, and duplication. Finally, GameLab supports GPU-based visualization of frames and model outputs for interactive inspection without costly GPU-to-CPU transfers. Shervin Shirmohammadi, Ihab Amer, Mohamed Hefeeda |
MMSys | 2 |
| 2026 | JNTD: Toward Just Noticeable Frame Rate-Based Temporal Difference for Perceptual Video CodingabstractJust Noticeable Difference (JND) refers to the maximum level of distortion in an image or video sequence that remains imperceptible to the Human Visual System (HVS). Current JND-based studies predominantly rely on existing datasets, developing models predicting JND levels in terms of Quantization Parameter (QP) or Quality Factor (QF). However, these solutions primarily focus on spatial-based Perceptual Video Coding (PVC) and neglect temporal-based optimization, which highly affects the video bitrate. This paper addresses this limitation by introducing Just Noticeable frame rate-based Temporal Difference (JNTD) to determine the optimal Frame Rate (FR) based on human perception. A novel dataset comprising 50 high frame rate video sequences is collected through subjective assessments. Subsequently, an ensemble method is proposed to predict the JNTD, by leveraging deep and hand-crafted features, for robust prediction. Experimental evaluations include the integration of the proposed method into several codecs (H.264, H.265, H.266, and a new learned codec), showcasing its ability to reduce bitrate without compromising visual quality. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | A Review of Player Engagement Estimation in Video Games: Challenges and OpportunitiesabstractThis article presents a review on the process of estimating player engagement in video gaming. To stay ahead of their competitors in entertainment, game developers need to understand, estimate, and maximize player engagement. We address the multidimensional nature of engagement, encompassing cognitive, emotional, and behavioral aspects across various gaming domains. We present a taxonomy of the diverse modalities for quantifying engagement, including physiological signals, observable behaviors, and gameplay data. We identify the challenges of conducting representative subjective studies in this domain and summarize various methods for establishing ground truth measurements. By synthesizing existing research, we provide insights into modeling techniques, highlight research gaps, and offer practical guidelines for implementing engagement measurement strategies. This review aims to aid researchers and industry professionals in navigating the complexities of player engagement estimation, ultimately contributing to enhanced game design, marketing, and user retention in the competitive gaming landscape. Ammar Rashed, Shervin Shirmohammadi, Ihab Amer, Mohamed Hefeeda |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | A Driver Activity Dataset with Multiple RGB-D Cameras and mmWave RadarsabstractDriver activity recognition has become crucial for intelligent transportation and automotive safety systems. However, existing studies mainly focus on fatigue-related behaviors while neglecting other activities for analyzing driver behavior and intent. In this work, we introduce a novel dataset for fine-grained driver activities, utilizing diverse sensors such as mmWave radars, RGB, and depth cameras, each of which includes three camera angles: body, face, and hands. This multi-modal and multi-angle approach allows for comprehensive driver behavior analysis, including hand gestures, head movement, and object interactions. Moreover, including mmWave radars provides significant privacy advantages, as the sparse dynamic point clouds prevent the identification of the driver's face and other personal information. This dataset is valuable for researchers and developers on driver activity recognition and behavior analysis. It enables the development and evaluation of robust, privacy-conscious solutions for improving road safety, driver assistance, and in-vehicle interaction. Furthermore, the multi-modal nature of the data enables the exploration of sensor fusion techniques, unlocking the full potential of diverse sensing modalities to understand complex driver behaviors. Guan-Hua Li, Hsin-Che Chiang, Yi-Chen Li 0005, Shervin Shirmohammadi, Cheng-Hsin Hsu |
MMSys | 4 |
| 2024 | Lightweight Multitask Learning for Robust JND Prediction Using Latent Space and Reconstructed FramesabstractThe Just Noticeable Difference (JND) refers to the smallest distortion in an image or video that can be perceived by Human Visual System (HVS), and is widely used in optimizing image/video compression. However, accurate JND modeling is very challenging due to its content dependence, and the complex nature of the HVS. Recent solutions train deep learning based JND prediction models, mainly based on a Quantization Parameter (QP) value, representing a single JND level, and train separate models to predict each JND level. We point out that a single QP-distance is insufficient to properly train a network with millions of parameters, for a complex content-dependent task. Inspired by recent advances in learned compression and multitask learning, we propose to address this problem by (1) learning to reconstruct the JND-quality frames, jointly with the QP prediction, and (2) jointly learning several JND levels to augment the learning performance. We propose a novel solution where first, an effective feature backbone is trained by learning to reconstruct JND-quality frames from the raw frames. Second, JND prediction models are trained based on features extracted from latent space (i.e., compressed domain), or reconstructed JND-quality frames. Third, a multi-JND model is designed, which jointly learns three JND levels, further reducing the prediction error. Extensive experimental results demonstrate that our multi-JND method outperforms the state-of-the-art and achieves an average JND1prediction error of only 1.57 in QP, and 0.72 dB in PSNR. Moreover, the multitask learning approach, and compressed domain prediction facilitate light-weight inference by significantly reducing the complexity and the number of parameters. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | MTJND: Multi-Task Deep Learning Framework for Improved JND PredictionabstractThe limitation of the Human Visual System (HVS) in perceiving small distortions allows us to lower the bitrate required to achieve a certain visual quality. Predicting and applying the Just Noticeable Distortion (JND), which is a threshold for maximum unperceived level of distortions, is among the popular ways to do so. Recently, machine learning based methods have been able to reduce bitrate even further by improving JND prediction accuracy. However, accurate modeling of JND is very challenging, as it is highly content dependent. Furthermore, existing datasets provide little information to learn the best parameters. To remedy this issue, we propose a multi-task deep learning framework that jointly learns various complementary visual information. We design three separate methods and training strategies that jointly learn: (1) three JND levels, (2) visual attention map and a JND level, and (3) three JND levels and the visual attention map. We show that accumulating information from multiple tasks leads to a more robust prediction of JND. Experimental results confirm the superiority of our framework compared to the state-of-the-art. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
ICIP | 4 |
| 2023 | A Dataset of Food Intake Activities Using Sensors with Heterogeneous Privacy Sensitivity LevelsabstractHuman activity recognition, which involves recognizing human activities from sensor data, has drawn a lot of interest from researchers and practitioners as a result of the advent of smart homes, smart cities, and smart systems. Existing studies on activity recognition mostly concentrate on coarse-grained activities like walking and jumping, while fine-grained activities like eating and drinking are understudied because it is more difficult to recognize fine-grained activities than coarse-grained ones. As such, food intake activity recognition in particular is under investigation in the literature despite its importance for human health and well-being, including telehealth and diet management. In order to determine sensors' practical recognition accuracy, preferably with the least amount of privacy intrusion, a dataset of food intake activities utilizing sensors with varying degrees of privacy sensitivity is required. In this study, we collected such a dataset by collecting fine-grained food intake activities using sensors of heterogeneous privacy sensitivity levels, namely a mmWave radar, an RGB camera, and a depth camera. Solutions to recognize food intake activities can be developed using this dataset, which may provide a more comprehensive picture of the accuracy and privacy trade-offs involved with heterogeneous sensors. Yi-Hung Wu, Hsin-Che Chiang, Shervin Shirmohammadi, Cheng-Hsin Hsu |
MMSys | 3 |
| 2023 | Multiple Description Coding for Best-Effort Delivery of Light Field Video Using GNN-Based CompressionabstractIn recent years, Light Field (LF) video has grabbed much attention as an emerging form of immersive media. LF collects, through a lens matrix, light information emanating in every direction, and obtains rich information about the scene, providing users with an immersive 6 Degrees of Freedom (DoF) experience. The visual content between different viewpoints is highly homogenized, suggesting the possibility of good compression and encoding. However, most fixed-structure LF coding schemes are difficult to adapt to the real-time requirements of different LF applications and best-effort network conditions causing packet loss. In this paper, we propose a dynamic adaptive LF video transmission scheme that can achieve high compression and yet provide near-distortion-free LF video when the network condition is stable. Additionally, for unstable network conditions a description scheduling algorithm is proposed, which can decode the LF video with the highest possible quality even if partial data cannot be received completely and/or timely. We achieve this by designing a Multiple Description Coding (MDC) based solution to transport the LF video compressed by a Graph Neural Network (GNN) model. Experimental results show that the scheduling algorithm can improve the quality of the decoding results by 3% to 15%. Compared with other similar schemes, our system greatly improves the reliability of the video streaming system against packet loss/error and supports heterogeneous receivers. Xinjue Hu, Yumei Wang, Lin Zhang 0013, Shervin Shirmohammadi |
IEEE Trans. Multim. | 5 |
| 2023 | BL-JUNIPER: A CNN-Assisted Framework for Perceptual Video Coding Leveraging Block-Level JNDabstractJust Noticeable Distortion (JND) finds the minimum distortion level perceivable by humans. This can be a natural solution for setting the compression for each video region in perceptual video coding. However, existing JND-based solutions estimate JND levels for each video frame and ignore the fact that different video regions have different perceptual importance. To address this issue, we propose a Block-Level Just Noticeable Distortion-based Perceptual (BL-JUNIPER) framework for video coding. The proposed four-stage framework combines different perceptual information to further improve the prediction accuracy. The JND mapping in the first stage derives block-level JNDs from frame-level information without the need to collect a new bock-level JND dataset. In the second stage, an efficient CNN-based model is proposed to predict JND levels for each block according to spatial and temporal characteristics. Unlike existing methods, BL-JUNIPER works on raw video frames and avoids re-encoding each frame several times, making it computationally practical. Third, the visual importance of each block is measured using a visual attention model. Finally, a proposed quantization control algorithm uses both JND levels and visual importance to adjust the Quantization Parameter (QP) for each block. The specific algorithm for each stage of the proposed framework can be changed, as long as the input and output formats of each block are followed, without the need to change other stages, based on any current or future methods, providing a flexible and robust solution. Extensive experimental results demonstrate that BL-JUNIPER achieves a mean bitrate reduction of 27.75% with a Delta Mean Opinion Score (DMOS) close to zero and BD-Rate gains of 25.44% based on MOS, compared to the baseline encoding, and also gains a better performance compared to competing methods. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
IEEE Trans. Multim. | 4 |
| 2022 | Using Compressive Sampling to Fill Interbatch Data Gap From Low-Cost IoT Vibration SensorabstractA low-cost wireless vibration sensor can be built using a 3-axis accelerometer, such as ADXL345, attached to a low-cost Wi-Fi microchip, such as ESP8266. In an Internet of Things (IoT) setting, a large number of such inexpensive sensor nodes can be setup with the widely used direct-read-and-send method which samples and sends individually acquired vibration data points from the sensor through the Internet to a server. In this work, we show that such a method is not effective. As the microcontroller alternates between sampling and sending the data, the micro delays of transmission will affect the sensor sampling rate and cause the data points to space unevenly, making the acquired data inaccurate. We propose that vibration should be sampled and transmitted in batches, as such data are acquired continuously without interruption and data points are more evenly spaced. However, the proposed batch-read-and-send will have interbatch gaps that need to be filled. Thus, the key contribution of this work is the novel use of compressive sampling (CS) technique to bridge those gaps. Experimental results show that the direct-read-and-send method loses more information and can only achieve a maximum sampling rate of 350 Hz with a standard uncertainty of 12.4, whereas the proposed solution can measure the vibration wirelessly and continuously up to 633 Hz. Gaps with up to 160 missing points can be filled using CS and achieve better accuracy, with a mean absolute error (MAE) of up to 0.048 and a standard uncertainty of 0.001, making the low-cost wireless vibration sensor a cost-effective solution in an IoT setting. Boon-Yaik Ooi, Woan Lin Beh, Xin-Yi Kh'Ng, Soung-Yue Liew, Shervin Shirmohammadi |
IEEE Internet Things J. | 5 |
| 2021 | 4DLFVD: A 4D Light Field Video DatasetabstractWe present a 4D Light Field (LF) video dataset, collected by a custom-made camera matrix, to be used for designing and testing algorithms and systems for LF video coding, processing, and streaming. Compared to existing LF datasets, ours provides LF videos, as opposed to only images, and at higher frame resolution, higher number of viewpoints, and/or higher framerate, offering the best visual quality LF video dataset. To achieve this, we built a 10 x 10 LF capture matrix composed of 100 cameras, each with a 1920 x 1056 resolution. We used this matrix to record videos in real and varying illumination and scene dynamics conditions. The dataset contains a total of nine groups of LF videos: eight groups collected with a fixed camera matrix position and orientation recording indoor potted plants, furniture, etc., and the last group collected by rotating around an outdoor environment with roadside vehicles, pedestrians, etc. Each group of LF videos consists of 100 video streams encoded with H.265/HEVC. Scene changes vary from static to slightly dynamic to highly dynamic, providing a good level of diversity. As an example, we present the results of a depth estimation method and show that our dataset can be used for applications such as objection detection, 3D modeling, and others. Xinjue Hu, Yunming Liu, Yumei Wang, Yu Liu 0001, Lin Zhang 0013, Shervin Shirmohammadi |
MMSys | 8 |
| 2021 | A review of temporal video error concealment techniques and their suitability for HEVC and VVC
Mohammad Kazemi 0002, Mohammed Ghanbari 0001, Shervin Shirmohammadi |
Multim. Tools Appl. | 3 |
| 2021 | A novel fast search method to find disparity vectors in multiview video coding
Ghane Zandi, Hoda Roodaki, Shervin Shirmohammadi |
Multim. Tools Appl. | 3 |
| 2021 | A Machine-Learning-Based Action Recommender for Network Operation CentersabstractFailure management and cost-aware traffic engineering are two important tasks done in Network Operation Centers (NOC). These are performed by expert technicians who must carefully analyze the network state and the flow of incoming alarms to decide how, where and when to take actions on the network. While based on implicit guiding principles, these network actions are very hard to automate with explicit rules due to the high complexity of the system; hence NOC action is essentially a manual process today. To automate part of that process, in this paper we introduce an Action Recommendation Engine (ARE) that can learn implicit NOC action rules with supervised machine learning from historical data. As a result, ARE can recommend suitable action(s) to remedy network faults and engineer the traffic to minimize costs, all while maximizing the users' Quality of Experience. To quantify the effectiveness of different NOC action scenarios, we introduce the QoE-OPEX metric which balances between users' quality of Experience and ISP's operational costs. After proper model training on 56,000 data points with 66 features, we demonstrate that ARE can effectively reproduce implicit action-taking logic of NOC technicians, thus moving us one step closer to reliable autonomous networks and fully-automated NOCs. Shady A. Mohammed, Ayse Rumeysa Mohammed, David Côté, Shervin Shirmohammadi |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2020 | Resource optimization through hierarchical SDN-enabled inter data center network for cloud gamingabstractGaming on demand is an emerging service that combines techniques from Cloud Computing and Online Gaming. This new paradigm is garnering prominence in the gaming industry and leading to a new "anywhere and anytime" online gaming model. Despite its advantages, cloud gaming's Quality of Experience (QoE) is challenged by high and varying end-to-end communication delay. Since the significant part of the computational processing, including game rendering and video compression, is performed on the cloud, properly allocating game requests to the geographically distributed data centers (DCs) can lead to QoE improvements resulting from lower delays. In this paper, we propose a hierarchical Software Defined Network (SDN) controller architecture to near-optimally allocate a gaming session to a DC while minimizing network delay and maximizing bandwidth utilization. To do so, we formulate an optimization problem, and propose the Online Convex Optimization (OCO) as a practical solution. Simulation results indicate that the proposed method can provide close-to-optimal solutions, and outperforms classic offline techniques e.g. Lagrangean relaxation. In addition, the proposed model improves the bandwidth utilization of DCs, and reduces end-to-end delay and delay variation by gamers. As a byproduct, our proposed method also achieves better fairness among multiple competing players in comparison with existing methods. Maryam Amiri, Hussein Al Osman, Shervin Shirmohammadi |
MMSys | 3 |
| 2020 | A Parameter-Free Vibration Analysis Solution for Legacy Manufacturing Machines' Operation TrackingabstractDespite the fact that the revolution of Industry 4.0 has started almost a decade ago, there are still many yesteryear's manufacturing machines that are still currently in operation in many small and medium enterprises (SME) factories. These legacy manufacturing machines are built without computing power and Internet connectivity. Therefore, the process of gathering operational information of such systems is often done manually. This article aims to automatically track these machines' operation status via the vibration produced by these machines, by using a retrofit Internet-of-Things (IoT) approach that attaches wireless vibration sensors onto legacy manufacturing machines to capture the vibration of the machines. One of the challenges of the proposed retrofit approach is to interpret the meaning of the vibration without any prior knowledge of the machine's vibration and also without the privilege to interrupt the manufacturing process to produce data sets with labels. Although there are many existing works that capture and analyze vibration, they very often only focus on fault diagnosis and prognosis. Also, many of these vibration analysis techniques are not parameter free; i.e., parameters need to be fine-tuned according to the data. The contribution of this article is the proposal of a parameter-free vibration analysis technique to cluster and classify the type of vibrations produced by a machine. Experiments, which were carried out in a limestone processing factory on real industrial machineries, show that the proposed technique is able to track the operation status of a 3-speed industrial exhaust fan with an average accuracy of 98.6% (worst case 95.5%) and standard uncertainty of 1.06%. Boon-Yaik Ooi, Woan Lin Beh, Wai-Kong Lee, Shervin Shirmohammadi |
IEEE Internet Things J. | 4 |
| 2020 | Bandwidth On-Demand for Multimedia Big Data Transfer Across Geo-Distributed Cloud Data CentersabstractMultimedia content is massively generated from various applications and devices, and processed in cloud data centers. Multimedia service providers prefer that their data are processed in data centers close to users in order to offer them high performance and reliable multimedia services that meet the requirements specified in the Service Level of Agreement (SLA). This requires transferring huge data sets of video streams, games content, images etc. across geographically distributed cloud data centers using underutilized bandwidth in backbone transport networks. As the amount of multimedia content increases, the demand to transfer big data sets across data centers increases as well. As such, the leftover bandwidth that appears at different times and for different durations in the backbone network becomes insufficient to satisfy the rapidly increasing demand for multimedia big data transfer. This challenge led to the creation of multi-rate Bandwidth on-Demand (BoD) service offerings for communication between geographically distributed cloud data centers. In this paper, we focus on BoD services which are offered by the Dense Wavelength Division Multiplexing (DWDM) layer because of its huge capacity. We propose a BoD broker which employs a scheduling algorithm that considers various deadlines of multimedia big data transfer requests. The broker in our model leverages the concept of standby wavelengths to minimize peak traffic and accommodate time requirements of delay-tolerant and delay-intolerant transfer requests. We also study strategies of routing and wavelength assignment using Mixed Integer Programming (MIP) optimization to rapidly handle volumes of multimedia big data transfer requests. Abdulsalam Yassine, Ali A. Nazari Shirehjini, Shervin Shirmohammadi |
IEEE Trans. Cloud Comput. | 3 |
| 2020 | Cooperative Tile-Based 360° Panoramic Streaming in Heterogeneous Networks Using Scalable Video CodingabstractThe use of high-quality 360° panoramic video is booming in the video industry. However, existing schemes for smartphones suffer from significant bandwidth consumption as they transmit the entire panoramic views in very high resolutions. This demand for bandwidth becomes even more problematic when multiple adjacent smartphones compete to access the same content, which further challenges a wireless network's capacity, and when the available bandwidth fluctuates much more than wired networks. In this paper, we propose a cooperative streaming scheme for tile-based 360° video using scalable video coding (SVC) to maximize a group of users' quality of experience. We formulate an optimization problem to choose optimal downloading and sharing subsets from a set of all requested SVC layers of tiles to maximize the effective quality of the users' viewport while meeting the feasibility of the bandwidth of heterogeneous networks. We then show that the problem is NP-hard and compose a heuristic approach. In the approach, we rank the SVC layers based on the aggregated group-level preference to guide the devices' downloading and sharing activities. A prototype on the Android platform is developed to test the approach's performance, and the real-world results show that our proposed scheme outperforms baseline alternatives. Xiaoyi Zhang 0001, Xinjue Hu, Shervin Shirmohammadi, Lin Zhang 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | An Empirical Approach to Modeling User-System Interaction Conflicts in Smart HomesabstractConflict is one of the important factors affecting user satisfaction and trust in smart environments, yet conflict modeling in mixed initiative smart environments has not been sufficiently explored. Most of the existing literature on conflict in smart homes are centered on conflicts between users. Although research has shown that about 75% of conflicts are between users and system [1], only a few studies have considered user-system conflicts in smart homes. The aim of this article is to empirically propose both a definition and a run-time detection method for conflicts between users and smart home systems. Our empirical study is based on conflict sample scenarios collected from 163 users. Using clustering on these scenarios, we form an empirical definition of user-system conflict in smart homes. We also propose two functions that characterize each class of the collected scenarios, and we detect conflicts from this characterization. Our conflict detection model could help users achieve a more satisfactory experience in smart homes. Moreover, the model can offer benefits for system developers to design and deploy more reliable smart homes. Fereshteh Jadidi Miandashti, Mohammad Izadi, Ali A. Nazari Shirehjini, Shervin Shirmohammadi |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2020 | The Effect of Room Complexity on Physical Object Selection Performance in 3-D Mobile User InterfacesabstractAn important challenge in smart environments is how to manipulate the smart objects. Although mobile applications are typically used for controlling a smart environment, no previous study has evaluated the users performance in manipulating smart objects under different environmental complexities. This article presents an experimental comparison between three different selection techniques 3-D, 2-D, and physical user interfaces (UIs). We evaluate these techniques across two levels of environment complexity measuring 51 participants timing data and errors. Our results indicate that the 3-D UI is superior for task completion time and error, and the 2-D UI is not a better solution than the physical UI when the environment is not complex. The results also show the importance of considering the environment complexity in choosing the proper UI. Maryam Rezaie, Morteza Malekmakan, Ali A. Nazari Shirehjini, Shervin Shirmohammadi |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2020 | The Performance of Quality Metrics in Assessing Error-Concealed Video QualityabstractIn highly-interactive video streaming applications such as video conferencing, tele-presence, or tele-operation, retransmission is typically not used, due to the tight deadline of the application. In such cases, the lost or erroneous data must be concealed. While various error concealment techniques exist, there is no defined rule to compare their perceived quality. In this paper, the performance of 16 existing image and video quality metrics (PSNR, SSIM, VQM, etc.) evaluating errorconcealed video quality is studied. The encoded video is subjected to packet loss and the loss is concealed using various error concealment techniques. We show that the subjective quality of the video cannot be necessarily predicted from the visual quality of the error-concealed frame alone. We then apply the metrics to the error-concealed images/videos and evaluate their success in predicting the scores reported by human subjects. The errorconcealed videos are judged by image quality metrics applied on the lossy frame, or by video quality metrics applied on the video clip containing that lossy frame; this way, the impact of error propagation is also considered by the objective metrics. The measurement and comparison of the results show that, mostly though not always, measuring the objective quality of the video is a better way to judge the error concealment performance. Moreover, our experiments show that when the objective quality metrics are used for the assessment of the performance of an error concealment technique, they do not behave as they would for general quality assessment. In fact, some newly developed metrics show the correct decision only about 60% of the time, leading to an unacceptable error rate of as much as 40%. Our analysis shows which specific quality metrics are relatively more suitable for error-concealed videos. Mohammad Kazemi 0002, Mohammed Ghanbari 0001, Shervin Shirmohammadi |
IEEE Trans. Image Process. | 3 |
| 2020 | Intra Coding Strategy for Video Error Resiliency: Behavioral AnalysisabstractOne challenge in video transmission is to deal with packet loss. Since the compressed video streams are sensitive to data loss, the error resiliency of the encoded video becomes important. When video data is lost and retransmission is not possible, the missed data should be concealed. But loss concealment causes distortion in the lossy frame which also propagates into the next frames even if their data are received correctly. One promising solution to mitigate this error propagation is intra coding. There are three approaches for intra coding: intra coding of a number of blocks selected randomly or regularly, intra coding of some specific blocks selected by an appropriate cost function, or intra coding of a whole frame. But Intra coding reduces the compression ratio; therefore, there exists a trade-off between bitrate and error resiliency achieved by intra coding. In this paper, we study and show the best strategy for getting the best rate-distortion performance. Considering the error propagation, an objective function is formulated, and with some approximations, this objective function is simplified and solved. The solution demonstrates that periodical I-frame coding is preferred over coding only a number of blocks as intra mode in P-frames. Through examination of various test sequences, it is shown that the best intra frame period depends on the coding bitrate as well as the packet loss rate. We then propose a scheme to estimate this period from curve fitting of the experimental results, and show that our proposed scheme outperforms other methods of intra coding especially for higher loss rates and coding bitrates. Mohammad Kazemi 0002, Mohammed Ghanbari 0001, Shervin Shirmohammadi |
IEEE Trans. Multim. | 3 |
| 2020 | QoE-Fair DASH Video Streaming Using Server-side Reinforcement LearningabstractTo design an optimal adaptive video streaming method, video service providers need to consider both the efficiency and the fairness of the Quality of Experience (QoE) of their users. In Reference [8], we proposed a server-side QoE-fair rate adaptation method that considers both efficiency and fairness of the QoE. The server uses Reinforcement Learning (RL) to select a bitrate for each client sharing the same bottleneck link to the server in a way that achieves fairness among concurrent DASH clients and imposes that bitrate by dynamically modifying the client’s Media Presentation Description (MPD) file. In this article, we extend that work to minimize the number of actions the server needs to take to keep the system in its equilibrium state. By incorporating a Recurrent Neural Network, specifically an LSTM model, we modify the server’s training algorithm to achieve improvements in both the quality and the quantity of actions the server takes to guide the client. Performance evaluation of the modified algorithm for clients running both homogeneous and heterogeneous adaptation algorithms showed that the number of server actions dropped by 14% and 22%, respectively, while QoE-fairness improved by at least 6% and 10%, respectively. Sa'di Altamimi, Shervin Shirmohammadi |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | An Adaptive Two-Layer Light Field Compression Scheme Using GNN-Based ReconstructionabstractAs a new form of volumetric media, Light Field (LF) can provide users with a true six degrees of freedom immersive experience because LF captures the scene with photo-realism, including aperture-limited changes in viewpoint. But uncompressed LF data is too large for network transmission, which is the reason why LF compression has become an important research topic. One of the more recent approaches for LF compression is to reduce the angular resolution of the input LF during compression and to use LF reconstruction to recover the discarded viewpoints during decompression. Following this approach, we propose a new LF reconstruction algorithm based on Graph Neural Networks; we show that it can achieve higher compression and better quality compared to existing reconstruction methods, although suffering from the same problem as those methods—the inability to deal effectively with high-frequency image components. To solve this problem, we propose an adaptive two-layer compression architecture that separates high-frequency and low-frequency components and compresses each with a different strategy so that the performance can become robust and controllable. Experiments with multiple datasets 1 show that our proposed scheme is capable of providing a decompression quality of above 40 dB, and can significantly improve compression efficiency compared with similar LF reconstruction schemes. Xinjue Hu, Jingming Shan, Yu Liu 0001, Lin Zhang 0013, Shervin Shirmohammadi |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2019 | Client-server cooperative and fair DASH video streamingabstractAdaptive video streaming over HTTP, such as the MPEG-DASH standard, is now widely used by video service provides to stream their videos to users. But DASH and similar methods are known to suffer from two practical challenges: on the one hand, clients use fixed heuristics that limit their ability to generalize across network conditions, making the clients unable to efficiently predict variations in new networking environments, in turn leading to more buffering. On the other hand, the absence of collaboration among DASH clients leads to unfair bandwidth allocation, and typically pushes the system to an unbalanced equilibrium point. In this paper, we propose a server-side rate adaptation method that significantly improves the fairness of network bandwidth allocation among concurrent DASH users. We formulate the problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) model, and use Reinforcement Learning (RL) to train two neural networks to find an optimal solution to the fairness problem. Since our solution is implemented at the server side, it requires no modifications to the widely-installed DASH clients, making our solution very practical. We show that our proposed method outperforms the state-of-the-art schemes in terms of QoE-efficiency, QoE-fairness, and social welfare by as much as 16%, 21%, and 24% respectively. Sa'di Altamimi, Shervin Shirmohammadi |
NOSSDAV | 2 |
| 2018 | Designing Trainer's Manual for the ISG for Competence Project
Szilvia Paxian, Veronika Szücs, Shervin Shirmohammadi, Boris Abersek, Petya Grudeva, Karel Van Isacker, Tibor Guzsvinecz, Cecilia Sik-Lányi |
ICCHP (1) | 3 |
| 2018 | Joint Intra and Multiple Description Coding for Packet Loss Resilient Video TransmissionabstractMultiple description coding (MDC) is a technique for video transmission over error prone networks where the descriptions are routed over multiple paths. Intra coding such as MDC provides error resiliency but coding in this mode must be decided with care since it degrades the compression ratio. In this paper, we present our investigation results for a new intra coding approach in MDC. We have found that, in MDC streams, the best policy is to encode selective frames as I-frame instead of coding some macroblocks of frames in intra mode. In order to find the most suitable I-frame positions within a given video stream, we developed a cost function based on which intra/inter frame type is decided. The MDC scheme with the proposed intra coding criterion, with and without redundancy optimization, is implemented in the H.264/AVC reference software, JM16.0. Based on the experimental performance evaluation, we show that our method achieves higher average PSNR compared to the other optimized MDCs found in the literature. Mohammad Kazemi 0002, Razib Iqbal, Shervin Shirmohammadi |
IEEE Trans. Multim. | 3 |
| 2017 | SDN-enabled Game-Aware Network Management for Residential GatewaysabstractResidential gateways play a key role in providing internet access to home consumers. Nowadays, users in the same home with heterogeneous applications share a common gateway. As such, the gateway becomes the bandwidth bottleneck, leading to impairments and negatively affecting users' Quality of Experience (QoE). In the case of delay sensitive applications like video streaming and online gaming, this impairment becomes more crucial. In this paper, we present an SDN-enabled optimization-based scheme for optimally sharing the bandwidth among network flows within a residential gateway. We target online game flows and try to provide them with a higher QoE while not starving other traffic flows. Our optimization model considers the nature of network flows, and aims to maximize the bandwidth utilization. Our experimental results show improvements in bandwidth efficiency by almost 15% and in fairness by almost 11.5% among active network flows compared to conventional methods. In addition, the proposed method minimizes the overall delay experienced by players by almost half, outperforming the commonly-used max-min fairness and Round Robin schemes. Maryam Amiri, Hussein Al Osman, Shervin Shirmohammadi |
ISM | 3 |
| 2017 | Priced-Based Fair Bandwidth Allocation for Networked MultimediaabstractThe high demand of bandwidth from multimedia applications, specially video applications which consume the great majority of the Internet bandwidth, has caused a challenge for service providers and network operators. On the one hand, the allocation of bandwidth in a fair manner for multimedia users is necessary, so that the total utility of all users is maximized for higher quality of experience. On the other hand, optimizing the utilization of network resources such as maximizing throughput is also important for network operators to reduce cost and/or maximize profits. These two requirements could potentially be conflicting; hence, achieving both at the same time is challenging, and the reason why very few previous efforts have targeted this problem. Examples include Traffic Management Using Multipath Protocol (TRUMP) and Logarithmic-Based Multipath Protocol (LBMP), both of which achieve good results but are not without shortcomings. In this paper, we propose a Price-Based Fair Bandwidth Allocation (PBFA) method by implementing an optimized sending rate adaptation technique and combining it with an intuitive investment method to optimize the feedback prices to achieve efficient and fair bandwidth allocation. The results of our performance tests, using different simulations under different network topologies, show that PBFA achieves improvements of as much as 90% in fairness, 207% in throughput, and 91% in utility compared to TRUMP Hamed Hamzeh, Mahdi Hemmati, Shervin Shirmohammadi |
ISM | 3 |
| 2017 | A Bitrate-Conservative Fast-Adjusting Rate Controller for Video ConferencingabstractWidely-used Rate Control (RC) algorithms, such as those in the H.264 encoder, have certain shortcomings for time-sensitive applications such as High Definition Video Conferencing (HDVC): they either respond too slowly to available bandwidth variations, causing degradation in the perceived quality of the video session, or do not optimize video quality for a given available bandwidth. To overcome these shortcomings, we propose Dynamic Rate Control (DRC) which: 1- can adjust the bitrate of the video within a fast 4 frames or so 2- is conservative and does not waste bandwidth by unnecessarily increasing the video quality, instead saving the bandwidth as bursts for future frames, and 3- uses a moving window to limit the effect of past bursts on current bitrate. We implemented DRC in the x264 codec and used it in an actual video conferencing product from Magor Corp. The results showed that, compared to the widely-used ABR and CRF rate controllers, DRC provides better video quality and user experience, while adjusting the video bitrate faster. Abbas Javadtalab, Mona Omidyeganeh, Shervin Shirmohammadi, Mojtaba Hosseini |
ISM | 3 |
| 2017 | A Cloud-Based Multi-threaded Implementation of View Synthesis SystemabstractIn multiview video applications, view synthesis is a computationally intensive task that needs to be done correctly and efficiently in order to deliver a seamless user experience. In order to provide fast and efficient view synthesis, in this paper, we present a cloud-based implementation that will be especially beneficial to mobile users whose devices may not be powerful enough for high quality view synthesis. Our proposed implementation balances the view synthesis algorithm's components across multiple threads and utilizes the computational capacity of modern CPUs for faster and higher quality view synthesis. For arbitrary view generation, we utilize the depth map of the scene from the cameras' viewpoint and estimate the depth information conceived from the virtual camera. The estimated depth is then used in a backward direction to warp the cameras' image onto the virtual view. Finally, we use a depth-aided inpainting strategy for the rendering step to reduce the effect of disocclusion regions (holes) and to paint the missing pixels. For our cloud implementation, we employed an automatic scaling feature to offer elasticity in order to adapt the service load according to the fluctuating user demands. Our performance results using 4 multiview videos over 2 different scenarios show that our proposed system achieves average improvement of around 1.7 times for speedup, 50% and 52% for efficiency and resource utilization, respectively. Parvaneh Pouladzadeh, Razib Iqbal, Shervin Shirmohammadi, Omid Fatemi |
ISM | 3 |
| 2017 | Sports VR Content Generation from Regular Camera FeedsabstractWith the recent availability of commodity Virtual Reality (VR) products, immersive video content is receiving a significant interest. However, producing high-quality VR content often requires upgrading the entire production pipeline, which is costly and time-consuming. In this work, we propose using video feeds from regular broadcasting cameras to generate immersive content. We utilize the motion of the main camera to generate a wide-angle panorama. Using various techniques, we remove the parallax and align all video feeds. We then overlay parts from each video feed on the main panorama using Poisson blending. We examined our technique on various sports including basketball, ice hockey and volleyball. Subjective studies show that most participants rated their immersive experience when viewing our generated content between Good to Excellent. In addition, most participants rated their sense of presence to be similar to ground-truth content captured using a GoPro Omni 360 camera rig. Kiana Calagari, Mohamed A. Elgharib, Shervin Shirmohammadi, Mohamed Hefeeda |
ACM Multimedia | 3 |
| 2017 | Towards QoE-aware HAS video streaming over LTEabstractLately, HTTP Adaptive Streaming (HAS) has become dominant among different video streaming techniques. However, owing to the unpredictable nature of the wireless radio channel and device mobility, using HAS over mobile wireless networks is still very challenging. In this paper, we propose a novel Quality of Experience (QoE) optimization mechanism for HAS in the context of mobile wireless networks. The proposed mechanism leverages recent advances in HAS specification, which includes new features for QoE measurements and reporting. First, we formulate a discrete optimization problem aiming at maximizing the overall average quality, and minimizing the negative impact of temporal video quality changes for all HAS users simultaneously. Second, in order to take advantage of well-known continuous optimization techniques and to decrease the computational complexity, we convert the formulated problem into a continuous form, and propose a gradient based algorithm to solve the continuous optimization problem. The results of our simulations demonstrate that our system attained better perceived video quality by almost 8% on average, while lowering the freezing period by 20% on average across HAS users when compared to other approaches where HAS users only rely on local adaptation logics. Ashkan Sobhani, Abdulsalam Yassine, Shervin Shirmohammadi |
PIMRC | 3 |
| 2017 | An intelligent cloud-based data processing broker for mobile e-health multimedia applications
Sri Vijay Bharat Peddi, Pallavi Kuhad, Abdulsalam Yassine, Parisa Pouladzadeh, Shervin Shirmohammadi, Ali A. Nazari Shirehjini |
Future Gener. Comput. Syst. | 5 |
| 2017 | Redundancy Allocation Based on the Weighted Mismatch-Rate Slope for Multiple Description Video CodingabstractMultiple description coding (MDC) is a robust coding technique for video transmission over error prone networks, whereby the video is encoded into multiple descriptions with some redundancy between the descriptions. This redundancy leads to error resiliency in the case of packet loss during the network transport. However, the amount of this redundancy has a critical role in MDC performance. Therefore, a crucial problem in MDC is to find what the optimum amount of redundancy budget is, and then how this redundancy budget can be optimally allocated to the frames. To solve this problem, we propose a scheme in which the redundancy budget is allocated to the frames based on the weighted mismatch-rate slopes so that this additional bitrate can attain maximum distortion reduction. The redundancy is added gradually so that fine tuning of the utilized bitrate is achievable. We have verified our proposed scheme by implementing it in H.264/AVC reference software JM16.0, and running experiments against two representative reference methods. Our experiments show that our scheme not only minimizes the end-to-end distortion with a rate-distortion performance that is better than the reference methods, especially for high PLRs, but also entirely uses the available bandwidth, unlike the reference methods. Mohammad Kazemi 0002, Razib Iqbal, Shervin Shirmohammadi |
IEEE Trans. Multim. | 3 |
| 2017 | Mobile Multi-Food Recognition Using Deep LearningabstractIn this article, we propose a mobile food recognition system that uses the picture of the food, taken by the user’s mobile device, to recognize multiple food items in the same meal, such as steak and potatoes on the same plate, to estimate the calorie and nutrition of the meal. To speed up and make the process more accurate, the user is asked to quickly identify the general area of the food by drawing a bounding circle on the food picture by touching the screen. The system then uses image processing and computational intelligence for food item recognition. The advantage of recognizing items, instead of the whole meal, is that the system can be trained with only single item food images. At the training stage, we first use region proposal algorithms to generate candidate regions and extract the convolutional neural network (CNN) features of all regions. Second, we perform region mining to select positive regions for each food category using maximum cover by our proposed submodular optimization method. At the testing stage, we first generate a set of candidate regions. For each region, a classification score is computed based on its extracted CNN features and predicted food names of the selected regions. Since fast response is one of the important parameters for the user who wants to eat the meal, certain heavy computational parts of the application are offloaded to the cloud. Hence, the processes of food recognition and calorie estimation are performed in cloud server. Our experiments, conducted with the FooDD dataset, show an average recall rate of 90.98%, precision rate of 93.05%, and accuracy of 94.11% compared to 50.8% to 88% accuracy of other existing food recognition systems. Parisa Pouladzadeh, Shervin Shirmohammadi |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | A Video Bitrate Adaptation and Prediction Mechanism for HTTP Adaptive StreamingabstractThe Hypertext Transfer Protocol (HTTP) Adaptive Streaming (HAS) has now become ubiquitous and accounts for a large amount of video delivery over the Internet. But since the Internet is prone to bandwidth variations, HAS's up and down switching between different video bitrates to keep up with bandwidth variations leads to a reduction in Quality of Experience (QoE). In this article, we propose a video bitrate adaptation and prediction mechanism based on Fuzzy logic for HAS players, which takes into consideration the estimate of available network bandwidth as well as the predicted buffer occupancy level in order to proactively and intelligently respond to current conditions. This leads to two contributions: First, it allows HAS players to take appropriate actions, sooner than existing methods, to prevent playback interruptions caused by buffer underrun, reducing the ON-OFF traffic phenomena associated with current approaches and increasing the QoE. Second, it facilitates fair sharing of bandwidth among competing players at the bottleneck link. We present the implementation of our proposed mechanism and provide both empirical/QoE analysis and performance comparison with existing work. Our results show that, compared to existing systems, our system has (1) better fairness among multiple competing players by almost 50% on average and as much as 80% as indicated by Jain's fairness index and (2) better perceived quality of video by almost 8% on average and as much as 17%, according to the estimate the Mean Opinion Score (eMOS) model. Ashkan Sobhani, Abdulsalam Yassine, Shervin Shirmohammadi |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2016 | Target Group Questionnaire in the "ISG for Competence" Project
Szilvia Paxian, Veronika Szücs, Shervin Shirmohammadi, Boris Abersek, Andrean Lazarov, Karel Van Isacker, Cecilia Sik-Lányi |
ICCHP (2) | 3 |
| 2016 | Game-Aware Resource Manager for Home GatewaysabstractOnline gaming, especially the new paradigm of Cloud Gaming, is advancing the state-of-the-art in interactive networked applications while creating significant business opportunities forgame providers. However, online gaming is fundamentallychallenged by network latency that impairs the interactive gamingexperience and negatively affects players' Quality of Experience(QoE). In this work, we propose a method to optimize bandwidthdistribution on the player's home network to maximize QoE without starving other concurrent applications. Our experimental resultsshow that the proposed model cuts the delay experienced by users by almost half and improves the user's QoE by enhancing game video quality. Maryam Amiri, Karun Pal Singh Malik, Hussein Al Osman, Shervin Shirmohammadi |
ISM | 4 |
| 2016 | Datacenter Traffic Shaping for Delay Reduction in Cloud GamingabstractCloud Gaming enables users to play games using a thin-client, regardless of their location or what platform they use (PCs, laptops, tablets, smartphones). Since the major computational parts of game processing are performed in datacenters, effectively assigning the resources (e.g. memory, bandwidth) to gaming sessions plays a key role in providing a high quality gaming experience to end-users. In this paper, we propose a traffic policing and shaping method using the Software Defined Networking (SDN) paradigm to solve the bandwidth allocation problem in cloud gaming datacenter networks. Our proposed method considers the current status of the datacenter paths in terms of remaining bandwidth and delay to achieve fair bandwidth allocation. The proposed scheme optimizes bandwidth utilization while ensuring compliance with the threshold of tolerable delay in cloud gaming systems. Our experimental results show that the proposed method improves bandwidth utilization and reduces end-to-end delay and delay variation (jitter) by almost 12% and 9%, respectively, without engendering additional packet loss compared to a representative conventional method: Equal Cost Multi-path (ECMP). These reductions lead to improvements in players' gaming experience. Maryam Amiri, Hussein Al Osman, Shervin Shirmohammadi |
ISM | 3 |
| 2016 | Fair and Efficient Bandwidth Allocation for Video Flows Using Sigmoidal ProgrammingabstractThe problem of bandwidth allocation in networks is traditionally solved using distributed rate allocation algorithms under the general framework of Network Utility Maximization(NUM). Despite many advances in solving the flow assignment problem in NUM, the common but unrealistic assumption of concavity of utility functions undermines the performance of existing systems in providing satisfactory QoE to the consumers of video traffic, the utility function of which is not concave, but sigmoidal. In this work, we model the bandwidth allocation problem as a nonconvex "Sigmoidal Programming" optimization problem and use an approximation algorithm to solve it while guaranteeing a suboptimal solution. Our simulation results for video streaming over two realistic network topologies indicate improvements of at least 50% in average utility, up to 29% in fairness, and on average 14% less network capacity usage, all compared to two existing representative methods: Proportional Fair and Max-Min Fair. Mahdi Hemmati, Bill McCormick, Shervin Shirmohammadi |
ISM | 3 |
| 2016 | A Selective Intra-Coding Approach for Multiple Description Video CodingabstractMultiple Description Coding (MDC) is a technique for video transmission over error prone networks where the descriptions are routed over multiple paths. Intra coding technique and MDC both provides error resiliency for real-time video transport over unreliable networks. In this paper, we present our investigation results for a new intra coding approach for MDC where selective frames are fully encoded in intra mode instead of coding selective macroblocks in intra mode in many frames. We implemented our scheme in the H.264/AVC reference software, JM16.0. Based on the experimental performance evaluation, we show that our method achieves higher average PSNR compared to the other optimized MDC schemes available in the literature. Mohammad Kazemi 0002, Razib Iqbal, Shervin Shirmohammadi |
ISM | 3 |
| 2016 | QoE-Driven Optimization for DASH Service in Wireless NetworksabstractIn this paper, we propose a novel QoE optimization mechanism for Dynamic Adaptive Streaming over HTTP (DASH) in the context of wireless mobile networks. The proposed mechanism leverages recent advances in 3GPP DASH specification, which includes new features for QoE measurements and reporting. The proposed optimization mechanism has two objective functions, the first function maximizes the overall average QoE among DASH clients while the second function minimizes the negative impact of temporal video quality changes, i.e. the up and down switching between different representation during playback. The results of our simulations demonstrate that the proposed method improves the overall QoE and outperform other approaches where DASH clients only rely on local adaptation logic. Ashkan Sobhani, Abdulsalam Yassine, Shervin Shirmohammadi |
ISM | 3 |
| 2016 | GSET somi: a game-specific eye tracking dataset for somiabstractIn this paper, we present an eye tracking dataset of computer game players who played the side-scrolling cloud game Somi. The game was streamed in the form of video from the cloud to the player. This dataset can be used for designing and testing game-specific visual attention models. The source code of the game is also available to facilitate further modifications and adjustments. For collecting this data, male and female candidates were asked to play the game in front of a remote eye-tracking device. For each player, we recorded gaze points, video frames of the gameplay, and mouse and keyboard commands. For each video frame, a list of its game objects with their locations and sizes was also recorded. This data, synchronized with eye-tracking data, allows one to calculate the amount of attention that each object or group of objects draw from each player. As a benchmark, we also show various attention patterns could be identified among players. Hamed Ahmadi, Saman Zad Tootaghaj, Sajad Mowlaei, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
MMSys | 5 |
| 2016 | Scalable multiview video coding for immersive video streaming systemsabstractImmersive video places the user inside the video scene, allowing the user to control the direction of the view. To achieve this, the view of every direction must be recorded using either a panoramic camera or multiple cameras placed at different positions with different angels. The size of the captured video can be quite large due to multiple video streams, one from each camera. Even with compression standards such as Multiview Video Coding (MVC), the transmission of the whole MVC video is still bandwidth-costly, especially for heterogeneous users whose bandwidths vary. In this paper, we present a new approach for immersive video streaming by using Scalable Multiview Video Coding (SMVC) to create multiple layers of the immersive video, supporting heterogeneous receivers more efficiently. Our method limits the number of views in its base layer, while it uses view scalability and free view-point scalability in the additional layers to synthesize more views at the receiver and provide high quality free view-point viewing to the user. Performance evaluations demonstrate that our method: 1-synthesizes missing views more accurately, as evident subjectively, and 2-achieves an average and maximum gain of 0.75 and 1.4 in Bjontegaard BD-Bitrate scale, respectively, compared to existing work which simply group adjacent views in the same layer. Hoda Roodaki, Shervin Shirmohammadi |
VCIP | 2 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 26 |
| 2016 | A View-Level Rate Distortion Model for Multi-View/3D VideoabstractMulti-view/3D video is currently available in games, entertainment, education, security, and surveillance applications . Since the amount of data in multi-view/3D increases proportionally with the number of cameras, and due to different bandwidth and playback capabilities of receivers, appropriate compression of multi-view/3D video to produce the correct bitrate while maintaining smooth video quality is crucial, a task that is mostly performed by the rate control module of the encoder. There are many existing rate control algorithms for single-view and multi-view video coding considering the specific features or aspects of these videos. In this paper, we introduce a novel view-level rate distortion (RD) model. We use a systematic methodology to derive this RD model by investigating the impact of multi-view/3D video characteristics on the bitrate of a compressed video. Our proposed RD model considers the concepts of intra-view and inter-view disparity as an effective feature of multi-view/3D video to estimate the overall bitrate of each view more accurately. Evaluation results indicate that our proposed view-level RD model outperforms existing linear models by a factor of 3 and can predict the rate of each view with relatively high precision and a low estimation error of 12% on average. Hoda Roodaki, Zahra Iravani, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
IEEE Trans. Multim. | 4 |
| 2016 | Toward Delay-Efficient Game-Aware Data Centers for Cloud GamingabstractGaming on demand is an emerging service that has recently started to garner prominence in the gaming industry. Cloud-based video games provide affordable, flexible, and high-performance solutions for end-users with constrained computing resources and enables them to play high-end graphic games on low-end thin clients. Despite its advantages, cloud gaming's Quality of Experience (QoE) suffers from high and varying end-to-end delay. Since the significant part of computational processing, including game rendering and video compression, is performed in data centers, controlling the transfer of information within the cloud has an important impact on the quality of cloud gaming services. In this article, a novel method for minimizing the end-to-end latency within a cloud gaming data center is proposed. We formulate an optimization problem for reducing delay, and propose a Lagrangian Relaxation (LR) time-efficient heuristic algorithm as a practical solution. Simulation results indicate that the heuristic method can provide close-to-optimal solutions. Also, the proposed model reduces end-to-end delay and delay variation by almost 11% and 13.5%, respectively, and outperforms the existing server-centric and network-centric models. As a byproduct, our proposed method also achieves better fairness among multiple competing players by almost 45%, on average, in comparison with existing methods. Maryam Amiri, Hussein Al Osman, Shervin Shirmohammadi, Maha Abdallah |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2015 | An Open Source Cloud Gaming Testbed Using DirectShowabstractDespite its challenges, cloud gaming is growing its share in the gaming market by attracting more players. This has led to an increasing number of researches trying to overcome cloud gaming's challenges, including the required high bandwidth and low latency, to make cloud gaming more practical and profitable. To perform this research, researchers need a testbed to evaluate their ideas and find the best solutions. Currently, GamingAnywhere is the only open source platform and testbed to serve this goal. However, it cannot be used to stream all video games, since it depends on hooking APIs which might be incompatible with some video games. In this paper, we introduce a new open source cloud gaming testbed. In this testbed, the screen capturing module is fundamentally a DirectShow filter and, hence, can be tuned for any DirectShow compatible video game. The testbed also facilitates the measurement of delay and quality as the video is processed through its modules. Hamed Ahmadi, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
CloudCom | 3 |
| 2015 | A Dynamic Alpha Congestion Controller for WebRTCabstractVideo conferencing applications have significantly changed the way people communicate over the Internet. Web Real-Time Communication (WebRTC), drafted by the World Wide Web Consortium (W3C) and Internet Engineering Task Force (IETF) working groups, has added new functionality to the web browsers, allowing audio/video calls between browsers without the need to install any video telephony applications. The Google Congestion Control (GCC) algorithm has been proposed as WebRTC's congestion control mechanism, but its performance is limited due to using a fixed incoming rate decrease factor, known as alpha (a). In this paper, we propose a dynamic alpha model to reduce the available receiving bandwidth estimate during overuse as indicated by the over-use detector. Experiments using our specific testbed show that our proposed model achieves a 33% higher incoming rate and a 16% lower round-trip time, while keeping a similar packet loss rate and video quality, compared to a fixed alpha model. Rasha Atwah, Razib Iqbal, Shervin Shirmohammadi, Abbas Javadtalab |
ISM | 3 |
| 2015 | An SDN Controller for Delay and Jitter Reduction in Cloud GamingabstractCloud gaming is an emerging service that has recently started to garner prominence in the gaming industry. Since the significant part of computational processing, including game rendering and video compression, is performed in data centers, controlling the transfer of information within the cloud has an important impact on the quality of cloud gaming services. In this paper, we make two contributions: we propose a design to apply the recent paradigm of Software Defined Networks (SDNs) to Cloud Gaming, and we propose an SDN controller that reduces end-to-end delay and delay variations experienced by players. Our SDN controller adaptively disperses the game traffic load among different network paths according to their corresponding end-to-end delays. Experimental results show that our proposed controller reduces end-to-end delay and delay variation by almost 9% and 50% respectively without engendering additional packet loss, compared to a representative conventional method: Open Shortest Path First (OSPF). These reductions lead to improvements in players' gaming experience. Maryam Amiri, Hussein Al Osman, Shervin Shirmohammadi, Maha Abdallah |
ACM Multimedia | 3 |
| 2015 | A fuzzy-based rate adaptation controller for DASHabstractAs dynamic delivery of video over HTTP becomes prominent, rate adaptation techniques become more challenging due to bandwidth variations. This paper presents a Fuzzy-based controller to dynamically adapt the video bitrate based on both the estimated throughput and the size of the playback buffer. The proposed Fuzzy Logic Controller (FLC) mechanism takes the observed throughput and buffer dynamics as inputs to change the policy of selecting the video bitrate and download scheduling to minimize the negative effect of ON-OFF switching. The experimental results show that our proposed mechanism generates a smoother stream compared to existing methods. Ashkan Sobhani, Abdulsalam Yassine, Shervin Shirmohammadi |
NOSSDAV | 3 |
| 2015 | Cloud-based SVM for food categorization
Parisa Pouladzadeh, Shervin Shirmohammadi, Aslan Bakirov, Ahmet Bulut, Abdulsalam Yassine |
Multim. Tools Appl. | 2 |
| 2015 | Erratum to: Cloud-based SVM for food categorization
Parisa Pouladzadeh, Shervin Shirmohammadi, Aslan Bakirov, Ahmet Bulut, Abdulsalam Yassine |
Multim. Tools Appl. | 2 |
| 2015 | A high capacity data hiding algorithm for H.264/AVC videoabstractAbstract This article presents an information hiding algorithm for H.264/AVC video stream. It utilizes position of the last nonzero level of quantized discrete cosine transform block to embed information bits. Because only the high‐frequency levels are changed, it can guarantee a high peak signal‐to‐noise ratio and slight increase in bit rate after the watermark embedding. The extraction process is not complex; thus, the proposed technique is an excellent solution for real‐time applications such as broadcasting. Experimental results on several test sequences demonstrate that the proposed approach can realize blind extraction with real‐time performance; it also provides very high capacity, low distortion and increase in bit rate by about 0.5%. Copyright © 2015 John Wiley & Sons, Ltd. Mehdi Fallahpour, Shervin Shirmohammadi, Mohammed Ghanbari 0001 |
Secur. Commun. Networks | 2 |
| 2015 | Rate/distortion optimization in multiple description video coding
Mohammad Kazemi 0002, Khosrow Haj Sadeghi, Shervin Shirmohammadi, Payman Moallem |
Signal Process. Image Commun. | 3 |
| 2015 | Video Encoding Acceleration in Cloud GamingabstractCloud computing provides reliable, affordable, and flexible resources for many applications and users with constrained computing resources and capabilities. The cloud computing concept is becoming an appealing paradigm for many industries including the gaming industry, leading to the introduction of cloud gaming architectures. Despite its advantages, cloud gaming suffers from unguaranteed end-to-end delay as well as server side's computational complexity. In this paper, a novel algorithm for reducing the computational complexity and hence speeding up the video encoding speed is proposed. Specifically, by performing minimum modifications in the game engine and the video codec, some information from the game engine is fed into the video encoder to bypass the motion estimation (ME) process. Our results show that the proposed method achieves up to 39% speedup in the ME process, leading to a 24% acceleration in the total encoding process. Mehdi Semsarzadeh, Abdulsalam Yassine, Shervin Shirmohammadi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Introduction to the Special Section on Visual Computing in the Cloud: Cloud Gaming and VirtualizationabstractCloud gaming, the newest entry in the online gaming world, leverages the well-known concept of cloud computing to provide real-time gaming services to players. The idea in cloud gaming is to capture the game events from players and transmit them to the cloud, process those events and run the game logic in the cloud, render the game scene as video in the cloud, and stream that video to the players. The advantage is that as long as the client can display video, which pretty much all smartphones, tablets, game consoles, desktops, laptops, and mobile devices today do, the user can play the game without installing it locally, and without needing to have a machine with high-grade 3-D graphics rendering and powerful computing hardware and software. This makes cloud gaming accessible to a huge market of mass consumers. While some variations of cloud gaming systems stream 3-D graphics, in addition to or instead of video, the great majority of cloud gaming implementations are video based. Using the well-known concept of software as a service, cloud gaming is also sometimes referred to as gaming as a service, which is already available as commercial products, such as Sony’s PlayStation Now, Ubitus’s GameNow, G-Cluster, Crytek’s GFACE, PlayGiga, and LiquidSky, to name a few. There are also many efforts concentrating specifically on the underlying technology behind cloud gaming, such as NVIDIA’s Grid, OTOY, CiiNow, Kalydo, and GamingAnywhere [4], the latter being the only open source and free technology. Microsoft is also exploring cloud gaming technologies, with recent successes such as its Kahawai project [1]. Shervin Shirmohammadi, Maha Abdallah, Dewan Tanvir Ahmed, Kuan-Ta Chen, Yan Lu 0001, Alex Snyatkov |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Decoder-Complexity-Aware Encoding of Motion Compensation for Multiple Heterogeneous ReceiversabstractFor mobile multimedia systems, advances in battery technology have been much slower than those in memory, graphics, and processing power, making power consumption a major concern in mobile systems. The computational complexity of video codecs, which consists of CPU operations and memory accesses, is one of the main factors affecting power consumption. In this article, we propose a method that achieves near-optimal video quality while respecting user-defined bounds on the complexity needed to decode a video. We specifically focus on the motion compensation process, including motion vector prediction and interpolation, because it is the single largest component of computation-based power consumption. We start by formulating a scenario with a single receiver as a rate-distortion optimization problem and we develop an efficient decoder-complexity-aware video encoding method to solve it. Then we extend our approach to handle multiple heterogeneous receivers, each with a different complexity requirement. We test our method experimentally using the H.264 standard for the single receiver scenario and the H.264 SVC extension for the multiple receiver scenario. Our experimental results show that our method can achieve up to 97% of the optimal solution value in the single receiver scenario, and an average of 97% of the optimal solution value in the multiple receiver scenario. Furthermore, our tests with actual power measurements show a power saving of up to 23% at the decoder when the complexity threshold is halved in the encoder. Mohsen Jamali Langroodi, Joseph G. Peters, Shervin Shirmohammadi |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2014 | Rate-distortion optimization for scalable multi-view video codingabstractIn recent years, multi-view/3D video applications, such as three-dimensional television (3DTV) and free-viewpoint television (FTV), have drawn increasing attention. Since the amount of data that has to be stored or transmitted increases proportionally with the number of cameras, efficient compression of multi-view/3D video is crucial. Scalable multi-view video coding is one of the methods to address this challenge. But, in streaming multi-view/3D video over a network to heterogeneous receivers, efficient video compression while maintaining a high quality of received video is very challenging. This paper presents a novel method for rate-distortion optimization in scalable multiview video. We apply the Karush-Kuhn-Tucker (KKT) conditions in minimizing the perceptual distortion of decoded video under the conditions that the sum of bits generated from different views is constrained within a given bit budget. Since the constraint-based optimization problem is usually computational intensive, our proposed approach considers the concept of disparity between layers and disparity between views to reduce this computational complexity. Simulation results indicate that the proposed approach is able to meet network bandwidth limitations with acceptable overall video quality. Hoda Roodaki, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
ICME | 3 |
| 2014 | Demo Paper: A Fast-Adjusting Rate Control Algorithm Using Network-Assisted Scheme for HD Video StreamingabstractWe present the integration of our Dynamic Rate Control (DRC) method with the Network Assisted Dynamic Adaptation (NADA) method, and their deployment in a High Definition Video Conferencing application over best effort networks (Internet). DRC offers the best frame quality at a given bitrate, and also jumps to new video bitrate on the fly, within 2-6 frames, which is much faster than other rate controllers, whereas NADA can detect congestion. By combining DRC with NADA, we have designed a complete solution that maintains the video bitrate according to the available and changing bandwidth quickly and with minimal effect on perceived quality. The demonstration shows that this solution provides a fast, efficient and adaptive video conferencing system. Abbas Javadtalab, Shervin Shirmohammadi |
ISM | 3 |
| 2014 | YawDD: a yawning detection datasetabstractIn this paper, we present two video datasets of drivers with various facial characteristics, to be used for designing and testing algorithms and models for yawning detection. For collecting these videos, male and female candidates were asked to sit in the driver's seat of a car. The videos are taken in real and varying illumination conditions. In the first dataset, the camera is installed under the front mirror of the car. Each participant has three or four videos and each video contains different mouth conditions such as normal, talking/singing, and yawning. In the second dataset, the camera is installed on the dash in front of the driver, and each participant has one video with the above-mentioned different mouth conditions all in the same video. The car was parked for both datasets to keep the environment safe for the participants. As a benchmark, we also present the results of our own yawning detection method, and show that we can achieve a much higher accuracy in the scenario with the camera installed on the dash in front of the driver. Shabnam Abtahi, Mona Omidyeganeh, Shervin Shirmohammadi, Behnoosh Hariri |
MMSys | 3 |
| 2014 | Complexity Aware Encoding of the Motion Compensation Process of the H.264/AVC Video Coding StandardabstractAdvances in battery technology have not kept pace with other recent advances in mobile multimedia systems with the result that power consumption is a major concern. The computational complexity of video codecs, which consists of CPU operations and memory accesses, is one of the main factors affecting power consumption. In this paper, we propose a method that achieves good video quality while at the same time guaranteeing that the complexity needed to decode the video does not exceed a specific threshold defined by a user. We focus on the motion compensation process, including motion vector prediction and interpolation, which is the biggest single component in computation-based power consumption. We formulate the rate-distortion optimization problem and present an efficient method for decoder complexity-aware video encoding in the H.264 video codec. Our results show that our method can achieve up to 95% of the optimal solution value. Mohsen Jamali Langroodi, Joseph G. Peters, Shervin Shirmohammadi |
NOSSDAV | 3 |
| 2014 | A generic, comprehensive and granular decoder complexity model for the H.264/AVC standard
Mehdi Semsarzadeh, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | A game attention model for efficient bit rate allocation in cloud gaming
Hamed Ahmadi, Saman Zad Tootaghaj, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
Multim. Syst. | 4 |
| 2014 | A review of multiple description coding techniques for error-resilient video delivery
Mohammad Kazemi 0002, Shervin Shirmohammadi, Khosrow Haj Sadeghi |
Multim. Syst. | 2 |
| 2014 | Utility based decision support engine for camera view selection in multimedia surveillance systems
Dewan Tanvir Ahmed, M. Anwar Hossain 0001, Shervin Shirmohammadi, Abdullah Sharaf Alghamdi, Pradeep K. Atrey, Abdulmotaleb El Saddik |
Multim. Tools Appl. | 3 |
| 2014 | Design and implementation of a system for body posture recognition
Ali A. Nazari Shirehjini, Abdulsalam Yassine, Shervin Shirmohammadi |
Multim. Tools Appl. | 3 |
| 2014 | ALP: Adaptive Loss Protection Scheme with Constant Overhead for Interactive Video ApplicationsabstractThere has been an increasing demand for interactive video transmission over the Internet for applications such as video conferencing, video calls, and telepresence applications. These applications are increasingly moving towards providing High Definition (HD) video quality to users. A key challenge in these applications is to preserve the quality of video when it is transported over best-effort networks that do not guarantee lossless transport of video packets. In such conditions, it is important to protect the transmitted video by using intelligent and adaptive protection schemes. Applications such as HD video conferencing require live interaction among participants, which limits the overall delay the system can tolerate. Therefore, the protection scheme should add little or no extra delay to video transport. We propose a novel Adaptive Loss Protection (ALP) scheme for interactive HD video applications such as video conferencing and video chats. This scheme adds negligible delay to the transmission process and is shown to achieve better quality than other schemes in lossy networks. The proposed ALP scheme adaptively applies four different protection modes to cope with the dynamic network conditions, which results in high video quality in all network conditions. Our ALP scheme consists of four protection modes ; each of these modes utilizes a protection method . Two of the modes rely on the state-of-the-art protection methods, and we propose a new Integrated Loss Protection (ILP) method for the other two modes. In the ILP method we integrate three factors for distributing the protection among packets. These three factors are error propagation, region of interest and header information. In order to decide when to switch between the protection modes, a new metric is proposed based on the effectiveness of each mode in performing protection, rather than just considering network statistics such as packet loss rate. Results show that by using this metric not only the overall quality will be improved but also the variance of quality will decrease. One of the main advantages of the proposed ALP scheme is that it does not increase the bit rate overhead in poor network conditions. Our results show a significant gain in video quality, up to 3dB PSNR improvement is achieved using our scheme, compared to protecting all packets equally with the same amount of overhead. Kiana Calagari, Mohammad Reza Pakravan, Shervin Shirmohammadi, Mohamed Hefeeda |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2013 | Energy-budget-compliant adaptive 3D texture streaming in mobile gamesabstractAdvances in computing hardware and novel multimedia applications have urged the development of handheld mobile devices such as smartphones and PDAs. Amongst the most used applications on handheld devices are mobile 3D graphics such as 3D games and 3D virtual environments. With this significant increase of mobile applications and games, one of the challenges is how to efficiently transmit the bulky 3D information to resource-constrained mobile devices. Despite the many attractive features, 3D graphics impose significant demands on the limited battery capacity of mobile devices. Thus the development of efficient approaches to decrease the amount of streamed data with the aim of increasing the battery lifetime has become a key research topic. Mohammad Hosseini 0002, Joseph G. Peters, Shervin Shirmohammadi |
MMSys | 3 |
| 2013 | Game as video: bit rate reduction through adaptive object encodingabstractWide-spread availability of broadband internet access and the ubiquity of thin smart end devices such as smart-phones and tablets have led to a trend of moving more services away from the end devices to data centers, commonly referred to as Cloud Computing. For gaming, the stringent network bandwidth requirements of Cloud Gaming and the rapid growth of massively multiplayer online games (MMOG) call for novel content encoding and scene customization schemes to control the bit rate of the streaming video of the game scene. In this paper, a first step has been taken towards this goal by presenting a selective object encoding method to reduce the required network bandwidth and processing power without much impact on the player's quality of experience. Mahdi Hemmati, Abbas Javadtalab, Ali A. Nazari Shirehjini, Shervin Shirmohammadi, Tarik Arici |
NOSSDAV | 4 |
| 2013 | Continuous one-way available bandwidth change detection in high definition video conferencingabstractIn order to deal, as fast as possible, with available bandwidth variations in High Definition Video Conferencing over best-effort networks, we propose a Bayesian instantaneous end-to-end bandwidth change discovery mechanism. Most other congestion detection mechanisms use network parameters such as packet loss probability, round trip time or jitter. On the other hand, our approach uses weighted inter-arrival time of video packets at the receiver in order to detect bandwidth variations more accurately and more quickly. Our approach is continuous, since it monitors available bandwidth with each incoming video packet, and therefore detects congestion occurrence in less than 200 ms, on average, which is significantly faster than existing RTCP-based approaches. It is also one-way, because it only takes into account the characteristics of the incoming path and not the outgoing path, as opposed to other approaches which use round trip time. Aziz Khanchi, Mehdi Semsarzadeh, Abbas Javadtalab, Shervin Shirmohammadi |
NOSSDAV | 4 |
| 2013 | Modeling and Evaluation of a Metadata-Based Adaptive P2P Video-Streaming SystemabstractIn this paper, we present a multi-parent adaptive video-streaming system (MAVSS). MAVSS is a cooperative video-streaming system, based on the peer-to-peer (P2P) content distribution concept, to simultaneously adapt and stream video contents to heterogeneous users. In order to ensure video codec independence, we emphasize on structured, metadata-based adaptation. Therefore, MPEG-21 generic Bitstream syntax description is chosen to describe the parts of the video contents selected for adaptation operations. Additionally, without an optimal solution, it is hard to judge the efficiency of such systems. Therefore, in this paper, we first present the fundamental properties of MAVSS. We then present a mathematical model of the adaptive streaming system, where we define efficient streaming as an optimization of a cost function that can be solved as an integer linear programming problem. The model illustrates the relation among all the parameters that affect the resource contribution, resource utilization, load balancing and service fairness. We use this model to analyze the trade-offs that exist between service fairness and system efficiency. Razib Iqbal, Shervin Shirmohammadi, Behnoosh Hariri |
Comput. J. | 2 |
| 2013 | Guest editorial for special issue on network and systems support for games
Shervin Shirmohammadi, Carsten Griwodz, Grenville J. Armitage |
Multim. Syst. | 1 |
| 2013 | Application of 3D-wavelet statistics to video analysis
Mona Omidyeganeh, Shahrokh Ghaemmaghami, Shervin Shirmohammadi |
Multim. Tools Appl. | 3 |
| 2013 | A fine-grain distortion and complexity aware parameter tuning model for the H.264/AVC encoder
Mehdi Semsarzadeh, Atieh Lotfi, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
Signal Process. Image Commun. | 4 |
| 2012 | Complexity Modeling of the Motion Compensation Process of the H.264/AVC Video Coding StandardabstractWith recent advances in computing and communication technologies, ubiquitous access to high quality multimedia content such as high definition video using smart phones, Net books, or tablets is a fact of our daily life. However, power is still a major concern for any mobile device, and requires optimization of power consumption using a power model for each multimedia application, such as a video decoder. In this paper, a generic decoding complexity model for the motion compensation (MC) process, which constitutes up to 25% of the computational complexity and hence power consumption of an H.264/AVC decoder, has been proposed. For the model to remain independent from a specific implementation or platform, it has been developed by analysing the MC algorithm as described in the standard. Simulation results indicate that the proposed model estimates MC complexity with an average accuracy of 95.63%, for a wide range of test sequences using both JM and x.264 software implementations of H.264/AVC. For a dedicated hardware implementation of the MC module the modeling accuracy is around 89.61%, according to our simulation results. It should be noted that in addition to power consumption control, the proposed model can be used for designing a receiver-aware H.264/AVC encoder, where the complexity constraints of the receiver side are taken into account during compression. Mehdi Semsarzadeh, Mohsen Jamali Langroodi, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
ICME | 4 |
| 2012 | ROI-based protection scheme for high definition interactive video applicationsabstractIn this work, first, a Region of Interest (ROI)-based Unequal Loss Protection (ULP) scheme with no delay is proposed for High Definition (HD) interactive video applications such as video calls. The proposed scheme uses Data Partitioning and Flexible Macro block Ordering (FMO) tools of H.264/AVC and reduces error propagation. The obtained data demonstrates that this scheme achieves better ROI quality than others in high Packet Loss Rates (PLR). Second, it is also shown that the efficient scheme for protecting ROI-based interactive video applications in wide PLR ranges is switching both the Forward Error Correction (FEC) distribution pattern and the codec configuration between various scenarios depending on network conditions. This is shown by using both PSNR and SSIM metrics for quality measurement. Kiana Calagari, Mohammad Reza Pakravan, Shervin Shirmohammadi |
ACM Multimedia | 3 |
| 2012 | Energy-aware adaptations in mobile 3d graphicsabstractSmartphone devices are becoming the de facto personal computing platform, rivaling the desktop, as the number of smartphone users is projected to reach 1.1 billion by 2013. Unlike the desktop, smartphones have a constrained energy budget, which is further challenged by increasingly sophisticated applications. Amongst the most popular applications on smartphone devices are games and virtual environments that rely on 3D graphics. Due to the computational intensity of geometry and rasterization, as well as the perpetually illuminated display, these applications are extremely power-hungry. To prolong the battery life of devices running these applications, we propose two new energy-aware adaptation schemes that can be employed in 3D graphics applications: lighting limitation and textural transformation. Our results show that we can conserve between 20% and 33% of energy with acceptable sacrifices to a user's visual experience. Mohammad Hosseini 0002, Alexandra Fedorova, Joseph G. Peters, Shervin Shirmohammadi |
ACM Multimedia | 4 |
| 2012 | Adaptive 3D texture streaming in M3G-based mobile gamesabstractWith the growing demand of mobile applications and games, one of the challenges is how to efficiently transmit the bulky 3D information to bandwidth- and computationally-limited mobile devices. In this paper, we propose two methods for improving the transmission delay of M3G-based 3D mobile game content over unreliable and congested networks. We introduce Object Mesh Similarity as a server side approach, in which we try to detect an alternative object with minimum complexities that is similar to the original object, and then transmit this reconstructed object, as well as Texture Stretching as a client-side approach, which leads to the efficient receipt of object textures. Our results show 35% to 70% improvement in the transmission delay of 3D textures. Mohammad Hosseini 0002, Dewan Tanvir Ahmed, Shervin Shirmohammadi |
MMSys | 3 |
| 2012 | Knowledge-empowered agent information system for privacy payoff in eCommerce
Abdulsalam Yassine, Ali A. Nazari Shirehjini, Shervin Shirmohammadi, Thomas T. Tran |
Knowl. Inf. Syst. | 3 |
| 2012 | Improving online gaming experience using location awareness and interaction details
Dewan Tanvir Ahmed, Shervin Shirmohammadi |
Multim. Tools Appl. | 2 |
| 2012 | A Mixed Layer Multiple Description Video Coding SchemeabstractMultiple description coding (MDC) is a technique where multiple streams from a source video are generated, each individually decodable and mutually refinable. MDC is a promising solution to overcome packet loss in video transmission over noisy channels, particularly for real-time applications in which retransmission of lost information is not practical. A problem with conventional MDC is that the achieved side distortion quality is considerably lower than single description coding (SDC) quality except at high redundancies which in turn leads to central quality degradation. In this paper, a new mixed layer MDC scheme is presented with no degradation in central quality, and providing better side quality (approximately as much as that of SDC) compared to conventional methods. Also, this property directly leads to higher average quality when delivering the video in lossy networks. For each discrete cosine transform coefficient, we generate two coefficients: base coefficient (BC) and enhancement coefficient which are combined together. When all descriptions are available, they are decomposed and decoded to achieve high quality video. When one description is not available, we use estimation to extract as much of the BC as possible from the received description. Simulation results show that the proposed scheme leads to an improved redundancy-rate-distortion performance compared to conventional methods. The algorithm is implemented in JM16.0 and its performance for two-description and four-description coding is verified by experiments. Mohammad Kazemi 0002, Khosrow Haj Sadeghi, Shervin Shirmohammadi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Equipment Location in Hospitals Using RFID-Based Positioning SystemabstractThroughout various complex processes within hospitals, context-aware services and applications can help to improve the quality of care and reduce costs. For example, sensors and radio frequency identification (RFID) technologies for e-health have been deployed to improve the flow of material, equipment, personal, and patient. Bed tracking, patient monitoring, real-time logistic analysis, and critical equipment tracking are famous applications of real-time location systems (RTLS) in hospitals. In fact, existing case studies show that RTLS can improve service quality and safety, and optimize emergency management and time critical processes. In this paper, we propose a robust system for position and orientation determination of equipment. Our system utilizes passive (RFID) technology mounted on flooring plates and several peripherals for sensor data interpretation. The system is implemented and tested through extensive experiments. The results show that our system's average positioning and orientation measurement outperforms existing systems in terms of accuracy. The details of the system as well as the experimental results are presented in this paper. Ali A. Nazari Shirehjini, Abdulsalam Yassine, Shervin Shirmohammadi |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2012 | A new methodology to derive objective quality assessment metrics for scalable multiview 3D video codingabstractWith the growing demand for 3D video, efforts are underway to incorporate it in the next generation of broadcast and streaming applications and standards. 3D video is currently available in games, entertainment, education, security, and surveillance applications. A typical scenario for multiview 3D consists of several 3D video sequences captured simultaneously from the same scene with the help of multiple cameras from different positions and through different angles. Multiview video coding provides a compact representation of these multiple views by exploiting the large amount of inter-view statistical dependencies. One of the major challenges in this field is how to transmit the large amount of data of a multiview sequence over error prone channels to heterogeneous mobile devices with different bandwidth, resolution, and processing/battery power, while maintaining a high visual quality. Scalable Multiview 3D Video Coding (SMVC) is one of the methods to address this challenge; however, the evaluation of the overall visual quality of the resulting scaled-down video requires a new objective perceptual quality measure specifically designed for scalable multiview 3D video. Although several subjective and objective quality assessment methods have been proposed for multiview 3D sequences, no comparable attempt has been made for quality assessment of scalable multiview 3D video. In this article, we propose a new methodology to build suitable objective quality assessment metrics for different scalable modalities in multiview 3D video. Our proposed methodology considers the importance of each layer and its content as a quality of experience factor in the overall quality. Furthermore, in addition to the quality of each layer, the concept of disparity between layers (inter-layer disparity) and disparity between the units of each layer (intra-layer disparity) is considered as an effective feature to evaluate overall perceived quality more accurately. Simulation results indicate that by using this methodology, more efficient objective quality assessment metrics can be introduced for each multiview 3D video scalable modalities. Hoda Roodaki, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | Introduction to special section on 3D mobile multimediaabstract10.1145/2348816.2348820 Shervin Shirmohammadi, Mohamed Hefeeda, Wei Tsang Ooi, Romulus Grigoras |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | A decision support engine for video surveillance systemsabstractDesign and implementation of an automated or a semi automated surveillance system is an active research area. Safety concerns of individuals have accelerated the research for identifying alarming human activities in public spaces like streets, shopping malls, airports, and others. But reliance on human operator for real-time actions can be inappropriate and expensive. On the other hand, demonstrating pragmatic scene of a surveillanced area with the help of a synthetic world can be effective that can guard operator's limitation. In this paper, we design a Decision Support Engine (DSE) coupled with a synthetic space to facilitate surveillance activities of operators. For this purpose, the detailed sensory data are processed and alarms are detected, classified and ranked according to the threat severity that follows some well-defined rules of the system. At the end, the identified alarming cases are marked and presented with the help of a synthetic environment in an elegant way so that the operators can take the right action in a specific security circumstance. Dewan Tanvir Ahmed, Shervin Shirmohammadi |
ICME | 2 |
| 2011 | A rate control algorithm for ×264 high definition video conferencingabstractCurrent rate control algorithms developed in the H.264 encoder are unable to address High Definition video streaming requirements over best effort networks such as the Internet. We propose a rate controller (RC) for H.264 high definition video conferencing (HDVC), specifically developed in the x264 codec: an open source and high performance H.264/AVC encoder. Called the Dynamic Constant Rate Factor (DCRF), our RC tries to provide the best possible quality in terms of the current bandwidth available at a given instance of video streaming. The proposed method is analyzed from different aspects including network bandwidth, frame size, PSNR and SSIM measurements. In addition to the objective tests, DCRF is evaluated by human observers for subjective performance evaluation. The results prove that DCRF provides better results than current RC methods for HDVC. Abbas Javadtalab, Mona Omidyeganeh, Shervin Shirmohammadi, Mojtaba Hosseini |
ICME | 3 |
| 2011 | A high video quality Multiple Description Coding scheme for lossy channelsabstractMultiple Description Coding (MDC) is a technique where multiple streams from a source are generated, each individually decodable and mutually refinable. In this paper, a new Mixed Layer MDC (MLMDC) scheme is presented which achieves a higher side quality compared to conventional MDCs. The improved side performance leads to higher average video quality at the receiver in lossy networks. For each DCT coefficient, we generate two coefficients: Base Coefficient (BC) and Enhancement Coefficient (EC) which are combined together. When all descriptions are available, they are decomposed and decoded to achieve high quality video. When one description is not available, we use estimation to extract as much of the BC as possible from the received description. The algorithm is implemented in JM16.0 and its performance for two-description and four-description coding is verified by experiments. Mohammad Kazemi 0002, Khosrow Haj Sadeghi, Shervin Shirmohammadi |
ICME | 3 |
| 2011 | A new Scalable Multi-View Video Coding configuration for mobile applicationsabstractTransmission of multi-view video content is not practical in most mobile environments due to the limited bandwidth and processing power of mobile devices. To support such environments, one can limit the number of views that are being transmitted, known as Scalable Multi-view Video Coding (SMVC). In this paper, we propose a new view selection method for view scalability in multi-view video coding in mobile environments, which uses inter and intra view dissimilarities to determine the most suitable views for the base layer corresponding to the prediction structure and user selected limited number of views. By selecting more correlated views for the base layer, the proposed method provides an improved performance, as confirmed by simulation results, even when all the enhancement layers are dropped due to network limitations. Hoda Roodaki, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
ICME | 3 |
| 2011 | Context-aware prioritized game streamingabstractIn this paper, we describe our proposed mechanism for a context-aware progressive 3D streaming for virtual environments (VEs) such as games running on mobile handheld devices. Our goal is to optimize the efficiency of streaming update messages in such environments by providing a new, dynamic, and context-aware method for selecting and prioritizing objects. Hesam Rahimi, Ali A. Nazari Shirehjini, Shervin Shirmohammadi |
ICME | 3 |
| 2011 | Transparent non-intrusive multimodal biometric system for video conference using the fusion of face and ear recognitionabstractMono-modal biometric systems face many limitations such as noisy data, intra-class variations, distinctiveness, spoof attacks, non-universality, and unacceptable error rates. Working on enhancing the performance of a mono-modal biometric system may not be highly efficient and effective. A multimodal biometric system combines two or more biometric features into a single identification system. It aims to improve several of the mono-modal biometric systems drawbacks and improve the recognition coverage and performance. In this paper, a transparent non-intrusive multimodal biometric system based on the fusion of the face and ear biometrics is proposed to identify individuals during a video conference environment with minimal explicit user involvement and hassle. The results of the experiment show that the performance of the transparent non-intrusive multimodal biometric system of the face and ear is higher than that of the mono-modal face or ear. Abbas Javadtalab, Laith Abbadi, Mona Omidyeganeh, Shervin Shirmohammadi, Carlisle M. Adams, Abdulmotaleb El Saddik |
PST | 4 |
| 2011 | Online information privacy: Agent-mediated payoffabstractWith the rapid development of applications in open distributed environments such as eCommerce, privacy of information is becoming a critical issue. Information about the preferences, activities, and demographic attributes of people using online shopping is very valuable to online businesses. Beyond the general anxieties with sharing personal information, people may more specifically have concerns about becoming increasingly identifiable; as increasing amounts of personal data are acquired. Not only that, but also it is widely known that information about consumers is often sold to online marketing and advertising companies, generally without the knowledge of consumers [4, 16]. In this paper, we introduce a model where people can opt out to share personal information for a payoff value that balances out the costs of privacy. Our model is based on agents working on behalf of consumers to maximize their benefit. The analysis of the model and a proof of concept implementation are presented in this paper. Abdulsalam Yassine, Ali A. Nazari Shirehjini, Shervin Shirmohammadi, Thomas T. Tran |
PST | 3 |
| 2011 | Video Keyframe Analysis Using a Segment-Based Statistical Metric in a Visually Sensitive Parametric SpaceabstractThis paper addresses a new approach to the keyframe extraction problem employing generalized Gaussian density (GGD) parameters of wavelet transform subbands along with Kullback-Leibler distance (KLD) measurement. Shot and cluster boundaries are selected using KLDs between GGD feature vectors, and then keyframes are located based on similarity and dissimilarity criteria. Objective and subjective evaluations show the high accuracy of this new approach compared with traditional methods. Mona Omidyeganeh, Shahrokh Ghaemmaghami, Shervin Shirmohammadi |
IEEE Trans. Image Process. | 3 |
| 2011 | Introduction to ACM multimedia 2010 best paper candidatesabstractNo abstract available. Shervin Shirmohammadi, Jiebo Luo 0001, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2009 | DCS: A Distributed Coordinate System for Network PositioningabstractPredicting latency between nodes on the Internet can have a significant impact on the performance of many services that use latency distances among nodes as a decision making input. Coordinate-based approaches are among the family of latency prediction techniques where latency between each pair of nodes is modeled as the virtual distance among those nodes over a virtual system. This article proposes the decentralized coordinate system (DCS) as a fully distributed system that does not rely on any infrastructure support from the underlying network. DCS uses a two-phase algorithm: first, each host is assigned rough coordinates, then, the rough estimation is refined in a way that it gradually converges to the accurate relative position of nodes. Simulation results demonstrate that the accuracy of DCS is competitive with existing network coordinate systems and that it can be considered as a good alternative for some applications. Negar Hariri, Jafar Habibi, Shervin Shirmohammadi, Behnoosh Hariri |
DS-RT | 3 |
| 2009 | A light-weight federated video adaptation system for P2P overlaysabstractThis paper presents a lightweight but effective mastersender-driven multiple parent approach for online video adaptation and streaming to heterogeneous clients. The proposed design facilitates the cooperation of participating clients for improving the utilization of the spare resources in an overlay. Our design uses standard H.264/AVC video streams. It enables peers to contribute both CPU and bandwidth in a multi-parent fashion, such that no dedicated adaptation server and streaming server is required, although we need a reliable rendezvous point for the video stream originators and the peers in the overlay. We present the live video processing paradigm using metadata followed by the video distribution overview. A brief performance evaluation supporting the design choices is also presented. Razib Iqbal, Shervin Shirmohammadi |
ICME | 2 |
| 2009 | A Cooperative video adaptation and streaming scheme for mobile and heterogeneous devices in a community networkabstractIn this paper, we aim to present the fundamental properties of a community-driven adaptive P2P streaming scheme. We show that if the participants in a community network agree to not only share their bandwidth, but also their computing resources according to the design principles mentioned here, then mobile and heterogeneous devices can be accommodated in the P2P paradigm ensuring adequate resource utilization, with respect to resilience to peer dynamics. We present simple design principles for a multimodal P2P system considering available bandwidth, computing power, and delay to build the video overlays. Some evaluation results supporting our design principles are also presented. Razib Iqbal, Shervin Shirmohammadi |
ICME | 2 |
| 2009 | MPEG-21 based temporal video adaptation for heterogeneous devices and mobile environmentsabstractIn the era of universal multimedia access (UMA), applications like video conferencing, surveillance and streaming are challenged by the multiplicity of devices. Since adaptation is the newly applied practice for digital video content customization, in this demo, we present a simple approach to adapt H.264 videos, in the compressed domain, for heterogeneous devices and mobile environments. Our approach is in contrast to most existing approaches that do adaptation in a series of cascaded decode/re-encode operations. We show that the adaptation operations can be expedited, using our approach, when adaptation systems are designed to adapt contents according to the encoding structure but in an intermediary node following a codec-independent technique. We present our adaptation system as a utility box where adaptation is performed on-demand based on its generic bitstream syntax description (gBSD), and in compressed domain. Razib Iqbal, Shervin Shirmohammadi |
ICME | 2 |
| 2009 | Compressed-domain temporal adaptation-resilient watermarking for H.264 video authenticationabstractIn this paper, we present a DCT domain watermarking approach for H.264/AVC video coding standard. This scheme is resilient to compressed-domain temporal adaptation. A cryptographic hash function is used to generate a semi-fragile watermark to provide content-based authentication. The embedded watermark can withstand frame-dropping due to temporal adaptation, yet it is able to detect malicious attacks such as content modification, transcoding etc. Simulation results demonstrate that the watermarking scheme is computationally efficient and suitable for practical use. Sharmeen Shahabuddin, Razib Iqbal, Shervin Shirmohammadi, Jiying Zhao |
ICME | 3 |
| 2009 | A compressed-domain spatio-temporal adaptation system for video deliveryabstractIn this demo, we present a working system that we have built for metadata-based compressed-domain spatio-temporal video adaptation of H.264/AVC videos. Our approach is in contrast to most existing approaches that do adaptation in a series of cascaded decode/re-encode procedures. We show that by applying compressed-domain operations, adaptation can be expedited for real-time application scenarios like news/sports broadcasting. Moreover, the system can be applied as a tool box for distributed adaptation within P2P video distribution applications to support heterogeneous devices. Razib Iqbal, Sharmeen Shahabuddin, Shervin Shirmohammadi |
ACM Multimedia | 3 |
| 2009 | Compressed domain spatial adaptation for H.264 videoabstractIn this paper, we present a metadata-based compressed-domain spatial adaptation scheme for H.264/AVC video. We have enhanced the H.264/AVC encoder with our proposed adaptation strategies in order to reduce video size by cropping individual frames in an intermediary node prior to transmitting that video to heterogeneous devices. In this regard, we exploit the sliced architecture of the video frames within the first version of the H.264/AVC specification and devise different slicing strategies. The compressed-domain bitstream modification is performed at the intermediary nodes, avoiding the need for any cascaded operations. Here, we briefly present our adaptation scheme as well as evaluation results showing the effectiveness of the slicing strategies on the bitrate reduction and processing time. A comparison of our approach with an existing cropping scheme is also presented. Sharmeen Shahabuddin, Razib Iqbal, Ali A. Nazari Shirehjini, Shervin Shirmohammadi |
ACM Multimedia | 4 |
| 2009 | A delaunay triangulation architecture supporting churn and user mobility in MMVEsabstractThis article proposes a new distributed architecture for update message exchange inmassively multi-user virtual environments (MMVE). MMVE applications require delivery of updates among various locations in the virtual environment. The proposed architecture here exploits the location addressing of geometrical routing in order to alleviate the need for IP-specific queries. However, the use of geometrical routing requires careful choice of overlay to achieve high performance in terms of minimizing the delay. At the same time, the MMVE is dynamic, in sense that users are constantly moving in the 3D virtual space. As such, our architecture uses a distributed topology control scheme that aims at maintaining the requires QoS to best support the greedy geometrical routing, despite user mobility or churn. We will further prove the functionality and performance of the proposed scheme through both theory and simulations. Mohsen Ghaffari 0001, Behnoosh Hariri, Shervin Shirmohammadi |
NOSSDAV | 3 |
| 2009 | An adaptive latency mitigation scheme for massively multiuser virtual environments
Behnoosh Hariri, Shervin Shirmohammadi, Mohammad Reza Pakravan, Mohammad Hossein Alavi |
J. Netw. Comput. Appl. | 2 |
| 2009 | A hybrid P2P communications architecture for zonal MMOGs
Dewan Tanvir Ahmed, Shervin Shirmohammadi, Jauvane Cavalcante de Oliveira |
Multim. Tools Appl. | 2 |
| 2009 | Using geometrical routing for overlay networking in MMOGs
Behnoosh Hariri, Mohammad Reza Pakravan, Shervin Shirmohammadi, Mohammad Hossein Alavi |
Multim. Tools Appl. | 3 |
| 2009 | Guest editorial for special issue on massively multiplayer online gaming systems and applications
Shervin Shirmohammadi, Mark Claypool |
Multim. Tools Appl. | 1 |
| 2009 | DAg-stream: Distributed video adaptation for overlay streaming to heterogeneous devices
Razib Iqbal, Shervin Shirmohammadi |
Peer-to-Peer Netw. Appl. | 2 |
| 2008 | Privacy and the market for private data: A negotiation model to capitalize on private dataabstractThe market for consumer information is already a lively market, where consumer information and consumer profile data are often among the most valuable assets owned by online retailers. The value of such commodity derives from the ability of firms to identify consumers and charge them personalized prices flj. We argue that if consumers' identity and personal information is such a valuable asset, should not consumers benefit from their asset as well? In this paper, we propose a negotiation process between an online consumer agent and an online seller. The online consumer agent acts on behalf of consumers to maximize their social welfare. In our model, the agent derives a quantified privacy risk for each private data and uses it to determine a cost premium value to make the bargaining process manageable. We also provide a computational example to evaluate the model. Abdulsalam Yassine, Shervin Shirmohammadi |
AICCSA | 2 |
| 2008 | A Dynamic Area of Interest Management and Collaboration Model for P2P MMOGsabstractIn this paper, we present a dynamic area of interest management for massively multiplayer online games (MMOG). Instead of mapping the virtual space to the area of interest (AOI), we scheme AOIs to the virtual space. This zoneless MMOG is the consequence of dynamic AOI that redeems the necessity of inter-AOI communication. In addition, the AOI maintenance cost is reduced significantly by assigning the maintenance responsibility to a subset of players for each AOI. Due to the integration of peer-to-peer communication model to the system, the scalability has improved. To satisfy the timing constraints, the projection of the underling network topology to the overlay network is more fruitful than building the overlay on the fly unintelligently. In response to this fact, the proposed communication model adapts a geometric algorithm which is usually used for minimax problem. The model is evaluated and justified through proper simulation. Dewan Tanvir Ahmed, Shervin Shirmohammadi |
DS-RT | 2 |
| 2008 | NL-DHT: A Non-uniform Locality Sensitive DHT Architecture for Massively Multi-user Virtual Environment ApplicationsabstractDHT networks offer a scalable structure for use in massively multi-user virtual environments (MMVEs). However, an issue with DHT structures is their use of uniform location-independent ID assignment. This conflicts with the locality-sensitive non-uniform ID assignment needed to achieve efficient latency-aware routing in MMVE applications. Our proposed solution is to use a modified version of the Hilbert space-filling curve in order to map users¿ locations from three dimensional VE space into single dimensional DHT ID space in the best locality preserving way. The need for such modified Hilbert curve is due to the fact that users are not typically homogeneously distributed in MMVEs while Hilbert curves fill space in a uniform manner. Generation of, and mapping to, an entire Hilbert curve can be very expensive in terms of computational complexity. Thus we present a fast dynamic mapping of MMVE locations to the modified Hilbert curve to reduce the mapping complexity. Saurabh Ratti, Behnoosh Hariri, Shervin Shirmohammadi |
ICPADS | 3 |
| 2008 | Online adaptation for video sharing applicationsabstractThe main concept of Peer-to-Peer (P2P) streaming is that viewers will contribute their bandwidth to the overlay and act as a relay for the video streams. In this paper, we introduce how a peer may implement an adaptive streaming scheme to serve peers in a P2P application. The technical contribution of this paper is to present the effectiveness and feasibility of utilizing the available computing power of the participating peers to serve mobile and heterogeneous clients by adapting the video content on the fly. The benefit is that there is no need for a dedicated adaptation or streaming server deployed in the system for video streaming/sharing applications. We emphasize on structured metadata-based adaptation and streaming utilizing MPEG-21 gBSD. Here, we briefly illustrate our scheme and present some experimental evaluations supporting our design choices. Copyright 2008 ACM. Razib Iqbal, Shervin Shirmohammadi |
ACM Multimedia | 2 |
| 2008 | Modeling and evaluation of overlay generation problem for peer-assisted video adaptation and streamingabstractIn this paper, we consider the problem of overlay generation for video adaptation and streaming applications in a way to efficiently utilize the bandwidth and computing power of the participating peers. Therefore, the proposed architecture performs regular streaming functions as well as video adaptation functions, moving the video contents adaptation computation load away from dedicated media-streaming/adaptation servers to the participating peers. To verify the performance of our design, we followed an analytical approach based on 0-1 Integer Linear Programming method to model the system and to calculate the optimum overlay. The performance of our scheme is evaluated by simulations. Preliminary results demonstrate that our design performance nearly follows the optimal boundary in terms of resource utilization. Razib Iqbal, Behnoosh Hariri, Shervin Shirmohammadi |
NOSSDAV | 3 |
| 2008 | Distributed Video Adaptation and Streaming for Heterogeneous DevicesabstractIn this paper, we present a new concept of distributed adaptation and P2P streaming supporting the expansion of context-aware mobile P2P systems. Here, we propose to adapt video contents in a distributed manner to address the peer heterogeneity. It also forms an overlay network and organizes video streaming sessions among the participating peers. We have followed a master-sender-driven approach (one-to-many), with the help of cooperative peers. The essential incentive behind this work is that with the modern computing capabilities, a single peer can carry out sufficient adaptation operations and stream media contents to other peers. Overall design and adaptation approach along with the performance evaluation is presented in brief. Razib Iqbal, Dewan Tanvir Ahmed, Shervin Shirmohammadi |
PerCom | 3 |
| 2008 | Touching beyond audio and video
Abdulmotaleb El Saddik, Shervin Shirmohammadi |
Multim. Tools Appl. | 2 |
| 2008 | Experiments in haptic-based authentication of humans
Mauricio Orozco Trujillo, Matthew Graydon, Shervin Shirmohammadi, Abdulmotaleb El Saddik |
Multim. Tools Appl. | 3 |
| 2008 | Experiments in haptic-based authentication of humans
Mauricio Orozco Trujillo, Matthew Graydon, Shervin Shirmohammadi, Abdulmotaleb El Saddik |
Multim. Tools Appl. | 3 |
| 2007 | Performance Enhancement in MMOGs Using Entity TypesabstractState management is a fundamental requirement of multi-user applications like distributed simulations or networked games. The quality of the experience of such applications heavily depends on synchronous communication. The movement of an entity from one logical zone into another initiates reorganization of the peer-to-peer structure and eventually ruins synchronous interactions. The frequent nature of such event in the distributed simulations has significant impact on the performance and cannot be ignored. In this paper, we present a performance enhancement mechanism considering the physical attributes of the entities. We demonstrate that clustering the entities based on their physical characteristics or behavior significantly reduces peer- to-peer structural reformation penalties. The effectiveness of such clustering mechanism is simulated and analyzed in the context MMOG. Dewan Tanvir Ahmed, Shervin Shirmohammadi, Jauvane Cavalcante de Oliveira |
DS-RT | 2 |
| 2007 | A Distributed Topology Control Algorithm for P2P Based SimulationsabstractAlthough collaborative distributed simulations and virtual environments (VE) have been an active area of research in the past few years, they have recently gained even more attention due to the emergence of online gaming, emergency simulation and planning systems, and disaster management applications. Such environments combine graphics, haptics, animations and networking to create interactive multimodal worlds that allows participants to collaborate in realtime. Massively Multiplayer Online Gaming (MMOG), perhaps the most widely deployed practical application of distributed virtual environments, allows players to act together concurrently in a virtual world over the Internet. IP Multicasting would be an optimal solution for the dissemination of updates among participants, but IP multicasting is not available to home users on the Internet, due to a number of technological, practical, and business reasons. In light of the lack availability of IP Multicasting on the global Internet, researchers have recently tended to shift multicasting from the networking layer to the application layer, known as Application Layer Multicasting, effectively constructing an overlay network among participants of the distributed simulation where end hosts themselves participate in the dissemination of update messages. In this paper, we propose a topology control architecture to support P2P based collaborative distributed simulations over the Internet by using AIM. We present our networking model and its rationale, theoretical proof, and simulation measurements in comparison with other methods as proof of concept. Behnoosh Hariri, Shervin Shirmohammadi, Mohammad Reza Pakravan |
DS-RT | 2 |
| 2007 | A Visibility-Driven Approach to Managing Interest in Distributed Simulations with Dynamic Load BalancingabstractDistributed simulations that support a massive number of users typically divide the virtual world into zones that are managed by separate servers to evenly distribute resources and achieve scalability. However, such zoning restricts cross-zonal interactions and exposes the division of the world to the participating parties. Problems such as crowding one zone among others defeats the very purpose of interest management and makes geographic partitioning inefficient for modeling interactions. In this work, we have designed and implemented a visibility-driven approach to make the partitioning transparent to users. The effectiveness of this distributed architecture is tested through a prototype implementation. We also introduce a novel idea to dynamic load balancing that can be achieved in real-time without modifying the communication architecture. By increasing the granularity of the partitioning and providing a layered approach to zoning, transient crowding can be handled by adoptively dispersing parts of the crowded zone to adjacent servers. Ihab Kazem, Dewan Tanvir Ahmed, Shervin Shirmohammadi |
DS-RT | 3 |
| 2007 | A Framework for Provisioning Overlay Network Based Multimedia Distribution ServicesabstractIn this paper, we present a graph-theoretic framework for provisioning overlay network based multimedia distribution services to a diverse set of receivers. Considering resource limitations and exploiting geographical positions, it greedily constructs degree-constrained minimum-cost connected graph to manipulate the topology to a significant extent by selecting mesh neighbors and changing the metrics. Data delivery paths and forwarding nodes are chosen using the dominating set. The minimal cardinality of the dominating set reduces the system's dependency on end-hosts. Simulation is used to demonstrate that the framework is more robust and responsive to tree partitions and suitable for multi-source multimedia applications. Dewan Tanvir Ahmed, Shervin Shirmohammadi |
ICME | 2 |
| 2007 | Hard Authentication of H.264 Video Applying MPEG-21 Generic Bitstream Syntax Description (gBSD)abstractWhile trivial research has been conducted in watermarking and authentication of H.264 video in recent years, most techniques require cascaded operations within a video adaptation scenario. In this paper, we propose an authentication scheme for adapted H.264 video content to detect integrity at the receiver's side without the need for cascaded operations. The proposed scheme utilizes MPEG-21 gBSD for hard authentication of H.264 video in the compressed domain and does not necessitate any cascaded decompression and recompression. The design uses content-based authentication which is derived from a hash value. The authentication data is embedded as a fragile watermark, and the marking space is selected during the adaptation process of the H.264 video by parsing the gBSD. The authentication information is embedded in already encoded videos during adaptation. Proof of concept and performance evaluation is also presented. Razib Iqbal, Shervin Shirmohammadi, Jiying Zhao |
ICME | 2 |
| 2007 | Improving gaming experience in zonal MMOGsabstractSynchronous communication is a primary concern for multi-user virtual environments like Massively Multiplayer Online Games (MMOGs). Most of the MMOGs offer discrete view for the avatars and follow logical zone layout for easy state management. The avatar movement, from one logical zone to another, causes reorganization at the P2P overlay structure. Its recurrent nature along with unintelligent zone crossing approaches eventually hampers synchronous communication. In this paper, we present performance enhancement mechanisms to reduce P2P overlay reorganization penalties based on avatars ’ physical characteristics. Avatars ’ unpredictable movement around the zone boundaries also incur repeated connections and disconnections either among the zone masters or among the multiple peers in the overlay networks. Interest driven zone crossing, dynamic shared region between adjacent zones, and clustering of the entities based on their attributes significantly alleviates these problems and also provides continuous view and seamless region crossing to players. The technical contribution of this paper is to present the effectiveness of object clustering mechanism and interest driven zone crossing with dynamic shared region between each adjacent zones to enhance the gaming experience of MMOG players or users of other types of networked virtual environments (VE). Dewan Tanvir Ahmed, Shervin Shirmohammadi, Jauvane Cavalcante de Oliveira |
ACM Multimedia | 2 |
| 2006 | Secured MPEG-21 Digital Item Adaptation for H.264 VideoabstractSeamless adaptation and transcoding techniques to adapt the digital content have achieved significant focus to serve the consumers with the desired content in a feasible way. With the succession of time we sense that secured adaptation should also be taken care of for not only serving sensitive digital contents but also to offer security as an embedded feature of the adaptation practice to ensure digital right management and confidentiality. In this paper, we propose an encryption framework for a transcoder while adapting H.264 video conforming to MPEG-21 DIA. Encryption mechanism is applied on the adapted video content thus reducing computational overhead compared to that on the original content Razib Iqbal, Shervin Shirmohammadi, Abdulmotaleb El Saddik |
ICME | 2 |
| 2006 | BM-ALM: An Application Layer Multicasting with Behavior Monitoring ApproachabstractIP multicasting is the most efficient way to perform group data distribution, as it eliminates traffic redundancy and improves bandwidth utilization. Application layer multicast (ALM) has been proposed to overcome some of the limitations in IP multicasting such as scalability and deployability. Limited computing power, scarcity of bandwidth and end-host's reluctance to share bandwidth make ALM difficult to spread. In this paper, we keep eye to those problems and present an ALM that scrutinizes the commitment of the ALM nodes. Failure to provide quality of service agreement triggers performance penalty for the node in concern. Thus, every node has an obligation to its descendants; as a result, a nice collaboration among the end-hosts is achieved for group communication. It has good performance for content distribution, as it reflects physical network topology onto the overlay network constructed by the end-hosts. Tree refinement and backup path strategies are taken to better satisfy heterogeneous QoS requirements Dewan Tanvir Ahmed, Shervin Shirmohammadi |
ISM | 2 |
| 2006 | Compressed-Domain Encryption of Adapted H.264 VideoabstractCommercial service providers and secret services yearn to employ the available environment for conveyance of their data in a secured way. In order to encrypt or to ensure personalized security of the video contents in an intermediary node, it is necessary to have the content structure conforming to an international standard. Moreover, pressure to satisfy user preferences and device requirements seamlessly are raising the need for content to be customized providing the best possible experience. In this paper, we present perceptual encryption scheme for video encryption that is incorporated with a dynamic temporal adaptation technique of the H.264 video conforming ISO/IEC MPEG-21 Digital Item Adaptation. Encryption is performed on demand directly from the adapted bitstream and its generic Bitstream Syntax Description (gBSD). Razib Iqbal, Shervin Shirmohammadi, Abdulmotaleb El Saddik |
ISM | 2 |
| 2006 | MPEG-21 Based Temporal Adaptation of Live H.264 VideoabstractThe diversity of devices in both wired and wireless networks via which multimedia contents are desired to be accessed and interacted with has grown significantly. Applications like video conferencing, surveillance and chatting is challenged by this diversity which requires live adaptation to meet user requirements and device specifications. In this paper, we present an architecture for temporal adaptation of ITU-T H.264 video conforming to ISO/IEC MPEG-21 DIA for live video stream along with the adaptation module implementation detail. Adaptation is performed on demand directly from the live bitstream and its generic bitstream syntax description (gBSD) avoiding conventional approaches seen in traditional transcoders. As a result, any MPEG-21 compliant host can adapt the stream without requiring the video codec. A prototype, based on the proposed architecture, and experimental evaluations of the system and its performance supporting the architecture are also presented Razib Iqbal, Shervin Shirmohammadi, Chris Joslin |
ISM | 2 |
| 2004 | Shared Object Manipulation with Decorators in Virtual EnvironmentsabstractOne of the known problems with shared object manipulation in virtual environments is the disruptive effect of network lag in collaboration sessions. Most solutions to this problem revolve around techniques to compensate for this lag at the network communication level. In this article, we examine a different approach: informing the user about the lag, using decorators, and allowing the user to react to it intuitively. Shervin Shirmohammadi, Nancy Ho Woo |
DS-RT | 1 |
| 2003 | An Approach for Recording Multimedia Collaborative Sessions: Design and Implementation
Shervin Shirmohammadi, Li Ding 0006, Nicolas D. Georganas |
Multim. Tools Appl. | 1 |
| 2003 | JASMINE: A Java Tool for Multimedia Collaboration on the Internet
Shervin Shirmohammadi, Abdulmotaleb El Saddik, Nicolas D. Georganas, Ralf Steinmetz |
Multim. Tools Appl. | 1 |
| 2001 | An end-to-end communication architecture for collaborative virtual environments
Shervin Shirmohammadi, Nicolas D. Georganas |
Comput. Networks | 1 |
| 2001 | Web-based multimedia tools for sharing educational resourcesabstractMany educational resources and objects have been developed as Java applets or applications, which can accessed by simply downloading them from various repositories. It is often necessary to share these resources in real time, for instance when an instructor teaches remote students how to use a certain resource explains the theory behind it. We have developed some tools for this purpose that emulate a virtual classroom, and are primarily designed for synchronous sharing of resources. They enable participants to share Java objects in real time and also allow the instructor to dynamically manage the telebearing session. Shervin Shirmohammadi, Abdulmotaleb El Saddik, Nicolas D. Georganas, Ralf Steinmetz |
ACM J. Educ. Resour. Comput. | 1 |
| 2000 | A Collaborative Virtual Environment for Industrial TrainingabstractThe concept of collaborative virtual environments has been deployed in many systems in the past few years. Applications of such technology have been used anywhere from military combat simulations to various civilian commercial applications. We present a specific system developed for industrial teletraining. The system, which aims to reduce the cost of training, consists of an interactive 3D environment with collaboration and management capabilities that allows a trainer to lead and control a session attended by many trainees. Jauvane Cavalcante de Oliveira, Shervin Shirmohammadi, Nicolas D. Georganas |
VR | 2 |
| 2000 | An Architecture for Collaboration in Virtual EnvironmentsabstractWe introduce an architecture for performing closely-coupled collaborative tasks in virtual environments. Our architecture consists of an application-layer model based on higher-level user interactions with shared objects, and a communication protocol for dissemination of collaborative update messages among participants. Shervin Shirmohammadi, Nicolas D. Georganas |
VR | 1 |