VLDB 2026 Research / reviewers in the wild / expert
Mahsa T. Pourazad
dblp:55/6295
· DBLP profile ↗
38ranked-venue papers
4as first author
8since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Federated Semi-Supervised Learning for Object Detection in Autonomous DrivingabstractOne of the main challenges in designing deep learning networks for autonomous driving is the lack of labeled data. Recent trends that address this problem involve the use of unlabeled data. In this paper, we propose a unified semi-supervised and federated learning (FL) approach that is designed to offer cost efficient and practical training of deep learning object detection models for autonomous driving. In our implementation, we assume that each vehicle is given some well-labeled image data which are coupled with unlabeled image data captured by its cameras. Each of the vehicles has a local object detection model, which will be trained leveraging a semi-supervised learning method with both labeled and unlabeled data. The local model parameters are uploaded to a cloud server and aggregated to update a global FL model which in turn is shared with all the vehicles involved. Performance evaluations showed that our proposed approach is a promising solution as it allows continuous training and thus improved performance in autonomous driving. Fangyuan Chi, Yixiao Wang 0001, Panos Nasiopoulos, Victor C. M. Leung, Mahsa T. Pourazad |
ICASSP | 5 |
| 2023 | A Spatial Calibrated and Colour Corrected Light Field Outdoor Video Dataset from a $5 \times 5$ Dense Camera ArrayabstractIn this paper, a new and calibrated light field (LF) video dataset is introduced, which focuses on outdoor scenes and objects. Each video stream is 10 seconds long and it is captured with a dense camera array that consists of$5\times 5$camera modules in$1640\times 1232$resolution at 40 frames per second. As multiple cameras in an array setup may suffer from various conditions of camera settings, lens structure, and lighting variations, the resulting images can be negatively affected by geometric distortion and colour difference. To address that, a unified calibration method involving both spatial calibration and colour correction is employed to correct inconsistences and achieve a better image quality with reduced image distortion. This video dataset would be suitable for further research and investigation of a variety LF applications, such as autonomous driving and immersive media. Yixiao Wang 0001, Nusrat Mehajabin, Hamid Reza Tohidypour, Jerry Song, Menghong Huang, Behnoosh Babaghorbani, Zuhao Chen, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung |
ISCAS | 8 |
| 2023 | A Novel No-Reference HD Video Quality Metric Based on Perceptual Temporal PoolingabstractImpressive advancements in capturing, display, and broadcasting technologies significantly elevate image and video quality, and with that the need for designing new reference and no-reference image and video quality metrics. One of the latest and perceptually accurate video quality metrics is the Video Multi-Method Assessment Fusion (VMAF) method. However, VMAF considers the temporal nature of video using basic average temporal pooling, an approach that falls short from human perception. In this paper, we introduce a new no-reference video quality metric that uses deep learning to extract spatial features and a unique temporal pooling approach to accurately predict the visual quality score. To this end, first we created a video quality dataset that consists of high-resolution, 20s-long test video clips compressed at several different bitrates. These videos were labeled based on subjective evaluations and were used to determine the perceptual importance of frames in our temporal pooling scheme. Evaluations showed that our proposed approach achieved correlation of 90.55% with human perception and outperformed the state-of-the-art VMAF approach by 15.63% accuracy. Hamid Reza Tohidypour, Yixiao Wang 0001, Panos Nasiopoulos, Mahsa T. Pourazad |
SMC | 5 |
| 2022 | A Security-Centric Deep Learning Enabled Camera Solution for Real-Time Human Fall DetectionabstractAutomatic human real-time fall detection is a challenging task in remote healthcare, demanding a non-intrusive, secure and affordable solution. In this paper, we present a real-time hardware system that uses a deep learning model for fall detection embedded in a color camera. To reduce the startup delay and achieve real-time performance for the inference phase, we optimized our model using TensorRT. In addition, we addressed the board memory limitation using virtual memory and linear memory allocation and garbage collection. Moreover, GStreamer was used to perform most of the video processing using Jetson's GPU. Our live evaluation shows that our system achieved the accuracy of 84.44% and real-time performance. Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos |
WiMob | 2 |
| 2022 | An Efficient Pseudo-Sequence-Based Light Field Video Coding Utilizing View Similarities for Prediction StructureabstractLight Field (LF) video technology is a step towards offering a better immersive experience through on-demand refocusing and perspective viewing. However, the significant increase in captured data makes the need for efficient compression of paramount importance. In this paper, we proposed two prediction structures and coding orders that efficiently compress LF video content using the existing HEVC standard. This is achieved by utilizing horizontal and vertical correlation among the views for better inter-view prediction. To assess the performance of the schemes, ten publicly available and widely used LF video sequences were used. Our first method is highly suitable for applications demanding high compression efficiency. It outperforms the best pseudo-sequence-based compression technique to date by up to 17% in bitrate reduction while being scalable in the number of views. The second method is proficient for low computational and random-access complexity to any arbitrary view in the light field video. It offers 10% faster decoding and 20% lower random-access complexity compared to the best existing technique. Nusrat Mehajabin, Mahsa T. Pourazad, Panos Nasiopoulos |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Learning-Based Light Field View Synthesis for Efficient Transmission and StorageabstractOne of the main advantages of Light Field (LF) technology is that it provides a truly immersive experience, critical for computer vision, autonomous driving and medical applications. However, one of the main problems of light field is the size of the data captured, which significantly increases bandwidth requirements. In this paper, we introduce a learning-based LF view synthesis approach for efficient transmission and storage, fundamental for performing remote surgery and storing data. This is achieved by dropping specific views at the transmitting end and then efficiently synthesizing them at the receiver end. Our deep learning approach uses the epipolar image plane (EPI) information to ensure smooth disparity between the generated and original views. We consider plenoptic, synthetic LF content and camera array implementations which support different baseline settings. Experimental results show that our proposed method outperforms state-of-the-art light field view synthesis techniques, offering improved visual quality for the generated views. Abrar Wafa, Mahsa T. Pourazad, Panos Nasiopoulos |
ICIP | 2 |
| 2021 | A Perception-Based Inverse Tone Mapping Operator for High Dynamic Range Video ApplicationsabstractThe drastic visual improvements introduced by High Dynamic Range (HDR) technologies open new markets for a wide range of industries. Among them, a significant opportunity is offered to owners of legacy Standard Dynamic Range (SDR) content that can be converted to the new standard to take advantage of the enhanced capabilities of the HDR displays. Similarly, since SDR broadcasting infrastructure will continue to be around for the time being, such conversion process is becoming an obvious necessity. To this end, different approaches have tried to efficiently convert SDR images and videos to HDR format, a procedure well known as inverse Tone Mapping. In this paper, we propose a novel high visual quality video inverse Tone Mapping Operator (iTMO) that addresses the inadequacies of the state-of-the-art methods, resulting in high visual quality HDR videos that match the capabilities of the HDR technology. Our approach is based on human visual perception and employs a segmentation method according to the Human Visual System (HVS) sensitivity to brightness changes in different regions of the frame and constructs the mapping curve using the brightness distribution information of these regions. Our iTMO uses a hybrid approach to achieve an optimal balance between the overall contrast and brightness of the output HDR frame by maximizing a weighted sum of contrast and brightness difference between input SDR and generated HDR frame. The proposed iTMO works equally well for all levels of brightness, eliminating any visual artifacts by dynamically maintaining changes at non-perceivable levels. Subjective and objective evaluations validated the superior visual performance of our proposed iTMO over state-of-the-art methods. Pedram Mohammadi, Mahsa T. Pourazad, Panos Nasiopoulos |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | A Fully Automatic Content Adaptive Inverse Tone Mapping Operator With Improved Color AccuracyabstractHigh Dynamic Range (HDR) technology offers a higher visual quality compared to its Standard Dynamic Range (SDR) counterpart, as it tries to imitate the way our eyes perceive brightness and color information. Converting SDR content to HDR format - using inverse Tone Mapping Operators (iTMOs) - to take advantage of the superior visual quality offered by HDR displays, is an attractive proposition to SDR content owners and real-time broadcasters. In this paper, we propose a novel content adaptive iTMO that works in the perceptual domain to model the sensitivity of the human eye to brightness changes in different areas of a scene. To preserve the overall visual impression, our proposed iTMO utilizes an entropy-based brightness segmentation, which also makes our method content adaptive. In addition, we propose a novel perception-based color adjustment method that can maintain the color accuracy between input SDR and generated HDR frames. By performing the color adjustment in the perceptual domain, our iTMO prevents hue shifts and generates HDR colors that closely follow their SDR counterparts. Our subjective evaluations indicate that our proposed method outperforms other state-of-the-art methods by an average of 81% in terms of visual quality, and 76% in terms of how closely the HDR colors match their SDR counterparts. In addition to subjective evaluations, we also performed objective evaluations using the HDR-VDP 2.2 and PU-SSIM metrics and concluded that, on average, our proposed iTMO outperforms the state-of-the-art methods in terms of these two metrics. Pedram Mohammadi, Mahsa T. Pourazad, Panos Nasiopoulos |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | A Novel Chroma Representation For Improved HDR Video Compression Efficiency Using The Hevc StandardabstractThe human visual system's sensitivity to changes in colors varies based on the perceived color. In this work, we propose a chroma processing scheme that assigns more code-words to colors that our eyes are most sensitive to, so that perceived color differences generated by the quantization processes in the HDR video delivery pipeline are reduced. Performance evaluations showed that the proposed method significantly reduces the number of pixels with visible color differences as well as the mean error of the frames when compared to the original method. The proposed scheme improves the compression efficiency of the existing 10-bit Y'CbCrby an average of 8.16% and 30.36%, in terms of tPSNR-XYZ and DE100, respectively. The proposed scheme is easily implementable in hardware and can be interpreted by the current HEVC video coding standard using existing supplemental enhancement information (SEI) messages. Maryam Azimi, Panos Nasiopoulos, Mahsa T. Pourazad |
ICIP | 3 |
| 2020 | A Color Adjustment Method for HDR Display of Video Content Received Over Wireless Multimedia NetworksabstractBandwidth limitations in wireless networks may be prohibitive for transmitting High Dynamic Range (HDR) video content to end users to take advantage of the capabilities of HDR displays. Instead, the Standard Dynamic Range (SDR) version of the content may be transmitted, which is inverse tone mapped to the visually rich HDR format at the receiver end. One of the challenges in this approach is that the mapping process causes color shifts. Failing to address this color change, degrades the overall visual quality of the generated HDR video. In this paper, we propose a perception-based color adjustment method that is capable of preserving the hue of colors and produces HDR colors that closely follow their SDR counterparts, while causing negligible luminance change. Performance evaluations show that our method outperforms existing state-of-the-art color adjustment methods. Pedram Mohammadi, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung |
WiMob | 2 |
| 2019 | An Efficient Random Access Light Field Video Compression Utilizing Diagonal Inter-View PredictionabstractA more realistic Virtual and Augmented Reality is promised by the emergence of Light Field (LF) technology. Considering LF technology benefits come with an exponential increase in the amount of data required as well as the complexity of view access, the need for efficient compression and random access schemes is essential. In this study, a novel and efficient pseudo-sequence based LF video compression scheme is proposed that offers the best trade-off between coding and access complexity. The philosophy behind our proposed prediction structure is maximally utilizing the low-level frames in hierarchy prediction structure to achieve better random access at a minimal expense of compression efficiency. Nusrat Mehajabin, Sichen Roger Luo, Hao Wei Yu, Joseph Khoury, Jashandeep Kaur, Mahsa T. Pourazad |
ICIP | 6 |
| 2019 | A High Contrast Video Inverse Tone Mapping Operator for High Dynamic Range ApplicationsabstractConversion of existing image and video content to High Dynamic Range (HDR) using inverse Tone Mapping Operators (iTMOs) is expected to enable the HDR market and open new market opportunities for studios and content owners. In this paper, we propose a high contrast video iTMO that addresses shortcomings of existing approaches, yielding HDR video quality worth the expectations of the emerging HDR technology. Our approach is content adaptive and is able to convert SDR videos to HDR videos of any target dynamic range. Our approach follows the Human Visual System (HVS) characteristic of being more sensitive to luminance changes in dark areas than bright and normal ones, mapping each of these areas accordingly. Our subjective evaluations demonstrate that our proposed method on average outperforms existing iTMOs in terms of overall HDR visual quality. Pedram Mohammadi, Mahsa T. Pourazad, Panos Nasiopoulos |
ISCAS | 2 |
| 2018 | Video-based Human Fall Detection in Smart Homes Using Deep LearningabstractAutomatic human fall detection is a challenging task of healthcare in smart homes, and video cameras have been proved to be efficient in addressing this problem. Although existing methods perform relatively well, they are all built upon "hand-crafted" features, thus constraining the performance of the model to some presumed conditions and scenarios, and making it vulnerable to any deviation from the assumed settings. In this paper, we propose a deep-learning-based approach for human fall detection, using long short-term memory neural network. Our model is not restricted to any specific circumstances, and performance evaluations show that it outperforms all the existing methods. Anahita Shojaei-Hashemi, Panos Nasiopoulos, James J. Little, Mahsa T. Pourazad |
ISCAS | 4 |
| 2017 | Optimizing Non Constant Luminance into Constant Luminance for High Dynamic Range Video DistributionabstractTo improve compression efficiency, pixels are traditionally represented using a luma and two chroma values. Such a representation aims at separating light from color information. Two methods are usually considered for computing luma values: Non-Constant Luminance (NCL) and Constant Luminance (CL). CL equations have been derived from the luminous efficacy of the used gamut color primaries in the light linear domain. NCL applies the same equations but on perceptually encoded values, thus resulting in lower compression efficiency and hue shifts. However, given the higher hardware complexity for implementing CL, the common operational practice in legacy television distribution is to use NCL. In this paper, our motivation is to derive a new set of equations that provides the compression benefits of CL with the lower complexity of NCL for High Dynamic Range video distribution. Results show that the proposed method increases compression efficiency significantly over the NCL approach while maintaining NCL's cost complexity. Fujun Xie, Ronan Boitard, Mahsa T. Pourazad, Panos Nasiopoulos |
ICASSP | 3 |
| 2017 | A color gamut mapping scheme for backward compatible UHD video distributionabstractThe new Ultra High Definition (UHD) standard digital imagery can represent much more color information than High Definition (HD) and Standard Definition (SD). Currently most manufactured displays support UHD colors while UHD is being deployed for content production. However not all service providers have updated their pipeline thoroughly. Thus, the enduser that buys a UHD display would not be able to benefit from the wider UHD color range. In this paper, we propose an invertible gamut mapping from UHD colors to HD colors so that UHD displays can reconstruct UHD colors, while HD displays are addressed directly using legacy video delivery pipeline. The proposed color mapping scheme allows the mapped signal to be converted back to the original signal with minimal perceptual error so that the viewers' quality of experience (QoE) is preserved. Our method includes a parameter that adjusts the trade-off between the quality of the HD content and that of the UHD content. Our experiment results provide a guideline on how to strike a balance between color errors in the mapped signal and the retrieved one. Maryam Azimi, Timothee-Florian Bronner, Panos Nasiopoulos, Mahsa T. Pourazad |
ICC | 4 |
| 2017 | A learning-based visual saliency prediction model for stereoscopic 3D video (LBVS-3D)
Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos |
Multim. Tools Appl. | 2 |
| 2017 | Online-Learning-Based Mode Prediction Method for Quality Scalable Extension of the High Efficiency Video Coding (HEVC) StandardabstractSHVC, the scalable extension of High Efficiency Video Coding (HEVC), uses advanced inter-layer prediction features in addition to the advanced compression tools of HEVC to improve the compression performance. Using combined features has brought us improved compression performance at the cost of huge computational complexity for the SHVC encoder. This complexity is mainly because of the the inter/intra-prediction mode search of the coding units. The focus of this study is on developing an efficient complexity reduction for quality scalability of SHVC encoder, with the intention to facilitate the adoption of SHVC for real-time applications. In this regard, first, we build a probabilistic model that uses the mode information and motion homogeneity of already encoded blocks in the enhancement layer (EL) and the base layer to predict the probabilities of all the available inter/intra modes of the to-be-coded block in the EL. Then, we propose an online-learning-based fast mode, assigning (FMA) method that uses the proposed probabilistic model to predict the mode of the to-be-coded block in the EL. Performance evaluation shows that our proposed FMA method reduces the total execution time of the SHVC encoder by 45.40% on average compared with unmodified SHVC codec while maintaining the overall video quality. Hamid Reza Tohidypour, Hossein Bashashati, Mahsa T. Pourazad, Panos Nasiopoulos |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Chroma scaling for high dynamic range video compressionabstractColor pixel encoding optimizes the conversion of linear physical values of light into integer values. The efficiency of such encoding methods depends on a trade-off between the bit-depth used and the visible distortion introduced by quantization. This efficiency for different color pixel encoding approaches has been evaluated in literature, without considering the fact that before transmission to the end-user, color encoded content needs to be compressed using a video codec. Thus, to be efficient, a color pixel encoding scheme needs not only to achieve the lowest bit-depth, but also to allow for efficient video compression ratio. Yet, when compressing dark video sequences, the most efficient color pixel encoding scheme known as Y'DuDv requires much higher bit-rates, hence negating its high encoding efficiency. In this article, we propose a chroma scaling technique that adaptively restricts the bit-depth of the chroma channels for optimized encoding of High Dynamic Range content. Results show that the proposed scaling reduces Y'DuDv bit-rate requirements for dark content while preserving its high color accuracy. Ronan Boitard, Mahsa T. Pourazad, Panos Nasiopoulos |
ICASSP | 2 |
| 2016 | An efficient human visual system based quality metric for 3D video
Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos |
Multim. Tools Appl. | 2 |
| 2016 | Online-Learning-Based Complexity Reduction Scheme for 3D-HEVCabstract3-D High Efficiency Video Coding (HEVC) is a new emerging video compression standard for multiview video applications. This standard utilizes advanced interview prediction characteristics in addition to the prediction features of the HEVC standard for efficient encoding of multiview video content. While using combined features improves the compression performance by utilizing the correlation between the views captured from slightly different angles of the same scene, they also increase coding complexity. The focus of this paper is on developing an efficient complexity reduction scheme for 3D-HEVC, with the intention to facilitate the adoption of this upcoming standard, especially for real-time applications. In this regard, first, we introduce two ways to decrease the complexity of the inter-/ intra-mode search process of the to-be-encoded blocks in the dependent texture views (${\mathrm {DV}}_{t}\text{s}$ ) of 3D-HEVC. Then, we propose a hybrid complexity reduction scheme that utilizes the two-mode prediction approaches, motion information of the base texture view (BVt), and the rate distortion cost of the already encoded blocks in the BVt and DVt. The performance of our proposed scheme is tested for the case with two views (i.e., base view + dependent view). The evaluations confirm that our proposed hybrid complexity reduction scheme reduces the 3D-HEVC codec complexity by 67.70% on average for the DVt compared with the unmodified 3D-HEVC encoder, while maintaining the overall video quality. Compared with the state-of-the-art method, it reduces complexity by 25.74% on average. Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Human Visual System-Based Saliency Detection for High Dynamic Range ContentabstractThe human visual system (HVS) attempts to select salient areas to reduce cognitive processing efforts. Computational models of visual attention try to predict the most relevant and important areas of videos or images viewed by the human eye. Such models, in turn, can be applied to areas such as computer graphics, video coding, and quality assessment. Although several models have been proposed, only one of them is applicable to high dynamic range (HDR) image content, and no work has been done for HDR videos. Moreover, the main shortcoming of the existing models is that they cannot simulate the characteristics of HVS under the wide luminous range found in HDR content. This paper addresses these issues by presenting a computational approach to model the bottom-up visual saliency for HDR input by combining spatial and temporal visual features. An analysis of eye movement data affirms the effectiveness of the proposed model. Comparisons employing three well-known quantitative metrics show that the proposed model substantially improves predictions of visual attention for HDR content. Mahsa T. Pourazad, Panos Nasiopoulos |
IEEE Trans. Multim. | 2 |
| 2016 | Probabilistic Approach for Predicting the Size of Coding Units in the Quad-Tree Structure of the Quality and Spatial Scalable HEVCabstractThe scalable extension of HEVC (known as SHVC), the recent scalable video coding standard, results in an improved compression performance at the cost of significant increase in computational coding complexity. One of the main factors that contribute to the SHVC encoder complexity is choosing the best partitioning structure for the coding tree units (CTUs). Our study focuses on developing a scheme for predicting the CTU structure in the quality and spatial scalable extension of HEVC. The proposed scheme uses the CTU partitioning structure of the already encoded CTUs in the enhancement layers (ELs) and base layer (BL) to predict the coding unit sizes of the to-be-encoded CTUs in the EL. Performance evaluations confirm that our proposed complexity reduction scheme significantly reduces the execution time of the SHVC encoder, while maintaining the overall quality of the coded streams. Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos |
IEEE Trans. Multim. | 2 |
| 2014 | A low complexity mode decision approach for HEVC-based 3D video coding using a Bayesian methodabstractThe 3D extension of High Efficiency Video Coding (HEVC) standard (3D-HEVC) aims at improving coding efficiency by introducing new and unique approaches for utilizing correlations between the different views of a scene. Reported coding efficiency, however, comes at the expense of increased computational complexity. For real-time applications, reducing the computational complexity of 3D-HEVC is very important. In this paper, we propose an adaptive fast mode assigning method based on a Bayesian classifier that reduces 3D-HEVC's coding complexity by up to 51.95%, while maintaining the overall quality and bitrate. Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos |
ICASSP | 2 |
| 2014 | A similarity measure for analyzing human activities using human-object interaction contextabstractUnderstanding the context of human-object interactions plays an important role in human activity recognition. Modeling the interaction context is a challenging problem due to the large number of possible objects in the scene and the large number of ways these objects may seem to relate to human activities taking place in the scene. In addition, providing labeling information of the object and human body parts is a very difficult and labor intense part of the training process. In this paper, we use a new class of kernels for image/video data as an extension of string kernels for 2 and 3 dimensional signals to model the human body parts and objects interaction context. In contrast to similar works, the proposed method does not require labeling of the human body parts and objects in the scene for the learning process, making it more practical when dealing with large datasets. Our experimental results show that the proposed kernel efficiently models the context of human-object interactions in image/video sequences and results in improved performance when compared to state-of-the-art methods. S. Mohsen Amiri, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung |
ICIP | 2 |
| 2014 | Effect of eye dominance on the perception of stereoscopic 3D videoabstractAsymmetric schemes have widespread applications in the 3D video transmission pipeline. The significance of eye dominance becomes a concern when designing such schemes. In this paper, in order to investigate the effect of eye dominance on the perceptual 3D video quality, a database of representative asymmetric stereoscopic sequences is prepared and the overall 3D quality of these sequences is evaluated through subjective experiments. Experiment results showed that viewers find an asymmetric video more pleasant when the view with higher quality is projected to their dominant eye. Moreover, the eye dominance changes the mean opinion quality score by 16 % at most, a result caused by slight asymmetric video compression. For all other representative types of asymmetry, the statistical difference is much lower and in some cases even negligible. Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos |
ICIP | 2 |
| 2014 | Content and network-aware multicast over wireless networksabstractThis paper proposes content and network-aware redundancy allocation algorithms for channel coding and network coding to optimally deliver data and video multicast services over error prone wireless mesh networks. Each network node allocates redundancies for channel coding and network coding taking in to account the content properties, channel bandwidth and channel status to improve the end-to-end performance of data and video multicast applications. For data multicast applications, redundancies are allocated at each network node in such a way that the total amount of redundant bits transmitted is minimised. As for video multicast applications, redundancies are allocated considering the priority of video packets such that the probability of delivering high priority video packets is increased. This not only ensures the continuous playback of a video but also increases the received video quality. Simulation results for bandwidth sensitive data multicast applications exhibit up to 10× reduction of the required amount of redundant bits compared to reference schemes to achieve a 100% packet delivery ratio. Similarly, for delay sensitive video multicast applications, simulation results exhibit up to 3.5dB PSNR gains in the received video quality. Chamitha de Alwis, Hemantha Kodikara Arachchi, Warnakulasuriya Anil Chandana Fernando, Mahsa T. Pourazad |
QSHINE | 4 |
| 2014 | Compression of high dynamic range video using the HEVC and H.264/AVC standardsabstractThe existing video coding standards such as H.264/AVC and High Efficiency Video Coding (HEVC) have been designed based on the statistical properties of Low Dynamic Range (LDR) videos and are not accustomed to the characteristics of High Dynamic Range (HDR) content. In this study, we investigate the performance of the latest LDR video compression standard, HEVC, as well as the recent widely commercially used video compression standard, H.264/AVC, on HDR content. Subjective evaluations of results on an HDR display show that viewers clearly prefer the videos coded via an HEVC-based encoder to the ones encoded using an H.264/AVC encoder. In particular, HEVC outperforms H.264/AVC by an average of 10.18% in terms of mean opinion score and 25.08% in terms of bit rate savings. Amin Banitalebi-Dehkordi, Mehran Azimi, Mahsa T. Pourazad, Panos Nasiopoulos |
QSHINE | 3 |
| 2014 | Perceptual quality-driven resource allocation in energy-aware wireless video multicastingabstractIn conjunction with the rising trend towards consumption of resource-hungry multimedia content over wireless medium, efficient radio resource allocation (eRRA) represents an ongoing challenge to network operators. This paper explores the application of multicasting in light of an eRRA that considers the perceptual quality aspect of wireless video transmission such that transmission energy is minimized. A multiuser orthogonal frequency-division multiplexing (OFDM) environment is considered. The resource allocation scheme presented is based on the genetic algorithm (GA) and offers fairness through the balance of average perceptual quality and minimal total energy among all subscribers in multicasting groups. In order to assist with this balance, a utility function is introduced as a fitness function for the GA. The allocation scheme relies on a well-established video quality model (VQM) for assessment of user-perceived quality. Simulation results show that the allocation scheme helps to preserve the energy requirement at low levels despite the proportionate rise in the number of subscribers. Moreover, perceived quality is well-secured and maintained at high levels, with a consistent increase of the utility value. Reflecting on today's consumption patterns of popular video content, adoption of the presented scheme in multicasting would help lessen the carbon footprint emitted by wireless communications, while consumers' quality of experience (QoE) is maintained. Emad Danish, Warnakulasuriya Anil Chandana Fernando, Omar Abdul-Hameed, Mahsa T. Pourazad |
QSHINE | 4 |
| 2014 | QoE aware resource allocation for video communications over LTE based mobile networksabstractAs the limits of video compression and usable wireless radio resources are exhausted, providing increased protection to critical data is regarded as a way forward to increase the effective capacity for delivering video data. This paper explores the provisioning of selective protection in the physical layer to critical video data and evaluates its effectiveness when transmitted through a wireless multipath fading channel. In this paper, the transmission of HEVC encoded video through an LTE-A wireless channel is considered. HEVC encoded video data is ranked based on how often each area of the picture is referenced by subsequent frames within a GOP in the sequence. The critical video data is allotted to the most robust OFDM resource blocks (RBs), which are the radio resources in the time-frequency domain of the LTE-A physical layer, to provide superior protection. The RBs are ranked based on a prediction for their robustness against noise. Simulation results show that the proposed content aware resource allocation scheme helps to improve the objective video quality up to 37dB at lower channel SNR levels when compared against the reference system, which treats video data uniformly. Alternatively, with the proposed technique the transmitted signal power can be lowered by 30% without sacrificing video quality at the receiver. Ryan Perera, Warnakulasuriya Anil Chandana Fernando, Thanuja Mallikarachchi, Hemantha Kodikara Arachchi, Mahsa T. Pourazad |
QSHINE | 5 |
| 2013 | Non-intrusive human activity monitoring in a smart home environmentabstractNon-intrusive activity monitoring of occupants in a home environment plays an important role in developing the next generation of smart environments and remote health monitoring systems. One important challenge in this research area is lack of a comprehensive dataset. In this paper, we introduce a new dataset related to a smart home environment, which can be used for human activity recognition. In addition, we benchmark our proposed human action recognition algorithm and some other state-of-the-art methods using our dataset. S. Mohsen Amiri, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung |
Healthcom | 2 |
| 2013 | 3D video quality metric for mobile applicationsabstractIn this paper, we propose a new full-reference quality metric for mobile 3D content. Our method is modeled around the Human Visual System, fusing the information of both left and right channels, considering color components, the cyclopean views of the two videos and disparity. Our method is assessing the quality of 3D videos displayed on a mobile 3DTV, taking into account the effect of resolution, distance from the viewers' eyes, and dimensions of the mobile display. Performance evaluations showed that our mobile 3D quality metric monitors the degradation of quality caused by several representative types of distortion with 82% correlation with results of subjective tests, an accuracy much better than that of the state-of-the-art mobile 3D quality metric. Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos |
ICASSP | 2 |
| 2013 | Content adaptive complexity reduction scheme for quality/fidelity scalable HEVCabstractThere has been significant interest in developing a scalable version of the High Efficiency Video Coding (HEVC) standard. As expected, the HEVC scalable video version increases the complexity of the codec compared to the non-scalable counterpart. In this paper, we propose an adaptive early-termination interlayer motion prediction mode search that significantly reduces HEVC/SVC's coding complexity by up to 85.77%, while maintaining the overall bitrate. Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos |
ICASSP | 2 |
| 2013 | Encoding and communication energy consumption trade-off in H.264/AVC based video sensor networkabstractVideo sensor networks (VSN) offer an interesting platform for a distributed and flexible surveillance system. In such a system, video compression and wireless transmission are the major operations on each video node. For a battery-powered wireless video sensor, it is essential to maximize the power efficiency of these two operations. Currently, H.264/AVC is the most widely used ITU-T and ISO/IEC video coding standard. Previous works on determining the trade-off between compression and transmission that minimizes energy consumption consider oversimplified coding configurations, thus not taking full advantage of the flexibility and advanced features of H.264/AVC. Choosing the right configuration and setting parameters that lead to optimal encoding performance is of prime importance for video sensor network (VSN) applications, especially since VSN is constrained in terms of bandwidth and energy resources. This paper studies the relationship between the picture quality, the transmission rate, and the complexity of the encoder to expound the energy consumption trade-off between encoding and transmission in VSN. The results of our study can be used as guidelines in optimizing the overall power consumption of a VSN system as it detailed in the paper. Bambang A. B. Sarif, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung |
WOWMOM | 2 |
| 2012 | Random Forests Based View Generation for Multiview TVabstractThe appearance of multiview display systems in the consumer market is not far from reality. With technical knowledge in this field constantly improving, production of multiview content is the only other key factor that will determine the successful adoption of this technology. Multiview content can be generated from two or three views and their associated depth maps. Estimating a high quality depth map is challenging. Moreover transmission of depth map information requires extra bandwidth. In this study, we propose an effective algorithm, which utilizes a 3D visual attention model, multiple monocular depth cues and a fraction of depth information for estimating the whole depth map of the scene using the Random Forests (RF) machine learning algorithm. Having the estimated depth maps and stereo videos, other views may be synthesized. Performance evaluations have shown that the proposed method estimates high quality depth maps for stereo sequences from limited depth information. Implementation of our proposed technique in the future multiview pipeline eliminates the need for estimating and transmitting the whole depth map for all the views, producing high quality multiview content while reducing the required bandwidth. Mahsa T. Pourazad, Di Xu 0001, Panos Nasiopoulos |
ICTAI | 1 |
| 2011 | Effect of brightness on the quality of visual 3D perceptionabstractOver the years, a consensus has been reached that the introduction of 3D entertainment can only be a lasting success if the perceived image quality and the viewing comfort are better than those of conventional 2D television. There are different factors that affect the perceived quality of 3D content. In this paper, our objective is to obtain a good understanding of the effect that brightness has on the visual quality of 3D videos and compare it to that of the 2D. We capture outdoor and indoor scenes with different exposures and we perform subjective evaluation to investigate how brightness affects the perceived quality of the 3D experience. Mahsa T. Pourazad, Zicong Mai, Panos Nasiopoulos, Konstantinos N. Plataniotis, Rabab K. Ward |
ICIP | 1 |
| 2010 | Correcting unsynchronized zoom in 3D videoabstractWhen capturing 3D video with a stereoscopic camera setup, it is important for the cameras to be precisely aligned and synchronized. This is particularly difficult in transitions such as zooming where the camera parameters must be changed in unison, or else the perceived 3D effect will be degraded. In this paper we study the problem of unsynchronized zooming in 3D video. First, we present a subjective study that shows that the perceived quality of stereo video is greatly reduced if the two views are zoomed by different amounts. Next, we present a method for correcting zoom mismatch by applying cropping and scaling to ones of the views. Our method involves finding matching points between the left and right views, and performing least-squares regressions to estimate the amount of scaling and cropping required to make the views consistent. Experiments were performed on videos with digitally introduced zoom mismatch and videos with optical unsynchronized zoom. In both cases the results show that our method is highly accurate and produces videos without size differences or vertical parallax between the two views. Colin Doutre, Mahsa T. Pourazad, Alexis M. Tourapis, Panos Nasiopoulos, Rabab K. Ward |
ISCAS | 2 |
| 2009 | An efficient low random-access delay panorama-based multiview video coding schemeabstractWe present an efficient low random delay scheme for multiview video coding (MVC). In the proposed scheme, inter-view prediction (disparity estimation), which introduces time-consuming computations and random access delay to MVC, is replaced with a residue-stream coding process. Our algorithm transforms the middle view to a panoramic view of the scene. Then the residue streams are created as the difference of the luma and chroma values of overlapping regions of each view and the panoramic view. Finally the panoramic stream and all residue streams are encoded separately (simulcast coding). The hierarchical B picture prediction structure is implemented for coding each stream. Performance evaluations show that our proposed coding method outperforms the recent multiview video coding standard by up to 2.13 dB PSNR and enhances the compression ratio by 24.6%, while reducing random-access delay by 50%. Mahsa T. Pourazad, Panos Nasiopoulos, Rabab K. Ward |
ICIP | 1 |
| 2009 | Converting H.264-Derived Motion Information into Depth Map
Mahsa T. Pourazad, Panos Nasiopoulos, Rabab K. Ward |
MMM | 1 |