Panos Nasiopoulos

dblp:55/3871 · DBLP profile ↗
← Back
115ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-2654-8096ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 76 · 4 first-author · 11 since 2021Computer networks · 12 · 2 first-author · 2 since 2021Systems, architecture and hardware · 11 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Security and privacy · 2Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 StereoMamba+: A Novel Stereo Image Super-Resolution Framework With Adaptive Dependency Capture and Enhanced Feature Fusion
abstract
Stereo image Super-Resolution (SR) aims to enhance image resolution by leveraging complementary information in stereo pairs. Convolutional Neural Networks (CNNs), widely used in stereo image SR for their strong local pattern extraction capabilities, often fail to capture long-range dependencies critical for stereo correspondence. On the other hand, Swin Transformers have demonstrated superior performance in modeling long-range dependencies for stereo image SR tasks. However, their computational complexity scales quadratically with the window size, leading to a trade-off between global receptive fields and computational efficiency. To tackle these challenges, we propose StereoMamba+, a novel stereo image SR method designed to adaptively capture both local and global dependencies in stereo pairs. Leveraging the Mamba architecture as its backbone, StereoMamba+ integrates an Adaptive State Space Module (ASSM) that efficiently extracts and fuses global and local features, maintaining linear computational complexity. Additionally, a Gated Enhanced Feed-Forward Network (GEFN) selectively amplifies essential features and depth cues, and a Residual Frequency Block (RFB) is employed to capture global features in the frequency domain. To further enhance stereo correspondence, we introduce a Stereo Bi-Directional Cross Attention Module (SBCAM), aligning unique features along both horizontal and vertical epipolar lines to improve stereo consistency. Extensive experiments demonstrate that our proposed StereoMamba+ method achieves state-of-the-art performance on 2× and 4× stereo image SR tasks, delivering PSNR improvements of up to 0.45dB, while maintaining competitive parameter efficiency compared to existing methods.
Zhenchao Ma, Hamid Reza Tohidypour, Panos Nasiopoulos, Victor C. M. Leung
IEEE Trans. Multim.3
2025 StereoMamba: Enhancing Stereo Image Super-Resolution with Structured State Space Models and Bi-Directional Cross Attention
abstract
Stereo image super-resolution (SR) aims to enhance image resolution by leveraging the complementary information from stereo image pairs. While convolutional neural network (CNN)-based methods have traditionally dominated this field, they struggle with capturing long-range dependencies. Transformer-based approaches have shown improvements by better modeling long-range dependencies, but their computational complexity scales quadratically with respect to the window length. To address these challenges, in this paper we propose StereoMamba, a new stereo image super-resolution method built on Structured State Space Models (SSMs). StereoMamba leverages the Mamba architecture to effectively capture long-range dependencies and inter-view correlations in stereo image pairs. Additionally, we introduce a Stereo Bi-directional Cross-Attention Module (SBCAM) to further improve stereo view correlation. Extensive experiments show that StereoMamba consistently surpasses state-of-the-art methods across several public datasets.
Zhenchao Ma, Hamid Reza Tohidypour, Panos Nasiopoulos, Victor C. M. Leung
ICASSP3
2025 Explainable Orthogonal Attention Networks for EEG-based Analysis: Leveraging Disentangled Representations to Enhance Diagnosis
abstract
The complexity of EEG data presents significant challenges for accurate diagnosis in neurological conditions such as Alzheimer’s disease. In this paper, we introduce Explainable Orthogonal Attention Networks, a novel approach for EEG-based analysis that decouples spatial and temporal features to more effectively capture disease-related neural patterns. By leveraging orthogonal attention mechanisms, our model independently processes spatial relationships across EEG channels and temporal dynamics, enhancing both explainability and predictive performance. Our approach outperforms baselines, achieving superior performance in objective metrics, with a 14% relative improvement, while offering insights into the neural mechanisms underlying Alzheimer’s disease. Using attention maps and spectral analysis, we identified critical parietal and frontal contributions, along with EEG markers like elevated theta and reduced alpha power, commonly associated with Alzheimer’s disease. This method represents a significant step forward in developing explainable and high-performing EEG-based diagnostic tools. We will make the code and model’s weights publicly available upon publication at anonymized.
Ailar Mahdizadeh, Puria Azadi Moghadam, Shahriar Mirabbasi, Panos Nasiopoulos
ICASSP4
2025 Parameter-Efficient Federated Cooperative Learning for 3-D Object Detection in Autonomous Driving
abstract
In the rapidly evolving field of autonomous driving, accurately detecting and understanding dynamic environments remains a challenge. Federated learning (FL) offers a promising approach by integrating decentralized models from multiple connected autonomous vehicles (CAVs) to enhance the performance of deep-learning (DL)-based object detection methods. However, traditional FL faces hurdles, such as extensive data synchronization requirements, limited data variance, and high communication costs. This article introduces a federated cooperative learning framework that addresses these challenges by combining data from both CAVs and roadside units. The framework combines local cooperative perception with global FL through a parameter-efficient FL adapter and a lazy communication strategy, improving DL-based object detection capabilities across diverse driving scenarios while significantly reducing bandwidth requirements. We also present a novel multiagent-multitown dataset Vehicle-to-Everything-Fed, specifically developed to validate the effectiveness of our approach under various conditions. Notably, our framework retains 97.51% of the detection accuracy achieved by full-model FL, while utilizing only 1.4% of the bandwidth typically required, demonstrating substantial improvements over conventional FL strategies. This study underscores the potential of our tailored approach to substantially enhance autonomous vehicle technologies with minimal resource utilization.
Fangyuan Chi, Yixiao Wang 0001, Panos Nasiopoulos, Victor C. M. Leung
IEEE Internet Things J.3
2025 Multiagent Collaborative Decision-Making Using Small Vision-Language Models for Autonomous Driving
Fangyuan Chi, Yixiao Wang 0001, Panos Nasiopoulos, Victor C. M. Leung
IEEE Internet Things J.3
2024 Multi-Modal GPT-4 Aided Action Planning and Reasoning for Self-driving Vehicles
abstract
Explainable decision-making is critical for building trust in autonomous vehicles. We investigate the use of a pre-trained large language model (LLM) to derive comprehensible driving decisions from multi-modal time-series data captured by a monocular camera on an autonomous vehicle. Leveraging a graph-of-thought structure, the LLM learns policies that perform robustly while generating natural language rationales. We generate a novel multi-modal dataset with sequential images, scene labels, and driving actions. Results demonstrate our method produces human- understandable explanations for its driving choices, providing transparency. Our work indicates incorporating language-based reasoning enables accountable and transparent decision-making for self-driving cars, making LLM a potential solution for autonomous driving.
Fangyuan Chi, Yixiao Wang 0001, Panos Nasiopoulos, Victor C. M. Leung
ICASSP3
2023 Federated Semi-Supervised Learning for Object Detection in Autonomous Driving
abstract
One of the main challenges in designing deep learning networks for autonomous driving is the lack of labeled data. Recent trends that address this problem involve the use of unlabeled data. In this paper, we propose a unified semi-supervised and federated learning (FL) approach that is designed to offer cost efficient and practical training of deep learning object detection models for autonomous driving. In our implementation, we assume that each vehicle is given some well-labeled image data which are coupled with unlabeled image data captured by its cameras. Each of the vehicles has a local object detection model, which will be trained leveraging a semi-supervised learning method with both labeled and unlabeled data. The local model parameters are uploaded to a cloud server and aggregated to update a global FL model which in turn is shared with all the vehicles involved. Performance evaluations showed that our proposed approach is a promising solution as it allows continuous training and thus improved performance in autonomous driving.
Fangyuan Chi, Yixiao Wang 0001, Panos Nasiopoulos, Victor C. M. Leung, Mahsa T. Pourazad
ICASSP3
2023 Detecting Stable Diffusion Generated Images Using Frequency Artifacts: A Case Study on Disney-Style Art
abstract
The use of Stable Diffusion models to generate realistic images has become a popular topic in recent years. However, this technology has also raised concerns about the potential harm it may cause to the copyright holders, particularly in the realm of art where these synthesized images can closely resemble the original work. As these synthesized images are hard for humans to distinguish from authentic ones, it is of great importance to develop methods that may identify them. In this paper, we propose a deep learning-based approach to detect synthesized images using information in the frequency domain. Since there exists no well-established dataset of images synthesized by stable diffusion models, in order to train and evaluate our network we generated a representative dataset consisting of carefully selecting human-created authentic images and synthesized animation images generated by the Stable Diffusion models. We chose to use Disney-style animated content for our case study, given its significance in the realm of intellectual property protection. Experimental results demonstrated that our proposed model outperforms humans and other state-of-the-art methods, achieving an accuracy rate of 99.46%.
Junbin Zhang 0002, Yixiao Wang 0001, Hamid Reza Tohidypour, Panos Nasiopoulos
ICIP4
2023 A Spatial Calibrated and Colour Corrected Light Field Outdoor Video Dataset from a $5 \times 5$ Dense Camera Array
abstract
In this paper, a new and calibrated light field (LF) video dataset is introduced, which focuses on outdoor scenes and objects. Each video stream is 10 seconds long and it is captured with a dense camera array that consists of$5\times 5$camera modules in$1640\times 1232$resolution at 40 frames per second. As multiple cameras in an array setup may suffer from various conditions of camera settings, lens structure, and lighting variations, the resulting images can be negatively affected by geometric distortion and colour difference. To address that, a unified calibration method involving both spatial calibration and colour correction is employed to correct inconsistences and achieve a better image quality with reduced image distortion. This video dataset would be suitable for further research and investigation of a variety LF applications, such as autonomous driving and immersive media.
Yixiao Wang 0001, Nusrat Mehajabin, Hamid Reza Tohidypour, Jerry Song, Menghong Huang, Behnoosh Babaghorbani, Zuhao Chen, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung
ISCAS9
2023 A Novel No-Reference HD Video Quality Metric Based on Perceptual Temporal Pooling
abstract
Impressive advancements in capturing, display, and broadcasting technologies significantly elevate image and video quality, and with that the need for designing new reference and no-reference image and video quality metrics. One of the latest and perceptually accurate video quality metrics is the Video Multi-Method Assessment Fusion (VMAF) method. However, VMAF considers the temporal nature of video using basic average temporal pooling, an approach that falls short from human perception. In this paper, we introduce a new no-reference video quality metric that uses deep learning to extract spatial features and a unique temporal pooling approach to accurately predict the visual quality score. To this end, first we created a video quality dataset that consists of high-resolution, 20s-long test video clips compressed at several different bitrates. These videos were labeled based on subjective evaluations and were used to determine the perceptual importance of frames in our temporal pooling scheme. Evaluations showed that our proposed approach achieved correlation of 90.55% with human perception and outperformed the state-of-the-art VMAF approach by 15.63% accuracy.
Hamid Reza Tohidypour, Yixiao Wang 0001, Panos Nasiopoulos, Mahsa T. Pourazad
SMC4
2022 Real-Time Deep Learning based Road Deterioration Detection for Smart Cities
abstract
Timely road condition inspection and maintenance are key components of infrastructure management for smart cities, as they reduce traffic congestion, accidents and repairing costs. Traditional road inspection methods that employ vibrations and/or laser scanning for detecting road deterioration use expensive equipment and dedicated municipality vehicles. Recently, computer vision techniques and artificial intelligence are emerging as alternative solutions to traditional approaches for road condition detection, offering more flexibility, higher accuracy, and overall lower cost. In this paper, we utilize convolutional neural network-based and vision transformer-based object detection models to accurately identify road deteriorations namely, potholes, cracks, and alligators. We compare four different state-of-the-art models in terms of detection accuracy and speed. Performance evaluations have shown that, on the same dataset the Swin Transformer model outperformed the other state-of-the-art methods by a substantial margin. With 74% detection accuracy, and 42 frames per second processing speed Swin Transformer exceled over EfficentDet, YOLOv4, and YOLOX. We also present a new comprehensive and balanced large-scale road condition dataset of 27,298 annotated images, captured by ordinary car cameras.
Nusrat Mehajabin, Zhenchao Ma, Yixiao Wang 0001, Hamid Reza Tohidypour, Panos Nasiopoulos
WiMob5
2022 A Security-Centric Deep Learning Enabled Camera Solution for Real-Time Human Fall Detection
abstract
Automatic human real-time fall detection is a challenging task in remote healthcare, demanding a non-intrusive, secure and affordable solution. In this paper, we present a real-time hardware system that uses a deep learning model for fall detection embedded in a color camera. To reduce the startup delay and achieve real-time performance for the inference phase, we optimized our model using TensorRT. In addition, we addressed the board memory limitation using virtual memory and linear memory allocation and garbage collection. Moreover, GStreamer was used to perform most of the video processing using Jetson's GPU. Our live evaluation shows that our system achieved the accuracy of 84.44% and real-time performance.
Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos
WiMob3
2022 An Efficient Pseudo-Sequence-Based Light Field Video Coding Utilizing View Similarities for Prediction Structure
abstract
Light Field (LF) video technology is a step towards offering a better immersive experience through on-demand refocusing and perspective viewing. However, the significant increase in captured data makes the need for efficient compression of paramount importance. In this paper, we proposed two prediction structures and coding orders that efficiently compress LF video content using the existing HEVC standard. This is achieved by utilizing horizontal and vertical correlation among the views for better inter-view prediction. To assess the performance of the schemes, ten publicly available and widely used LF video sequences were used. Our first method is highly suitable for applications demanding high compression efficiency. It outperforms the best pseudo-sequence-based compression technique to date by up to 17% in bitrate reduction while being scalable in the number of views. The second method is proficient for low computational and random-access complexity to any arbitrary view in the light field video. It offers 10% faster decoding and 20% lower random-access complexity compared to the best existing technique.
Nusrat Mehajabin, Mahsa T. Pourazad, Panos Nasiopoulos
IEEE Trans. Circuits Syst. Video Technol.3
2021 Learning-Based Light Field View Synthesis for Efficient Transmission and Storage
abstract
One of the main advantages of Light Field (LF) technology is that it provides a truly immersive experience, critical for computer vision, autonomous driving and medical applications. However, one of the main problems of light field is the size of the data captured, which significantly increases bandwidth requirements. In this paper, we introduce a learning-based LF view synthesis approach for efficient transmission and storage, fundamental for performing remote surgery and storing data. This is achieved by dropping specific views at the transmitting end and then efficiently synthesizing them at the receiver end. Our deep learning approach uses the epipolar image plane (EPI) information to ensure smooth disparity between the generated and original views. We consider plenoptic, synthetic LF content and camera array implementations which support different baseline settings. Experimental results show that our proposed method outperforms state-of-the-art light field view synthesis techniques, offering improved visual quality for the generated views.
Abrar Wafa, Mahsa T. Pourazad, Panos Nasiopoulos
ICIP3
2021 Improving compression efficiency of HEVC using perceptual coding
Sima Valizadeh, Panos Nasiopoulos, Rabab K. Ward
Multim. Tools Appl.2
2021 A Perception-Based Inverse Tone Mapping Operator for High Dynamic Range Video Applications
abstract
The drastic visual improvements introduced by High Dynamic Range (HDR) technologies open new markets for a wide range of industries. Among them, a significant opportunity is offered to owners of legacy Standard Dynamic Range (SDR) content that can be converted to the new standard to take advantage of the enhanced capabilities of the HDR displays. Similarly, since SDR broadcasting infrastructure will continue to be around for the time being, such conversion process is becoming an obvious necessity. To this end, different approaches have tried to efficiently convert SDR images and videos to HDR format, a procedure well known as inverse Tone Mapping. In this paper, we propose a novel high visual quality video inverse Tone Mapping Operator (iTMO) that addresses the inadequacies of the state-of-the-art methods, resulting in high visual quality HDR videos that match the capabilities of the HDR technology. Our approach is based on human visual perception and employs a segmentation method according to the Human Visual System (HVS) sensitivity to brightness changes in different regions of the frame and constructs the mapping curve using the brightness distribution information of these regions. Our iTMO uses a hybrid approach to achieve an optimal balance between the overall contrast and brightness of the output HDR frame by maximizing a weighted sum of contrast and brightness difference between input SDR and generated HDR frame. The proposed iTMO works equally well for all levels of brightness, eliminating any visual artifacts by dynamically maintaining changes at non-perceivable levels. Subjective and objective evaluations validated the superior visual performance of our proposed iTMO over state-of-the-art methods.
Pedram Mohammadi, Mahsa T. Pourazad, Panos Nasiopoulos
IEEE Trans. Circuits Syst. Video Technol.3
2021 A Fully Automatic Content Adaptive Inverse Tone Mapping Operator With Improved Color Accuracy
abstract
High Dynamic Range (HDR) technology offers a higher visual quality compared to its Standard Dynamic Range (SDR) counterpart, as it tries to imitate the way our eyes perceive brightness and color information. Converting SDR content to HDR format - using inverse Tone Mapping Operators (iTMOs) - to take advantage of the superior visual quality offered by HDR displays, is an attractive proposition to SDR content owners and real-time broadcasters. In this paper, we propose a novel content adaptive iTMO that works in the perceptual domain to model the sensitivity of the human eye to brightness changes in different areas of a scene. To preserve the overall visual impression, our proposed iTMO utilizes an entropy-based brightness segmentation, which also makes our method content adaptive. In addition, we propose a novel perception-based color adjustment method that can maintain the color accuracy between input SDR and generated HDR frames. By performing the color adjustment in the perceptual domain, our iTMO prevents hue shifts and generates HDR colors that closely follow their SDR counterparts. Our subjective evaluations indicate that our proposed method outperforms other state-of-the-art methods by an average of 81% in terms of visual quality, and 76% in terms of how closely the HDR colors match their SDR counterparts. In addition to subjective evaluations, we also performed objective evaluations using the HDR-VDP 2.2 and PU-SSIM metrics and concluded that, on average, our proposed iTMO outperforms the state-of-the-art methods in terms of these two metrics.
Pedram Mohammadi, Mahsa T. Pourazad, Panos Nasiopoulos
IEEE Trans. Circuits Syst. Video Technol.3
2020 A Novel Chroma Representation For Improved HDR Video Compression Efficiency Using The Hevc Standard
abstract
The human visual system's sensitivity to changes in colors varies based on the perceived color. In this work, we propose a chroma processing scheme that assigns more code-words to colors that our eyes are most sensitive to, so that perceived color differences generated by the quantization processes in the HDR video delivery pipeline are reduced. Performance evaluations showed that the proposed method significantly reduces the number of pixels with visible color differences as well as the mean error of the frames when compared to the original method. The proposed scheme improves the compression efficiency of the existing 10-bit Y'CbCrby an average of 8.16% and 30.36%, in terms of tPSNR-XYZ and DE100, respectively. The proposed scheme is easily implementable in hardware and can be interpreted by the current HEVC video coding standard using existing supplemental enhancement information (SEI) messages.
Maryam Azimi, Panos Nasiopoulos, Mahsa T. Pourazad
ICIP2
2020 A Color Adjustment Method for HDR Display of Video Content Received Over Wireless Multimedia Networks
abstract
Bandwidth limitations in wireless networks may be prohibitive for transmitting High Dynamic Range (HDR) video content to end users to take advantage of the capabilities of HDR displays. Instead, the Standard Dynamic Range (SDR) version of the content may be transmitted, which is inverse tone mapped to the visually rich HDR format at the receiver end. One of the challenges in this approach is that the mapping process causes color shifts. Failing to address this color change, degrades the overall visual quality of the generated HDR video. In this paper, we propose a perception-based color adjustment method that is capable of preserving the hue of colors and produces HDR colors that closely follow their SDR counterparts, while causing negligible luminance change. Performance evaluations show that our method outperforms existing state-of-the-art color adjustment methods.
Pedram Mohammadi, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung
WiMob3
2019 A High Contrast Video Inverse Tone Mapping Operator for High Dynamic Range Applications
abstract
Conversion of existing image and video content to High Dynamic Range (HDR) using inverse Tone Mapping Operators (iTMOs) is expected to enable the HDR market and open new market opportunities for studios and content owners. In this paper, we propose a high contrast video iTMO that addresses shortcomings of existing approaches, yielding HDR video quality worth the expectations of the emerging HDR technology. Our approach is content adaptive and is able to convert SDR videos to HDR videos of any target dynamic range. Our approach follows the Human Visual System (HVS) characteristic of being more sensitive to luminance changes in dark areas than bright and normal ones, mapping each of these areas accordingly. Our subjective evaluations demonstrate that our proposed method on average outperforms existing iTMOs in terms of overall HDR visual quality.
Pedram Mohammadi, Mahsa T. Pourazad, Panos Nasiopoulos
ISCAS3
2018 Virtual View Color Estimation for Free Viewpoint TV Applications Using Gaussian Mixture Model
abstract
Free Viewpoint TV applications extensively use view synthesis technologies to generate virtual views. Mismatch in color and brightness between the original views captured by far apart cameras in this setting, is a real challenge as they result in visual discontinuities in the synthesized views and affect the overall viewer's quality of experience. In this paper, we present a novel color estimation algorithm specifically designed for FTV applications. Our method is based on Gaussian Mixture Model (GMM) histogram approximation and estimates the color of virtual views based on their relative position in space and the color of the real views. Subjective tests confirm that our approach significantly reduces the color related artifacts in the virtual views, which results in increased overall visual quality.
Ilya Ganelin, Panos Nasiopoulos
ICIP2
2018 Video-based Human Fall Detection in Smart Homes Using Deep Learning
abstract
Automatic human fall detection is a challenging task of healthcare in smart homes, and video cameras have been proved to be efficient in addressing this problem. Although existing methods perform relatively well, they are all built upon "hand-crafted" features, thus constraining the performance of the model to some presumed conditions and scenarios, and making it vulnerable to any deviation from the assumed settings. In this paper, we propose a deep-learning-based approach for human fall detection, using long short-term memory neural network. Our model is not restricted to any specific circumstances, and performance evaluations show that it outperforms all the existing methods.
Anahita Shojaei-Hashemi, Panos Nasiopoulos, James J. Little, Mahsa T. Pourazad
ISCAS2
2018 Saliency inspired quality assessment of stereoscopic 3D video
Amin Banitalebi-Dehkordi, Panos Nasiopoulos
Multim. Tools Appl.2
2018 Perceptual rate distortion optimization of 3D-HEVC using PSNR-HVS
Sima Valizadeh, Panos Nasiopoulos, Rabab K. Ward
Multim. Tools Appl.2
2017 Optimizing Non Constant Luminance into Constant Luminance for High Dynamic Range Video Distribution
abstract
To improve compression efficiency, pixels are traditionally represented using a luma and two chroma values. Such a representation aims at separating light from color information. Two methods are usually considered for computing luma values: Non-Constant Luminance (NCL) and Constant Luminance (CL). CL equations have been derived from the luminous efficacy of the used gamut color primaries in the light linear domain. NCL applies the same equations but on perceptually encoded values, thus resulting in lower compression efficiency and hue shifts. However, given the higher hardware complexity for implementing CL, the common operational practice in legacy television distribution is to use NCL. In this paper, our motivation is to derive a new set of equations that provides the compression benefits of CL with the lower complexity of NCL for High Dynamic Range video distribution. Results show that the proposed method increases compression efficiency significantly over the NCL approach while maintaining NCL's cost complexity.
Fujun Xie, Ronan Boitard, Mahsa T. Pourazad, Panos Nasiopoulos
ICASSP4
2017 A color gamut mapping scheme for backward compatible UHD video distribution
abstract
The new Ultra High Definition (UHD) standard digital imagery can represent much more color information than High Definition (HD) and Standard Definition (SD). Currently most manufactured displays support UHD colors while UHD is being deployed for content production. However not all service providers have updated their pipeline thoroughly. Thus, the enduser that buys a UHD display would not be able to benefit from the wider UHD color range. In this paper, we propose an invertible gamut mapping from UHD colors to HD colors so that UHD displays can reconstruct UHD colors, while HD displays are addressed directly using legacy video delivery pipeline. The proposed color mapping scheme allows the mapped signal to be converted back to the original signal with minimal perceptual error so that the viewers' quality of experience (QoE) is preserved. Our method includes a parameter that adjusts the trade-off between the quality of the HD content and that of the UHD content. Our experiment results provide a guideline on how to strike a balance between color errors in the mapped signal and the retrieved one.
Maryam Azimi, Timothee-Florian Bronner, Panos Nasiopoulos, Mahsa T. Pourazad
ICC3
2017 A learning-based visual saliency prediction model for stereoscopic 3D video (LBVS-3D)
Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos
Multim. Tools Appl.3
2017 Online-Learning-Based Mode Prediction Method for Quality Scalable Extension of the High Efficiency Video Coding (HEVC) Standard
abstract
SHVC, the scalable extension of High Efficiency Video Coding (HEVC), uses advanced inter-layer prediction features in addition to the advanced compression tools of HEVC to improve the compression performance. Using combined features has brought us improved compression performance at the cost of huge computational complexity for the SHVC encoder. This complexity is mainly because of the the inter/intra-prediction mode search of the coding units. The focus of this study is on developing an efficient complexity reduction for quality scalability of SHVC encoder, with the intention to facilitate the adoption of SHVC for real-time applications. In this regard, first, we build a probabilistic model that uses the mode information and motion homogeneity of already encoded blocks in the enhancement layer (EL) and the base layer to predict the probabilities of all the available inter/intra modes of the to-be-coded block in the EL. Then, we propose an online-learning-based fast mode, assigning (FMA) method that uses the proposed probabilistic model to predict the mode of the to-be-coded block in the EL. Performance evaluation shows that our proposed FMA method reduces the total execution time of the SHVC encoder by 45.40% on average compared with unmodified SHVC codec while maintaining the overall video quality.
Hamid Reza Tohidypour, Hossein Bashashati, Mahsa T. Pourazad, Panos Nasiopoulos
IEEE Trans. Circuits Syst. Video Technol.4
2016 Chroma scaling for high dynamic range video compression
abstract
Color pixel encoding optimizes the conversion of linear physical values of light into integer values. The efficiency of such encoding methods depends on a trade-off between the bit-depth used and the visible distortion introduced by quantization. This efficiency for different color pixel encoding approaches has been evaluated in literature, without considering the fact that before transmission to the end-user, color encoded content needs to be compressed using a video codec. Thus, to be efficient, a color pixel encoding scheme needs not only to achieve the lowest bit-depth, but also to allow for efficient video compression ratio. Yet, when compressing dark video sequences, the most efficient color pixel encoding scheme known as Y'DuDv requires much higher bit-rates, hence negating its high encoding efficiency. In this article, we propose a chroma scaling technique that adaptively restricts the bit-depth of the chroma channels for optimized encoding of High Dynamic Range content. Results show that the proposed scaling reduces Y'DuDv bit-rate requirements for dark content while preserving its high color accuracy.
Ronan Boitard, Mahsa T. Pourazad, Panos Nasiopoulos
ICASSP3
2016 An efficient human visual system based quality metric for 3D video
Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos
Multim. Tools Appl.3
2016 Online-Learning-Based Complexity Reduction Scheme for 3D-HEVC
abstract
3-D High Efficiency Video Coding (HEVC) is a new emerging video compression standard for multiview video applications. This standard utilizes advanced interview prediction characteristics in addition to the prediction features of the HEVC standard for efficient encoding of multiview video content. While using combined features improves the compression performance by utilizing the correlation between the views captured from slightly different angles of the same scene, they also increase coding complexity. The focus of this paper is on developing an efficient complexity reduction scheme for 3D-HEVC, with the intention to facilitate the adoption of this upcoming standard, especially for real-time applications. In this regard, first, we introduce two ways to decrease the complexity of the inter-/ intra-mode search process of the to-be-encoded blocks in the dependent texture views (${\mathrm {DV}}_{t}\text{s}$ ) of 3D-HEVC. Then, we propose a hybrid complexity reduction scheme that utilizes the two-mode prediction approaches, motion information of the base texture view (BVt), and the rate distortion cost of the already encoded blocks in the BVt and DVt. The performance of our proposed scheme is tested for the case with two views (i.e., base view + dependent view). The evaluations confirm that our proposed hybrid complexity reduction scheme reduces the 3D-HEVC codec complexity by 67.70% on average for the DVt compared with the unmodified 3D-HEVC encoder, while maintaining the overall video quality. Compared with the state-of-the-art method, it reduces complexity by 25.74% on average.
Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos
IEEE Trans. Circuits Syst. Video Technol.3
2016 Human Visual System-Based Saliency Detection for High Dynamic Range Content
abstract
The human visual system (HVS) attempts to select salient areas to reduce cognitive processing efforts. Computational models of visual attention try to predict the most relevant and important areas of videos or images viewed by the human eye. Such models, in turn, can be applied to areas such as computer graphics, video coding, and quality assessment. Although several models have been proposed, only one of them is applicable to high dynamic range (HDR) image content, and no work has been done for HDR videos. Moreover, the main shortcoming of the existing models is that they cannot simulate the characteristics of HVS under the wide luminous range found in HDR content. This paper addresses these issues by presenting a computational approach to model the bottom-up visual saliency for HDR input by combining spatial and temporal visual features. An analysis of eye movement data affirms the effectiveness of the proposed model. Comparisons employing three well-known quantitative metrics show that the proposed model substantially improves predictions of visual attention for HDR content.
Mahsa T. Pourazad, Panos Nasiopoulos
IEEE Trans. Multim.3
2016 Probabilistic Approach for Predicting the Size of Coding Units in the Quad-Tree Structure of the Quality and Spatial Scalable HEVC
abstract
The scalable extension of HEVC (known as SHVC), the recent scalable video coding standard, results in an improved compression performance at the cost of significant increase in computational coding complexity. One of the main factors that contribute to the SHVC encoder complexity is choosing the best partitioning structure for the coding tree units (CTUs). Our study focuses on developing a scheme for predicting the CTU structure in the quality and spatial scalable extension of HEVC. The proposed scheme uses the CTU partitioning structure of the already encoded CTUs in the enhancement layers (ELs) and base layer (BL) to predict the coding unit sizes of the to-be-encoded CTUs in the EL. Performance evaluations confirm that our proposed complexity reduction scheme significantly reduces the execution time of the SHVC encoder, while maintaining the overall quality of the coded streams.
Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos
IEEE Trans. Multim.3
2014 A low complexity mode decision approach for HEVC-based 3D video coding using a Bayesian method
abstract
The 3D extension of High Efficiency Video Coding (HEVC) standard (3D-HEVC) aims at improving coding efficiency by introducing new and unique approaches for utilizing correlations between the different views of a scene. Reported coding efficiency, however, comes at the expense of increased computational complexity. For real-time applications, reducing the computational complexity of 3D-HEVC is very important. In this paper, we propose an adaptive fast mode assigning method based on a Bayesian classifier that reduces 3D-HEVC's coding complexity by up to 51.95%, while maintaining the overall quality and bitrate.
Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos
ICASSP3
2014 A similarity measure for analyzing human activities using human-object interaction context
abstract
Understanding the context of human-object interactions plays an important role in human activity recognition. Modeling the interaction context is a challenging problem due to the large number of possible objects in the scene and the large number of ways these objects may seem to relate to human activities taking place in the scene. In addition, providing labeling information of the object and human body parts is a very difficult and labor intense part of the training process. In this paper, we use a new class of kernels for image/video data as an extension of string kernels for 2 and 3 dimensional signals to model the human body parts and objects interaction context. In contrast to similar works, the proposed method does not require labeling of the human body parts and objects in the scene for the learning process, making it more practical when dealing with large datasets. Our experimental results show that the proposed kernel efficiently models the context of human-object interactions in image/video sequences and results in improved performance when compared to state-of-the-art methods.
S. Mohsen Amiri, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung
ICIP3
2014 Effect of eye dominance on the perception of stereoscopic 3D video
abstract
Asymmetric schemes have widespread applications in the 3D video transmission pipeline. The significance of eye dominance becomes a concern when designing such schemes. In this paper, in order to investigate the effect of eye dominance on the perceptual 3D video quality, a database of representative asymmetric stereoscopic sequences is prepared and the overall 3D quality of these sequences is evaluated through subjective experiments. Experiment results showed that viewers find an asymmetric video more pleasant when the view with higher quality is projected to their dominant eye. Moreover, the eye dominance changes the mean opinion quality score by 16 % at most, a result caused by slight asymmetric video compression. For all other representative types of asymmetry, the statistical difference is much lower and in some cases even negligible.
Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos
ICIP3
2014 Compression of high dynamic range video using the HEVC and H.264/AVC standards
abstract
The existing video coding standards such as H.264/AVC and High Efficiency Video Coding (HEVC) have been designed based on the statistical properties of Low Dynamic Range (LDR) videos and are not accustomed to the characteristics of High Dynamic Range (HDR) content. In this study, we investigate the performance of the latest LDR video compression standard, HEVC, as well as the recent widely commercially used video compression standard, H.264/AVC, on HDR content. Subjective evaluations of results on an HDR display show that viewers clearly prefer the videos coded via an HEVC-based encoder to the ones encoded using an H.264/AVC encoder. In particular, HEVC outperforms H.264/AVC by an average of 10.18% in terms of mean opinion score and 25.08% in terms of bit rate savings.
Amin Banitalebi-Dehkordi, Mehran Azimi, Mahsa T. Pourazad, Panos Nasiopoulos
QSHINE4
2013 Non-intrusive human activity monitoring in a smart home environment
abstract
Non-intrusive activity monitoring of occupants in a home environment plays an important role in developing the next generation of smart environments and remote health monitoring systems. One important challenge in this research area is lack of a comprehensive dataset. In this paper, we introduce a new dataset related to a smart home environment, which can be used for human activity recognition. In addition, we benchmark our proposed human action recognition algorithm and some other state-of-the-art methods using our dataset.
S. Mohsen Amiri, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung
Healthcom3
2013 3D video quality metric for mobile applications
abstract
In this paper, we propose a new full-reference quality metric for mobile 3D content. Our method is modeled around the Human Visual System, fusing the information of both left and right channels, considering color components, the cyclopean views of the two videos and disparity. Our method is assessing the quality of 3D videos displayed on a mobile 3DTV, taking into account the effect of resolution, distance from the viewers' eyes, and dimensions of the mobile display. Performance evaluations showed that our mobile 3D quality metric monitors the degradation of quality caused by several representative types of distortion with 82% correlation with results of subjective tests, an accuracy much better than that of the state-of-the-art mobile 3D quality metric.
Amin Banitalebi-Dehkordi, Mahsa T. Pourazad, Panos Nasiopoulos
ICASSP3
2013 Content adaptive complexity reduction scheme for quality/fidelity scalable HEVC
abstract
There has been significant interest in developing a scalable version of the High Efficiency Video Coding (HEVC) standard. As expected, the HEVC scalable video version increases the complexity of the codec compared to the non-scalable counterpart. In this paper, we propose an adaptive early-termination interlayer motion prediction mode search that significantly reduces HEVC/SVC's coding complexity by up to 85.77%, while maintaining the overall bitrate.
Hamid Reza Tohidypour, Mahsa T. Pourazad, Panos Nasiopoulos
ICASSP3
2013 Encoding and communication energy consumption trade-off in H.264/AVC based video sensor network
abstract
Video sensor networks (VSN) offer an interesting platform for a distributed and flexible surveillance system. In such a system, video compression and wireless transmission are the major operations on each video node. For a battery-powered wireless video sensor, it is essential to maximize the power efficiency of these two operations. Currently, H.264/AVC is the most widely used ITU-T and ISO/IEC video coding standard. Previous works on determining the trade-off between compression and transmission that minimizes energy consumption consider oversimplified coding configurations, thus not taking full advantage of the flexibility and advanced features of H.264/AVC. Choosing the right configuration and setting parameters that lead to optimal encoding performance is of prime importance for video sensor network (VSN) applications, especially since VSN is constrained in terms of bandwidth and energy resources. This paper studies the relationship between the picture quality, the transmission rate, and the complexity of the encoder to expound the energy consumption trade-off between encoding and transmission in VSN. The results of our study can be used as guidelines in optimizing the overall power consumption of a VSN system as it detailed in the paper.
Bambang A. B. Sarif, Mahsa T. Pourazad, Panos Nasiopoulos, Victor C. M. Leung
WOWMOM3
2013 Visually Favorable Tone-Mapping With High Compression Performance in Bit-Depth Scalable Video Coding
abstract
In bit-depth scalable video coding, the tone-mapping scheme used to convert high-bit-depth to eight-bit videos is an essential yet very often ignored component. In this paper, we demonstrate that an appropriate choice of a tone-mapping operator can improve the coding efficiency of bit-depth scalable encoders. We present a new tone-mapping scheme that delivers superior compression efficiency while adhering to a predefined base layer perceptual quality. We develop numerical models that estimate the base layer bit-rate (Rb), the enhancement layer bitrate (Re), and the mismatch (QL) between the resulting low dynamic range (LDR) base-layer signal and the predefined base layer representation. Our proposed tone curve is given by the solution of an optimization problem which minimizes a weighted sum of Rb, Re, and QL. The problem formulation also considers the temporal effect of tone-mapping by adding a constraint to the optimization problem that suppresses flickering artifacts. We also propose a technique with which to tone-map a high-bit-depth video directly in a compression-friendly color space (e.g., one luma and two chroma channels) without converting to the RGB domain. Experimental results show that we can save up to 40% of the total bit-rate (or 3.5 dB PSNR improvement for the same bitrate), and, in general, about 20% bit-rate savings can be achieved.
Zicong Mai, Hassan Mansour, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Multim.3
2012 Computationally efficient tone-mapping of high-bit-depth video in the YCbCr domain
abstract
High dynamic range (HDR) video content is able to provide superior picture quality. This is because the representation of HDR signals requires more bits than the 8-bit low dynamic range (LDR) video. Tone-mapping is the process that converts HDR to LDR signals. Most tone-mapping methods are derived only for the luminance component. This mapping function is then used in each of the R, G and B components to generate the LDR color image. This color tone mapping correction approach, however, cannot be directly applied to most videos since they are usually encoded in the YCbCr color space. This paper addresses this problem and proposes a tone-mapping method that is applied directly on the YCbCr signals. Experimental results show that the Cb and the Cr signals generated by our method are almost identical to those produced with the conventional pipeline up to round-off errors, with average PSNR at about 55 dB and average SSIM at 0.991. By avoiding all the round-off errors introduced in the conventional method, our approach provides a more accurate LDR picture. Moreover, the proposed solution has significantly lower complexity because it bypasses the processes such as color space transformation and up-sampling which are required by the conventional method.
Zicong Mai, Panos Nasiopoulos, Rabab K. Ward
ICASSP2
2012 Non-negative sparse coding for human action recognition
abstract
We consider the problem of human action recognition using non-negative sparse representation of extracted features from spatiotemporal video patches. Our algorithm trains dictionaries for the calculation of a non-negative sparse representation for feature vectors and uses a linear Support Vector Machine (SVM) to distinguish between different actions. We evaluate the performance of the proposed techniques by using two human action datasets (KTH and IXMAS). In both cases, the proposed technique outperforms state-of-the-art techniques, achieving 100% accuracy on the KTH dataset.
S. Mohsen Amiri, Panos Nasiopoulos, Victor C. M. Leung
ICIP2
2012 Random Forests Based View Generation for Multiview TV
abstract
The appearance of multiview display systems in the consumer market is not far from reality. With technical knowledge in this field constantly improving, production of multiview content is the only other key factor that will determine the successful adoption of this technology. Multiview content can be generated from two or three views and their associated depth maps. Estimating a high quality depth map is challenging. Moreover transmission of depth map information requires extra bandwidth. In this study, we propose an effective algorithm, which utilizes a 3D visual attention model, multiple monocular depth cues and a fraction of depth information for estimating the whole depth map of the scene using the Random Forests (RF) machine learning algorithm. Having the estimated depth maps and stereo videos, other views may be synthesized. Performance evaluations have shown that the proposed method estimates high quality depth maps for stereo sequences from limited depth information. Implementation of our proposed technique in the future multiview pipeline eliminates the need for estimating and transmitting the whole depth map for all the views, producing high quality multiview content while reducing the required bandwidth.
Mahsa T. Pourazad, Di Xu 0001, Panos Nasiopoulos
ICTAI3
2012 Distance based heuristic for power and rate allocation of video sensor networks
abstract
Video sensor networks (VSNs) offer an alternative to several existing surveillance technologies. However, unlike in conventional sensor network, video processing and transmission requires large amount of resources both in signal processing, i.e., encoding, and transmission of the encoded data. For such networks, an optimal encoding power and rate allocation method based on a power-rate-distortion (PRD) analysis has previously been proposed, where the power consumption of video encoding can be controlled by managing some encoding parameters. However, these parameters are currently obtained offline by examining the stored video, an approach which may not be suitable for surveillance applications. In this paper, a distance based heuristic for encoding power and rate allocation of VSNs is proposed. The proposed technique is a practical solution since the video coding parameter is controlled by the node's location in the network. Although the proposed technique offers a sub-optimal solution, in some scenarios it achieves performance up to 94% of the optimal solution in terms of network lifetime.
Bambang A. B. Sarif, Victor C. M. Leung, Panos Nasiopoulos
WCNC3
2011 3D medical image coding with optimal channel protection for wireless transmission
abstract
We propose a 3D medical image coding method with optimal channel protection for wireless transmission. The proposed method employs the 3D integer wavelet transform and a modified EBCOT with 3D contexts to create a scalable layered bit-stream. Optimal channel protection is attained by assigning protection bits to the different sections of the compressed bit-stream according to their mean energy content. The robustness of the proposed method is evaluated over a Rayleigh-fading channel with a concatenation of a cyclic redundancy check code and a rate-compatible convolutional code. Comparisons are made with the cases of equal channel protection and unequal channel protection. Simulation results show a significant improvement in reconstruction quality of the received 3D images.
Victor Sanchez, Panos Nasiopoulos
ICASSP2
2011 Effect of brightness on the quality of visual 3D perception
abstract
Over the years, a consensus has been reached that the introduction of 3D entertainment can only be a lasting success if the perceived image quality and the viewing comfort are better than those of conventional 2D television. There are different factors that affect the perceived quality of 3D content. In this paper, our objective is to obtain a good understanding of the effect that brightness has on the visual quality of 3D videos and compare it to that of the 2D. We capture outdoor and indoor scenes with different exposures and we perform subjective evaluation to investigate how brightness affects the perceived quality of the 3D experience.
Mahsa T. Pourazad, Zicong Mai, Panos Nasiopoulos, Konstantinos N. Plataniotis, Rabab K. Ward
ICIP3
2011 Collaborative routing and camera selection for visual wireless sensor networks
abstract
The authors propose a new approach for network lifetime maximisation in visual wireless sensor networks (VWSN). The existence of redundant multiple sensing for some parts of the field results in additional computational complexity and requires more bandwidth and battery power in the nodes, hence reducing the network lifetime. In a VWSN, efficient node collaboration for data sensing and gathering is a key factor to determine network lifetime. The proposed approach provides a jointly optimised traffic routing and camera selection strategy to enhance the lifetime of the whole network. In this approach, different nodes can collaborate with each other to prevent unnecessary multiple sensing of different areas in the network, and at the same time collaborate in the routing of the generated traffics to the network sink.
S. Mohsen Amiri, Panos Nasiopoulos, Victor C. M. Leung
IET Commun.2
2011 Optimizing a Tone Curve for Backward-Compatible High Dynamic Range Image and Video Compression
abstract
For backward compatible high dynamic range (HDR) video compression, the HDR sequence is reconstructed by inverse tone-mapping a compressed low dynamic range (LDR) version of the original HDR content. In this paper, we show that the appropriate choice of a tone-mapping operator (TMO) can significantly improve the reconstructed HDR quality. We develop a statistical model that approximates the distortion resulting from the combined processes of tone-mapping and compression. Using this model, we formulate a numerical optimization problem to find the tone-curve that minimizes the expected mean square error (MSE) in the reconstructed HDR sequence. We also develop a simplified model that reduces the computational complexity of the optimization problem to a closed-form solution. Performance evaluations show that the proposed methods provide superior performance in terms of HDR MSE and SSIM compared to existing tone-mapping schemes. It is also shown that the LDR image quality resulting from the proposed methods matches that produced by perceptually-based TMOs.
Zicong Mai, Hassan Mansour, Rafal Mantiuk, Panos Nasiopoulos, Rabab K. Ward, Wolfgang Heidrich
IEEE Trans. Image Process.4
2011 Rate and Distortion Modeling of CGS Coded Scalable Video Content
abstract
In this paper, we derive single layer and scalable video rate and distortion models for video bitstreams encoded using the coarse grain quality scalability (CGS) feature of the scalable extension of H.264/AVC. In these models, we assume the source is Laplacian distributed and compensate for errors in the distribution assumption by linearly scaling the Laplacian parameter . Moreover, we present simplified approximations of the derived models that allow for a run-time calculation of sequence dependent model constants. Our models use the mean absolute difference (MAD) of the prediction residual signal and the encoder quantization parameter (QP) as input parameters. Consequently, we are able to estimate the residual MAD, bitrate, and distortion of a future video frame at any QP value and for both base-layer and CGS layer packets. We also present simulation results that demonstrate the accuracy of the proposed models.
Hassan Mansour, Panos Nasiopoulos, Vikram Krishnamurthy
IEEE Trans. Multim.2
2011 Correction of Clipped Pixels in Color Images
abstract
Conventional images store a very limited dynamic range of brightness. The true luma in the bright area of such images is often lost due to clipping. When clipping changes the R, G, B color ratios of a pixel, color distortion also occurs. In this paper, we propose an algorithm to enhance both the luma and chroma of the clipped pixels. Our method is based on the strong chroma spatial correlation between clipped pixels and their surrounding unclipped area. After identifying the clipped areas in the image, we partition the clipped areas into regions with similar chroma, and estimate the chroma of each clipped region based on the chroma of its surrounding unclipped region. We correct the clipped R, G, or B color channels based on the estimated chroma and the unclipped color channel(s) of the current pixel. The last step involves smoothing of the boundaries between regions of different clipping scenarios. Both objective and subjective experimental results show that our algorithm is very effective in restoring the color of clipped pixels.
Di Xu 0001, Colin Doutre, Panos Nasiopoulos
IEEE Trans. Vis. Comput. Graph.3
2010 Color image desaturation using sparse reconstruction
abstract
In this paper, we propose an algorithm to estimate the true values of saturated pixels in color images. Pixel saturation occurs when at least one color channel is clipped at some value below the full dynamic range of the scene, resulting in a loss in image fidelity. The proposed algorithm is based on the assumptions that images are nearly sparse in an appropriate transform domain, and that saturated pixels can be inferred from the structure of non-saturated neighboring pixels. Consequently, we use a hierarchical windowing algorithm which selects image regions containing relatively few saturated pixels for processing. Starting with small sized regions, and progressively increasing the size, we solve a sparsity promoting constrained ℓ1minimization problem for each selected region to recover the saturated pixels. Moreover, we provide simulation results to show the effectiveness of our algorithm.
Hassan Mansour, Rayan Saab, Panos Nasiopoulos, Rabab K. Ward
ICASSP3
2010 A stereo matching data cost robust to blurring
abstract
Most modern stereo matching algorithms involve solving an optimization problem where the objective function includes a data cost term and a smoothness term. The data cost term measures how well corresponding pixels match between the left and right images. In this paper a new stereo matching data cost is proposed which is robust to variations in blurring between the images caused by camera focus. In our method, each image is blurred once with a large filter. By comparing the original and blurred versions of each image we obtain a range of possible values each pixel could take on for different levels of blurring. Based on this range we construct a blur robust data cost for comparing pixels between two images. Experimental results show our proposed method greatly improves stereo matching accuracy when the left and right images in a stereo pair are focused differently.
Colin Doutre, Panos Nasiopoulos
ICIP2
2010 Visually-favorable tone-mapping with high compression performance
abstract
We develop a tone-mapping operator (TMO) that considers the perceptual quality of the tone-mapped image together with the compression efficiency. The proposed TMO is formulated as an optimization problem that incorporates statistical models of i) the quality of the tone-mapped image given a desired TMO, ii) the base layer bit-rate and iii) the enhancement layer bit-rate. The results show that our method achieves high coding gain while maintaining good quality tone-mapped images.
Zicong Mai, Hassan Mansour, Panos Nasiopoulos, Rabab K. Ward
ICIP3
2010 An improved Bayesian algorithm for color image desaturation
abstract
Current digital imaging systems are unable to capture the entire dynamic range of the visible luminance, causing saturation in the very bright parts of a scene. Color distortion occurs when the amounts of saturation are different in the red (R), green (G), and blue (B) color channels. A Bayesian algorithm was developed in the past to correct the saturated pixels in raw images. For each image, it estimates the distributions of the R, G, and B color channels based on the unsaturated pixels, and then corrects the saturated pixels based on this prior distribution. In this paper, we improve this Bayesian algorithm by incorporating spatial information in the correction process. We utilize the strong spatial correlation of images as well as the correlation between the R, G, and B channels of each individual pixel to estimate the prior distributions of the R, G, and B color channels. The prior distribution of each saturated region is modeled individually based on its surrounding region, which is determined by morphological dilation. Experimental results show that our modified algorithm greatly outperforms the original Bayesian algorithm for fixing saturated pixels in color images.
Di Xu 0001, Colin Doutre, Panos Nasiopoulos
ICIP3
2010 Correcting unsynchronized zoom in 3D video
abstract
When capturing 3D video with a stereoscopic camera setup, it is important for the cameras to be precisely aligned and synchronized. This is particularly difficult in transitions such as zooming where the camera parameters must be changed in unison, or else the perceived 3D effect will be degraded. In this paper we study the problem of unsynchronized zooming in 3D video. First, we present a subjective study that shows that the perceived quality of stereo video is greatly reduced if the two views are zoomed by different amounts. Next, we present a method for correcting zoom mismatch by applying cropping and scaling to ones of the views. Our method involves finding matching points between the left and right views, and performing least-squares regressions to estimate the amount of scaling and cropping required to make the views consistent. Experiments were performed on videos with digitally introduced zoom mismatch and videos with optical unsynchronized zoom. In both cases the results show that our method is highly accurate and produces videos without size differences or vertical parallax between the two views.
Colin Doutre, Mahsa T. Pourazad, Alexis M. Tourapis, Panos Nasiopoulos, Rabab K. Ward
ISCAS4
2010 On-the-fly tone mapping for backward-compatible high dynamic range image/video compression
abstract
In this paper, we propose a real-time tone-mapping scheme for backward compatible high dynamic range (HDR) video compression. The appropriate choice of a tone-mapping operator (TMO) can significantly improve the HDR quality reconstructed from a low dynamic range (LDR) version. We develop a statistical model that approximates the mean square error (MSE) distortion resulting from the combined processes of tone-mapping and compression. Using this model, we formulate a numerical optimization problem to find the tone-curve that minimizes the expected MSE in the reconstructed HDR sequence. We then simplify the developed model in order to reduce the computational complexity of the optimization problem to a closed-form solution. Performance evaluations show that the proposed methods provide superior performance in terms of HDR MSE and SSIM compared to existing tone-mapping schemes. It is also shown that the LDR image quality resulting from the proposed methods matches that produced by perceptually-based TMOs.
Zicong Mai, Hassan Mansour, Rafal Mantiuk, Panos Nasiopoulos, Rabab K. Ward, Wolfgang Heidrich
ISCAS4
2010 Fast block-size partitioning using empirical rate-distortion models for MPEG-2 to H.264/AVC transcoding
abstract
We present an efficient H.264/AVC block-size partitioning prediction method, which is based on our proposed empirical rate and distortion models. Compared to other state-of-the-art transcoding methods, and for the same rate-distortion performance, our proposed algorithm requires the least computational complexity, reaching a 73% reduction in variable block-size motion estimation for SDTV sequences, and 71% reduction for CIF sequences.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ISCAS2
2010 Saturated-pixel enhancement for color images
abstract
We propose an algorithm to correct both luma and chroma of the saturated pixels in an overexposed image. Our method is based on the strong chroma spatial correlation between saturated pixels and their surrounding unsaturated area. We first identify the saturated areas in the image. Then, we partition these areas into regions with similar chroma, and estimate the chroma of each saturated region based on the chroma of its surrounding unsaturated region. Next, we correct the saturated R, G, or B color channels according to the estimated chroma and the unsaturated color channel(s) of the pixel. The last step involves smoothing of the boundaries between regions of different saturation scenarios. Both objective and subjective experimental results show that our algorithm is very effective in restoring the color of saturated pixels.
Di Xu 0001, Colin Doutre, Panos Nasiopoulos
ISCAS3
2010 Fair scheduling for real-time multimedia support in IEEE 802.16 wireless access networks
abstract
Successful deployment of Broadband Wireless Access Networks such as WiMAX (IEEE 802.16) will be contingent on provisions for supporting multimedia traffic. In this paper, we review the quality of service features of access networks such as the 802.16 standard, and identify algorithms and schemes that are needed for supporting multimedia traffic in such networks. The 802.16 standard only specifies the features that should be implemented and leaves the design of a quality of service solution to developers. This includes the design of a mandatory scheduling framework. We present a comprehensive multimedia support framework based on the standard features. The framework specifies the architectures for the base station and the subscriber station, and contributes a number of algorithms for different service provisioning objectives. We use the concept of virtual packets to provide fair packet based centralized scheduling of uplink and downlink packets. The presented solution also provides algorithms for temporal and throughput fair scheduling in multirate physical layer of the 802.16 networks. An important part of the presented design is a multi-class fair scheduling scheme which is proposed for providing better delay performance for real time applications, while maintaining slightly longer term fairness.
Yaser P. Fallah, Panos Nasiopoulos, Raja Sengupta 0002
WOWMOM2
2010 Efficient Motion Re-Estimation With Rate-Distortion Optimization for MPEG-2 to H.264/AVC Transcoding
abstract
One objective in MPEG-2 to H.264/advanced video coding transcoding is to improve the H.264/AVC compression ratio by using more advanced macroblock encoding modes. The motion re-estimation process is by far the most time-consuming process in this type of video transcoding. In this paper, we present an efficient H.264/AVC block size partitioning prediction algorithm for MPEG-2 to H.264/AVC transcoding applications. Our algorithm uses rate-distortion optimization techniques and predicted initial motion vectors to estimate block size partitioning. It is also shown that using block size partitioning smaller than 8 × 8 (i.e., 8 × 4, 4 × 8, and 4 × 4) results in negligible compression improvements, and thus these sizes should be avoided in transcoding. Experimental results show that, compared to the state-of-the-art transcoding scheme, our transcoder yields similar rate-distortion performance, while the computational complexity is significantly reduced, requiring an average of 29% of the computations. Compared to the full-search scheme, our proposed algorithm reduces the computational complexity by about 99.47% for standard-definition television sequences and 98.66% for common intermediate format sequences. Compared to UMHexagonS, the fast motion estimation algorithm used in H.264/AVC, the experimental results show that our proposed algorithm is a better trade-off between computational complexity and picture quality.
Qiang Tang 0002, Panos Nasiopoulos
IEEE Trans. Circuits Syst. Video Technol.2
2010 3-D Scalable Medical Image Compression With Optimized Volume of Interest Coding
abstract
We present a novel 3-D scalable compression method for medical images with optimized volume of interest (VOI) coding. The method is presented within the framework of interactive telemedicine applications, where different remote clients may access the compressed 3-D medical imaging data stored on a central server and request the transmission of different VOIs from an initial lossy to a final lossless representation. The method employs the 3-D integer wavelet transform and a modified EBCOT with 3-D contexts to create a scalable bit-stream. Optimized VOI coding is attained by an optimization technique that reorders the output bit-stream after encoding, so that those bits belonging to a VOI are decoded at the highest quality possible at any bit-rate, while allowing for the decoding of background information with peripherally increasing quality around the VOI. The bit-stream reordering procedure is based on a weighting model that incorporates the position of the VOI and the mean energy of the wavelet coefficients. The background information with peripherally increasing quality around the VOI allows for placement of the VOI into the context of the 3-D image. Performance evaluations based on real 3-D medical imaging data showed that the proposed method achieves a higher reconstruction quality, in terms of the peak signal-to-noise ratio, than that achieved by 3D-JPEG2000 with VOI coding, when using the MAXSHIFT and general scaling-based methods.
Victor Sanchez, Rafeef Abugharbieh, Panos Nasiopoulos
IEEE Trans. Medical Imaging3
2009 Efficient Utilization of Error Protection Techniques for Transmission of Data-Partitioned H.264 Video in a Capacity Constrained Network
abstract
We propose an efficient error protection technique for data-partitioned H.264 video in a capacity constrained network. Our scheme maximizes video quality by choosing the optimal point in the application layer and medium access control (MAC) layer redundancy. We have shown that, in a capacity constrained network and highly lossy environment, neither forward error correction (FEC) nor retransmissions alone can result in optimum performance. Instead, it is the combination of these two techniques that effectively reduces the overall loss.
Ashfiqua T. Connie, Yaser P. Fallah, Panos Nasiopoulos, Victor C. M. Leung
ICC3
2009 Fast vignetting correction and color matching for panoramic image stitching
abstract
When images are stitched together to form a panorama there is often color mismatch between the source images due to vignetting and differences in exposure and white balance between images. In this paper a low complexity method is proposed to correct vignetting and differences in color between images, producing panoramas that look consistent across all source images. Unlike most previous methods which require complex non-linear optimization to solve for correction parameters, our method requires only linear regressions with a low number of parameters, resulting in a fast, computationally efficient method. Experimental results show the proposed method effectively removes vignetting effects and produces images that are highly visually consistent in color and brightness.
Colin Doutre, Panos Nasiopoulos
ICIP2
2009 Modified H.264 intra prediction for compression of video and images captured with a color filter array
abstract
Most consumer digital cameras capture color information with a single light sensor and a color filter array (CFA). In these cameras, only one color sample (red, green or blue) is captured at each pixel location. This paper presents a modified H.264 intra prediction scheme for compressing image and video data captured with a color filter array. The H.264 intra prediction modes are modified for the green channel to account for the fact that the green data is not sampled in a rectangular manner in the Bayer pattern, the most popular CFA design. The proposed method increases the compression efficiency of I frames on the green channel by up to 1.2 dB.
Colin Doutre, Panos Nasiopoulos
ICIP2
2009 An efficient low random-access delay panorama-based multiview video coding scheme
abstract
We present an efficient low random delay scheme for multiview video coding (MVC). In the proposed scheme, inter-view prediction (disparity estimation), which introduces time-consuming computations and random access delay to MVC, is replaced with a residue-stream coding process. Our algorithm transforms the middle view to a panoramic view of the scene. Then the residue streams are created as the difference of the luma and chroma values of overlapping regions of each view and the panoramic view. Finally the panoramic stream and all residue streams are encoded separately (simulcast coding). The hierarchical B picture prediction structure is implemented for coding each stream. Performance evaluations show that our proposed coding method outperforms the recent multiview video coding standard by up to 2.13 dB PSNR and enhances the compression ratio by 24.6%, while reducing random-access delay by 50%.
Mahsa T. Pourazad, Panos Nasiopoulos, Rabab K. Ward
ICIP2
2009 3D scalable lossless compression of medical images based on global and local symmetries
abstract
We recently proposed a symmetry-based scalable lossless compression method for 3D medical images using the 2D integer wavelet transform and the embedded block coder with optimized truncation (EBCOT). In this paper, we present two major contributions that enhance our early work: 1) a new block-based intra-band prediction method that exploits the global and local symmetries of the wavelet-transform sub-bands based on the main axis of symmetry as detected using the analytical Fourier-Mellin transform; and 2) a new inter-slice DPCM prediction method that exploits the correlation between slices. Performance evaluations on real 3D medical images show an average improvement of up to 17% in lossless compression ratios when compared to the state-of-the-art compression methods including 3D-JPEG2000, JPEG2000 and H.264 intra-coding.
Victor Sanchez, Rafeef Abugharbieh, Panos Nasiopoulos
ICIP3
2009 Efficient motion vector re-estimation for MPEG-2 TO H.264/AVC transcoding with arbitrary down-sizing ratios
abstract
As for down-sizing MPEG-2 to H.264/AVC transcoding, an efficient algorithm of estimating initial H.264/AVC motion vectors is proposed. By using the estimated initial motion vectors, only a small range of motion vector refinement is sufficient to find the final motion vector for each partition. Experimental results show that our proposed algorithm achieves average 0.08 dB improvement (maximum 0.24 dB) in the picture quality compared to the other state-of-art method. At the same time, the computational complexity of estimating the initial motion vectors is less than that of the other state-of-art technique (average 36% reduction).
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ICIP2
2009 Logo insertion transcoding for H.264/AVC compressed video
abstract
H.264/AVC quickly gains ground in many aspects of video applications, due to its superior coding performance. Inserting a company logo into H.264/AVC compressed video streams has been a highly desirable application in the TV telecasting industry. In this paper, we propose a novel and efficient logo-insertion scheme for H.264/AVC compressed videos. Our proposed scheme overcomes the numerous coding dependencies, and minimizes the changes to the original compressed videos. Experimental results show that our proposed transcoding scheme achieves extraordinary video quality and significantly reduces the bit rate and computational cost. Compared with the cascaded transcoding scheme, our proposed logo-insertion method achieves an average of 1.16 dB PSNR increase, or a 68.6% bit-rate reduction. Our scheme also dramatically reduces the total transcoding time and the motion-estimation time by at least 67.2% and 97.4%, respectively.
Di Xu 0001, Panos Nasiopoulos
ICIP2
2009 Color Correction of Multiview Video with Average Color as Reference
abstract
When capturing multiview video, there can be significant variations in the color of views captured with different cameras. This negatively affects compression efficiency when multiview video is coded using inter-view prediction. In this paper we propose a method for correcting the color of multiview video sets as a preprocessing step to compression. Unlike previous work where one of the captured views is used as the color reference, we correct all views to match the average color of the set of views. Block based disparity estimation is used to find matching points between all views in the video set, and the average color is calculated for these matching points. Least squares regressions are used to find functions that will make each view match the average color. Experimental results show that the proposed method results in video sets that closely match in subjective color. Furthermore, when multiview video is compressed with JMVM, the proposed method increases compression efficiency by up to 1.0 dB compared to compressing the original uncorrected video.
Colin Doutre, Panos Nasiopoulos
ISCAS2
2009 Converting H.264-Derived Motion Information into Depth Map
Mahsa T. Pourazad, Panos Nasiopoulos, Rabab K. Ward
MMM2
2009 Color Correction Preprocessing for Multiview Video Coding
abstract
In multiview video, a number of cameras capture the same scene from different viewpoints. There can be significant variations in the color of views captured with different cameras, which negatively affects performance when the videos are compressed with inter-view prediction. In this letter, a method is proposed for correcting the color of multiview video sets as a preprocessing step to compression. Unlike previous work, where one of the captured views is used as the color reference, we correct all views to match the average color of the set of views. Block-based disparity estimation is used to find matching points between all views in the video set, and the average color is calculated for these matching points. A least-squares regression is performed for each view to find a function that will make the view most closely match the average color. Experimental results show that when multiview video is compressed with joint multiview video model, the proposed method increases compression efficiency by up to 1.0 dB in luma peak signal-to-noise ratio (PSNR) compared to compressing the original uncorrected video.
Colin Doutre, Panos Nasiopoulos
IEEE Trans. Circuits Syst. Video Technol.2
2009 Novel Lossless fMRI Image Compression Based on Motion Compensation and Customized Entropy Coding
abstract
We recently proposed a method for lossless compression of 4-D medical images based on the advanced video coding standard (H.264/AVC). In this paper, we present two major contributions that enhance our previous work for compression of functional MRI (fMRI) data: 1) a new multiframe motion compensation process that employs 4-D search, variable-size block matching, and bidirectional prediction; and 2) a new context-based adaptive binary arithmetic coder designed for lossless compression of the residual and motion vector data. We validate our method on real fMRI sequences of various resolutions and compare the performance to two state-of-the-art methods: 4D-JPEG2000 and H.264/AVC. Quantitative results demonstrate that our proposed technique significantly outperforms current state of the art with an average compression ratio improvement of 13%.
Victor Sanchez, Panos Nasiopoulos, Rafeef Abugharbieh
IEEE Trans. Inf. Technol. Biomed.2
2009 Symmetry-Based Scalable Lossless Compression of 3D Medical Image Data
abstract
We propose a novel symmetry-based technique for scalable lossless compression of 3D medical image data. The proposed method employs the 2D integer wavelet transform to decorrelate the data and an intraband prediction method to reduce the energy of the sub-bands by exploiting the anatomical symmetries typically present in structural medical images. A modified version of the embedded block coder with optimized truncation (EBCOT), tailored according to the characteristics of the data, encodes the residual data generated after prediction to provide resolution and quality scalability. Performance evaluations on a wide range of real 3D medical images show an average improvement of 15% in lossless compression ratios when compared to other state-of-the art lossless compression methods that also provide resolution and quality scalability including 3D-JPEG2000, JPEG2000, and H.264/AVC intra-coding.
Victor Sanchez, Rafeef Abugharbieh, Panos Nasiopoulos
IEEE Trans. Medical Imaging3
2009 Dynamic Resource Allocation for MGS H.264/AVC Video Transmission Over Link-Adaptive Networks
abstract
In this paper, we address the problem of efficiently allocating network resources to support multiple scalable video streams over a constrained wireless channel. We present a resource allocation framework that jointly optimizes the operation of the link adaptation scheme in the physical layer (PHY), and that of a traffic control module in the network or medium access control (MAC) layer in multirate wireless networks, while satisfying bandwidth/capacity constraints. Multirate networks, such as IEEE 802.16 or IEEE 802.11, adjust the PHY coding and modulation schemes to maintain the reliability of transmission under varying channel conditions. Higher reliability is achieved at the cost of reduced PHY bit-rate which in turn necessitates a reduction in video stream bit-rates. The rate reduction for scalable video is implemented using a traffic control module. Conventional solutions operate unaware of the importance and loss tolerance of data and drop the higher layers of scalable video altogether. In this paper, we consider medium grain scalable (MGS) extension of H.264/AVC video and develop new rate and distortion models that characterize the coded bitstream. Performance evaluations show that our proposed framework results in significant gains over existing schemes in terms of average video PSNR that can reach 3 dB in some cases for different channel SNRs and different bandwidth budgets.
Hussein Mansour, Yaser P. Fallah, Panos Nasiopoulos, Vikram Krishnamurthy
IEEE Trans. Multim.3
2008 Video Packetization Techniques for Enhancing H.264 Video Transmission over 3G Networks
abstract
Transmission of video over a wireless network is a challenging task due to the error prone characteristics of the wireless link. This application requires the video encoding process to offer error resilience features in order to deal with the high packet loss probability of wireless networks. H.264, the newest and most efficient video compression standard, offers several error resiliency features. H.264 also offers network support by utilizing a network abstraction layer (NAL) and encapsulating individually decodable video slices in NAL units. The size of the video slices, or packets, has a significant effect on the overall performance of a video application in wireless networks. In this paper we study the possibility of determining the optimal packet size of encoded video that minimizes the packet loss rate and maximizes the video quality during transmission over 3G networks. Through simulation experiments, we show that smaller slices are more favorable, but encoding inefficiency and increase in overhead set a limit on the minimum acceptable packet size.
Ashfiqua T. Connie, Panos Nasiopoulos, Victor C. M. Leung, Yaser P. Fallah
CCNC2
2008 Motion vector prediction for improving one bit transform based motion estimation
abstract
One bit transforms (1BT) have been proposed for lowering the complexity of motion estimation (ME) in video coding. These transforms generate a one bit representation of each pixel in the video that is used in the motion search. This approach can greatly reduce the silicon area and power required for hardware based video encoding. However 1BT methods under-perform traditional Sum of absolute differences (SAD) based motion estimation, particularly for smaller block sizes. In this paper, it is proposed to improve 1BT based ME by predicting the motion vector for each block based on the vectors from previous blocks and modifying the cost function to favor motion vectors close to the predicted one. This takes advantage of the spatial correlation between motion vectors and produces a more uniform motion field. Simulation results show the proposed method can improve the PSNR of frames reconstructed through motion compensation by up to 1 dB and substantially improve the subjective video quality by reducing blocking artifacts.
Colin Doutre, Panos Nasiopoulos
ICASSP2
2008 Joint media-channel aware unequal error protection for wireless scalable video streaming
abstract
In this paper, we propose a joint source-channel unequal error protection scheme for scalable video streaming over capacity constrained high speed packet access (HSPA) networks. Conventional link adaptation schemes in HSPA networks use the modulation and coding scheme (MCS) that achieves a preset channel frame error rate. Our scheme utilizes video priority information along with channel quality information to set the channel coding rate that maximizes the cumulative coding rate of channel coding and application layer unequal error protection. Performance evaluations show that that under the same constraints our scheme results in an average performance improvement of 0.5dB in video PSNR for different channel conditions and different video sequences.
Hassan Mansour, Panos Nasiopoulos, Vikram Krishnamurthy
ICASSP2
2008 Efficient 4D motion compensated lossless compression of dynamic volumetric medical image data
abstract
Dynamic volumetric (four dimensional- 4D) medical images are typically huge in file size and require a vast amount of resources for storage and transmission purposes. In this paper, we propose an efficient lossless compression method for 4D medical images that is based on a multi-frame motion compensation process employing a 4D search, variable block- sizes and bi-directional prediction. Data redundancies are reduced by recursively applying multi-frame motion compensation in the spatial and temporal dimensions. The proposed method also uses a novel differential coding algorithm to reduce redundancies in motion vectors and a new context-based adaptive binary arithmetic coder (CABAC) for compression of the residual data. Performance evaluations on real medical images of varying modality resulted in lossless compression ratios of up to 16:1.
Victor Sanchez, Panos Nasiopoulos, Rafeef Abugharbieh
ICASSP2
2008 Fast block size prediction for MPEG-2 to H.264/AVC transcoding
abstract
One objective in MPEG-2 to H.264 transcoding is to improve the H.264 compression ratio by using more accurate H.264 motion vectors. Motion re-estimation is by far the most time consuming process in video transcoding, and improving the searching speed is a challenging problem. We introduce a new transcoding scheme that uses the MPEG-2 DCT coefficients to predict the block size partitioning for H.264. Performance evaluations have shown that, for the same rate-distortion performance, our proposed scheme achieves an impressive reduction in the computational complexity of more than 82% compared to the full range motion estimation used by H.264.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ICASSP2
2008 Hybrid OFDMA/CSMA Based Medium Access Control for Next-Generation Wireless LANs
abstract
Existing medium access control (MAC) schemes for wireless local area networks (WLAN) have been shown to lack scalability in crowded networks, and efficiency in supporting heterogeneous traffic types. These issues are mostly due to the use of random multiple access techniques in the MAC layer. The design of these techniques is highly linked to the choice of the underlying physical (PHY) layer technology. The advent of new PHY schemes that are based on orthogonal frequency division multiple access (OFDMA) provides new opportunities for devising more efficient MAC protocols. We propose a new adaptive MAC design based on OFDMA technology. The design uses OFDMA to reduce collision during transmission request phases, and makes channel access more predictable. To improve efficiency, we combine the OFDMA access with a carrier sense multiple access (CSMA) scheme. Data transmission opportunities are assigned through an access point that can schedule traffic streams in both time and frequency (subchannels) domains. We demonstrate the effectiveness of the proposed MAC and compare it to existing mechanisms through simulation experiments and by deriving an analytical model for the operation of the MAC in saturation mode.
Yaser P. Fallah, Panos Nasiopoulos, Hussein M. Alnuweiri
ICC3
2008 Rate and distortion modeling of medium grain scalable video coding
abstract
Scalability in video coding is becoming the primary choice for providing quality of service (QoS) guarantees in wireless video communication. In this paper, we develop real-time rate and distortion prediction models for medium grained scalable (MGS) coded video streams. These models allow mobile video encoders to predict the packet size and corresponding distortion of a video frame using only the mean absolute difference (MAD) of the motion prediction and the quantization parameter (QP). The prediction of rate and distortion measures can be used in devices with cross layer optimization capabilities to choose the combination of base and enhancement layer packets that deliver the best picture quality given channel quality information. Performance evaluations demonstrate that our models accurately predict the size and distortion of base and enhancement layer MGS packets.
Hassan Mansour, Vikram Krishnamurthy, Panos Nasiopoulos
ICIP3
2008 An optimized link adaptation scheme for efficient delivery of scalable H.264 Video over IEEE 802.11n
abstract
In this paper, we propose a cross-layer optimization scheme for delivery of scalable video over variable bit-rate wireless networks, in particular 802.11 based wireless local area networks (WLAN). For scalable video streaming applications, the conventional solution to reduced throughput due to channel distortions is to reduce the video bitrate by dropping the higher enhancement layers of the scalable video. We show that video quality can be improved, without adding to traffic load, when the WLAN link adaptation scheme uses a temporal fairness criterion along with scalable video distortion estimates to adjust its physical (PHY) layer modulation and coding parameters used for delivering each video layer. We formulate the problem as an optimization problem for assigning different PHY modes to different layers of scalable video under temporal fairness constrains; the solution to this problem provides a set of PHY configuration parameters that achieve the highest possible video quality while meeting the admission control constraints. Performance evaluations demonstrate the effectiveness of our method and the accuracy of the models.
Yaser P. Fallah, Hassan Mansour, Panos Nasiopoulos, Hussein M. Alnuweiri
ISCAS4
2008 Real-time joint rate and protection allocation for multi-user scalable video streaming
abstract
In this paper, we present a real-time joint bit-rate and error protection allocation scheme for multiple scalable video streams sharing a single downlink channel. High speed downlink packet access (HSDPA) systems allow for multiple live video streams to share a common downlink channel among multiple mobile users. However, the unreliable nature of the wireless link results in packet losses and fluctuations in the available channel capacity. This calls for flexible error protection and rate control strategies implemented at the video encoders that can respond to the variation in channel conditions. In this paper, we formulate a global optimization problem, which is solved at every frame transmission instant and minimizes the expected sum of the video frame distortions of all users by adjusting the encoding quality at the base- and enhancement-layers as well as the application layer error protection overhead used to combat packet losses. We consider frame-level unequal erasure protection (UXP) as the application layer forward error correction scheme. Performance evaluations show that compared with existing schemes our proposed scheme delivers far superior decoded video quality averaging 1.2 dB in PSNR.
Hassan Mansour, Panos Nasiopoulos, Vikram Krishnamurthy
PIMRC2
2008 H.264-Based Compression of Bayer Pattern Video Sequences
abstract
Most consumer digital cameras use a single light sensor which captures color information using a color filter array (CFA). This produces a mosaic image, where each pixel location contains a sample of only one of three colors, either red, green or blue. The two missing colors at each pixel location must be interpolated from the surrounding samples in a process called demosaicking. The conventional approach to compressing video captured with these devices is to first perform demosaicking and then compress the resulting full-color video using standard methods. In this paper two methods for compressing CFA video prior to demosaicking are proposed. In our first method, the CFA video is directly compressed with the H.264 video coding standard in 4:2:2 sampling mode. Our second method uses a modified version of H.264, where motion compensation is altered to take advantage of the properties of CFA data. Simulations show both proposed methods give better compression efficiency than the demosaick-first approach at high bit rates, and thus are suitable for applications, such as digital camcorders, where high quality video is required.
Colin Doutre, Panos Nasiopoulos, Konstantinos N. Plataniotis
IEEE Trans. Circuits Syst. Video Technol.2
2008 A Link Adaptation Scheme for Efficient Transmission of H.264 Scalable Video Over Multirate WLANs
abstract
In this paper, we propose a cross-layer optimization scheme for delivery of scalable video over multirate wireless networks, in particular the popular 802.11 based wireless local area network (WLAN). The 802.11 based networks use a link adaptation mechanism in the physical layer (PHY) to maintain the reliability of transmission under varying channel conditions. When channel condition worsens, the reliability is maintained by employing more robust modulation and coding schemes, at the cost of reduced PHY bit rate. The reduced bit rate will result in lower available throughput for applications. For scalable video streaming applications, the conventional solution to this problem is to reduce the video bit rate by dropping the higher enhancement layers of the scalable video. We show in this article that the video quality can be improved, if the link adaptation scheme uses more intelligent reliability criteria and adjusts the PHY parameters used for delivering each video layer, according to the relative importance of that layer. Our scheme achieves better video quality without increasing the traffic load of the WLAN. For this purpose we present temporal fairness constraints and formulate an optimization problem for assigning different PHY modes to different layers of scalable video; the solution to this problem provides a set of PHY configuration parameters that achieve the highest possible video quality while meeting the admission control constraints in the network. Performance evaluations demonstrate that our method outperforms the existing mechanisms.
Yaser P. Fallah, Hassan Mansour, Panos Nasiopoulos, Hussein M. Alnuweiri
IEEE Trans. Circuits Syst. Video Technol.4
2008 Compensation of Requantization and Interpolation Errors in MPEG-2 to H.264 Transcoding
abstract
Implementing MPEG-2 to H.264 transcoding schemes in the pixel domain introduces a high degree of computational complexity. In the transform domain, this transcoding is more computationally efficient, and several methods have been developed to address that approach. However, incompatibilities between the two standards, such as the mismatches between the MPEG-2 and H.264 motion compensation processes, cause several distortions that may affect the overall picture quality. In this study, we address the main distortions that result from requantization errors: luminance half-pixel and chrominance quarter/three-quarter interpolation errors. Then, we propose algorithms that compensate for these errors. The traditional requantization error compensation algorithm for DCT coefficients is updated so that it can be applied to the H.264 integer transform coefficients. Equations that compensate for the luminance half-pixel and chrominance quarter/three-quarter pixel interpolation errors are derived. To remove the interpolation errors, the previous H.264 frame is needed. Thus, the compensation scheme includes a closed-loop H.264 motion compensation process, which is implemented in the pixel domain. To evaluate the performance of the proposed compensation algorithms in terms of picture quality, our scheme is compared with two different cascaded pixel-domain transcoding structures. The first structure reuses the MPEG-2 motion vectors, and the other implements plusmn2 pixels motion vector refinement, but each one has an H.264 deblocking filter. The experimental results show that the proposed compensation algorithms achieve 5-dB quality improvement over the open-loop transform-domain-based transcoding and almost the same picture quality (0.3-0.6 dB) as the cascaded structures. An additional advantage is the reduction in computational complexity that ranges from 13% to 69% compared with the two cascaded methods.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Circuits Syst. Video Technol.2
2008 A Video Watermarking Scheme Based on the Dual-Tree Complex Wavelet Transform
abstract
A watermarking scheme that discourages theater camcorder piracy through the enforcement of playback control is presented. In this method, the video is watermarked so that its display is not permitted if a compliant video player detects the watermark. A watermark that is robust to geometric distortions (rotation, scaling, cropping) and lossy compression is required in order to block access to media content that has been re-recorded with a camera inside a movie theater. We introduce a new video watermarking algorithm for playback control that takes advantage of the properties of the dual-tree complex wavelet transform. This transform offers the advantages of the regular and the complex wavelets (perfect reconstruction, shift invariance, and good directional selectivity). Our method relies on these characteristics to create a watermark that is robust to geometric distortions and lossy compression. The proposed scheme is simple to implement and outperforms comparable methods when tested against geometric distortions.
Lino Coria-Mendoza, Mark R. Pickering, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Inf. Forensics Secur.3
2008 Efficient Lossless Compression of 4-D Medical Images Based on the Advanced Video Coding Scheme
abstract
This paper presents an efficient lossless compression method for 4-D medical images based on the advanced video coding scheme (H.264/AVC). The proposed method efficiently reduces data redundancies in all four dimensions by recursively applying multiframe motion compensation. Performance evaluations on real 4-D medical images of varying modalities including functional magnetic resonance show an improvement in compression efficiency of up to three times that of other state-of-the-art compression methods such as 3D-JPEG2000.
Victor Sanchez, Panos Nasiopoulos, Rafeef Abugharbieh
IEEE Trans. Inf. Technol. Biomed.2
2008 Channel Aware Multiuser Scalable Video Streaming Over Lossy Under-Provisioned Channels: Modeling and Analysis
abstract
In this paper, we analyze the performance of media-aware multiuser video streaming strategies in capacity limited wireless channels suffering from latency problems and packet losses. Wireless video streaming applications are characterized by their bandwidth-intensity, delay-sensitivity, and loss-tolerance. Our main contributions include (i) a rate-minimized unequal erasure protection (UXP) scheme, (ii) an analytical expression for packet delay and play-out deadline of UXP protected scalable video, (iii) a loss-distortion model for hierarchical predictive video coders with picture copy concealment, (iv) an analysis of the performance and complexity of delay-aware, capacity-aware, and optimized UXP streaming scenarios, and (v) we show that the use of unequal error protection causes a rate-constrained optimization problem to be nonconvex. Performance evaluations using a 3GPP network simulator show that, for different channel capacities and packet loss rates, delay-aware nonstationary rate-allocation streaming policies deliver significant gains which range between 1.65 dB to 2 dB in average Y-PSNR of the received video streams over delay-unaware strategies. These gains come at a cost of increasedofflinecomputation which is performed prior to the start of the streaming session or in batches during transmission and therefore, do not affect the run-time performance of the streaming system.
Hassan Mansour, Vikram Krishnamurthy, Panos Nasiopoulos
IEEE Trans. Multim.3
2007 A Cross Layer Optimization Mechanism to Improve H.264 Video Transmission over WLANs
abstract
Supporting video applications over 802.11 wireless local area networks is a challenging task due to the constant fluctuations in channel error rates and the inefficiency of the MAC layer. New video compression technologies, such as H.264, provide a network adaptation layer for adapting the output of the video encoder to the characteristics of the underlying transport network. In this article we demonstrate that it is possible to improve the performance of H.264 video applications over 802.11 WLANs through a cross-layer design that optimizes the encoded H.264 packet sizes. We propose the use of aggregation and fragmentation mechanisms to create the optimal frame lengths. We also investigate several application layer error
Yaser P. Fallah, Darrell Koskinen, Avideh Shahabi, Faizal Karim, Panos Nasiopoulos
CCNC5
2007 Scheduled and Contention Access Transmission of Partitioned H.264 Video Over WLANs
abstract
Supporting Multimedia applications, such as video, over 802.11 wireless local area networks (WLAN) is a challenging task due to the inefficiency of the 802.11 MAC layer and the constant fluctuations in channel error rates. Therefore, specific measures must be taken in both application and delivery layers in order to achieve efficient and satisfactory quality for multimedia applications. Advanced video compression technologies, such as H.264, provide new data partitioning error resiliency feature that allows delivery of the video data in different streams with different levels of importance. There are several possible schemes for mapping these streams to the services of the IEEE 802.1 le MAC. In this article we examine the existing schemes and propose new mechanisms that improve the performance of the video communications system. We propose to use scheduled access schemes in MAC, along with aggregation of packets in the application layer to achieve the best results. We evaluate our proposed mechanisms using simulation experiments.
Yaser P. Fallah, Panos Nasiopoulos, Hussein M. Alnuweiri
GLOBECOM2
2007 Efficient Chrominance Compensation for MPEG2 to H.264 Transcoding
abstract
Although open-loop transcoding is known as the most computational efficient transcoding structure, it is also known to introduce many distortions in the transcoded video. This paper addresses the chrominance distortions resulting from the open-loop MPEG2 to H.264 transcoding structure and proposes algorithms to compensate for the chrominance distortions. The open-loop structure is replaced by a closed-loop transcoding structure, which provides high-quality video by removing the chrominance distortions, resulting in an average of 6 dB picture quality improvement.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ICASSP (1)2
2006 A robust content-dependent algorithm for video watermarking
abstract
A watermarking method that relies on informed coding and informed embedding is presented. Our method uses a subset of various codewords to represent the 0 and 1 message bits to be embedded. We propose a codeword generation scheme that keeps control of the distance between codewords in order to secure fidelity and robustness of the watermark. When compared to existing video watermarking schemes, our method yields superior robustness to video compression.
Lino Coria-Mendoza, Panos Nasiopoulos, Rabab K. Ward
Digital Rights Management Workshop2
2006 Lossless Compression of 4D Medical Images using H.264/AVC
abstract
Four dimensional (4D) medical data are sequences of volumetric images captured in time. These data sets are typically very large in size and demand a great amount of resources for storage and transmission. In this paper, we present a lossless compression technique for 4D medical images which is based on the H.264/AVC video coding standard. Our lossless compression technique efficiently exploits spatial and temporal redundancies between 2D image slices and 3D images in 4D medical images and eliminates any concerns regarding the effects of compression on image quality for diagnostic purposes. Performance evaluations have shown that the proposed compression technique outperforms current 4D compression methods by 70%
Victor Sanchez, Panos Nasiopoulos, Rafeef Abugharbieh
ICASSP (2)2
2006 An Efficient MPEG2 to H.264 Half-Pixel Motion Compensation Transcoding
abstract
An efficient MPEG2 to H.264 half-pixel motion compensation transcoding method is proposed. The Inter macroblock transcoding is implemented in the transform domain. An algorithm is designed to compensate for the errors which arise because of the different half-pixel interpolation procedures used by MPEG2 and H.264/AVC. The experimental results show that the PSNR values of the transcoded H.264 streams result in significant improvement (average 5.5 dB) after we reduce the half-pixel interpolation errors.
Qiang Tang 0002, Rabab K. Ward, Panos Nasiopoulos
ICIP3
2006 A Fast Video Motion Estimation Algorithm for the H.264 Standard
abstract
Video applications are becoming an essential component for mobile devices. H.264, the latest video-coding standard, shows significant potential in terms of bandwidth savings at the cost of substantially increased complexity compared to former standards. The computing power currently available on mobile devices is not sufficient to allow high quality real-time encoding using H.264. Our algorithm uses on average only 0.41% of the computational complexity of the full search method used by H.264, leading to a significant reduction in computational requirements and enabling real-time applications for mobile devices with the efficiency of H.264
Panos Nasiopoulos, Matthias von dem Knesebeck
ICME1
2006 An improved scalar quantization-based digital video watermarking scheme for H.264/AVC
abstract
Digital video watermarking has attracted a great deal of research interest in the past few years in applications such as digital fingerprinting and owner identification. H.264/AVC is the latest and most advanced video coding standard, but to this date, there are very few watermarking schemes designed for it. This is mainly due to its complexity and compression efficiency which presents a major challenge for any video watermarking approach. We developed a quantization-based video watermarking scheme, which is designed to work with H.264/AVC. Our scheme offers constant robustness at all compression rates without affecting the overall bite rate and quality of the video stream. Experimental results show that compared to existing methods, our scheme significantly outperforms existing methods under compression, transcoding, filtering, scaling, rotation and collusion attacks
Adarsh Golikeri, Panos Nasiopoulos, Z. Jane Wang 0001
ISCAS2
2006 An Efficient Compression Scheme for Colour Filter Array Video Sequences
abstract
Most consumer digital cameras use a single light sensor which captures colour information using a colour filter array (CFA). This produces a mosaic image, where each pixel contains either a red, green or blue sample. The two missing colours at each pixel location must be interpolated from the surrounding samples in a process called demosaicking. The conventional approach to compressing video in these devices is to first perform demosaicking and then compress the resulting full-colour video using standard methods. In this paper a method for compressing CFA video prior to demosaicking is proposed, using a modified version of the H.264 video coding standard. Simulations show the proposed method gives better compression efficiency than the demosaick-first approach at high bit-rates
Colin Doutre, Panos Nasiopoulos
MMSP2
2006 Data Transmission Schemes for DVD-Like Interactive TV
abstract
Current interactive services for digital TV are limited. They basically display a Web page alongside the TV program, which enhances the viewer's experience by providing extra information about the TV program. We define new interactive services for digital TV, which provide DVD-like interactivity to TV viewers. These services enable viewers to control the content and final presentation of a TV program. Some of the attractive applications of our services include parental management, multilingual audio, multiangle video, video in video, etc. The challenge in implementing these services is in transmitting an extra audio or video stream (called incidental) along with the main streams of the TV program. In the first part of this paper, we present a framework for adding the incidental streams to the original transmission stream without increasing the required bandwidth, degrading the picture quality of the main streams, or violating the compatibility of the transmitted stream with standard TV receivers. In the second part of this paper, we explore the two basic mechanisms of the presented framework: traffic characterization and admission control. We present methods for implementing these mechanisms. Using our methods, one can determine whether a TV transmission network has the capability of sending an incidental stream or not. Simulations were conducted to test the validity of our method. The results verify that our method successfully transmits the incidental streams without any discrepancy and without affecting the quality of the main streams
Mehran Azimi, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Multim.2
2005 A robust watermarking scheme based on informed coding and informed embedding
abstract
A watermarking algorithm that relies on informed coding and informed embedding is presented. This method embeds one bit of the watermark in every image block, but offers higher watermarked image quality as well as higher robustness to image processing operation attacks than other known methods using informed coding and informed embedding. The method is shown to withstand high values of added Gaussian noise, valumetric scaling, low-pass filtering as well as lossy JPEG compression. Each bit (0 or 1) to be embedded is represented by a subset of codewords. For every image block, a vector is extracted. This vector is modified so that its correlation with the codewords related to the bit to be embedded in it has higher probability than those of the codewords representing the other bit even if the image is later modified by image processing operations. The vector modification is also carried so that the change in the image fidelity is minimal.
Lino Coria-Mendoza, Panos Nasiopoulos, Rabab K. Ward
ICIP (1)2
2005 Fast Video Motion Estimation Algorithm for Mobile Devices
abstract
The power consumption of video enabled mobile devices is an ongoing challenge. The computational burden caused by the motion estimation process involved in video compression is the main cause for this consumption. We developed a new algorithm which drastically reduces the computation costs of the motion estimation process over the existing techniques. Our algorithm uses only 0.5% of the computational complexity of the full search method, making it the fastest presently available motion compensation method
Ho-Jae Lee, Panos Nasiopoulos, Victor C. M. Leung
ICME2
2004 A new signal model and identification algorithm for hidden semi-Markov signals
abstract
Markovian models form a powerful tool for modelling physical signals. In this approach, a signal generation model is employed, and its parameters are estimated from signal samples. We present a novel signal generation model for hidden semi-Markov models, HSMMs. Our model results in a significantly easier and more efficient parameter identification method. Instead of the constant probabilities presently used for modelling state transitions, we use state transition probabilities that are state-duration dependant. We then develop a parameter identification algorithm based on the maximum likelihood criterion. Our numerical results show that our parameter identification algorithm can successfully, and more efficiently, estimate the actual values of the model parameters of an HSMM signal.
Mehran Azimi, Panos Nasiopoulos, Rabab K. Ward
ICASSP (2)2
2002 A scheduling scheme for multiplexing of VBR sources in digital TV systems
abstract
Digital TV transmission systems allow a transmission channel to be shared by a number of sources. In order to improve the bandwidth utilization, variable bit rate encoding and statistical multiplexing techniques are usually used. However, the channel sharing requires a careful scheduling method for multiplexing. This is because the video and audio materials have to be presented at the receivers at specific points in time. In this paper, we present a novel scheduling scheme for statistical multiplexing of VBR sources. Our method is sensitive to the timing requirements of the sources and sends the packets as close to their transmission deadlines as possible. The advantages of our method are: (1) it decreases the broadcast deadline violation probability (or improves the bandwidth utilization), (2) it minimizes the delay and delay jitter of packets and (3) it generates a transport stream compliant with all the standard TV receivers. Simulations were conducted to compare our algorithm with the first-come-first-serve scheduling method. The results show that our algorithm significantly reduces both the percentage of dropped packets (by 35%-50%) and the average packet delay.
Mehran Azimi, Panos Nasiopoulos, Rabab K. Ward
ICIP (3)2
2001 New interactive services for digital TV
abstract
Current interactive TV services are limited. They mainly consist of accessing the World Wide Web. These services basically display a Web page beside a TV program that is related to the contents of TV program. We define new interactive services for digital TV, which can be used in many attractive applications such as parental management, multilingual audio and video in video. These services have also the advantages that they do not require an Internet return path and that they preserve the compatibility of the transmitted stream with conventional TV systems and receivers. In this paper we evaluate the feasibility of the proposed interactive services for digital TV and specify the challenges and problems in implementing these services. We present methods for overcoming these problems. We also specify the minimum decoder buffer required and the average random access delay for multilingual support in a typical standard definition TV program.
Mehran Azimi, Panos Nasiopoulos, Rabab K. Ward
ICIP (1)2
2000 DVD: Redefining Multimedia (Abstract)
Panos Nasiopoulos
ICIP1
2000 Interactive DVD Programming Using Next Generation Content-Based Encoded Multimedia Data
abstract
In this paper we propose a method of expanding the existing DVD-Video standard by incorporating MPEG-4 encoded content and its associated object-based interactivity into DVD. The proposed integration of MPEG-4 into DVD will enrich the already proven and successful DVD technology with the advanced interactive capabilities of MPEG-4, while maintaining backward capability with the existing DVD standard.
Katerina Pronina, Rabab K. Ward, Panos Nasiopoulos
ICIP3
1999 Joint MPEG-2 coding for multi-program broadcasting of pre-recorded video
abstract
We developed a cost-effective operational system suitable for digital TV, video on demand, and high definition TV broadcast over satellite networks with limited bandwidth. This MPEG-2 based system is easy to implement and allows the joint video coding of multiple video programs. Compared to present broadcast operations and for the same level of picture quality, our system greatly increases the number of video streams transmitted in each channel. As a result, either a large number of transponders can be freed up to carry real-time broadcasting or the level of the transmitted picture quality can be significantly increased, By switching from tape storage to video server technology, the need for numerous (expensive) playback VTR systems at the headend is eliminated. In addition, the majority of the complete MPEG-2 encoders are replaced by much less complex MPEG-2 transcoders. All this means significant savings for the broadcast stations. In addition to the gain in bandwidth and the reduction in cost, our system speeds up the encoding process by six fold.
Irene Koo, Panos Nasiopoulos, Rabab K. Ward
ICASSP2
1998 MPEG-2 video coding with image partitioning
abstract
In this paper we present an original motion compensation strategy based on frame partitioning. The proposed method uses different temporal resolutions within a frame to improve compression. We present a new bit allocation and rate control algorithm complementing our motion compensation technique. This unique approach to bit allocation ensures the consistency of quality throughout a single frame and a GOP. For the same picture quality, frame partitioning alone yields an additional increase of up to 20 percent or more of the encoding efficiency.
Ekaterina G. Barzykina, Panos Nasiopoulos, Rabab K. Ward
ICASSP2
1996 HDTV picture quality performance in the presence of random errors, analysis and measures for improvement
Panos Nasiopoulos, Rabab K. Ward, Dimitrios P. Bouras, P. Takis Mathiopoulos
Signal Process. Image Commun.1
1995 A Hybrid Coding Method for Digital HDTV Signals
abstract
Digitally transmitted video images suffer from sudden degradation in picture quality at high channel noise levels. This is due to the variable length coding and the large synchronization blocks used. To remedy that, we propose encoding the DCT coefficients by a hybrid fixed and variable lengths compression scheme. This method improves the compression ratio by approximately 20% and increases the system's resistance to channel errors. We then combine this hybrid algorithm with an error protected synchronization method which uses small blocks of size 32/spl times/16 pixels. The resulting combined schemes 1) significantly improve the error resistance characteristic of the system, 2) eliminate the abrupt picture degradation, and 3) does not alter the original data transmission rate.
Panos Nasiopoulos, Rabab K. Ward
ISCAS1
1995 A high-quality fixed-length compression scheme for color images
abstract
We present a new compression method which compresses 8/spl times/8 picture blocks by fixed-length codewords. The compression operation is performed on the discrete cosine transforms, DCT, of each block. As a result, our method combines the distinct advantage of being fixed-length with the high image quality obtained by the DCT based compression methods. Our method has excellent error-resistance characteristics since it does not have the synchronization and error propagation problems inherent in variable-length coding methods.
Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Commun.1
1993 Improving the HDTV picture performance under noisy transmission conditions
Panos Nasiopoulos, Rabab K. Ward
ICASSP (5)1
1991 Adaptive compression coding
abstract
A compression technique which preserves edges in compressed pictures is developed. The proposed compression algorithm adapts itself to the local nature of the image. Smooth regions are represented by their averages and edges are preserved using quad trees. Textured regions are encoded using BTC (block truncation coding) and a modification of BTC using look-up tables. A threshold using a range which is the difference between the maximum and the minimum grey levels in a 4*4 pixel quadrant is used. At the recommended value of the threshold (equal to 18), the quality of the compressed texture regions is very high, the same as that of AMBTC (absolute moment block truncation coding), but the edge preservation quality is far superior to that of AMBTC. Compression levels below 0.5-0.8 b/pixel may be achieved.>
Panos Nasiopoulos, Rabab K. Ward, Daryl J. Morse
IEEE Trans. Commun.1