Ioannis Katsavounidis

dblp:66/3977 · DBLP profile ↗
← Back
41ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0002-9072-4250ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 2Systems, architecture and hardware · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 HoloQA: Full Reference Video Quality Assessor of Rendered Human Avatars in Virtual Reality
abstract
We present HoloQA, a new state-of-the-art Full Reference Video Quality Assessment (VQA) model that was designed using principles of visual neuroscience, information theory, and self-supervised deep learning to accurately predict the quality of rendered digital human avatars in Virtual Reality (VR) and Augmented Reality (AR) systems. The growing adoption of VR/AR applications that aim to transmit digital human avatars over bandwidth-limited video networks has driven the need for VQA algorithms that better account for the kinds of distortions that reduce the quality of rendered and viewed avatars. As we will show, standard VQA models often fail to capture distortions unique to the rendering, transmission, and compression of videos containing human avatars. Towards solving this difficult problem, we adopt a multi-level Mixture-of-Experts approach. This involves computing distortion-aware perceptual features and high-level content-aware deep features that capture semantic attributes of human body avatars. The high-level features are computed using a self-supervised, pre-trained deep learning network. We show that HoloQA is able to achieve state-of-the-art performance on the recently introduced LIVE-Meta Rendered Human Avatar VQA database, demonstrating its efficacy in predicting the quality of rendered human avatars in VR. Furthermore, we demonstrate the competitive performance of HoloQA on other digital human avatar databases and on another synthetically generated video quality use case: cloud gaming. The code associated with this work will be made available on https://github.com/avinabsaha/HologramQAGitHub.
Avinab Saha, Yu-Chih Chen, Christian Häne, Jean-Charles Bazin, Ioannis Katsavounidis, Alexandre Chapiro, Alan C. Bovik
IEEE Trans. Image Process.5
2026 Joint Quality Assessment and Example-Guided Tone Mapping by Disentangling Picture Appearance From Content
abstract
The deep learning revolution has strongly impacted low-level image processing tasks such as style/domain transfer, enhancement/restoration, and visual quality assessments. Despite often being treated separately, the aforementioned tasks share a common theme of understanding, editing, or enhancing the appearance of input images without modifying the underlying content. We leverage this observation to develop a novel disentangled representation learning method that decomposes inputs into content and appearance features. The model is trained in a self-supervised manner and we use the learned features to develop a new quality prediction model named DisQUE. We demonstrate through extensive evaluations that DisQUE achieves state-of-the-art accuracy across quality prediction tasks and distortion types. Moreover, we demonstrate that the same features may also be used for image processing tasks such as HDR tone mapping, where the desired output characteristics may be tuned using example input-output pairs.
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Hassene Tmar, Alan C. Bovik
IEEE Trans. Image Process.3
2025 Cut-FUNQUE: An objective quality model for compressed tone-mapped High Dynamic Range videos
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Hassene Tmar, Alan C. Bovik
Signal Process. Image Commun.3
2024 Encoding Time and Energy Model for SVT-AV1 Based on Video Complexity
abstract
The share of online video traffic in global carbon dioxide emissions is growing steadily. To comply with the demand for video media, dedicated compression techniques are continuously optimized, but at the expense of increasingly higher computational demands and thus rising energy consumption at the video encoder side. In order to find the best trade-off between compression and energy consumption, modeling encoding energy for a wide range of encoding parameters is crucial. We propose an encoding time and energy model for SVT-AV1 based on empirical relations between the encoding time and video parameters as well as encoder configurations. Furthermore, we model the influence of video content by established content descriptors such as spatial and temporal information. We then use the predicted encoding time to estimate the required energy demand and achieve a prediction error of 19.6% for encoding time and 20.9% for encoding energy.
Lena Eichermüller, Gaurang Chaudhari, Ioannis Katsavounidis, Zhijun Lei, Hassene Tmar, Christian Herglotz, André Kaup
ICASSP3
2024 Comparison of Crowdsourcing And Laboratory Settings for Subjective Assessment of Video Quality and Acceptability & Annoyance
abstract
User satisfaction is significantly influenced by their expectations of video quality. Even when users are presented with identical video stimuli, the Quality of Experience (QoE) can vary based on the context. The acceptability and annoyance paradigm serves as a tool to understand this relationship by measuring QoE as a function of user expectations and video quality. Traditionally, subjective experiments assessing QoE have been conducted in controlled laboratory settings. While the extension of traditional video quality experiments to crowdsourcing settings is well-explored, the impact of crowdsourcing on QoE studies has not been thoroughly examined. This study explore the potential use of crowdsourcing platforms for acceptability & annoyance experiments. To this end, video quality and acceptability & annoyance experiments were conducted in both laboratory and crowdsourcing settings. The findings reveal a more linear relationship between video quality and QoE in crowdsourcing settings. Subjects in crowdsourcing settings tend to have higher expectations of video quality, resulting in a slight increase in acceptability & annoyance thresholds compared to laboratory experiments. Analyses suggest that extending acceptability & annoyance experiments to crowdsourcing is not as straightforward as extending traditional video quality experiments. In crowdsourcing settings, priming subject expectations with instructions is not as effective as it is in laboratory conditions.
Ali Ak, Abhishek Gera, Denise Noyes, Hassene Tmar, Ioannis Katsavounidis, Patrick Le Callet
ICIP5
2024 Bitrate Ladder Construction Using Visual Information Fidelity
abstract
Recently proposed perceptually optimized per-title video encoding methods provide better BD-rate savings than fixed bitrate-ladder approaches that have been employed in the past. However, a disadvantage of per-title encoding is that it requires significant time and energy to compute bitrate ladders. Over the past few years, a variety of methods have been proposed to construct optimal bitrate ladders including using low-level features to predict cross-over bitrates, optimal resolutions for each bitrate, predicting visual quality, etc. Here, we deploy features drawn from Visual Information Fidelity (VIF) (VIF features) extracted from uncompressed videos to predict the visual quality (VMAF) of compressed videos. We present multiple VIF feature sets extracted from different scales and subbands of a video to tackle the problem of bitrate ladder construction. Comparisons are made against a fixed bitrate ladder and a bitrate ladder obtained from exhaustive encoding using Bjontegaard delta metrics.
Krishna Srikar Durbha, Hassene Tmar, Cosmin Stejerean, Ioannis Katsavounidis, Alan C. Bovik
PCS4
2024 "Discriminability-Experimental Cost" Tradeoff in Subjective Video Quality Assessment of Codec: DCR with EVP Rating Scale Versus ACR-HR
abstract
This work uses naive observers to compare two subjective studies conducted in a controlled laboratory environment on SDR HD, UHD, and HDR UHD contents. These tests aim to compare the precision and accuracy of a modified Degradation Category Rating (DCR) and Absolute Category Rating with Hidden Reference (ACR-HR) subjective methods for video quality assessment. The modified version of the DCR method includes a repetition of both reference and distorted stimuli; and utilizes an 11-grade rating scale from Expert Viewing Protocol (EVP) of ITU-R BT.500-15 standards. In the second subjective protocol, ACR-HR operates without repetition and with the 5-grade quality scale from ITU standards. We extensively analyze the scale usage and compare Mean Opinion Score (MOS) discriminability in both subjective studies. We show that both methods can retrieve accurate MOS. However, the ACR-HR method achieves better discriminability among MOS than DCR with the EVP rating scale while reducing the experimental effort by a factor of two, i.e., the cost of the experiment. The findings of this work give new insight into how to perform cost-efficient subjective tests for video quality estimation with naive observers and how to retrieve good MOS estimates.
Andreas Pastor, Ioannis Katsavounidis, Lukas Krasula, Andrey Norkin, Hassene Tmar, Patrick Le Callet
PCS3
2024 A FUNQUE Approach to the Quality Assessment of Compressed HDR Videos
abstract
Recent years have seen steady growth in the popularity and availability of High Dynamic Range (HDR) content, particularly videos, streamed over the internet. As a result, assessing the subjective quality of HDR videos, which are generally subjected to compression, is of increasing importance. In particular, we target the task of full-reference quality assessment of compressed HDR videos. The state-of-the-art (SOTA) approach HDRMAX involves augmenting off-the-shelf video quality models, such as VMAF, with features computed on nonlinearly transformed video frames. However, HDRMAX increases the computational complexity of models like VMAF. Here, we show that an efficient class of video quality prediction models named FUNQUE+ achieves SOTA accuracy. This shows that the FUNQUE+ models are flexible alternatives to VMAF that achieve higher HDR video quality prediction accuracy at lower computational cost.
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Alan C. Bovik
PCS3
2024 Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual Reality
abstract
We study the visual quality judgments of human subjects on digital human avatars (sometimes referred to as "holograms" in the parlance of virtual reality [VR] and augmented reality [AR] systems) that have been subjected to distortions. We also study the ability of video quality models to predict human judgments. As streaming human avatar videos in VR or AR become increasingly common, the need for more advanced human avatar video compression protocols will be required to address the tradeoffs between faithfully transmitting high-quality visual representations while adjusting to changeable bandwidth scenarios. During transmission over the internet, the perceived quality of compressed human avatar videos can be severely impaired by visual artifacts. To optimize trade-offs between perceptual quality and data volume in practical workflows, video quality assessment (VQA) models are essential tools. However, there are very few VQA algorithms developed specifically to analyze human body avatar videos, due, at least in part, to the dearth of appropriate and comprehensive datasets of adequate size. Towards filling this gap, we introduce the LIVE-Meta Rendered Human Avatar VQA Database, which contains 720 human avatar videos processed using 20 different combinations of encoding parameters, labeled by corresponding human perceptual quality judgments that were collected in six degrees of freedom VR headsets. To demonstrate the usefulness of this new and unique video resource, we use it to study and compare the performances of a variety of state-of-the-art Full Reference and No Reference video quality prediction models, including a new model called HoloQA. As a service to the research community, we publicly releases the metadata of the new database at https://live.ece.utexas.edu/research/LIVE-Meta-rendered-human-avatar/index.html.
Yu-Chih Chen, Avinab Saha, Alexandre Chapiro, Christian Häne, Jean-Charles Bazin, Stefano Zanetti, Ioannis Katsavounidis, Alan C. Bovik
IEEE Trans. Image Process.8
2024 One Transform to Compute Them All: Efficient Fusion-Based Full-Reference Video Quality Assessment
abstract
The Visual Multimethod Assessment Fusion (VMAF) algorithm has recently emerged as a state-of-the-art approach to video quality prediction, that now pervades the streaming and social media industry. However, since VMAF requires the evaluation of a heterogeneous set of quality models, it is computationally expensive. Given other advances in hardware-accelerated encoding, quality assessment is emerging as a significant bottleneck in video compression pipelines. Towards alleviating this burden, we propose a novel Fusion of Unified Quality Evaluators (FUNQUE) framework, by enabling computation sharing and by using a transform that is sensitive to visual perception to boost accuracy. Further, we expand the FUNQUE framework to define a collection of improved low-complexity fused-feature models that advance the state-of-the-art of video quality performance with respect to both accuracy, by 4.2% to 5.3%, and computational efficiency, by factors of 3.8 to 11 times!.
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Alan C. Bovik
IEEE Trans. Image Process.3
2023 Estimating Uncertainty On Video Quality Metrics
abstract
Video Quality Metrics (VQM) are models used to predict the score that a user would give to the quality of a video visualization. They are widely used in video processing systems, for monitoring end-to-end quality or system troubleshooting for example. In these scenarios, the improvement is quantified based on a certain enhancement of a VQM score and trouble-detection is done based on a certain drop or threshold computed based on a VQM. Yet, whether such improvement or fault-detection is worth a significant increase in power consumption is questionable. Therefore, the goal of this work is to propose a method to predict the uncertainty of the quality metric. In this paper, we propose a framework to evaluate the confidence interval of a VQM for a given content using simple video features. We assess the performance of the framework by using the confidence intervals to predict if two videos are of similar or different quality and show that in most cases our approach performs better than just using a constant confidence interval.
Patrick Le Callet, Suiyi Ling, Haixiong Wang, Ioannis Katsavounidis, Zafar Shahid, Cosmin Stejerean
ICASSP5
2023 Encoder Complexity Control in SVT-AV1 by Speed-Adaptive Preset Switching
abstract
Current developments in video encoding technology lead to continuously improving compression performance but at the expense of increasingly higher computational demands. Regarding the online video traffic increases during the last years and the concomitant need for video encoding, encoder complexity control mechanisms are required to restrict the processing time to a sufficient extent in order to find a reasonable trade-off between performance and complexity. We present a complexity control mechanism in SVT-AV1 by using speed-adaptive preset switching to comply with the remaining time budget. This method enables encoding with a user-defined time constraint within the complete preset range with an average precision of 8.9 % without introducing any additional latencies.
Lena Eichermüller, Gaurang Chaudhari, Ioannis Katsavounidis, Zhijun Lei, Hassene Tmar, André Kaup, Christian Herglotz
ICIP3
2023 Video Consumption in Context: Influence of Data Plan Consumption on QoE
abstract
User expectations are one of the main factors on providing satisfactory QoE for streaming service providers. Measuring acceptability and annoyance of video content, therefore, provide a valuable insight when measured under a given context. In this ongoing work, we measure video QoE in terms of acceptability and annoyance for the remaining data in a mobile data plan context.. We show that simple logos can be used during the experiment to prompt the context to subjects and the different context levels may impact the user expectations and consequently their satisfactions. Finally, we show that objective metrics can be used to determine the acceptability and annoyance thresholds for a given context.
Ali Ak, Anne-Flore Perrin, Denise Noyes, Ioannis Katsavounidis, Patrick Le Callet
IMX4
2023 GAMIVAL: Video Quality Prediction on Mobile Cloud Gaming Content
abstract
The mobile cloud gaming industry has been rapidly growing over the last decade. When streaming gaming videos are transmitted to customers' client devices from cloud servers, algorithms that can monitor distorted video quality without having any reference video available are desirable tools. However, creating No-Reference Video Quality Assessment (NR VQA) models that can accurately predict the quality of streaming gaming videos rendered by computer graphics engines is a challenging problem, since gaming content generally differs statistically from naturalistic videos, often lacks detail, and contains many smooth regions. Until recently, the problem has been further complicated by the lack of adequate subjective quality databases of mobile gaming content. We have created a new gaming-specific NR VQA model called the Gaming Video Quality Evaluator (GAMIVAL), which combines and leverages the advantages of spatial and temporal gaming distorted scene statistics models, a neural noise model, and deep semantic features. Using a support vector regression (SVR) as a regressor, GAMIVAL achieves superior performance on the new LIVE-Meta Mobile Cloud Gaming (LIVE-Meta MCG) video quality database.
Yu-Chih Chen, Avinab Saha, Chase Davis, Rahul Gowda, Ioannis Katsavounidis, Alan C. Bovik
IEEE Signal Process. Lett.7
2023 Study of Subjective and Objective Quality Assessment of Mobile Cloud Gaming Videos
abstract
We present the outcomes of a recent large-scale subjective study of Mobile Cloud Gaming Video Quality Assessment (MCG-VQA) on a diverse set of gaming videos. Rapid advancements in cloud services, faster video encoding technologies, and increased access to high-speed, low-latency wireless internet have all contributed to the exponential growth of the Mobile Cloud Gaming industry. Consequently, the development of methods to assess the quality of real-time video feeds to end-users of cloud gaming platforms has become increasingly important. However, due to the lack of a large-scale public Mobile Cloud Gaming Video dataset containing a diverse set of distorted videos with corresponding subjective scores, there has been limited work on the development of MCG-VQA models. Towards accelerating progress towards these goals, we created a new dataset, named the LIVE-Meta Mobile Cloud Gaming (LIVE-Meta-MCG) video quality database, composed of 600 landscape and portrait gaming videos, on which we collected 14,400 subjective quality ratings from an in-lab subjective study. Additionally, to demonstrate the usefulness of the new resource, we benchmarked multiple state-of-the-art VQA algorithms on the database. The new database will be made publicly available on our website: https://live.ece.utexas.edu/research/LIVE-Meta-Mobile-Cloud-Gaming/index.html.
Avinab Saha, Yu-Chih Chen, Chase Davis, Rahul Gowda, Ioannis Katsavounidis, Alan C. Bovik
IEEE Trans. Image Process.7
2022 Domain-Specific Fusion Of Objective Video Quality Metrics
abstract
Video processing algorithms like video upscaling, denoising, and compression are now increasingly optimized for perceptual quality metrics instead of signal distortion. This means that they may score well for metrics like video multi-method assessment fusion (VMAF), but this may be because of metric overfitting. This imposes the need for costly subjective quality assessments that cannot scale to large datasets and large parameter explorations. We propose a methodology that fuses multiple quality metrics based on small scale subjective testing in order to unlock their use at scale for specific application domains of interest. This is achieved by employing pseudo-random sampling of the resolution, quality range and test video content available, which is initially guided by quality metrics in order to cover the quality range useful to each application. The selected samples then undergo a subjective test, such as ITU-T P.910 absolute categorical rating, with the results of the test postprocessed and used as the means to derive the best combination of multiple objective metrics using support vector regression. We showcase the benefits of this approach in two applications: video encoding with and without perceptual preprocessing, and deep video denoising & upscaling of compressed content. For both applications, the derived fusion of metrics allows for a more robust alignment to mean opinion scores than a perceptually-uninformed combination of the original metrics themselves. The dataset and code is available at https://github.com/isize-tech/VideoQualityFusion.
Aaron Chadha, Ioannis Katsavounidis, Ayan Kumar Bhunia, Cosmin Stejerean, Muhammad Umar Karim Khan, Yiannis Andreopoulos
ACM Multimedia2
2022 PTR-CNN for in-loop filtering in video coding
Tong Shao, Dapeng Oliver Wu, Chia-Yang Tsai, Zhijun Lei, Ioannis Katsavounidis
J. Vis. Commun. Image Represent.6
2021 Encoding Parameters Prediction for Convex Hull Video Encoding
abstract
Fast encoding parameter selection technique have been proposed in the past. Leveraging the power of convex hull video encoding framework, an encoder with a faster speed setting, or faster encoder such as hardware encoder, can be used to determine the optimal encoding parameters. It has been shown that one can speed up 3 to 5 times, while still achieve 20–40% BD-rate savings, compared to the traditional fixed-QP encodings. Such approach presents two problems. First, although the speedup is impressive, there is still ~3% loss in BD-rate. Secondly, the previous approach only works with encoders implementing standards that have the same quantization scheme and QP range, such as between VP9 and AV1, and not in the scenario where one might want to use a much faster H.264 encoder, to determine and predict the encoding parameters for VP9 or AV1 encoder. In this work we propose and present new additions to address these issues. First, we show that the predictive model can be used to gain back the quality loss incurred from using a faster encoder. Secondly, we also demonstrate that, by introducing a mapping layer, it is possible to predict between two arbitrary encoders and reduce the computational complexity even further. Such degree of complexity reduction is made possible by the fact that different generations of video codecs are usually orders of magnitude apart in terms of complexity and thus one can use an encoder from an earlier generation to predict the encoding parameters for an encoder from later generations.
Ping-Hao Wu, Volodymyr Kondratenko, Gaurang Chaudhari, Ioannis Katsavounidis
PCS4
2021 Towards Perceptually Optimized Adaptive Video Streaming-A Realistic Quality of Experience Database
abstract
Measuring Quality of Experience (QoE) and integrating these measurements into video streaming algorithms is a multi-faceted problem that fundamentally requires the design of comprehensive subjective QoE databases and objective QoE prediction models. To achieve this goal, we have recently designed the LIVE-NFLX-II database, a highly-realistic database which contains subjective QoE responses to various design dimensions, such as bitrate adaptation algorithms, network conditions and video content. Our database builds on recent advancements in content-adaptive encoding and incorporates actual network traces to capture realistic network variations on the client device. The new database focuses on low bandwidth conditions which are more challenging for bitrate adaptation algorithms, which often must navigate tradeoffs between rebuffering and video quality. Using our database, we study the effects of multiple streaming dimensions on user experience and evaluate video quality and quality of experience models and analyze their strengths and weaknesses. We believe that the tools introduced here will help inspire further progress on the development of perceptually-optimized client adaptation and video streaming strategies. The database is publicly available at http://live.ece.utexas.edu/research/LIVE_NFLX_II/live_nflx_plus.html.
Christos G. Bampis, Zhi Li 0001, Ioannis Katsavounidis, Te-Yuan Huang, Chaitanya Ekanadham, Alan C. Bovik
IEEE Trans. Image Process.3
2020 Towards Perceptually-Optimized Compression Of User Generated Content (UGC): Prediction Of UGC Rate-Distortion Category
abstract
How to best evaluate the perceptual quality, and efficiently optimize the compression of User Generated Content (UGC) within an adaptive streaming system is becoming one of the most intractable challenges in the community. Rate-Distortion (R-D) characteristic based content analyses, which could be applied on the non-pristine originals, is inevitable to provide guidance in developing quality metrics and efficient compression system. To this end, we present a novel complete R-D category prediction system through the identification of discriminate features. To better understand the Rate-Distortion (R-D) behaviors of UGC, we first propose a Bjontegaard Delta (BD)-Rate, BD-Quality-based algorithm to categorize UGC. By using the predicted R-D related categories as ground-truth labels, we further identify features that characterize the R-D behaviors of UGC via a hierarchical feature selection framework. Finally, selected features are employed to predict the R-D category of under-test UGC. Comprehensive observations and results are summarized through extensive experiments.
Suiyi Ling, Yoann Baveye, Patrick Le Callet, Jim Skinner, Ioannis Katsavounidis
ICME5
2020 Image Coding With Data-Driven Transforms: Methodology, Performance and Potential
abstract
Image compression has always been an important topic in the last decades due to the explosive increase of images. The popular image compression formats are based on different transforms which convert images from the spatial domain into compact frequency domain to remove the spatial correlation. In this paper, we focus on the exploration of data-driven transform, Karhunen-Loéve transform (KLT), the kernels of which are derived from specific images via Principal Component Analysis (PCA), and design a high efficient KLT based image compression algorithm with variable transform sizes. To explore the optimal compression performance, the multiple transform sizes and categories are utilized and determined adaptively according to their rate-distortion (RD) costs. Moreover, comprehensive analyses on the transform coefficients are provided and a band-adaptive quantization scheme is proposed based on the coefficient RD performance. Extensive experiments are performed on several class-specific images as well as general images, and the proposed method achieves significant coding gain over the popular image compression standards including JPEG, JPEG 2000, and the state-of-the-art dictionary learning based methods.
Xinfeng Zhang 0001, Chao Yang 0021, Shan Liu 0001, Haitao Yang 0001, Ioannis Katsavounidis, Shawmin Lei, C.-C. Jay Kuo
IEEE Trans. Image Process.6
2018 Prediction of Satisfied User Ratio for Compressed Video
abstract
A large-scale video quality dataset called the VideoSet has been constructed recently to measure human subjective experience of H.264 coded video in terms of the just-noticeable-difference (JND). It measures the first three JND points of 5-second video of resolution 1080p, 720p, 540p and 360p. Based on the VideoSet, we propose a method to predict the satisfied-user-ratio (SUR) curves using a machine learning framework. First, we partition a video clip into local spatial-temporal segments and evaluate the quality of each segment using the VMAF quality index. Then, we aggregate these local VMAF measures to derive a global one. Finally, the masking effect is incorporated and the support vector regression (SVR) is used to predict the SUR curves, from which the JND points can be derived. Experimental results are given to demonstrate the performance of the proposed SUR prediction method.
Haiqiang Wang, Ioannis Katsavounidis, Qin Huang 0006, Xin Zhou 0001, C.-C. Jay Kuo
ICASSP2
2018 Recurrent and Dynamic Models for Predicting Streaming Video Quality of Experience
abstract
Streaming video services represent a very large fraction of global bandwidth consumption. Due to the exploding demands of mobile video streaming services, coupled with limited bandwidth availability, video streams are often transmitted through unreliable, low-bandwidth networks. This unavoidably leads to two types of major streaming-related impairments: compression artifacts and/or rebuffering events. In streaming video applications, the end-user is a human observer; hence being able to predict the subjective Quality of Experience (QoE) associated with streamed videos could lead to the creation of perceptually optimized resource allocation strategies driving higher quality video streaming services. We propose a variety of recurrent dynamic neural networks that conduct continuous-time subjective QoE prediction. By formulating the problem as one of time-series forecasting, we train a variety of recurrent neural networks and non-linear autoregressive models to predict QoE using several recently developed subjective QoE databases. These models combine multiple, diverse neural network inputs, such as predicted video quality scores, rebuffering measurements, and data related to memory and its effects on human behavioral responses, using them to predict QoE on video streams impaired by both compression artifacts and rebuffering events. Instead of finding a single time-series prediction model, we propose and evaluate ways of aggregating different models into a forecasting ensemble that delivers improved results with reduced forecasting variance. We also deploy appropriate new evaluation metrics for comparing time-series predictions in streaming applications. Our experimental results demonstrate improved prediction performance that approaches human performance. An implementation of this work can be found at https://github.com/christosbampis/NARX_QoE_release.
Christos G. Bampis, Zhi Li 0001, Ioannis Katsavounidis, Alan C. Bovik
IEEE Trans. Image Process.3
2017 VideoSet: A large-scale compressed video quality dataset based on JND measurement
abstract
• A large-scale JND-based coded video quality dataset is presented. • The VideoSet contains 220 5-s sequences in four resolutions coded by H.264/AVC. • The subjective test procedure, JND data cleaning and properties are described. • The significance and implications of the VideoSet are discussed. • This work points out a clear path to data-driven perceptual coding. A new methodology to measure coded image/video quality using the just-noticeable-difference (JND) idea was proposed in Lin et al. (2015). Several small JND-based image/video quality datasets were released by the Media Communications Lab at the University of Southern California in Jin et al. (2016) and Wang et al. (2016) [3]. In this work, we present an effort to build a large-scale JND-based coded video quality dataset. The dataset consists of 220 5-s sequences in four resolutions (i.e., 1920 × 1080 , 1280 × 720 , 960 × 540 and 640 × 360 ). For each of the 880 video clips, we encode it using the H.264/AVC codec with QP = 1 , … , 51 and measure the first three JND points with 30 + subjects. The dataset is called the “VideoSet”, which is an acronym for “Video Subject Evaluation Test (SET)”. This work describes the subjective test procedure, detection and removal of outlying measured data, and the properties of collected JND data. Finally, the significance and implications of the VideoSet to future video coding research and standardization efforts are pointed out. All source/coded video clips as well as measured JND data included in the VideoSet are available to the public in the IEEE DataPort (Wang et al., 2016 [4]).
Haiqiang Wang, Ioannis Katsavounidis, Jiantong Zhou, Jeong-Hoon Park, Shawmin Lei, Xin Zhou 0001, Man-On Pun, Xin Jin 0002, Ronggang Wang, Xu Wang 0006, Yun Zhang 0002, Jiwu Huang, Sam Kwong, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.2
2017 Study of Temporal Effects on Subjective Video Quality of Experience
abstract
HTTP adaptive streaming is being increasingly deployed by network content providers, such as Netflix and YouTube. By dividing video content into data chunks encoded at different bitrates, a client is able to request the appropriate bitrate for the segment to be played next based on the estimated network conditions. However, this can introduce a number of impairments, including compression artifacts and rebuffering events, which can severely impact an end-user's quality of experience (QoE). We have recently created a new video quality database, which simulates a typical video streaming application, using long video sequences and interesting Netflix content. Going beyond previous efforts, the new database contains highly diverse and contemporary content, and it includes the subjective opinions of a sizable number of human subjects regarding the effects on QoE of both rebuffering and compression distortions. We observed that rebuffering is always obvious and unpleasant to subjects, while bitrate changes may be less obvious due to content-related dependencies. Transient bitrate drops were preferable over rebuffering only on low complexity video content, while consistently low bitrates were poorly tolerated. We evaluated different objective video quality assessment algorithms on our database and found that objective video quality models are unreliable for QoE prediction on videos suffering from both rebuffering events and bitrate changes. This implies the need for more general QoE models that take into account objective quality models, rebuffering-aware information, and memory. The publicly available video content as well as metadata for all of the videos in the new database can be found at http://live.ece.utexas.edu/research/LIVE_NFLXStudy/nflx_index.html.
Christos G. Bampis, Zhi Li 0001, Anush K. Moorthy, Ioannis Katsavounidis, Anne Aaron, Alan C. Bovik
IEEE Trans. Image Process.4
2016 MCL-JCV: A JND-based H.264/AVC video quality assessment dataset
abstract
A compressed video quality assessment dataset based on the just noticeable difference (JND) model, called MCL-JCV, is recently constructed and released. In this work, we explain its design objectives, selected video content and subject test procedures. Then, we conduct statistical analysis on collected JND data. We compute the difference between every two adjacent JND points and propose an outlier detection algorithm to remove unreliable data. We also show that each JND difference group can be well approximated by a normal distribution so that we can adopt the Gaussian mixture model (GMM) to characterize the distribution of multiple JND points. Finally, it is demonstrated by experimental results that the proposed JND analysis performed in the difference domain, called the D-method, achieves a lower BIC (Bayesian information criteria) value than the previously proposed G-method.
Haiqiang Wang, Weihao Gan, Sudeng Hu, Joe Yuchieh Lin, Lina Jin, Longguang Song, Ioannis Katsavounidis, Anne Aaron, C.-C. Jay Kuo
ICIP8
2016 Initialization of dynamic time warping using tree-based fast Nearest Neighbor
Stergios Poularakis, Ioannis Katsavounidis
Pattern Recognit. Lett.2
2016 Blind Picture Upscaling Ratio Prediction
abstract
Natural scene statistics are well studied in the context of picture quality assessment and have been used in a wide variety of top-performing picture quality prediction models. Upscaling artifacts have been measured with regards to quality impairment using these kinds of models. However, the assessment and classification of subtle, less discriminable upscaling artifacts remains an unsolved problem. The nearly imperceptible artifacts pertaining to the extent and type of upscaling have not been predicted using natural scene statistics (NSS)-based models. We develop an accurate model for predicting the upscaling ratio applied to any natural image. By decomposing an input image frame using an orthogonal filter bank and locally normalizing the resulting responses, we show that the local energy terms can be used to predict the upscaling ratio. In fact, a simple linear regressor can be trained on these energy measurements; hence, no hyperparameter tuning is necessary. We compare the proposed model with other no-reference models using real-world data contained in the Netflix collection.
Todd Richard Goodall, Ioannis Katsavounidis, Zhi Li 0001, Anne Aaron, Alan C. Bovik
IEEE Signal Process. Lett.2
2016 Low-Complexity Hand Gesture Recognition System for Continuous Streams of Digits and Letters
abstract
In this paper, we propose a complete gesture recognition framework based on maximum cosine similarity and fast nearest neighbor (NN) techniques, which offers high-recognition accuracy and great computational advantages for three fundamental problems of gesture recognition: 1) isolated recognition; 2) gesture verification; and 3) gesture spotting on continuous data streams. To support our arguments, we provide a thorough evaluation on three large publicly available databases, examining various scenarios, such as noisy environments, limited number of training examples, and time delay in system's response. Our experimental results suggest that this simple NN-based approach is quite accurate for trajectory classification of digits and letters and could become a promising approach for implementations on low-power embedded systems.
Stergios Poularakis, Ioannis Katsavounidis
IEEE Trans. Cybern.2
2015 HEVC decoder optimization in low power configurable architecture for wireless devices
abstract
High Efficiency Video Coding (HEVC) is the new video compression standard, reducing bitrates nearly at half compared to H.264, offering potentially significant power savings for wireless video transmission at the network interface. This reduction in bitrate is achieved by a series of computationally expensive algorithms, thus making imperative to optimize HEVC decoding in order to provide a low-power implementation that can be used in mobile devices. Extending the Instruction Set Architecture (ISA) of a configurable microprocessor with new instructions for a target application can reduce the total effort of the application, thus reducing operating frequency and eventually power. The flexibility and relatively low design effort of such microprocessors - compared to hardwired Application-Specific-Integrated-Circuit (ASIC) designs - reduces the time space for adoption of HEVC and makes them an efficient alternative for wireless devices. We propose an efficient quarter-pixel interpolation filter implementation for HEVC using new custom-made instructions and other techniques for optimization of motion compensation, implemented on a configurable microprocessor architecture. Simulation results show a four times acceleration on average of the interpolation filter module over the reference HEVC software and an overall doubling in decoder performance.
Vasileios Magoulianitis, Ioannis Katsavounidis
WOWMOM2
2014 Finger detection and hand posture recognition based on depth information
abstract
In this work, we propose a novel framework for automatic finger detection and hand posture recognition, based mainly on depth information. Our method locates apex-shaped structures in a hand contour and deals efficiently with the challenging problem of partially merged fingers. Hand posture recognition is achieved using Fourier Descriptors of the contour, while global information about the fingers helps reducing the size of the search space. Our experiments on a dataset obtained from a Kinect device confirm the high recognition accuracy of our approach.
Stergios Poularakis, Ioannis Katsavounidis
ICASSP2
2013 Sparse representations for hand gesture recognition
abstract
Dynamic recognition of gestures from video sequences is a challenging task due to the high variability in the characteristics of each gesture with respect to different individuals. In this work, we propose a novel representation of gestures as linear combinations of the elements of an overcomplete dictionary, based on the emerging theory of sparse representations. We evaluate our approach on a publicly available gesture dataset of Palm Grafti Digits and compare it with other state-of-the-art methods, such as Hidden Markov Models, Dynamic Time Warping and the recently proposed distance metric termed Move-Split-Merge. Our experimental results suggest that the proposed recognition scheme offers high recognition accuracy in isolated gesture recognition and a satisfying robustness to noisy data, thus indicating that sparse representations can be successfully applied in the field of gesture recognition.
Stergios Poularakis, Grigorios Tsagkatakis, Panagiotis Tsakalides, Ioannis Katsavounidis
ICASSP4
2012 A Multiscale Error Diffusion Technique for Digital Multitoning
abstract
Multitoning is the representation of digital pictures using a given set of available color intensities, which are also known as tones or quantization levels. It can be viewed as the generalization of halftoning, where only two such quantization levels are available. Its main application is for printing and, similar to halftoning, can be applied to both colored and grayscale images. In this paper, we present a method to produce multitones based on the multiscale error diffusion technique. Key characteristics of this technique are: 1) the use of an image quadtree; 2) the quantization order of the pixels being determined through "maximum intensity guidance" on the image quadtree; and 3) noncausal error diffusion. Special care has been given to the problem of banding, which is one of the inherent limitations in error diffusion when applied to multitoning. Banding is evident in areas of the image with values close to one of the available quantization levels; our approach is to apply a preprocessing step to alleviate part of the problem. Our results are evaluated both in terms of visual appearance and using a set of standard metrics, with the latter demonstrating the blue-noise characteristics and very low anisotropy of the proposed method.
Giorgos Sarailidis, Ioannis Katsavounidis
IEEE Trans. Image Process.2
2009 A high performance and low power hardware architecture for the transform & quantization stages in H.264
abstract
In this work, we present a hardware architecture prototype for the various types of transforms and the accompanying quantization, supported in H.264 baseline profile video encoding standard. The proposed architecture achieves high performance and can satisfy quad full high definition (QFHD) (3840middot2160@150Hz) coding. The transforms are implemented using only add and shift operations, which reduces the computation overhead. A modification in the quantization equations representation is suggested to remove the absolute value and resign operation stages overhead. Additionally, a post-scale Hadamard transform computation is presented. The architecture can achieve a reduction of about 20% in power consumption, compared to existing implementations.
Muhsen Owaida, Maria G. Koziri, Ioannis Katsavounidis, Georgios I. Stamoulis
ICME3
2007 A Novel Low-Power Motion Estimation Design for H.264
abstract
The H.264 video coding standard can achieve considerably higher coding efficiency than previous video coding standards. The keys to this high coding efficiency are the two prediction modes (Intra & Inter) provided by H.264 which adopt many new features such as variable block size searching, motion vector prediction etc. However, these result in a considerably higher encoder complexity that adversely affects speed and power, which are both significant for the mobile multimedia applications targeted by the standard. Therefore, it is of high importance to design architectures that minimize the speed and power overhead of the prediction modes. In this paper we present a new algorithm, and the architecture that implements it, that can replace the standard sum of absolute differences (SAD) approach in the two main prediction modes, supports the variable block size motion estimation (VBSME) as it is defined in the standard and provide a power efficient hardware implementation without perceivable degradation in coding efficiency or video quality.
Maria G. Koziri, Antonios N. Dadaliaris, Georgios I. Stamoulis, Ioannis Katsavounidis
ASAP4
2006 Power reduction in an H.264 encoder through algorithmic and logic transformations
abstract
The H.264 video coding standard can achieve considerably higher coding efficiency than previous video coding standards. The keys to this high coding efficiency are the two prediction modes (Intra & Inter) provided by H.264. Unfortunately, these result in a considerably higher encoder complexity that adversely affects speed and power, which are both significant for the mobile multimedia applications targeted by the standard. Therefore, it is of high importance to design architectures that minimize the speed and power overhead of the prediction modes. In this paper we present a new algorithm, and the logic transformations that enable it, that can replace the standard Sum of Absolute Differences (SAD) approach in the two main prediction modes, and provide a power efficient hardware implementation without perceivable degradation in coding efficiency or video quality.
Maria G. Koziri, Georgios I. Stamoulis, Ioannis Katsavounidis
ISLPED3
2005 Robust MMSE video decoding: theory and practical implementations
Chang-Su Kim 0001, Jongwon Kim 0001, Ioannis Katsavounidis, C.-C. Jay Kuo
IEEE Trans. Circuits Syst. Video Technol.3
1997 A multiscale error diffusion technique for digital halftoning
abstract
A new digital halftoning technique based on multiscale error diffusion is examined. We use an image quadtree to represent the difference image between the input gray-level image and the output halftone image. In iterative algorithm is developed that searches the brightest region of a given image via "maximum intensity guidance" for assigning dots and diffuses the quantization error noncausally at each iteration. To measure the quality of halftone images, we adopt a new criterion based on hierarchical intensity distribution. The proposed method provides very good results both visually and in terms of the hierarchical intensity quality measure.
Ioannis Katsavounidis, C.-C. Jay Kuo
IEEE Trans. Image Process.1
1996 Fast tree-structured nearest neighbor encoding for vector quantization
abstract
This work examines the nearest neighbor encoding problem with an unstructured codebook of arbitrary size and vector dimension. We propose a new tree-structured nearest neighbor encoding method that significantly reduces the complexity of the full-search method without any performance degradation in terms of distortion. Our method consists of efficient algorithms for constructing a binary tree for the codebook and nearest neighbor encoding by using this tree. Numerical experiments are given to demonstrate the performance of the proposed method.
Ioannis Katsavounidis, C.-C. Jay Kuo, Zhen Zhang 0010
IEEE Trans. Image Process.1
1994 A new initialization technique for generalized Lloyd iteration
abstract
The generalized Lloyd algorithm plays an important role in the design of vector quantizers (VQ) and in feature clustering for pattern recognition. In the VQ context, this algorithm provides a procedure to iteratively improve a codebook and results in a local minimum that minimizes the average distortion function. We propose an efficient method to obtain a good initial codebook that can accelerate the convergence of the generalized Lloyd algorithm and achieve a better local minimum as well.>
Ioannis Katsavounidis, C.-C. Jay Kuo, Zhen Zhang 0010
IEEE Signal Process. Lett.1
1993 Recursive multiscale error-diffusion technique for digital halftoning
abstract
The technique of mapping a given gray level to some arrangement of dots such that it renders the desired gray level is called halftoning. In this research, we propose a new digital halftoning algorithm to achieve this goal based on an approach called the recursive multiscale error diffusion. Our main assumption is that the resulting intensity from a raster of dots is in proportion to the number of dots on that raster. In analogy, the intensity of the corresponding region of the input image is simply the integral of the (normalized) gray level over the region. The two intensities should be matched as much as possible. It is shown that the area of integration plays an important role to how successful the matching of the two intensities can be, and since the area of integration corresponds to different resolutions (therefore to different viewing distances), we address the problem of matching the intensities, as much as possible for every resolution. Advantages of our method include very good performance, versatility and ease of hardware implementation.
Ioannis Katsavounidis, C.-C. Jay Kuo
VCIP1