VLDB 2026 Research / reviewers in the wild / expert
Balu Adsumilli
dblp:182/7075
· DBLP profile ↗
50ranked-venue papers
1as first author
38since 2021 · last 2026
0000-0002-5187-6331ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 1 first-author · 37 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BrightRate: Quality Assessment for User-Generated HDR VideosabstractHigh Dynamic Range (HDR) videos offer superior luminance and color fidelity as compared to Standard Dynamic Range (SDR) content. The rapid growth of User-Generated Content (UGC) on platforms such as YouTube, Instagram, and TikTok has brought a significant increase in the volumes of streamed and shared UGC videos. This newer category of videos brings new challenges to the development of effective No-Reference (NR) video quality assessment (VQA) models specialized to HDR UGC, because of the extreme variety and severities of distortions, arising from diverse capture, editing, and processing outcomes. Towards addressing this issue, we introduce BrightVQ, a sizeable new psychometric data resource. It is the first large-scale subjective video quality database dedicated to the quality modelling of HDR UGC videos. BrightVQ comprises 2,100 videos, on which we collected 73,794 perceptual quality ratings. Using this dataset, we also developed BrightRate, a novel video quality prediction model designed to capture both UGC-specific distortions coexisting with HDR-specific artifacts. Extensive experimental results demonstrate that BrightRate achieves state-of-the-art performance across HDR databases. Shreshth Saini, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
WACV | 5 |
| 2026 | Quality Prediction of Embedded and Overlaid Text in User-Generated Visual ContentabstractUser-generated visual content (UGC) now occupies a significant fraction of internet traffic, and billions of UGC videos and pictures are uploaded daily. Among these, short-form video content now accounts for most of the videos consumed by online users. Given the popularity of short-form UGC content, being able to control the perceptual quality of UGC videos has emerged as an important problem. Visual UGC is subject to myriad types, severity, and combinations of distortions. While UGC video quality has been closely studied, the quality and legibility of text that is overlaid or embedded in short-form UGC videos has received relatively low attention. However, being able to accurately predict text quality in images is important, since it both impacts the overall perception of the content it is embedded in, as well as the messages being conveyed. It is also beneficial for applications involving image or video text recognition which can affect visual search and content identification. Analyzing the quality of text embedded in pictures or videos is a hard problem, since perception of it is commingled with the surrounding visual content. Our work, which greatly extends our early report on text legibility prediction, contributes to both the psychophysics of embedded text quality as well as to computational models of its perception. We have created two subjective datasets-designated as the LIVE-COCO Text Legibility (LIVE-COCO-TL) Database (a modification of COCO-Text), and the LIVE-YouTube Text-in-Video Quality (LIVE-YT-TVQ) Database. LIVE-COCO-TL contains 74,440 text patches with legibility annotations, while LIVE-YT-TVQ contains $\sim ~19$ K subjective quality ratings on 405 videos and 641 text patches extracted from them. We build models that predict embedded or overlaid text legibility and text quality, as well as a multi-task model that simultaneously predicts the overall quality of videos with embedded or overlaid and local text quality. We are making the databases and all models freely available at https://live.ece.utexas.edu/research/LIVE_YouTube_Text_Quality_Assessment/index.html. Maniratnam Mandal, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2025 | An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLMabstractThe rise of short-form videos, characterized by diverse content, editing styles, and artifacts, poses substantial challenges for learning-based blind video quality assessment (BVQA) models. Multimodal large language models (MLLMs), renowned for their superior generalization capabilities, present a promising solution. This paper focuses on effectively leveraging a pretrained MLLM for short-form video quality assessment, regarding the impacts of pre-processing and response variability, and insights on combining the MLLM with BVQA models. We first investigated how frame pre-processing and sampling techniques influence the MLLM’s performance. Then, we introduced a lightweight learning-based ensemble method that adaptively integrates predictions from the MLLM and state-of-the-art BVQA models. Our results demonstrated superior generalization performance with the proposed ensemble approach. Furthermore, the analysis of content-aware ensemble weights highlighted that some video characteristics are not fully represented by existing BVQA models, revealing potential directions to improve BVQA models further. Wen Wen 0007, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli |
ICASSP | 4 |
| 2025 | Rate-Distortion Optimization with Non-Reference Metrics for UGC CompressionabstractService providers must encode a large volume of noisy videos to meet the demand for user-generated content (UGC) in online video-sharing platforms. However, low-quality UGC challenges conventional codecs based on rate-distortion optimization (RDO) with full-reference metrics (FRMs). While effective for pristine videos, FRMs drive codecs to preserve artifacts when the input is degraded, resulting in suboptimal compression. A more suitable approach used to assess UGC quality is based on non-reference metrics (NRMs). However, RDO with NRMs as a measure of distortion requires an iterative workflow of encoding, decoding, and metric evaluation, which is computationally impractical. This paper overcomes this limitation by linearizing the NRM around the uncompressed video. The resulting cost function enables block-wise bit allocation in the transform domain by estimating the alignment of the quantization error with the gradient of the NRM. To avoid large deviations from the input, we add sum of squared errors (SSE) regularization. We derive expressions for both the SSE regularization parameter and the Lagrangian, akin to the relationship used for SSE-RDO. Experiments with images and videos show bitrate savings of more than 30% over SSE-RDO using the target NRM, with no decoder complexity overhead and minimal encoder complexity increase. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega, Neil Birkbeck, Balu Adsumilli |
ICIP | 6 |
| 2025 | CHUG: Crowdsourced User-Generated HDR Video Quality DatasetabstractHigh Dynamic Range (HDR) videos enhance visual experiences with superior brightness, contrast, and color depth. The surge of User-Generated Content (UGC) on platforms like YouTube and TikTok introduces unique challenges for HDR video quality assessment (VQA) due to diverse capture conditions, editing artifacts, and compression distortions. Existing HDR-VQA datasets primarily focus on professionally generated content (PGC), leaving a gap in understanding real-world UGC-HDR degradations. To address this, we introduce CHUG: Crowdsourced User-Generated HDR Video Quality Dataset, the first large-scale subjective study on UGC-HDR quality. CHUG comprises 856 UGC-HDR source videos, transcoded across multiple resolutions and bitrates to simulate real-world scenarios, totaling 5,992 videos. A large-scale study via Amazon Mechanical Turk collected 211,848 perceptual ratings. CHUG provides a benchmark for analyzing UGC-specific distortions in HDR videos. We anticipate CHUG will advance No-Reference (NR) HDR-VQA research by offering a large-scale, diverse, and real-world UGC dataset. The dataset is publicly available at: https://shreshthsaini.github.io/CHUG/. Shreshth Saini, Alan C. Bovik, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli |
ICIP | 5 |
| 2025 | Google Industry Seminar: Video Processing in the New Age of AIabstractVideo processing and compression are being re-envisioned in the age of AI. Traditional video codecs, which rely on rigid, pre-defined rules, are being augmented and, in some cases, replaced by AI-driven approaches. These new methods leverage machine learning to intelligently analyze video content, allowing for more adaptive and efficient compression. We are going to discuss AOM's new codec AV2, its low- and high-level features that enable significantly smaller file sizes with no perceptible loss in quality, a crucial development for streaming and storage. The shift to AI has also transformed how we evaluate video quality. Traditional metrics, while directionally useful, don't always align with human perception, especially for user generated content (UGC). They fail to capture what's most important for machine vision tasks. We will talk about new AI-based quality metrics that are being developed. They correlate better with a human's subjective experience and a machine's ability to perform tasks like object recognition. Along the way, we'll cover large scale industrial infrastructure challenges and the ways to achieve high reliability and accuracy. Balu Adsumilli, Jianle Chen, In Suk Chong, Yilin Wang 0001 |
ACM Multimedia | 1 |
| 2025 | An Innovative Industry Program on Multimedia in A New AI EraabstractThe ACM Multimedia 2025 Industry Program presents a comprehensive overview of how multimodal AI is revolutionizing real-world applications. The program features contributions from over twenty industry leaders, covering a spectrum of domains from content creation and distribution to healthcare, manufacturing, and foundational technology. Keynotes by leaders from NEC and Google DeepMind address critical challenges in business transformation and media integrity, respectively. A dedicated seminar from Google explores advancements in video codecs like AV2 and the development of next-generation quality metrics using Large Language Models (LLMs). The program also includes twelve expert talks and seven demonstrations that showcase practical innovations, such as LLM-driven recommendation systems at Meta, generative AI for industrial optimization by Mitsubishi Electric, and a Siemens-developed protocol for semantic interoperability in manufacturing. These components collectively highlight the profound societal and industrial impact of multimedia research and bridge the gap between academic theory and real-world deployment. Jianquan Liu, Balu Adsumilli, Yukiko Yanagawa, Haiwei Dong 0001 |
ACM Multimedia | 2 |
| 2025 | Study on content-dependency of acceptability/annoyance (AccAnn) scale in User-Generated Content (UGC) videosabstractInternational audience Pierre R. Lebreton, Patrick Le Callet, Neil Birkbeck, Yilin Wang 0001, Zeina Sinno, Balu Adsumilli |
PCS | 6 |
| 2025 | Understanding, detecting, and removing perceptual banding artifacts in compressed videosabstractBanding artifacts, or false contouring, are a common compression impairment that often appears on large smooth regions of encoded videos and images. These staircase-like color bands can be very noticeable and annoying, even on otherwise high-quality videos, especially when displayed on high-definition screens. Yet, relatively little attention has been applied to this problem. Here we study this artifact, by first analyzing the perceptual and encoding aspects of banding artifacts, then propose a new distortion-specific no-reference video quality algorithm for predicting banding artifacts, inspired by perceptual models. The proposed banding detector can generate a pixel-wise banding visibility map, and output overall banding severity scores at both the frame and video levels. Furthermore, we propose a deep learning based approach to improve the overall perceptual quality of compressed videos by joint debanding and compression artifact removal. Our experimental results show that the proposed banding detector delivers better consistency with subjective evaluations, and is able to detect different perceptual severity levels of bands. The debanding experiments also show that the proposed algorithm outperforms recent debanding models both visually and quantitatively. The code is available at https://github.com/google/bband-adaband and https://github.com/vztu/DebandingNet . Zhengzhong Tu, Chia-Ju Chen, Jessie Lin, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
Signal Process. Image Commun. | 6 |
| 2025 | Subjective and Objective Quality Assessment of Banding Artifacts on Compressed VideosabstractAlthough there have been notable advancements in video compression technologies in recent years, banding artifacts remain a serious issue affecting the quality of compressed videos, particularly on smooth regions of high-definition videos. Noticeable banding artifacts can severely impact the perceptual quality of videos viewed on a high-end HDTV or high-resolution screen. Hence, there is a pressing need for a systematic investigation of the banding video quality assessment problem for advanced video codecs. Given that the existing publicly available datasets for studying banding artifacts are limited to still picture data only, which cannot account for temporal banding dynamics, we have created a first-of-a-kind open video dataset, dubbed LIVE-YT-Banding, which consists of 160 videos generated by four different compression parameters using the AV1 video codec. A total of 7,200 subjective opinions are collected from a cohort of 45 human subjects. To demonstrate the value of this new resources, we tested and compared a variety of models that detect banding occurrences, and measure their impact on perceived quality. Among these, we introduce an effective and efficient new no-reference (NR) video quality evaluator which we call CBAND. CBAND leverages the properties of the learned statistics of natural images expressed in the embeddings of deep neural networks. Our experimental results show that the perceptual banding prediction performance of CBAND significantly exceeds that of previous state-of-the-art models, and is also orders of magnitude faster. Moreover, CBAND can be employed as a differentiable loss function to optimize video debanding models. The LIVE-YT-Banding database, code, and pre-trained model are all publically available at https://github.com/uniqzheng/CBAND. Qi Zheng 0004, Li-Heng Chen, Chenlong He, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik, Yibo Fan, Zhengzhong Tu |
IEEE Trans. Image Process. | 6 |
| 2024 | Youtube SFV+HDR Quality DatasetabstractThe popularity of Short form videos (SFV) has grown dramatically in the past few years, and has become a phenomenal video category with billions of viewers. Meanwhile, High Dynamic Range (HDR) as an advanced feature also becomes more and more popular on video sharing platforms. As a hot topic with huge impact, SFV and HDR bring new questions to video quality research: 1) is SFV+HDR quality assessment significantly different from traditional User Generated Content (UGC) quality assessment? 2) do objective quality metrics designed for traditional UGC still work well for SFV+HDR? To answer the above questions, we created the first large scale SFV+HDR dataset with reliable subjective quality scores, covering 10 popular content categories. Further, we also introduce a general sampling framework to maximize the representativeness of the dataset. We provided a comprehensive analysis of subjective quality scores for Short form SDR and HDR videos, and discuss the reliability of state-of-the-art UGC quality metrics and potential improvements. Yilin Wang 0001, Joong Gon Yim, Neil Birkbeck, Balu Adsumilli |
ICIP | 4 |
| 2024 | A Dataset for Understanding Open UGC Video DatasetsabstractUser Generated Content (UGC) video streaming is a major application on the Internet. Even small bitrate savings can have large network impacts at this scale. In order to achieve improvements without sacrificing experience, the quality of UGC videos needs to be better understood. In recent years video quality evaluation models designed for the evaluation of UGC videos have received a lot of attention. However, considering that these models are learning-based models, they heavily depend on the training data that has been used. In this paper, a new dataset is introduced that allows studying the differences in characteristics between existing UGC video datasets. It reveals the range of quality that was covered by existing UGC video datasets, and the implication of these quality ranges on training and validation performance of UGC video quality prediction models. Furthermore, this work demonstrates that dataset alignment enables existing UGC models to achieve higher performance. This alignment dataset can be found openly available on Zenodo (https://zenodo.org/doi/10.5281/zenodo.12155934). Pierre R. Lebreton, Patrick Le Callet, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli |
ICIP | 5 |
| 2024 | Subjective Portrait Region Cropping On Landscape Video StudyabstractWith the rise of mobile video consumption, adapting videos to non-traditional aspect ratios poses challenges for existing content. The use of static cropping and border padding often compromises visual quality, while warping may distort a video’s intended meaning. Here we advocate for a more effective approach - cropping significant regions within video frames in a temporal manner, while minimizing distortion and preserving essential content. However, the lack of a large-scale database devoted to informing these tasks impedes progress in this direction. Addressing this gap, we introduce the LIVE-YouTube Video Cropping (LIVE-YT VC) database, featuring 1800 videos labeled by 90 human subjects. Sourced from the YouTube-UGC and LSVQ databases, this collection serves as the largest subjective video portrait region cropping database. We evaluate our methodology using the SmartVidCrop [1] algorithm, establishing a benchmark for future research. Our contributions offer a crucial resource for advancing video aspect ratio transformation, ensuring that mobile-friendly video content retains its quality and meaning. The details of accessing the dataset have been provided in the supplementary material. Cheng-Han Lee, Maniratnam Mandal, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
ICIP | 5 |
| 2024 | Legit: Text Legibility For User-Generated MediaabstractUser-generated content (UGC) is ubiquitous across the internet as a result of billions of videos and images being uploaded each day. All kinds of UGC media are affected by natural distortions, occurring both during and after capture, which are inherently diverse and commingled. These distortions have different perceptual effects based on the media content. Given recent dramatic increases in the consumption of short-form content, the analysis and control of their perceptual quality has become an important problem. Regardless of the content, many UGC videos have overlaid and embedded texts in them, which are visually salient. Hence text quality has a significant impact on the global perception of video or image quality and needs to be studied. One of the most important factors in perceptual text quality in user-generated media is legibility, which has been studied very little in the context of computer vision. Predicting text legibility can also help in text recognition applications such as image search or document identification. This work aims at modeling text legibility using computer vision techniques and thus studying the relationship between text quality and legibility. We propose a modified dataset variant of COCO-Text [1] and a model for predicting text legibility for both handwritten and machine-generated texts. We also demonstrate how models trained to predict text legibility can help in the prediction of text (perceptual) quality. The dataset and models can be accessed here https://live.ece.utexas.edu/research/Quality/index.htm. Maniratnam Mandal, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 3 |
| 2024 | An Innovative Industry Program in A New Era of Multimedia with Generative AIabstractThe ACM Multimedia 2024 industry program offers a unique platform for fostering collaboration between academia and industry. This year's program features a diverse range of industry keynotes, expert talks, seminars, and demonstrations, showcasing the latest advancements in multimedia technology. Renowned experts from industry and academia will share their insights on topics such as generative AI, automotive design, computer vision, spatial experience, healthcare, and more. Attendees will have the opportunity to network with industry leaders, learn about cutting-edge technologies, and explore potential collaborations. The industry program highlights the growing importance of multimedia technology in various domains and demonstrates the innovative ways in which AI and other emerging technologies are transforming industries. By participating in this program, attendees can gain valuable knowledge, expand their professional networks, and contribute to the advancement of the field. Jianquan Liu, Balu Adsumilli, Yukiko Yanagawa, Haiwei Dong 0001 |
ACM Multimedia | 2 |
| 2024 | Message from the MMSP 2024 General and Technical Program ChairsabstractThe 26th IEEE International Workshop on Multimedia Signal Processing (MMSP 2024), organized by the Multimedia Signal Processing Technical Committee (MMSP-TC) of IEEE Signal Processing Society (SPS), was held at Purdue University, West Lafayette, Indiana, U.S.A., from October 2 - 4, 2024. Fengqing Zhu 0001, Nikolaos Thomos, Balu Adsumilli, Enrico Magli |
MMSP | 4 |
| 2024 | Subjective and Objective Analysis of Streamed Gaming VideosabstractThe rising popularity of online User-Generated-Content (UGC) in the form of streamed and shared videos, has hastened the development of perceptual Video Quality Assessment (VQA) models, which can be used to help optimize their delivery. Gaming videos, which are a relatively new type of UGC videos, are created when skilled and casual gamers post videos of their gameplay. These kinds of screenshots of UGC gameplay videos have become extremely popular on major streaming platforms like YouTube and Twitch. Synthetically-generated gaming content presents challenges to existing VQA algorithms, including those based on natural scene/video statistics models. Synthetically generated gaming content presents different statistical behavior than naturalistic videos. A number of studies have been directed towards understanding the perceptual characteristics of professionally generated gaming videos arising in gaming video streaming, online gaming, and cloud gaming. However, little work has been done on understanding the quality of UGC gaming videos, and how it can be characterized and predicted. Towards boosting the progress of gaming video VQA model development, we conducted a comprehensive study of subjective and objective VQA models on UGC gaming videos. To do this, we created a novel UGC gaming video resource, called the LIVE-YouTube Gaming video quality (LIVE-YT-Gaming) database, comprised of 600 real UGC gaming videos. We conducted a subjective human study on this data, yielding 18,600 human quality ratings recorded by 61 human subjects. We also evaluated a number of state-of-the-art (SOTA) VQA models on the new database, including a new one, called GAME-VQP, based on both natural video statistics and CNN-learned features. To help support work in this field, we are making the new LIVE-YT-Gaming Database, along with code for GAME-VQP, publicly available through the link:https://live.ece.utexas.edu/research/LIVE-YT-Gaming/index.html. Xiangxu Yu, Zhenqiang Ying, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Games | 5 |
| 2023 | Rate-Distortion Optimization with Alternative References for UGC Video CompressionabstractUser generated content (UGC) refers to videos that are uploaded by users and shared over the Internet. UGC may have low quality due to noise and previous compression. When re-encoding UGC for streaming or downloading, a traditional video coding pipeline will perform rate-distortion (RD) optimization to choose coding parameters. However, in the UGC video coding case, since the input is not pristine, quality “saturation” (or even degradation) can be observed, i.e., increased bitrate only leads to improved representation of coding artifacts and noise present in the UGC input. In this paper, we study the saturation problem in UGC compression, where the goal is to identify and avoid during encoding, the coding parameters and rates that lead to quality saturation. We proposed a geometric criterion for saturation detection that works with rate-distortion optimization, and only requires a few frames from the UGC video. In addition, we show how to combine the proposed saturation detection method with existing video coding systems that implement rate-distortion optimization for efficient compression of UGC videos. Eduardo Pavez, Antonio Ortega, Balu Adsumilli |
ICASSP | 4 |
| 2023 | Comparison of HDR quality metrics in Per-Clip Lagrangian multiplier optimisation with AV1abstractThe complexity of modern codecs along with the increased need of delivering high-quality videos at low bitrates has reinforced the idea of a per-clip tailoring of parameters for optimised rate-distortion performance. While the objective quality metrics used for Standard Dynamic Range (SDR) videos have been well studied, the transitioning of consumer displays to support High Dynamic Range (HDR) videos, poses a new challenge to rate-distortion optimisation. In this paper, we review the popular HDR metrics DeltaE100 (DE100), PSNRL100, wPSNR, and HDR-VQM. We measure the impact of employing these metrics in per-clip direct search optimisation of the rate-distortion Lagrange multiplier in AV1. We report, on 35 HDR videos, average Bjontegaard Delta Rate (BD-Rate) gains of 4.675%, 2.226%, and 7.253% in terms of DE100, PSNRL100, and HDR-VQM. We also show that the inclusion of chroma in the quality metrics has a significant impact on optimisation, which can only be partially addressed by the use of chroma offsets. Vibhoothi, François Pitié, Angeliki V. Katsenou, Yeping Su, Balu Adsumilli, Anil C. Kokaram |
ICME | 5 |
| 2023 | CONVIQT: Contrastive Video Quality EstimatorabstractPerceptual video quality assessment (VQA) is an integral component of many streaming and video sharing platforms. Here we consider the problem of learning perceptually relevant video quality representations in a self-supervised manner. Distortion type identification and degradation level determination is employed as an auxiliary task to train a deep learning model containing a deep Convolutional Neural Network (CNN) that extracts spatial features, as well as a recurrent unit that captures temporal information. The model is trained using a contrastive loss and we therefore refer to this training framework and resulting model as CONtrastive VIdeo Quality EstimaTor (CONVIQT). During testing, the weights of the trained model are frozen, and a linear regressor maps the learned features to quality scores in a no-reference (NR) setting. We conduct comprehensive evaluations of the proposed model against leading algorithms on multiple VQA databases containing wide ranges of spatial and temporal distortions. We analyze the correlations between model predictions and ground-truth quality ratings, and show that CONVIQT achieves competitive performance when compared to state-of-the-art NR-VQA models, even though it is not trained on those databases. Our ablation experiments demonstrate that the learned representations are highly robust and generalize well across synthetic and realistic distortions. Our results indicate that compelling representations with perceptual bearing can be obtained using self-supervised learning. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2022 | An Empirical Approach for Optimising the Impact of a Preprocessor in a Transcoding PipelineabstractThe volume of User Generated Content (UGC) on the internet has exploded throughout the pandemic. The relatively low quality of that content generally implies an increased bitrate and much reduced quality after transcoding. Preprocessing e.g. using a noise reducer, is one approach for reducing bitrate and increasing quality. The impact of the noise reducer is however affected by the target bitrate of the encoder. That relationship is known but not previously quantitatively examined. In this paper we present a methodology and new metric for measuring this impact based on the Rate-Distortion curves before and after pre-processing. The metric is used as a cost function for estimating the optimal filter parameter for our chosen denoiser. Our experiments show that optimising the filter parameter in this way yields as much as 4-5dB improvement in PSNR at 3 Mbps. Varoun Hanooman, Anil C. Kokaram, Yeping Su, Neil Birkbeck, Balu Adsumilli |
ICIP | 5 |
| 2022 | Compression of User Generated Content Using Denoised ReferencesabstractVideo shared over the internet is commonly referred to as user generated content (UGC). UGC video may have low quality due to various factors including previous compression. UGC video is uploaded by users, and then it is re-encoded to be made available at various levels of quality. In a traditional video coding pipeline the encoder parameters are optimized to minimize a rate-distortion criterion, but when the input signal has low quality, this results in sub-optimal coding parameters optimized to preserve undesirable artifacts. In this paper we formulate the UGC compression problem as that of compression of a noisy/corrupted source. The noisy source coding theorem reveals that an optimal UGC compression system is comprised of optimal denoising of the UGC signal, followed by compression of the denoised signal. Since optimal denoising is unattainable and users may be against modification of their content, we propose encoding the UGC signal, and using denoised references only to compute distortion, so the encoding process can be guided towards perceptually better solutions. We demonstrate the effectiveness of the proposed strategy for JPEG compression of UGC images and videos. Eduardo Pavez, Enrique Perez, Antonio Ortega, Balu Adsumilli |
ICIP | 5 |
| 2022 | When is the Cleaning of Subjective Data Relevant to Train UGC Video Quality Metrics?abstractOutlier analysis and spammer detection recently gained momentum in order to reduce uncertainty of subjective ratings in image & video quality assessment tasks. The large proportion of unreliable ratings from online crowdsourcing experiments and the need for qualitative and quantitative large-scale studies in the deep-learning ecosystem played a role in this event. We study the effect that data cleaning has on trainable models predicting the visual quality for videos, and present results demonstrating when cleaning is necessary to reach higher efficiency. To this end, we present and analyze a benchmark on clean and noisy User Generated Content (UGC) large-scale datasets on which we re-trained models, followed by an empirical exploration of the constraint of data removal. Our results show that a dataset presenting between 7 and 30% of outliers benefits from cleaning before training. Anne-Flore Perrin, Charles Dormeval, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Patrick Le Callet |
ICIP | 5 |
| 2022 | Revisiting the Efficiency of UGC Video Quality AssessmentabstractUGC video quality assessment (UGC-VQA) is a challenging research topic due to the high video diversity and limited public UGC quality datasets. State-of-the-art (SOTA) UGC quality models tend to use high complexity models, and rarely discuss the trade-off among complexity, accuracy, and generalizability. We propose a new perspective on UGC-VQA, and show that model complexity may not be critical to the performance, whereas a more diverse dataset is essential to train a better model. We illustrate this by using a light weight model, UVQ-lite, which has higher efficiency and better generalizability (less overfitting) than baseline SOTA models. We also propose a new way to analyze the sufficiency of the training set, by leveraging UVQ’s comprehensive features. Our results motivate a new perspective about the future of UGC-VQA research, which we believe is headed toward more efficient models and more diverse datasets. Yilin Wang 0001, Joong Gon Yim, Neil Birkbeck, Junjie Ke, Hossein Talebi, Feng Yang 0008, Balu Adsumilli |
ICIP | 8 |
| 2022 | A Deep Learning post-processor with a perceptual loss function for video compression artifact removalabstractWhile video compression is necessary for large scale video streaming services, compression at low bitrate can degrade the original video and negatively affect the end user’s quality of experience. Deep Neural Networks (DNNs) are actively researched with respect to artifact removal, however the loss functions that are typically employed follows a derivation of a pixel-wise Lpnorm. In this paper we consider a DNN as a post-processor for video compression artifact removal. The DNN is trained using a composite perceptual loss that combines a traditional Lpnorm loss and a VMAF proxy network based on the Video Multimethod Assessment Function (VMAF). Results show an improvement in VMAF score over both the training and testing sets. Darren Ramsook, Anil C. Kokaram, Neil Birkbeck, Yeping Su, Balu Adsumilli |
PCS | 5 |
| 2022 | Making Video Quality Assessment Models Sensitive to Frame Rate DistortionsabstractWe consider the problem of capturing distortions arising from changes in frame rate as part of Video Quality Assessment (VQA). Variable frame rate (VFR) videos have become much more common, and streamed videos commonly range from 30 frames per second (fps) up to 120 fps. VFR-VQA offers unique challenges in terms of distortion types as well as in making non-uniform comparisons of reference and distorted videos having different frame rates. The majority of current VQA models require compared videos to be of the same frame rate, but are unable to adequately account for frame rate artifacts. The recently proposed Generalized Entropic Difference (GREED) VQA model succeeds at this task, using natural video statistics models of entropic differences of temporal band-pass coefficients, delivering superior performance on predicting video quality changes arising from frame rate distortions. Here we propose a simple fusion framework, whereby temporal features from GREED are combined with existing VQA models, towards improving model sensitivity towards frame rate distortions. We find through extensive experiments that this feature fusion significantly boosts model performance on both HFR/VFR datasets as well as fixed frame rate (FFR) VQA databases. Our results suggest that employing efficient temporal representations can result much more robust and accurate VQA models when frame rate variations can occur. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Signal Process. Lett. | 4 |
| 2022 | Image Quality Assessment Using Contrastive LearningabstractWe consider the problem of obtaining image quality representations in a self-supervised manner. We use prediction of distortion type and degree as an auxiliary task to learn features from an unlabeled image dataset containing a mixture of synthetic and realistic distortions. We then train a deep Convolutional Neural Network (CNN) using a contrastive pairwise objective to solve the auxiliary problem. We refer to the proposed training framework and resulting deep IQA model as the CONTRastive Image QUality Evaluator (CONTRIQUE). During evaluation, the CNN weights are frozen and a linear regressor maps the learned representations to quality scores in a No-Reference (NR) setting. We show through extensive experiments that CONTRIQUE achieves competitive performance when compared to state-of-the-art NR image quality models, even without any additional fine-tuning of the CNN backbone. The learned representations are highly robust and generalize well across images afflicted by either synthetic or authentic distortions. Our results suggest that powerful quality representations with perceptual relevance can be obtained without requiring large labeled subjective image quality datasets. The implementations used in this paper are available at https://github.com/pavancm/CONTRIQUE. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2021 | Rich Features for Perceptual Quality Assessment of UGC VideosabstractVideo quality assessment for User Generated Content (UGC) is an important topic in both industry and academia. Most existing methods only focus on one aspect of the perceptual quality assessment, such as technical quality or compression artifacts. In this paper, we create a large scale dataset to comprehensively investigate characteristics of generic UGC video quality. Besides the subjective ratings and content labels of the dataset, we also propose a DNN-based framework to thoroughly analyze importance of content, technical quality, and compression level in perceptual quality. Our model is able to provide quality scores as well as human-friendly quality indicators, to bridge the gap between low level video signals to human perceptual quality. Experimental results show that our model achieves state-of-the-art correlation with Mean Opinion Scores (MOS). Yilin Wang 0001, Junjie Ke, Hossein Talebi, Joong Gon Yim, Neil Birkbeck, Balu Adsumilli, Peyman Milanfar, Feng Yang 0008 |
CVPR | 6 |
| 2021 | Regression or classification? New methods to evaluate no-reference picture and video quality modelsabstractVideo and image quality assessment has long been projected as a regression problem, which requires predicting a continuous quality score given an input stimulus. However, recent efforts have shown that accurate quality score regression on real-world user-generated content (UGC) is a very challenging task. To make the problem more tractable, we propose two new methods - binary, and ordinal classification - as alternatives to evaluate and compare no-reference quality models at coarser levels. Moreover, the proposed new tasks convey more practical meaning on perceptually optimized UGC transcoding, or for preprocessing on media processing platforms. We conduct a comprehensive benchmark experiment of popular no-reference quality models on recent in-the-wild picture and video quality datasets, providing reliable baselines for both evaluation methods to support further studies. We hope this work promotes coarse-grained perceptual modeling and its applications to efficient UGC processing. Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICASSP | 6 |
| 2021 | Video Quality Assessment of User Generated Content: A Benchmark Study and a New ModelabstractRecent years have witnessed an explosion of user-generated content (UGC) shared and streamed over the Internet. Accordingly, there is a great need for accurate video quality assessment (VQA) models for consumer or UGC videos to monitor, control, and optimize this vast content. Here we contribute to advancing the UGC-VQA problem by conducting a comprehensive evaluation of leading blind VQA (BVQA) models. Besides, we also created a new fusion-based BVQA model, which we dub the VIDeo quality EVALuator (VIDEVAL), that effectively balances the trade-off between performance and efficiency. Our experimental results show that VIDEVAL achieves state-of-the-art performance at a lower computational cost. We believe our reliable and reproducible benchmark will facilitate further research on deep learning-based BVQA modeling. An implementation of VIDEVAL has been made available online1.1https://github.com/vztu/VIDEVAL_release Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 5 |
| 2021 | A Temporal Statistics Model For UGC Video Quality PredictionabstractBlind video quality assessment of user-generated content (UGC) has become a trending and challenging problem. Previous studies have shown the efficacy of natural scene statistics for capturing spatial distortions. The exploration of temporal video statistics on UGC, however, is relatively limited. Here we propose the first general, effective and efficient temporal statistics model accounting for temporal- or motion-related distortions for UGC video quality assessment, by analyzing regularities in the temporal bandpass domain. The proposed temporal model can serve as a plug-in module to boost existing no-reference video quality predictors that lack motion-relevant features. Our experimental results on recent large-scale UGC video databases show that the proposed model can significantly improve the performances of existing methods, at a very reasonable computational expense. Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 5 |
| 2021 | High Frame Rate Video Quality Assessment using VMAF and Entropic DifferencesabstractThe popularity of streaming videos with live, high-action content has led to an increased interest in High Frame Rate (HFR) videos. In this work we address the problem of frame rate dependent Video Quality Assessment (VQA) when the videos to be compared have different frame rate and compression factor. The current VQA models such as VMAF have superior correlation with perceptual judgments when videos to be compared have same frame rates and contain conventional distortions such as compression, scaling etc. However this framework requires additional pre-processing step when videos with different frame rates need to be compared, which can potentially limit its overall performance. Recently, Generalized Entropic Difference (GREED) VQA model was proposed to account for artifacts that arise due to changes in frame rate, and showed superior performance on the LIVE-YT-HFR database which contains frame rate dependent artifacts such as judder, strobing etc. In this paper we propose a simple extension, where the features from VMAF and GREED are fused in order to exploit the advantages of both models. We show through various experiments that the proposed fusion framework results in more efficient features for predicting frame rate dependent video quality. We also evaluate the fused feature set on standard non-HFR VQA databases and obtain superior performance than both GREED and VMAF, indicating the combined feature set captures complimentary perceptual quality information. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
PCS | 4 |
| 2021 | A differentiable estimator of VMAF for VideoabstractModern Perceptual Visual Quality Metrics (PVQMs) for video are generally complex and non-differentiable. This makes them difficult to use as loss functions in restoration and compression tuning. Traditional metrics such as PSNR/MSE which are differentiable remain important but do not capture perceptual visual criteria. In this paper we present a DNN which models a popular perceptual video metric VMAF. In so doing, we introduce a differentiable loss function that closely matches the behaviour of a perceptual metric. Employing degradation generated with H.265 compression, our model achieves a 4.41% RMSE in predicting VMAF. This can now be deployed as a video based loss function in video enhancement and compression tasks. Darren Ramsook, Anil C. Kokaram, Noel E. O'Connor, Neil Birkbeck, Yeping Su, Balu Adsumilli |
PCS | 6 |
| 2021 | Efficient User-Generated Video Quality PredictionabstractBlind video quality assessment of user-generated content (UGC) has become a trending, challenging, unsolved problem. Accurate and efficient video quality predictors suitable for this content are thus in great demand to achieve intelligent analysis and processing of UGC videos. However, previous video quality models are either incapable or inefficient for predicting the quality of complex, diverse UGC videos in practical applications. Here we introduce an effective and efficient video quality model for UGC content, which we dub the Rapid and Accurate Video Quality Evaluator (RAPIQUE), which we show performs comparably to state-of-the-art models but with orders-of-magnitude faster runtime. Our experimental results on recent large-scale UGC video quality databases show that RAPIQUE delivers top performances on all datasets at a considerably lower computational expense. An implementation of RAPIQUE is online: https://github.com/vztu/RAPIQUE. Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
PCS | 5 |
| 2021 | Block-based Learned Image Coding with Convolutional Autoencoder and Intra-Prediction Aided Entropy CodingabstractRecent works on learned image coding using autoencoder models have achieved promising results in rate-distortion performance. Typically, an autoencoder is used to transform an image into a latent tensor, which is then quantized and entropy coded. Based on a work by Ballé et al., we adapted the autoencoder with a hyperprior model to code images in a block-based approach. When the autoencoder model is directly applied to code small image blocks, spatial redundancy in the larger image cannot be fully utilized, resulting in a decrease in ratedistortion performance. We propose a method to utilize border information in the entropy coding of latent and hyper-latent tensors, which has achieved promising results. We show that using intra-prediction to help entropy coding is more effective than applying a convolutional autoencoder with hyper priors to intra-prediction residual blocks. Zhongzheng Yuan, Debargha Mukherjee, Balu Adsumilli, Yao Wang 0001 |
PCS | 4 |
| 2021 | ST-GREED: Space-Time Generalized Entropic Differences for Frame Rate Dependent Video Quality PredictionabstractWe consider the problem of conducting frame rate dependent video quality assessment (VQA) on videos of diverse frame rates, including high frame rate (HFR) videos. More generally, we study how perceptual quality is affected by frame rate, and how frame rate and compression combine to affect perceived quality. We devise an objective VQA model called Space-Time GeneRalized Entropic Difference (GREED) which analyzes the statistics of spatial and temporal band-pass video coefficients. A generalized Gaussian distribution (GGD) is used to model band-pass responses, while entropy variations between reference and distorted videos under the GGD model are used to capture video quality variations arising from frame rate changes. The entropic differences are calculated across multiple temporal and spatial subbands, and merged using a learned regressor. We show through extensive experiments that GREED achieves state-of-the-art performance on the LIVE-YT-HFR Database when compared with existing VQA models. The features used in GREED are highly generalizable and obtain competitive performance even on standard, non-HFR VQA databases. The implementation of GREED has been made available online: https://github.com/pavancm/GREED. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2021 | UGC-VQA: Benchmarking Blind Video Quality Assessment for User Generated ContentabstractRecent years have witnessed an explosion of user-generated content (UGC) videos shared and streamed over the Internet, thanks to the evolution of affordable and reliable consumer capture devices, and the tremendous popularity of social media platforms. Accordingly, there is a great need for accurate video quality assessment (VQA) models for UGC/consumer videos to monitor, control, and optimize this vast content. Blind quality prediction of in-the-wild videos is quite challenging, since the quality degradations of UGC videos are unpredictable, complicated, and often commingled. Here we contribute to advancing the UGC-VQA problem by conducting a comprehensive evaluation of leading no-reference/blind VQA (BVQA) features and models on a fixed evaluation architecture, yielding new empirical insights on both subjective video quality studies and objective VQA model design. By employing a feature selection strategy on top of efficient BVQA models, we are able to extract 60 out of 763 statistical features used in existing methods to create a new fusion-based model, which we dub the VIDeo quality EVALuator (VIDEVAL), that effectively balances the trade-off between VQA performance and efficiency. Our experimental results show that VIDEVAL achieves state-of-the-art performance at considerably lower computational cost than other leading models. Our study protocol also defines a reliable benchmark for the UGC-VQA problem, which we believe will facilitate further research on deep learning-based VQA modeling, as well as perceptually-optimized efficient UGC video processing, transcoding, and streaming. To promote reproducible research and public evaluation, an implementation of VIDEVAL has been made available online: https://github.com/vztu/VIDEVAL. Zhengzhong Tu, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2021 | Predicting the Quality of Compressed Videos With Pre-Existing DistortionsabstractBecause of the increasing ease of video capture, many millions of consumers create and upload large volumes of User-Generated-Content (UGC) videos to social and streaming media sites over the Internet. UGC videos are commonly captured by naive users having limited skills and imperfect techniques, and tend to be afflicted by mixtures of highly diverse in-capture distortions. These UGC videos are then often uploaded for sharing onto cloud servers, where they are further compressed for storage and transmission. Our paper tackles the highly practical problem of predicting the quality of compressed videos (perhaps during the process of compression, to help guide it), with only (possibly severely) distorted UGC videos as references. To address this problem, we have developed a novel Video Quality Assessment (VQA) framework that we call 1stepVQA (to distinguish it from two-step methods that we discuss). 1stepVQA overcomes limitations of Full-Reference, Reduced-Reference and No-Reference VQA models by exploiting the statistical regularities of both natural videos and distorted videos. We also describe a new dedicated video database, which was created by applying a realistic VMAF-Guided perceptual rate distortion optimization (RDO) criterion to create realistically compressed versions of UGC source videos, which typically have pre-existing distortions. We show that 1stepVQA is able to more accurately predict the quality of compressed videos, given imperfect reference videos, and outperforms other VQA models in this scenario. Xiangxu Yu, Neil Birkbeck, Yilin Wang 0001, Christos G. Bampis, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 5 |
| 2020 | BBAND INDEX: A NO-REFERENCE BANDING ARTIFACT PREDICTORabstractBanding artifact, or false contouring, is a common video compression impairment that tends to appear on large flat regions in encoded videos. These staircase-shaped color bands can be very noticeable in high-definition videos. Here we study this artifact, and propose a new distortion-specific no-reference video quality model for predicting banding artifacts, called the Blind BANding Detector (BBAND index). BBAND is inspired by human visual models. The proposed detector can generate a pixel-wise banding visibility map and output a banding severity score at both the frame and video levels. Experimental results show that our proposed method outperforms state-of-the-art banding detection algorithms and delivers better consistency with subjective evaluations. Zhengzhong Tu, Jessie Lin, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
ICASSP | 4 |
| 2020 | Rate Distortion Optimization Over Large Scale Video Corpus With Machine LearningabstractWe present an efficient codec-agnostic method for bitrate allocation over a large scale video corpus with the goal of minimizing the average bitrate subject to constraints on average and minimum quality. Our method clusters the videos in the corpus such that videos within one cluster have similar rate-distortion (R-D) characteristics. We train a support vector machine classifier to predict the R-D cluster of a video using simple video complexity features that are computationally easy to obtain. The model allows us to classify a large sample of the corpus in order to estimate the distribution of the number of videos in each of the clusters. We use this distribution to find the optimal encoder operating point for each R-D cluster. Experiments with AV1 encoder show that our method can achieve the same average quality over the corpus with 22% less average bitrate. Sam John, Akshay Gadde, Balu Adsumilli |
ICIP | 3 |
| 2020 | A Comparative Evaluation Of Temporal Pooling Methods For Blind Video Quality AssessmentabstractMany objective video quality assessment (VQA) algorithms include a key step of temporal pooling of frame-level quality scores. However, less attention has been paid to studying the relative efficiencies of different pooling methods on noreference (blind) VQA. Here we conduct a large-scale comparative evaluation to assess the capabilities and limitations of multiple temporal pooling strategies on blind VQA of usergenerated videos. The study yields insights and general guidance regarding the application and selection of temporal pooling models. In addition, we also propose an ensemble pooling model built on top of high-performing temporal pooling models. Our experimental results demonstrate the relative efficacies of the evaluated temporal pooling models, using several popular VQA algorithms evaluated on two recent largescale natural video quality databases. Conclusively, we also provide an empirical recipe for applying temporal pooling of frame-based quality predictions. Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 5 |
| 2020 | Subjective Quality Assessment For Youtube Ugc DatasetabstractDue to the scale of social video sharing, User Generated Content (UGC) is getting more attention from academia and industry. To facilitate compression-related research on UGC, YouTube has released a large-scale dataset [1]. The initial dataset only provided videos, limiting its use in quality assessment. We used a crowd-sourcing platform to collect subjective quality scores for this dataset. We analyzed the distribution of Mean Opinion Score (MOS) in various dimensions, and investigated some fundamental questions in video quality assessment, like the correlation between full video MOS and corresponding chunk MOS, and the influence of chunk variation in quality score aggregation. Joong Gon Yim, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli |
ICIP | 4 |
| 2020 | A Viewport-Driven Multi-Metric Fusion Approach for 360-Degree Video Quality AssessmentabstractWe propose a new viewport-based multi-metric fusion (MMF) approach for visual quality assessment of 360-degree (omnidirectional) videos. Our method is based on computing multiple spatio-temporal objective quality metrics (features) on viewports extracted from 360-degree videos, and learning a model that combines these features into a metric, which closely matches subjective quality scores. The main motivations for the proposed method are that: 1) quality metrics computed on viewports better captures the user experience than metrics computed on the projection domain; 2) no individual objective image quality metric always performs best for all types of visual distortions, while a learned combination of them is able to adapt to different conditions and produce better results overall. Experimental results, based on the largest available 360-degree videos quality dataset, demonstrate that the proposed metric outperforms state-of-the-art 360-degree and 2D video quality metrics. Roberto Gerson De Albuquerque Azevedo, Neil Birkbeck, Ivan Janatra, Balu Adsumilli, Pascal Frossard |
ICME | 4 |
| 2020 | Translation of Perceived Video Quality Across DisplaysabstractDisplay devices can affect the perceived quality of a video significantly. In this paper, we focus on the scenario where video resolution does not exceed screen resolution, and investigate the relationship of perceived video quality on mobile, laptop and TV. A novel transformation of Mean Opinion Scores (MOS) among different devices is proposed and is shown to be effective at normalizing ratings across user devices for in lab and crowd sourced subjective studies. The model allows us to perform more focused in lab subjective studies as we can reduce the number of test devices and helps us reduce noise during crowd-sourcing subjective video quality tests. It is also more effective than utilizing existing device dependent objective metrics for translating MOS ratings across devices. Jessie Lin, Neil Birkbeck, Balu Adsumilli |
MMSP | 3 |
| 2020 | Capturing Video Frame Rate Variations via Entropic DifferencingabstractHigh frame rate videos are increasingly getting popular in recent years, driven by the strong requirements of the entertainment and streaming industries to provide high quality of experiences to consumers. To achieve the best trade-offs between the bandwidth requirements and video quality in terms of frame rate adaptation, it is imperative to understand the effects of frame rate on video quality. In this direction, we devise a novel statistical entropic differencing method based on a Generalized Gaussian Distribution model expressed in the spatial and temporal band-pass domains, which measures the difference in quality between reference and distorted videos. The proposed design is highly generalizable and can be employed when the reference and distorted sequences have different frame rates. Our proposed model correlates very well with subjective scores in the recently proposed LIVE-YT-HFR database and achieves state of the art performance when compared with existing methodologies. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Signal Process. Lett. | 4 |
| 2020 | Adaptive Debanding FilterabstractBanding artifacts, which manifest as staircase-like color bands on pictures or video frames, is a common distortion caused by compression of low-textured smooth regions. These false contours can be very noticeable even on high-quality videos, especially when displayed on high-definition screens. Yet, relatively little attention has been applied to this problem. Here we consider banding artifact removal as a visual enhancement problem, and accordingly, we solve it by applying a form of content-adaptive smoothing filtering followed by dithered quantization, as a post-processing module. The proposed debanding filter is able to adaptively smooth banded regions while preserving image edges and details, yielding perceptually enhanced gradient rendering with limited bit-depths. Experimental results show that our proposed debanding filter outperforms state-of-the-art false contour removing algorithms both visually and quantitatively. Zhengzhong Tu, Jessie Lin, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Signal Process. Lett. | 4 |
| 2020 | Visual Distortions in 360° VideosabstractOmnidirectional (or 360°) images and videos are emergent signals being used in many areas, such as robotics and virtual/augmented reality. In particular, for virtual reality applications, they allow an immersive experience in which the user can interactively navigate through a scene with three degrees of freedom, wearing a head-mounted display. Current approaches for capturing, processing, delivering, and displaying 360° content, however, present many open technical challenges and introduce several types of distortions in the visual signal. Some of the distortions are specific to the nature of 360° images and often differ from those encountered in classical visual communication frameworks. This paper provides a first comprehensive review of the most common visual distortions that alter 360° signals going through the different processing elements of the visual communication pipeline. While their impact on viewers' visual perception and the immersive experience at large is still unknown-thus, it is an open research topic-this review serves the purpose of proposing a taxonomy of the visual distortions that can be encountered in 360° signals. Their underlying causes in the end-to-end 360° content distribution pipeline are identified. This taxonomy is essential as a basis for comparing different processing techniques, such as visual enhancement, encoding, and streaming strategies, and allowing the effective design of new algorithms and applications. It is also a useful resource for the design of psycho-visual studies aiming to characterize human perception of 360° content in interactive and immersive applications. Roberto Gerson De Albuquerque Azevedo, Neil Birkbeck, Francesca De Simone, Ivan Janatra, Balu Adsumilli, Pascal Frossard |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Mutual Noise Estimation Algorithm for Video DenoisingabstractThis paper presents a novel algorithm to estimate spatio-temporal noise variance in videos. The algorithm uses mutual information from the spatial and temporal noise statistics to detect homogeneous blocks in a video frame and estimate the noise. The experimental results show the accuracy and robustness of the algorithm over videos with low to high spatial and temporal complexities, and varying levels of noise power and noise correlations. Mohammad Izadi, Neil Birkbeck, Balu Adsumilli |
ICIP | 3 |
| 2019 | YouTube UGC Dataset for Video Compression ResearchabstractNon-professional video, commonly known as User Generated Content (UGC) has become very popular in today's video sharing applications. However, traditional metrics used in compression and quality assessment, like BD-Rate and PSNR, are designed for pristine originals. Thus, their accuracy drops significantly when being applied on non-pristine originals (the majority of UGC). Understanding difficulties for compression and quality assessment in the scenario of UGC is important, but there are few public UGC datasets available for research. This paper introduces a large scale UGC dataset (1500 20 sec video clips) sampled from millions of YouTube videos. The dataset covers popular categories like Gaming, Sports, and new features like High Dynamic Range (HDR). Besides a novel sampling method based on features extracted from encoding, challenges for UGC compression and quality evaluation are also discussed. Shortcomings of traditional reference-based metrics on UGC are addressed. We demonstrate a promising way to evaluate UGC quality by no-reference objective quality metrics, and evaluate the current dataset with three no-reference metrics (Noise, Banding, and SLEEQ). Yilin Wang 0001, Sasi Inguva, Balu Adsumilli |
MMSP | 3 |
| 2017 | Deformable block-based motion estimation in omnidirectional image sequencesabstractThis paper presents an extension of block-based motion estimation for omnidirectional videos, based on a translational object motion model that accounts for the spherical geometry of the imaging system. We use this model to design a new algorithm to perform block matching in sequences of panoramic frames that are the result of the equirectangular projection. Experimental results demonstrate that significant gains can be achieved with respect to the classical exhaustive block matching algorithm in terms of accuracy of motion prediction. In particular, average quality improvements up to approximately 6 dB in terms of Peak Signal to Noise Ratio (PSNR), 0.043 in terms of Structural SIMilarity index (SSIM), and 2 dB in terms of spherical PSNR, can be achieved on the predicted frames. Francesca De Simone, Pascal Frossard, Neil Birkbeck, Balu Adsumilli |
MMSP | 4 |