VLDB 2026 Research / reviewers in the wild / expert
Neil Birkbeck
dblp:79/494
· DBLP profile ↗
55ranked-venue papers
6as first author
30since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 5 first-author · 29 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-authorSystems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BrightRate: Quality Assessment for User-Generated HDR VideosabstractHigh Dynamic Range (HDR) videos offer superior luminance and color fidelity as compared to Standard Dynamic Range (SDR) content. The rapid growth of User-Generated Content (UGC) on platforms such as YouTube, Instagram, and TikTok has brought a significant increase in the volumes of streamed and shared UGC videos. This newer category of videos brings new challenges to the development of effective No-Reference (NR) video quality assessment (VQA) models specialized to HDR UGC, because of the extreme variety and severities of distortions, arising from diverse capture, editing, and processing outcomes. Towards addressing this issue, we introduce BrightVQ, a sizeable new psychometric data resource. It is the first large-scale subjective video quality database dedicated to the quality modelling of HDR UGC videos. BrightVQ comprises 2,100 videos, on which we collected 73,794 perceptual quality ratings. Using this dataset, we also developed BrightRate, a novel video quality prediction model designed to capture both UGC-specific distortions coexisting with HDR-specific artifacts. Extensive experimental results demonstrate that BrightRate achieves state-of-the-art performance across HDR databases. Shreshth Saini, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
WACV | 4 |
| 2026 | Quality Prediction of Embedded and Overlaid Text in User-Generated Visual ContentabstractUser-generated visual content (UGC) now occupies a significant fraction of internet traffic, and billions of UGC videos and pictures are uploaded daily. Among these, short-form video content now accounts for most of the videos consumed by online users. Given the popularity of short-form UGC content, being able to control the perceptual quality of UGC videos has emerged as an important problem. Visual UGC is subject to myriad types, severity, and combinations of distortions. While UGC video quality has been closely studied, the quality and legibility of text that is overlaid or embedded in short-form UGC videos has received relatively low attention. However, being able to accurately predict text quality in images is important, since it both impacts the overall perception of the content it is embedded in, as well as the messages being conveyed. It is also beneficial for applications involving image or video text recognition which can affect visual search and content identification. Analyzing the quality of text embedded in pictures or videos is a hard problem, since perception of it is commingled with the surrounding visual content. Our work, which greatly extends our early report on text legibility prediction, contributes to both the psychophysics of embedded text quality as well as to computational models of its perception. We have created two subjective datasets-designated as the LIVE-COCO Text Legibility (LIVE-COCO-TL) Database (a modification of COCO-Text), and the LIVE-YouTube Text-in-Video Quality (LIVE-YT-TVQ) Database. LIVE-COCO-TL contains 74,440 text patches with legibility annotations, while LIVE-YT-TVQ contains $\sim ~19$ K subjective quality ratings on 405 videos and 641 text patches extracted from them. We build models that predict embedded or overlaid text legibility and text quality, as well as a multi-task model that simultaneously predicts the overall quality of videos with embedded or overlaid and local text quality. We are making the databases and all models freely available at https://live.ece.utexas.edu/research/LIVE_YouTube_Text_Quality_Assessment/index.html. Maniratnam Mandal, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2025 | An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLMabstractThe rise of short-form videos, characterized by diverse content, editing styles, and artifacts, poses substantial challenges for learning-based blind video quality assessment (BVQA) models. Multimodal large language models (MLLMs), renowned for their superior generalization capabilities, present a promising solution. This paper focuses on effectively leveraging a pretrained MLLM for short-form video quality assessment, regarding the impacts of pre-processing and response variability, and insights on combining the MLLM with BVQA models. We first investigated how frame pre-processing and sampling techniques influence the MLLM’s performance. Then, we introduced a lightweight learning-based ensemble method that adaptively integrates predictions from the MLLM and state-of-the-art BVQA models. Our results demonstrated superior generalization performance with the proposed ensemble approach. Furthermore, the analysis of content-aware ensemble weights highlighted that some video characteristics are not fully represented by existing BVQA models, revealing potential directions to improve BVQA models further. Wen Wen 0007, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli |
ICASSP | 3 |
| 2025 | Rate-Distortion Optimization with Non-Reference Metrics for UGC CompressionabstractService providers must encode a large volume of noisy videos to meet the demand for user-generated content (UGC) in online video-sharing platforms. However, low-quality UGC challenges conventional codecs based on rate-distortion optimization (RDO) with full-reference metrics (FRMs). While effective for pristine videos, FRMs drive codecs to preserve artifacts when the input is degraded, resulting in suboptimal compression. A more suitable approach used to assess UGC quality is based on non-reference metrics (NRMs). However, RDO with NRMs as a measure of distortion requires an iterative workflow of encoding, decoding, and metric evaluation, which is computationally impractical. This paper overcomes this limitation by linearizing the NRM around the uncompressed video. The resulting cost function enables block-wise bit allocation in the transform domain by estimating the alignment of the quantization error with the gradient of the NRM. To avoid large deviations from the input, we add sum of squared errors (SSE) regularization. We derive expressions for both the SSE regularization parameter and the Lagrangian, akin to the relationship used for SSE-RDO. Experiments with images and videos show bitrate savings of more than 30% over SSE-RDO using the target NRM, with no decoder complexity overhead and minimal encoder complexity increase. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega, Neil Birkbeck, Balu Adsumilli |
ICIP | 5 |
| 2025 | CHUG: Crowdsourced User-Generated HDR Video Quality DatasetabstractHigh Dynamic Range (HDR) videos enhance visual experiences with superior brightness, contrast, and color depth. The surge of User-Generated Content (UGC) on platforms like YouTube and TikTok introduces unique challenges for HDR video quality assessment (VQA) due to diverse capture conditions, editing artifacts, and compression distortions. Existing HDR-VQA datasets primarily focus on professionally generated content (PGC), leaving a gap in understanding real-world UGC-HDR degradations. To address this, we introduce CHUG: Crowdsourced User-Generated HDR Video Quality Dataset, the first large-scale subjective study on UGC-HDR quality. CHUG comprises 856 UGC-HDR source videos, transcoded across multiple resolutions and bitrates to simulate real-world scenarios, totaling 5,992 videos. A large-scale study via Amazon Mechanical Turk collected 211,848 perceptual ratings. CHUG provides a benchmark for analyzing UGC-specific distortions in HDR videos. We anticipate CHUG will advance No-Reference (NR) HDR-VQA research by offering a large-scale, diverse, and real-world UGC dataset. The dataset is publicly available at: https://shreshthsaini.github.io/CHUG/. Shreshth Saini, Alan C. Bovik, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli |
ICIP | 3 |
| 2025 | Study on content-dependency of acceptability/annoyance (AccAnn) scale in User-Generated Content (UGC) videosabstractInternational audience Pierre R. Lebreton, Patrick Le Callet, Neil Birkbeck, Yilin Wang 0001, Zeina Sinno, Balu Adsumilli |
PCS | 3 |
| 2025 | Understanding, detecting, and removing perceptual banding artifacts in compressed videosabstractBanding artifacts, or false contouring, are a common compression impairment that often appears on large smooth regions of encoded videos and images. These staircase-like color bands can be very noticeable and annoying, even on otherwise high-quality videos, especially when displayed on high-definition screens. Yet, relatively little attention has been applied to this problem. Here we study this artifact, by first analyzing the perceptual and encoding aspects of banding artifacts, then propose a new distortion-specific no-reference video quality algorithm for predicting banding artifacts, inspired by perceptual models. The proposed banding detector can generate a pixel-wise banding visibility map, and output overall banding severity scores at both the frame and video levels. Furthermore, we propose a deep learning based approach to improve the overall perceptual quality of compressed videos by joint debanding and compression artifact removal. Our experimental results show that the proposed banding detector delivers better consistency with subjective evaluations, and is able to detect different perceptual severity levels of bands. The debanding experiments also show that the proposed algorithm outperforms recent debanding models both visually and quantitatively. The code is available at https://github.com/google/bband-adaband and https://github.com/vztu/DebandingNet . Zhengzhong Tu, Chia-Ju Chen, Jessie Lin, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
Signal Process. Image Commun. | 5 |
| 2025 | Subjective and Objective Quality Assessment of Banding Artifacts on Compressed VideosabstractAlthough there have been notable advancements in video compression technologies in recent years, banding artifacts remain a serious issue affecting the quality of compressed videos, particularly on smooth regions of high-definition videos. Noticeable banding artifacts can severely impact the perceptual quality of videos viewed on a high-end HDTV or high-resolution screen. Hence, there is a pressing need for a systematic investigation of the banding video quality assessment problem for advanced video codecs. Given that the existing publicly available datasets for studying banding artifacts are limited to still picture data only, which cannot account for temporal banding dynamics, we have created a first-of-a-kind open video dataset, dubbed LIVE-YT-Banding, which consists of 160 videos generated by four different compression parameters using the AV1 video codec. A total of 7,200 subjective opinions are collected from a cohort of 45 human subjects. To demonstrate the value of this new resources, we tested and compared a variety of models that detect banding occurrences, and measure their impact on perceived quality. Among these, we introduce an effective and efficient new no-reference (NR) video quality evaluator which we call CBAND. CBAND leverages the properties of the learned statistics of natural images expressed in the embeddings of deep neural networks. Our experimental results show that the perceptual banding prediction performance of CBAND significantly exceeds that of previous state-of-the-art models, and is also orders of magnitude faster. Moreover, CBAND can be employed as a differentiable loss function to optimize video debanding models. The LIVE-YT-Banding database, code, and pre-trained model are all publically available at https://github.com/uniqzheng/CBAND. Qi Zheng 0004, Li-Heng Chen, Chenlong He, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik, Yibo Fan, Zhengzhong Tu |
IEEE Trans. Image Process. | 4 |
| 2024 | Youtube SFV+HDR Quality DatasetabstractThe popularity of Short form videos (SFV) has grown dramatically in the past few years, and has become a phenomenal video category with billions of viewers. Meanwhile, High Dynamic Range (HDR) as an advanced feature also becomes more and more popular on video sharing platforms. As a hot topic with huge impact, SFV and HDR bring new questions to video quality research: 1) is SFV+HDR quality assessment significantly different from traditional User Generated Content (UGC) quality assessment? 2) do objective quality metrics designed for traditional UGC still work well for SFV+HDR? To answer the above questions, we created the first large scale SFV+HDR dataset with reliable subjective quality scores, covering 10 popular content categories. Further, we also introduce a general sampling framework to maximize the representativeness of the dataset. We provided a comprehensive analysis of subjective quality scores for Short form SDR and HDR videos, and discuss the reliability of state-of-the-art UGC quality metrics and potential improvements. Yilin Wang 0001, Joong Gon Yim, Neil Birkbeck, Balu Adsumilli |
ICIP | 3 |
| 2024 | A Dataset for Understanding Open UGC Video DatasetsabstractUser Generated Content (UGC) video streaming is a major application on the Internet. Even small bitrate savings can have large network impacts at this scale. In order to achieve improvements without sacrificing experience, the quality of UGC videos needs to be better understood. In recent years video quality evaluation models designed for the evaluation of UGC videos have received a lot of attention. However, considering that these models are learning-based models, they heavily depend on the training data that has been used. In this paper, a new dataset is introduced that allows studying the differences in characteristics between existing UGC video datasets. It reveals the range of quality that was covered by existing UGC video datasets, and the implication of these quality ranges on training and validation performance of UGC video quality prediction models. Furthermore, this work demonstrates that dataset alignment enables existing UGC models to achieve higher performance. This alignment dataset can be found openly available on Zenodo (https://zenodo.org/doi/10.5281/zenodo.12155934). Pierre R. Lebreton, Patrick Le Callet, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli |
ICIP | 3 |
| 2024 | Subjective Portrait Region Cropping On Landscape Video StudyabstractWith the rise of mobile video consumption, adapting videos to non-traditional aspect ratios poses challenges for existing content. The use of static cropping and border padding often compromises visual quality, while warping may distort a video’s intended meaning. Here we advocate for a more effective approach - cropping significant regions within video frames in a temporal manner, while minimizing distortion and preserving essential content. However, the lack of a large-scale database devoted to informing these tasks impedes progress in this direction. Addressing this gap, we introduce the LIVE-YouTube Video Cropping (LIVE-YT VC) database, featuring 1800 videos labeled by 90 human subjects. Sourced from the YouTube-UGC and LSVQ databases, this collection serves as the largest subjective video portrait region cropping database. We evaluate our methodology using the SmartVidCrop [1] algorithm, establishing a benchmark for future research. Our contributions offer a crucial resource for advancing video aspect ratio transformation, ensuring that mobile-friendly video content retains its quality and meaning. The details of accessing the dataset have been provided in the supplementary material. Cheng-Han Lee, Maniratnam Mandal, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
ICIP | 3 |
| 2024 | Legit: Text Legibility For User-Generated MediaabstractUser-generated content (UGC) is ubiquitous across the internet as a result of billions of videos and images being uploaded each day. All kinds of UGC media are affected by natural distortions, occurring both during and after capture, which are inherently diverse and commingled. These distortions have different perceptual effects based on the media content. Given recent dramatic increases in the consumption of short-form content, the analysis and control of their perceptual quality has become an important problem. Regardless of the content, many UGC videos have overlaid and embedded texts in them, which are visually salient. Hence text quality has a significant impact on the global perception of video or image quality and needs to be studied. One of the most important factors in perceptual text quality in user-generated media is legibility, which has been studied very little in the context of computer vision. Predicting text legibility can also help in text recognition applications such as image search or document identification. This work aims at modeling text legibility using computer vision techniques and thus studying the relationship between text quality and legibility. We propose a modified dataset variant of COCO-Text [1] and a model for predicting text legibility for both handwritten and machine-generated texts. We also demonstrate how models trained to predict text legibility can help in the prediction of text (perceptual) quality. The dataset and models can be accessed here https://live.ece.utexas.edu/research/Quality/index.htm. Maniratnam Mandal, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 2 |
| 2024 | Subjective and Objective Analysis of Streamed Gaming VideosabstractThe rising popularity of online User-Generated-Content (UGC) in the form of streamed and shared videos, has hastened the development of perceptual Video Quality Assessment (VQA) models, which can be used to help optimize their delivery. Gaming videos, which are a relatively new type of UGC videos, are created when skilled and casual gamers post videos of their gameplay. These kinds of screenshots of UGC gameplay videos have become extremely popular on major streaming platforms like YouTube and Twitch. Synthetically-generated gaming content presents challenges to existing VQA algorithms, including those based on natural scene/video statistics models. Synthetically generated gaming content presents different statistical behavior than naturalistic videos. A number of studies have been directed towards understanding the perceptual characteristics of professionally generated gaming videos arising in gaming video streaming, online gaming, and cloud gaming. However, little work has been done on understanding the quality of UGC gaming videos, and how it can be characterized and predicted. Towards boosting the progress of gaming video VQA model development, we conducted a comprehensive study of subjective and objective VQA models on UGC gaming videos. To do this, we created a novel UGC gaming video resource, called the LIVE-YouTube Gaming video quality (LIVE-YT-Gaming) database, comprised of 600 real UGC gaming videos. We conducted a subjective human study on this data, yielding 18,600 human quality ratings recorded by 61 human subjects. We also evaluated a number of state-of-the-art (SOTA) VQA models on the new database, including a new one, called GAME-VQP, based on both natural video statistics and CNN-learned features. To help support work in this field, we are making the new LIVE-YT-Gaming Database, along with code for GAME-VQP, publicly available through the link:https://live.ece.utexas.edu/research/LIVE-YT-Gaming/index.html. Xiangxu Yu, Zhenqiang Ying, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Games | 3 |
| 2023 | CONVIQT: Contrastive Video Quality EstimatorabstractPerceptual video quality assessment (VQA) is an integral component of many streaming and video sharing platforms. Here we consider the problem of learning perceptually relevant video quality representations in a self-supervised manner. Distortion type identification and degradation level determination is employed as an auxiliary task to train a deep learning model containing a deep Convolutional Neural Network (CNN) that extracts spatial features, as well as a recurrent unit that captures temporal information. The model is trained using a contrastive loss and we therefore refer to this training framework and resulting model as CONtrastive VIdeo Quality EstimaTor (CONVIQT). During testing, the weights of the trained model are frozen, and a linear regressor maps the learned features to quality scores in a no-reference (NR) setting. We conduct comprehensive evaluations of the proposed model against leading algorithms on multiple VQA databases containing wide ranges of spatial and temporal distortions. We analyze the correlations between model predictions and ground-truth quality ratings, and show that CONVIQT achieves competitive performance when compared to state-of-the-art NR-VQA models, even though it is not trained on those databases. Our ablation experiments demonstrate that the learned representations are highly robust and generalize well across synthetic and realistic distortions. Our results indicate that compelling representations with perceptual bearing can be obtained using self-supervised learning. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2022 | An Empirical Approach for Optimising the Impact of a Preprocessor in a Transcoding PipelineabstractThe volume of User Generated Content (UGC) on the internet has exploded throughout the pandemic. The relatively low quality of that content generally implies an increased bitrate and much reduced quality after transcoding. Preprocessing e.g. using a noise reducer, is one approach for reducing bitrate and increasing quality. The impact of the noise reducer is however affected by the target bitrate of the encoder. That relationship is known but not previously quantitatively examined. In this paper we present a methodology and new metric for measuring this impact based on the Rate-Distortion curves before and after pre-processing. The metric is used as a cost function for estimating the optimal filter parameter for our chosen denoiser. Our experiments show that optimising the filter parameter in this way yields as much as 4-5dB improvement in PSNR at 3 Mbps. Varoun Hanooman, Anil C. Kokaram, Yeping Su, Neil Birkbeck, Balu Adsumilli |
ICIP | 4 |
| 2022 | When is the Cleaning of Subjective Data Relevant to Train UGC Video Quality Metrics?abstractOutlier analysis and spammer detection recently gained momentum in order to reduce uncertainty of subjective ratings in image & video quality assessment tasks. The large proportion of unreliable ratings from online crowdsourcing experiments and the need for qualitative and quantitative large-scale studies in the deep-learning ecosystem played a role in this event. We study the effect that data cleaning has on trainable models predicting the visual quality for videos, and present results demonstrating when cleaning is necessary to reach higher efficiency. To this end, we present and analyze a benchmark on clean and noisy User Generated Content (UGC) large-scale datasets on which we re-trained models, followed by an empirical exploration of the constraint of data removal. Our results show that a dataset presenting between 7 and 30% of outliers benefits from cleaning before training. Anne-Flore Perrin, Charles Dormeval, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Patrick Le Callet |
ICIP | 4 |
| 2022 | Revisiting the Efficiency of UGC Video Quality AssessmentabstractUGC video quality assessment (UGC-VQA) is a challenging research topic due to the high video diversity and limited public UGC quality datasets. State-of-the-art (SOTA) UGC quality models tend to use high complexity models, and rarely discuss the trade-off among complexity, accuracy, and generalizability. We propose a new perspective on UGC-VQA, and show that model complexity may not be critical to the performance, whereas a more diverse dataset is essential to train a better model. We illustrate this by using a light weight model, UVQ-lite, which has higher efficiency and better generalizability (less overfitting) than baseline SOTA models. We also propose a new way to analyze the sufficiency of the training set, by leveraging UVQ’s comprehensive features. Our results motivate a new perspective about the future of UGC-VQA research, which we believe is headed toward more efficient models and more diverse datasets. Yilin Wang 0001, Joong Gon Yim, Neil Birkbeck, Junjie Ke, Hossein Talebi, Feng Yang 0008, Balu Adsumilli |
ICIP | 3 |
| 2022 | A Deep Learning post-processor with a perceptual loss function for video compression artifact removalabstractWhile video compression is necessary for large scale video streaming services, compression at low bitrate can degrade the original video and negatively affect the end user’s quality of experience. Deep Neural Networks (DNNs) are actively researched with respect to artifact removal, however the loss functions that are typically employed follows a derivation of a pixel-wise Lpnorm. In this paper we consider a DNN as a post-processor for video compression artifact removal. The DNN is trained using a composite perceptual loss that combines a traditional Lpnorm loss and a VMAF proxy network based on the Video Multimethod Assessment Function (VMAF). Results show an improvement in VMAF score over both the training and testing sets. Darren Ramsook, Anil C. Kokaram, Neil Birkbeck, Yeping Su, Balu Adsumilli |
PCS | 3 |
| 2022 | Making Video Quality Assessment Models Sensitive to Frame Rate DistortionsabstractWe consider the problem of capturing distortions arising from changes in frame rate as part of Video Quality Assessment (VQA). Variable frame rate (VFR) videos have become much more common, and streamed videos commonly range from 30 frames per second (fps) up to 120 fps. VFR-VQA offers unique challenges in terms of distortion types as well as in making non-uniform comparisons of reference and distorted videos having different frame rates. The majority of current VQA models require compared videos to be of the same frame rate, but are unable to adequately account for frame rate artifacts. The recently proposed Generalized Entropic Difference (GREED) VQA model succeeds at this task, using natural video statistics models of entropic differences of temporal band-pass coefficients, delivering superior performance on predicting video quality changes arising from frame rate distortions. Here we propose a simple fusion framework, whereby temporal features from GREED are combined with existing VQA models, towards improving model sensitivity towards frame rate distortions. We find through extensive experiments that this feature fusion significantly boosts model performance on both HFR/VFR datasets as well as fixed frame rate (FFR) VQA databases. Our results suggest that employing efficient temporal representations can result much more robust and accurate VQA models when frame rate variations can occur. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Signal Process. Lett. | 2 |
| 2022 | Image Quality Assessment Using Contrastive LearningabstractWe consider the problem of obtaining image quality representations in a self-supervised manner. We use prediction of distortion type and degree as an auxiliary task to learn features from an unlabeled image dataset containing a mixture of synthetic and realistic distortions. We then train a deep Convolutional Neural Network (CNN) using a contrastive pairwise objective to solve the auxiliary problem. We refer to the proposed training framework and resulting deep IQA model as the CONTRastive Image QUality Evaluator (CONTRIQUE). During evaluation, the CNN weights are frozen and a linear regressor maps the learned representations to quality scores in a No-Reference (NR) setting. We show through extensive experiments that CONTRIQUE achieves competitive performance when compared to state-of-the-art NR image quality models, even without any additional fine-tuning of the CNN backbone. The learned representations are highly robust and generalize well across images afflicted by either synthetic or authentic distortions. Our results suggest that powerful quality representations with perceptual relevance can be obtained without requiring large labeled subjective image quality datasets. The implementations used in this paper are available at https://github.com/pavancm/CONTRIQUE. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2021 | Rich Features for Perceptual Quality Assessment of UGC VideosabstractVideo quality assessment for User Generated Content (UGC) is an important topic in both industry and academia. Most existing methods only focus on one aspect of the perceptual quality assessment, such as technical quality or compression artifacts. In this paper, we create a large scale dataset to comprehensively investigate characteristics of generic UGC video quality. Besides the subjective ratings and content labels of the dataset, we also propose a DNN-based framework to thoroughly analyze importance of content, technical quality, and compression level in perceptual quality. Our model is able to provide quality scores as well as human-friendly quality indicators, to bridge the gap between low level video signals to human perceptual quality. Experimental results show that our model achieves state-of-the-art correlation with Mean Opinion Scores (MOS). Yilin Wang 0001, Junjie Ke, Hossein Talebi, Joong Gon Yim, Neil Birkbeck, Balu Adsumilli, Peyman Milanfar, Feng Yang 0008 |
CVPR | 5 |
| 2021 | Regression or classification? New methods to evaluate no-reference picture and video quality modelsabstractVideo and image quality assessment has long been projected as a regression problem, which requires predicting a continuous quality score given an input stimulus. However, recent efforts have shown that accurate quality score regression on real-world user-generated content (UGC) is a very challenging task. To make the problem more tractable, we propose two new methods - binary, and ordinal classification - as alternatives to evaluate and compare no-reference quality models at coarser levels. Moreover, the proposed new tasks convey more practical meaning on perceptually optimized UGC transcoding, or for preprocessing on media processing platforms. We conduct a comprehensive benchmark experiment of popular no-reference quality models on recent in-the-wild picture and video quality datasets, providing reliable baselines for both evaluation methods to support further studies. We hope this work promotes coarse-grained perceptual modeling and its applications to efficient UGC processing. Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICASSP | 5 |
| 2021 | Video Quality Assessment of User Generated Content: A Benchmark Study and a New ModelabstractRecent years have witnessed an explosion of user-generated content (UGC) shared and streamed over the Internet. Accordingly, there is a great need for accurate video quality assessment (VQA) models for consumer or UGC videos to monitor, control, and optimize this vast content. Here we contribute to advancing the UGC-VQA problem by conducting a comprehensive evaluation of leading blind VQA (BVQA) models. Besides, we also created a new fusion-based BVQA model, which we dub the VIDeo quality EVALuator (VIDEVAL), that effectively balances the trade-off between performance and efficiency. Our experimental results show that VIDEVAL achieves state-of-the-art performance at a lower computational cost. We believe our reliable and reproducible benchmark will facilitate further research on deep learning-based BVQA modeling. An implementation of VIDEVAL has been made available online1.1https://github.com/vztu/VIDEVAL_release Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 4 |
| 2021 | A Temporal Statistics Model For UGC Video Quality PredictionabstractBlind video quality assessment of user-generated content (UGC) has become a trending and challenging problem. Previous studies have shown the efficacy of natural scene statistics for capturing spatial distortions. The exploration of temporal video statistics on UGC, however, is relatively limited. Here we propose the first general, effective and efficient temporal statistics model accounting for temporal- or motion-related distortions for UGC video quality assessment, by analyzing regularities in the temporal bandpass domain. The proposed temporal model can serve as a plug-in module to boost existing no-reference video quality predictors that lack motion-relevant features. Our experimental results on recent large-scale UGC video databases show that the proposed model can significantly improve the performances of existing methods, at a very reasonable computational expense. Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 4 |
| 2021 | High Frame Rate Video Quality Assessment using VMAF and Entropic DifferencesabstractThe popularity of streaming videos with live, high-action content has led to an increased interest in High Frame Rate (HFR) videos. In this work we address the problem of frame rate dependent Video Quality Assessment (VQA) when the videos to be compared have different frame rate and compression factor. The current VQA models such as VMAF have superior correlation with perceptual judgments when videos to be compared have same frame rates and contain conventional distortions such as compression, scaling etc. However this framework requires additional pre-processing step when videos with different frame rates need to be compared, which can potentially limit its overall performance. Recently, Generalized Entropic Difference (GREED) VQA model was proposed to account for artifacts that arise due to changes in frame rate, and showed superior performance on the LIVE-YT-HFR database which contains frame rate dependent artifacts such as judder, strobing etc. In this paper we propose a simple extension, where the features from VMAF and GREED are fused in order to exploit the advantages of both models. We show through various experiments that the proposed fusion framework results in more efficient features for predicting frame rate dependent video quality. We also evaluate the fused feature set on standard non-HFR VQA databases and obtain superior performance than both GREED and VMAF, indicating the combined feature set captures complimentary perceptual quality information. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
PCS | 2 |
| 2021 | A differentiable estimator of VMAF for VideoabstractModern Perceptual Visual Quality Metrics (PVQMs) for video are generally complex and non-differentiable. This makes them difficult to use as loss functions in restoration and compression tuning. Traditional metrics such as PSNR/MSE which are differentiable remain important but do not capture perceptual visual criteria. In this paper we present a DNN which models a popular perceptual video metric VMAF. In so doing, we introduce a differentiable loss function that closely matches the behaviour of a perceptual metric. Employing degradation generated with H.265 compression, our model achieves a 4.41% RMSE in predicting VMAF. This can now be deployed as a video based loss function in video enhancement and compression tasks. Darren Ramsook, Anil C. Kokaram, Noel E. O'Connor, Neil Birkbeck, Yeping Su, Balu Adsumilli |
PCS | 4 |
| 2021 | Efficient User-Generated Video Quality PredictionabstractBlind video quality assessment of user-generated content (UGC) has become a trending, challenging, unsolved problem. Accurate and efficient video quality predictors suitable for this content are thus in great demand to achieve intelligent analysis and processing of UGC videos. However, previous video quality models are either incapable or inefficient for predicting the quality of complex, diverse UGC videos in practical applications. Here we introduce an effective and efficient video quality model for UGC content, which we dub the Rapid and Accurate Video Quality Evaluator (RAPIQUE), which we show performs comparably to state-of-the-art models but with orders-of-magnitude faster runtime. Our experimental results on recent large-scale UGC video quality databases show that RAPIQUE delivers top performances on all datasets at a considerably lower computational expense. An implementation of RAPIQUE is online: https://github.com/vztu/RAPIQUE. Zhengzhong Tu, Chia-Ju Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
PCS | 4 |
| 2021 | ST-GREED: Space-Time Generalized Entropic Differences for Frame Rate Dependent Video Quality PredictionabstractWe consider the problem of conducting frame rate dependent video quality assessment (VQA) on videos of diverse frame rates, including high frame rate (HFR) videos. More generally, we study how perceptual quality is affected by frame rate, and how frame rate and compression combine to affect perceived quality. We devise an objective VQA model called Space-Time GeneRalized Entropic Difference (GREED) which analyzes the statistics of spatial and temporal band-pass video coefficients. A generalized Gaussian distribution (GGD) is used to model band-pass responses, while entropy variations between reference and distorted videos under the GGD model are used to capture video quality variations arising from frame rate changes. The entropic differences are calculated across multiple temporal and spatial subbands, and merged using a learned regressor. We show through extensive experiments that GREED achieves state-of-the-art performance on the LIVE-YT-HFR Database when compared with existing VQA models. The features used in GREED are highly generalizable and obtain competitive performance even on standard, non-HFR VQA databases. The implementation of GREED has been made available online: https://github.com/pavancm/GREED. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2021 | UGC-VQA: Benchmarking Blind Video Quality Assessment for User Generated ContentabstractRecent years have witnessed an explosion of user-generated content (UGC) videos shared and streamed over the Internet, thanks to the evolution of affordable and reliable consumer capture devices, and the tremendous popularity of social media platforms. Accordingly, there is a great need for accurate video quality assessment (VQA) models for UGC/consumer videos to monitor, control, and optimize this vast content. Blind quality prediction of in-the-wild videos is quite challenging, since the quality degradations of UGC videos are unpredictable, complicated, and often commingled. Here we contribute to advancing the UGC-VQA problem by conducting a comprehensive evaluation of leading no-reference/blind VQA (BVQA) features and models on a fixed evaluation architecture, yielding new empirical insights on both subjective video quality studies and objective VQA model design. By employing a feature selection strategy on top of efficient BVQA models, we are able to extract 60 out of 763 statistical features used in existing methods to create a new fusion-based model, which we dub the VIDeo quality EVALuator (VIDEVAL), that effectively balances the trade-off between VQA performance and efficiency. Our experimental results show that VIDEVAL achieves state-of-the-art performance at considerably lower computational cost than other leading models. Our study protocol also defines a reliable benchmark for the UGC-VQA problem, which we believe will facilitate further research on deep learning-based VQA modeling, as well as perceptually-optimized efficient UGC video processing, transcoding, and streaming. To promote reproducible research and public evaluation, an implementation of VIDEVAL has been made available online: https://github.com/vztu/VIDEVAL. Zhengzhong Tu, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2021 | Predicting the Quality of Compressed Videos With Pre-Existing DistortionsabstractBecause of the increasing ease of video capture, many millions of consumers create and upload large volumes of User-Generated-Content (UGC) videos to social and streaming media sites over the Internet. UGC videos are commonly captured by naive users having limited skills and imperfect techniques, and tend to be afflicted by mixtures of highly diverse in-capture distortions. These UGC videos are then often uploaded for sharing onto cloud servers, where they are further compressed for storage and transmission. Our paper tackles the highly practical problem of predicting the quality of compressed videos (perhaps during the process of compression, to help guide it), with only (possibly severely) distorted UGC videos as references. To address this problem, we have developed a novel Video Quality Assessment (VQA) framework that we call 1stepVQA (to distinguish it from two-step methods that we discuss). 1stepVQA overcomes limitations of Full-Reference, Reduced-Reference and No-Reference VQA models by exploiting the statistical regularities of both natural videos and distorted videos. We also describe a new dedicated video database, which was created by applying a realistic VMAF-Guided perceptual rate distortion optimization (RDO) criterion to create realistically compressed versions of UGC source videos, which typically have pre-existing distortions. We show that 1stepVQA is able to more accurately predict the quality of compressed videos, given imperfect reference videos, and outperforms other VQA models in this scenario. Xiangxu Yu, Neil Birkbeck, Yilin Wang 0001, Christos G. Bampis, Balu Adsumilli, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2020 | A Comparative Evaluation Of Temporal Pooling Methods For Blind Video Quality AssessmentabstractMany objective video quality assessment (VQA) algorithms include a key step of temporal pooling of frame-level quality scores. However, less attention has been paid to studying the relative efficiencies of different pooling methods on noreference (blind) VQA. Here we conduct a large-scale comparative evaluation to assess the capabilities and limitations of multiple temporal pooling strategies on blind VQA of usergenerated videos. The study yields insights and general guidance regarding the application and selection of temporal pooling models. In addition, we also propose an ensemble pooling model built on top of high-performing temporal pooling models. Our experimental results demonstrate the relative efficacies of the evaluated temporal pooling models, using several popular VQA algorithms evaluated on two recent largescale natural video quality databases. Conclusively, we also provide an empirical recipe for applying temporal pooling of frame-based quality predictions. Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik |
ICIP | 4 |
| 2020 | Subjective Quality Assessment For Youtube Ugc DatasetabstractDue to the scale of social video sharing, User Generated Content (UGC) is getting more attention from academia and industry. To facilitate compression-related research on UGC, YouTube has released a large-scale dataset [1]. The initial dataset only provided videos, limiting its use in quality assessment. We used a crowd-sourcing platform to collect subjective quality scores for this dataset. We analyzed the distribution of Mean Opinion Score (MOS) in various dimensions, and investigated some fundamental questions in video quality assessment, like the correlation between full video MOS and corresponding chunk MOS, and the influence of chunk variation in quality score aggregation. Joong Gon Yim, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli |
ICIP | 3 |
| 2020 | A Viewport-Driven Multi-Metric Fusion Approach for 360-Degree Video Quality AssessmentabstractWe propose a new viewport-based multi-metric fusion (MMF) approach for visual quality assessment of 360-degree (omnidirectional) videos. Our method is based on computing multiple spatio-temporal objective quality metrics (features) on viewports extracted from 360-degree videos, and learning a model that combines these features into a metric, which closely matches subjective quality scores. The main motivations for the proposed method are that: 1) quality metrics computed on viewports better captures the user experience than metrics computed on the projection domain; 2) no individual objective image quality metric always performs best for all types of visual distortions, while a learned combination of them is able to adapt to different conditions and produce better results overall. Experimental results, based on the largest available 360-degree videos quality dataset, demonstrate that the proposed metric outperforms state-of-the-art 360-degree and 2D video quality metrics. Roberto Gerson De Albuquerque Azevedo, Neil Birkbeck, Ivan Janatra, Balu Adsumilli, Pascal Frossard |
ICME | 2 |
| 2020 | Translation of Perceived Video Quality Across DisplaysabstractDisplay devices can affect the perceived quality of a video significantly. In this paper, we focus on the scenario where video resolution does not exceed screen resolution, and investigate the relationship of perceived video quality on mobile, laptop and TV. A novel transformation of Mean Opinion Scores (MOS) among different devices is proposed and is shown to be effective at normalizing ratings across user devices for in lab and crowd sourced subjective studies. The model allows us to perform more focused in lab subjective studies as we can reduce the number of test devices and helps us reduce noise during crowd-sourcing subjective video quality tests. It is also more effective than utilizing existing device dependent objective metrics for translating MOS ratings across devices. Jessie Lin, Neil Birkbeck, Balu Adsumilli |
MMSP | 2 |
| 2020 | Capturing Video Frame Rate Variations via Entropic DifferencingabstractHigh frame rate videos are increasingly getting popular in recent years, driven by the strong requirements of the entertainment and streaming industries to provide high quality of experiences to consumers. To achieve the best trade-offs between the bandwidth requirements and video quality in terms of frame rate adaptation, it is imperative to understand the effects of frame rate on video quality. In this direction, we devise a novel statistical entropic differencing method based on a Generalized Gaussian Distribution model expressed in the spatial and temporal band-pass domains, which measures the difference in quality between reference and distorted videos. The proposed design is highly generalizable and can be employed when the reference and distorted sequences have different frame rates. Our proposed model correlates very well with subjective scores in the recently proposed LIVE-YT-HFR database and achieves state of the art performance when compared with existing methodologies. Pavan C. Madhusudana, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik |
IEEE Signal Process. Lett. | 2 |
| 2020 | Visual Distortions in 360° VideosabstractOmnidirectional (or 360°) images and videos are emergent signals being used in many areas, such as robotics and virtual/augmented reality. In particular, for virtual reality applications, they allow an immersive experience in which the user can interactively navigate through a scene with three degrees of freedom, wearing a head-mounted display. Current approaches for capturing, processing, delivering, and displaying 360° content, however, present many open technical challenges and introduce several types of distortions in the visual signal. Some of the distortions are specific to the nature of 360° images and often differ from those encountered in classical visual communication frameworks. This paper provides a first comprehensive review of the most common visual distortions that alter 360° signals going through the different processing elements of the visual communication pipeline. While their impact on viewers' visual perception and the immersive experience at large is still unknown-thus, it is an open research topic-this review serves the purpose of proposing a taxonomy of the visual distortions that can be encountered in 360° signals. Their underlying causes in the end-to-end 360° content distribution pipeline are identified. This taxonomy is essential as a basis for comparing different processing techniques, such as visual enhancement, encoding, and streaming strategies, and allowing the effective design of new algorithms and applications. It is also a useful resource for the design of psycho-visual studies aiming to characterize human perception of 360° content in interactive and immersive applications. Roberto Gerson De Albuquerque Azevedo, Neil Birkbeck, Francesca De Simone, Ivan Janatra, Balu Adsumilli, Pascal Frossard |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Mutual Noise Estimation Algorithm for Video DenoisingabstractThis paper presents a novel algorithm to estimate spatio-temporal noise variance in videos. The algorithm uses mutual information from the spatial and temporal noise statistics to detect homogeneous blocks in a video frame and estimate the noise. The experimental results show the accuracy and robustness of the algorithm over videos with low to high spatial and temporal complexities, and varying levels of noise power and noise correlations. Mohammad Izadi, Neil Birkbeck, Balu Adsumilli |
ICIP | 2 |
| 2018 | Film Grain Synthesis for AV1 Video CodecabstractFilm grain is abundant in TV and movie content. It is often part of the creative intent and needs to be preserved while encoding. However, the random nature of film grain is difficult to compress using traditional coding tools. This paper describes a film grain modeling and synthesis algorithm proposed for the AV1 video codec. At the encoder, an autoregressive model of film grain is transmitted relative to a denoised signal, and the film grain strength is modeled as a function of intensity. The corresponding renoising at the decoder is implemented using an efficient block-based approach suitable for use in consumer electronic devices. Preliminary results indicate that the approach can give significant bitrate savings (up to 50%) on sequences with heavy film grain. Andrey Norkin, Neil Birkbeck |
DCC | 2 |
| 2017 | Deformable block-based motion estimation in omnidirectional image sequencesabstractThis paper presents an extension of block-based motion estimation for omnidirectional videos, based on a translational object motion model that accounts for the spherical geometry of the imaging system. We use this model to design a new algorithm to perform block matching in sequences of panoramic frames that are the result of the equirectangular projection. Experimental results demonstrate that significant gains can be achieved with respect to the classical exhaustive block matching algorithm in terms of accuracy of motion prediction. In particular, average quality improvements up to approximately 6 dB in terms of Peak Signal to Noise Ratio (PSNR), 0.043 in terms of Structural SIMilarity index (SSIM), and 2 dB in terms of spherical PSNR, can be achieved on the predicted frames. Francesca De Simone, Pascal Frossard, Neil Birkbeck, Balu Adsumilli |
MMSP | 3 |
| 2017 | Quantitative evaluation of omnidirectional video qualityabstractOmnidirectional video encoding and delivery are rapidly evolving fields, where choosing an efficient representation for storage and transmission of pixel data is critical. Given that there are a number of projections (pixel representations), a projection independent measure is needed to evaluate the merits of different options. We present a technique to evaluate projection quality by rendering virtual views and use this to evaluate three projections in common use: Equirectangular, Cubemap, and Equi-Angular Cubemap1. Through evaluation on dozens of videos, our metrics rank the projection types consistently with pixel density computations and small scale user studies. Neil Birkbeck, Chip Brown, Robert Suderman |
QoMEX | 1 |
| 2016 | Geometry-driven quantization for omnidirectional image codingabstractIn this paper we propose a method to adapt the quantization tables of typical block-based transform codecs when the input to the encoder is a panoramic image resulting from equirectangular projection of a spherical image. When the visual content is projected from the panorama to the viewport, a frequency shift is occurring. The quantization can be adapted accordingly: the quantization step sizes that would be optimal to quantize the transform coefficients of the viewport image block, can be used to quantize the coefficients of the panoramic block. As a proof of concept, the proposed quantization strategy has been used in JPEG compression. Results show that a rate reduction up to 2.99% can be achieved for the same perceptual quality of the spherical signal with respect to a standard quantization. Francesca De Simone, Pascal Frossard, Paul Wilkins, Neil Birkbeck, Anil C. Kokaram |
PCS | 4 |
| 2014 | Temporal synchronization of multiple audio signalsabstractGiven the proliferation of consumer media recording devices, events often give rise to a large number of recordings. These recordings are taken from different spatial positions and do not have reliable timestamp information. In this paper, we present two robust graph-based approaches for synchronizing multiple audio signals. The graphs are constructed atop the over-determined system resulting from pairwise signal comparison using cross-correlation of audio features. The first approach uses a Minimum Spanning Tree (MST) technique, while the second uses Belief Propagation (BP) to solve the system. Both approaches can provide excellent solutions and robustness to pairwise outliers, however the MST approach is much less complex than BP. In addition, an experimental comparison of audio features-based synchronization shows that spectral flatness outperforms the zero-crossing rate and signal energy. Julius Kammerl, Neil Birkbeck, Sasi Inguva, Damien Kelly, Andrew J. Crawford, Hugh Denman, Anil C. Kokaram, Caroline Pantofaru |
ICASSP | 2 |
| 2014 | Lung Segmentation from CT with Severe Pathologies Using Anatomical Constraints
Neil Birkbeck, Timo Kohlberger, Jingdan Zhang, Michal Sofka, Jens N. Kaftan, Dorin Comaniciu, Shaohua Kevin Zhou |
MICCAI (1) | 1 |
| 2014 | Segmentation of Multiple Knee Bones from CT for Orthopedic Knee Surgery Planning
Dijia Wu, Michal Sofka, Neil Birkbeck, Shaohua Kevin Zhou |
MICCAI (1) | 3 |
| 2013 | IntellEditS: Intelligent Learning-Based Editor of Segmentations
Adam P. Harrison, Neil Birkbeck, Michal Sofka |
MICCAI (3) | 2 |
| 2012 | Precise Segmentation of Multiple Organs in CT Volumes Using Learning-Based Approach and Information Theory
Chao Lu 0011, Yefeng Zheng 0001, Neil Birkbeck, Jingdan Zhang, Timo Kohlberger, Christian Tietjen, Thomas Böttger, James S. Duncan, Shaohua Kevin Zhou |
MICCAI (2) | 3 |
| 2011 | Basis constrained 3D scene flow on a dynamic proxyabstractExisting scene flow approaches mainly focus on two-frame stereo-pair configurations and reconstruct an image-based representation of scene flow. Instead, we propose a variational formulation of scene flow relative to a coarse proxy geometry, which is better suited for many views. Furthermore, a linear basis is used to represent temporal surface flow, allowing for longer-range temporal correspondence with fewer variables. Our formulation takes known proxy motion into account (e.g, if the proxy is a tracked human subject), which enables 3D trajectory reconstruction when only a single view is available. Additionally, through the appropriate proxy and basis, our framework generalizes existing approaches for scene flow, optic-flow, and two-frame stereo. We illustrate results on real-data for both static and moving proxy surfaces over several frames. Neil Birkbeck, Dana Cobzas, Martin Jägersand |
ICCV | 1 |
| 2011 | Automatic Multi-organ Segmentation Using Learning-Based Segmentation and Level Set Optimization
Timo Kohlberger, Michal Sofka, Jingdan Zhang, Neil Birkbeck, Jens Wetzl, Jens N. Kaftan, Jérôme Declerck, Shaohua Kevin Zhou |
MICCAI (3) | 4 |
| 2011 | Multi-stage Learning for Robust Lung Segmentation in Challenging CT Volumes
Michal Sofka, Jens Wetzl, Neil Birkbeck, Jingdan Zhang, Timo Kohlberger, Jens N. Kaftan, Jérôme Declerck, Shaohua Kevin Zhou |
MICCAI (3) | 3 |
| 2010 | Performance evaluation of monocular predictive displayabstractIn teleoperation systems, operator performance is negatively affected by time-delayed visual feedback. Predictive display (PD) compensates for delays by providing synthesized visual feedback. While most existing PD methods rely on a priori models (e.g., from laser range finding or stereo vision), recent work on monocular SLAM and SFM makes it possible to acquire PD models in single camera applications. In this work, we evaluate operator performance of PD visual feedback based on a coarse 3D model. We report the experimental results of 12 human tele-operators each performing 96 visual alignment tasks with a 300ms delay. Four operating modes are considered: delayed video (no PD), video-based PD using a stabilizing plane (homography), 3D model-based PD, and no delay (ground truth). The results indicate that vision-based PD (both plane and 3D model-based) is significantly better than delayed video. It reduced task completion time 40% and is nearly as good as the no delay condition. PD based on a sparse a 3D model was somewhat better than the simpler plane based method. Adam Rachmielowski, Neil Birkbeck, Martin Jägersand |
ICRA | 2 |
| 2010 | Predictive display for mobile manipulators in unknown environments using online vision-based monocular modeling and localizationabstractTo tele-operate a robot, visual feedback is critical. However, communication channel latency can delay feedback to the point where the operator is impeded in performing his task. This work presents a vision-based “predictive display” system that compensates for visual delay. The approach is online and relatively uncalibrated, thus it has the advantage of being useful in unknown environments and many applications. From monocular eye-in-hand video, we incrementally compute a 3D graphics model of the robot site in real time using our new technique. The method exploits free-space/occlusion constraints on the scene to produce a physically consistent mesh. Novel vantage points are immediately rendered in response to the operator's control commands, without waiting for delayed video. We implement a full prototype tele-operation system where the operator controls, via a PHANTOM Omni device, a Barrett WAM robot mounted on a mobile Segway. Experiments with this setup validate the efficacy of the proposed approach. We demonstrate significant improvement in task completion time with predictive display on a real robot, while our previous related results were established only in simulation. David Lovi, Neil Birkbeck, Alejandro Hernandez Herdocia, Adam Rachmielowski, Martin Jägersand, Dana Cobzas |
IROS | 2 |
| 2009 | An interactive graph cut method for brain tumor segmentationabstractTumor segmentation from MRI data is an important but time consuming task performed manually by medical experts. Automating this process is challenging due to the high diversity in appearance of tumor tissue among different patients and, in many cases, similarity between tumor and normal tissue. We propose a semi-automatic interactive brain tumor segmentation system that incorporates 2D interactive and 3D automatic tools with the ability to adjust operator control. The provided methods are based on an energy that incorporates region statistics computed on available MRI modalities and the usual regularization term. The energy is efficiently minimized on-line using graph cut. Experiments with radiation oncologists testing the semi-automatic tool vs. a manual tool show that the proposed system improves both segmentation time and repeatability. Neil Birkbeck, Dana Cobzas, Martin Jägersand, Albert Murtha, Tibor Kesztyues |
WACV | 1 |
| 2007 | A Dimension Abstraction Approach to Vectorization in MatlababstractMatlab is a matrix-processing language that offers very efficient built-in operations for data organized in arrays. However Matlab operation is slow when the program accesses data through interpreted loops. Often during the development of a Matlab application writing loop-based code is more intuitive than crafting the data organization into arrays. Furthermore, many Matlab users do not command the linear algebra expertise necessary to write efficient code. Thus loop-based Matlab coding is a fairly common practice. This paper presents a tool that automatically converts loop-based Matlab code into equivalent array-based form and built-in Matlab constructs. Array-based code is produced by checking the input and output dimensions of equations within loops, and by transposing terms when necessary to generate correct code. This paper also describes an extensible loop pattern database that allows user-defined patterns to be discovered and replaced by more efficient Matlab routines that perform the same computation. The safe conversion of loop-based into more efficient array-based code is made possible by the introduction of a new abstract representation for dimensions Neil Birkbeck, Jonathan Levesque, José Nelson Amaral |
CGO | 1 |
| 2007 | 3D Variational Brain Tumor Segmentation using a High Dimensional Feature SetabstractTumor segmentation from MRI data is an important but time consuming task performed manually by medical experts. Automating this process is challenging due to the high diversity in appearance of tumor tissue, among different patients and, in many cases, similarity between tumor and normal tissue. One other challenge is how to make use of prior information about the appearance of normal brain. In this paper we propose a variational brain tumor segmentation algorithm that extends current approaches from texture segmentation by using a high dimensional feature set calculated from MRI data and registered atlases. Using manually segmented data we learn a statistical model for tumor and normal tissue. We show that using a conditional model to discriminate between normal and abnormal regions significantly improves the segmentation results compared to traditional generative models. Validation is performed by testing the method on several cancer patient MRI scans. Dana Cobzas, Neil Birkbeck, Mark Schmidt 0001, Martin Jägersand, Albert Murtha |
ICCV | 2 |
| 2006 | Variational Shape and Reflectance Estimation Under Changing Light and Viewpoints
Neil Birkbeck, Dana Cobzas, Peter F. Sturm, Martin Jägersand |
ECCV (1) | 1 |