VLDB 2026 Research / reviewers in the wild / expert
Anil C. Kokaram
dblp:47/2143
· DBLP profile ↗
92ranked-venue papers
15as first author
20since 2021 · last 2025
0000-0001-5304-6238ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 85 · 12 first-author · 20 since 2021Artificial intelligence and machine learning · 9 · 3 first-authorHuman-computer interaction and ubiquitous computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LiteVPNet: A Lightweight Network for Video Encoding Control in Quality-Critical Applications
Vibhoothi Vibhoothi, François Pitié, Anil C. Kokaram |
PCS | 3 |
| 2025 | An Empirical Study of Reducing AV1 Decoder Complexity and Energy Consumption via Encoder Parameter Tuning
Vibhoothi Vibhoothi, Julien Zouein, Shanker Shreejith, Jean-Baptiste Kempf, Anil C. Kokaram |
PCS | 5 |
| 2025 | AV1 Motion Vector Fidelity and Application for Efficient Optical Flow
Julien Zouein, Vibhoothi Vibhoothi, Anil C. Kokaram |
PCS | 3 |
| 2025 | An Efficient Quality Metric for Video Frame Interpolation Based on Motion-Field DivergenceabstractVideo frame interpolation is a fundamental tool for temporal video enhancement, but existing quality metrics struggle to evaluate the perceptual impact of interpolation artefacts effectively. Metrics like PSNR, SSIM and LPIPS ignore temporal coherence. State-of-the-art quality metrics tailored towards video frame interpolation, like FloLPIPS, have been developed but suffer from computational inefficiency that limits their practical application. We present PSNRDIV, a novel full-reference quality metric that enhances PSNR through motion divergence weighting, a technique adapted from archival film restoration where it was developed to detect temporal inconsistencies. Our approach highlights singularities in motion fields which is then used to weight image errors. Evaluation on the BVI-VFI dataset (180 sequences across multiple frame rates, resolutions and interpolation methods) shows PSNRDIVachieves statistically significant improvements: +0.09 Pearson Linear Correlation Coefficient over FloLPIPS, while being 2.5× faster and using 4× less memory. Performance remains consistent across all content categories and are robust to the motion estimator used. The efficiency and accuracy of PSNRDIVenables fast quality evaluation and practical use as a loss function for training neural networks for video frame interpolation tasks. An implementation of our metric is available at www.github.com/conalld/psnr-div. Conall Daly, Darren Ramsook, Anil C. Kokaram |
QoMEX | 3 |
| 2024 | A Neural Enhancement Post-Processor with a Dynamic AV1 Encoder Configuration Strategy for CLIC 2024abstractAt practical streaming bitrates, traditional video compression pipelines frequently lead to visible artifacts that degrade perceptual quality. This submission couples the effectiveness of a neural post-processor with a different dynamic optimsation strategy for achieving an improved bitrate/quality compromise. The neural post-processor is refined via adversarial training and employs perceptual loss functions. By optimising the post-processor and encoder directly our method demonstrates significant improvement in video fidelity. The neural post-processor achieves substantial VMAF score increases of +6.72 and +1.81 at bitrates of 50 kb/s and 500 kb/s respectively. Darren Ramsook, Anil C. Kokaram |
DCC | 2 |
| 2024 | A Dictionary Based Approach for Removing Out-of-Focus BlurabstractThe field of image deblurring has seen tremendous progress with the rise of deep learning models. These models, albeit efficient, are computationally expensive and energy consuming. Dictionary based learning approaches have shown promising results in image denoising and Single Image Super-Resolution. We propose an extension of the Rapid and Accurate Image Super-Resolution (RAISR) algorithm introduced by Isidoro, Romano and Milanfar for the task of out-of-focus blur removal. We define a sharpness quality measure which aligns well with the perceptual quality of an image. A metric based blending strategy based on asset allocation management is also proposed. Our method demonstrates an average increase of approximately 13%(PSNR) and 10% (SSIM) compared to popular deblurring methods. Furthermore, our blending scheme curtails ringing artefacts post restoration. Uditangshu Aurangabadkar, Anil C. Kokaram |
ICIP | 2 |
| 2024 | A Sharpness Based Loss Function for Removing Out-of-Focus BlurabstractThe success of modern Deep Neural Network (DNN) approaches can be attributed to the use of complex optimization criteria beyond standard losses such as mean absolute error (MAE) or mean squared error (MSE). In this work, we propose a novel method of utilising a no-reference sharpness metric$Q$introduced by Zhu and Milanfar for removing out-of-focus blur from images. We also introduce a novel dataset of real-world out-of-focus images for assessing restoration models. Our fine-tuned method produces images with a 7.5% increase in perceptual quality (LPIPS) as compared to a standard model trained only on MAE. Furthermore, we observe a 6.7% increase in$Q$(reflecting sharper restorations) and 7.25% increase in PSNR over most state-of-the-art (SOTA) methods. Uditangshu Aurangabadkar, Darren Ramsook, Anil C. Kokaram |
MMSP | 3 |
| 2024 | Comparative Analysis of Subjective Evaluations for Traditional and Neural-Based Video Enhancement TechniquesabstractThis work evaluates the effectiveness of modern video restoration methods, contrasting neural network-based techniques with traditional statistical algorithms to improve perceived video quality. Our analysis focused on three distinct methods: VBM4D, CVEGAN, and Ramsook, assessing their performance using pairwise subjective assessments with a compressed baseline. Results indicate a significant disparity between objective and subjective evaluations, with traditional methods like VBM4D showing limited improvements in perceptual quality, as demonstrated by a statistically non-significant increase in Mean-Opinion-Score (MOS). In contrast, the neural-based methods, CVEGAN and Ramsook, showed statistically significant improvements in subjective video quality. The findings highlight the superior capability of neural approaches to enhance perceptual quality, suggesting that current objective metrics may not fully capture quality as perceived by human observers. This study also contributes the results of the comparative analysis and the dataset to the research community. Darren Ramsook, Vibhoothi, Anil C. Kokaram, Angeliki V. Katsenou, David Bull 0001 |
QoMEX | 3 |
| 2024 | Predicting total time to compress a video corpus using online inference systemsabstractPredicting the computational cost of compressing/transcoding clips in a video corpus is important for resource management of cloud services and VOD (Video On Demand) providers. Currently, customers of cloud video services are unaware of the cost of transcoding their files until the task is completed. Previous work concentrated on predicting perclip compression time, and thus estimating the cost of video compression. In this work, we propose new Machine Learning (ML) systems which predict cost for the entire corpus instead. This is a more appropriate goal since users are not interested in per-clip cost but instead the cost for the whole corpus. In this work, we evaluate our systems with respect to two video codecs (x264, x265) and a novel high-quality video corpus. We find that the accuracy of aggregate time prediction for a video corpus is more than two times better than using per-clip predictions. Furthermore, we present an online inference framework in which we update the ML models as files are processed. A consideration of video compute overhead and appropriate choice of ML predictor for each fraction of corpus completed yields a prediction error of less than 5%. This is approximately two times better than previous work which proposed generalised predictors. Vibhoothi Vibhoothi, Anil C. Kokaram |
VCIP | 3 |
| 2023 | Learnt Deep Hyperparameter Selection in Adversarial Training for Compressed Video Enhancement with a Perceptual CriticabstractImage based Deep Feature Quality Metrics (DFQMs) have been shown to better correlate with subjective perceptual scores over traditional metrics. The fundamental focus of these DFQMs is to exploit internal representations from a large scale classification network as the metric feature space. Previously, no attention has been given to the problem of identifying which layers are most perceptually relevant. In this paper we present a new method for selecting perceptually relevant layers from such a network, based on a neuroscience interpretation of layer behaviour. The selected layers are treated as a hyperparameter to the critic network in a W-GAN. The critic uses the output from these layers in the preliminary stages to extract perceptual information. A video enhancement network is trained adversarially with this critic. Our results show that the introduction of these selected features into the critic yields up to 10% (FID) and 15% (KID) performance increase against other critic networks that do not exploit the idea of optimised feature selection. Darren Ramsook, Anil C. Kokaram |
ICIP | 2 |
| 2023 | Subjective Assessment of the Impact of a Content Adaptive Optimiser for Compressing 4K HDR Content With AV1abstractSince 2015 video dimensionality has expanded to higher spatial and temporal resolutions and a wider colour gamut. This High Dynamic Range (HDR) content has gained traction in the consumer space as it delivers an enhanced quality of experience. At the same time, the complexity of codecs is growing. This has driven the development of tools for content-adaptive optimisation that achieve optimal rate-distortion performance for HDR video at 4K resolution. While improvements of just a few percentage points in BD-Rate (1-5%) are significant for the streaming media industry, the impact on subjective quality has been less studied especially for HDR/AV1. In this paper, we conduct a subjective quality assessment (42 subjects) of 4K HDR content with a per-clip optimisation strategy. We correlate these subjective scores with existing popular objective metrics used in standard development and show that some perceptual metrics correlate surprisingly well even though they are not tuned for HDR. We find that the DSQCS protocol is too insensitive to categorically compare the methods but the data allows us to make recommendations about the use of experts vs non-experts in HDR studies, and explain the subjective impact of film grain in HDR content under compression. Vibhoothi, Angeliki V. Katsenou, François Pitié, Katarina Domijan, Anil C. Kokaram |
ICIP | 5 |
| 2023 | Comparison of HDR quality metrics in Per-Clip Lagrangian multiplier optimisation with AV1abstractThe complexity of modern codecs along with the increased need of delivering high-quality videos at low bitrates has reinforced the idea of a per-clip tailoring of parameters for optimised rate-distortion performance. While the objective quality metrics used for Standard Dynamic Range (SDR) videos have been well studied, the transitioning of consumer displays to support High Dynamic Range (HDR) videos, poses a new challenge to rate-distortion optimisation. In this paper, we review the popular HDR metrics DeltaE100 (DE100), PSNRL100, wPSNR, and HDR-VQM. We measure the impact of employing these metrics in per-clip direct search optimisation of the rate-distortion Lagrange multiplier in AV1. We report, on 35 HDR videos, average Bjontegaard Delta Rate (BD-Rate) gains of 4.675%, 2.226%, and 7.253% in terms of DE100, PSNRL100, and HDR-VQM. We also show that the inclusion of chroma in the quality metrics has a significant impact on optimisation, which can only be partially addressed by the use of chroma offsets. Vibhoothi, François Pitié, Angeliki V. Katsenou, Yeping Su, Balu Adsumilli, Anil C. Kokaram |
ICME | 6 |
| 2023 | Recommendations for Verifying HDR Subjective Testing WorkflowsabstractOver the past few years, there has been an increase in the demand and availability of High Dynamic Range (HDR) displays and content. To ensure the production of high-quality materials, human evaluation is required. However, ascertaining whether the full playback pipeline is indeed HDR-compliant can be challenging. In this paper, we present a set of recommendations for conformance testing to validate various aspects of the testing workflow, including playback, displays, brightness, colours, and viewing environment. We assessed the effectiveness of HDR conversion techniques used in current standards development (3GPP) for making source materials. Additionally, we evaluate HDR display technologies, including OLED and LCD, using both consumer television and a reference monitor. Vibhoothi, Angeliki V. Katsenou, John Squires, François Pitié, Anil C. Kokaram |
QoMEX | 5 |
| 2022 | Impact of Video Compression on the Performance of Object Detection Systems for Surveillance ApplicationsabstractThis study examines the relationship between H.264 video compression and the performance of an object detection network (YOLOv5). We curated a set of 50 surveillance videos and annotated targets of interest (people, bikes, and vehicles). Videos were encoded at 5 quality levels using Constant Rate Factor (CRF) values in the set {22,32,37,42,47}. YOLOv5 was applied to compressed videos and detection performance was analyzed at each CRF level. Test results indicate that the detection performance is generally robust to moderate levels of compression; using a CRF value of 37 instead of 22 leads to significantly reduced bitrates/file sizes without adversely affecting detection performance. However, detection performance degrades appreciably at higher compression levels, especially in complex scenes with poor lighting and fast-moving targets. Finally, retraining YOLOv5 on compressed imagery gives up to a 1% improvement in F1score when applied to highly compressed footage. Michael O'Byrne, Mark Sugrue, Vibhoothi, Anil C. Kokaram |
AVSS | 4 |
| 2022 | An Empirical Approach for Optimising the Impact of a Preprocessor in a Transcoding PipelineabstractThe volume of User Generated Content (UGC) on the internet has exploded throughout the pandemic. The relatively low quality of that content generally implies an increased bitrate and much reduced quality after transcoding. Preprocessing e.g. using a noise reducer, is one approach for reducing bitrate and increasing quality. The impact of the noise reducer is however affected by the target bitrate of the encoder. That relationship is known but not previously quantitatively examined. In this paper we present a methodology and new metric for measuring this impact based on the Rate-Distortion curves before and after pre-processing. The metric is used as a cost function for estimating the optimal filter parameter for our chosen denoiser. Our experiments show that optimising the filter parameter in this way yields as much as 4-5dB improvement in PSNR at 3 Mbps. Varoun Hanooman, Anil C. Kokaram, Yeping Su, Neil Birkbeck, Balu Adsumilli |
ICIP | 2 |
| 2022 | Frame-Type Sensitive RDO Control for Content-Adaptive EncodingabstractVideo transcoding is an increasingly important application in the streaming media industry. It has become important to investigate the optimisation of transcoder parameters for a single clip simply because of the immense number of playbacks for popular clips. In this paper, we explore the use of a canned optimiser to estimate the optimal Rate-Distortion (RD) tradeoff achievable for a particular clip. We show that by adjusting the Lagrange multiplier in RD optimisation on keyframes alone we can achieve more than 10× the previous BD-Rate gains possible without affecting quality for any operating point. Vibhoothi, François Pitié, Anil C. Kokaram |
ICIP | 3 |
| 2022 | A Deep Learning post-processor with a perceptual loss function for video compression artifact removalabstractWhile video compression is necessary for large scale video streaming services, compression at low bitrate can degrade the original video and negatively affect the end user’s quality of experience. Deep Neural Networks (DNNs) are actively researched with respect to artifact removal, however the loss functions that are typically employed follows a derivation of a pixel-wise Lpnorm. In this paper we consider a DNN as a post-processor for video compression artifact removal. The DNN is trained using a composite perceptual loss that combines a traditional Lpnorm loss and a VMAF proxy network based on the Video Multimethod Assessment Function (VMAF). Results show an improvement in VMAF score over both the training and testing sets. Darren Ramsook, Anil C. Kokaram, Neil Birkbeck, Yeping Su, Balu Adsumilli |
PCS | 2 |
| 2021 | CNN-Based Video Codec Classifier For Multimedia ForensicsabstractIn video forensics, identification of codec type is complicated by a lack of standards compliant compressed bitstreams. Previous work is unable to identify codec types without actually decoding the file successfully. This paper presents a CNN classifier derived from the AlexNet architecture that can detect codec types without decoding the bitstream. It is based on classification of the raw bitstream data itself without decoding. The algorithm is tested on real data in a video forensics setting as well as user generated content supplied in the YouTube test set. Our results show better than 96.73% accuracy with over 43 combinations of codec/containers and, at least, 88.59% accuracy at 20% data corruption across both test sets. Rodrigo Pessoa, Anil C. Kokaram, François Pitié, Mark Sugrue |
ICIP | 2 |
| 2021 | A differentiable estimator of VMAF for VideoabstractModern Perceptual Visual Quality Metrics (PVQMs) for video are generally complex and non-differentiable. This makes them difficult to use as loss functions in restoration and compression tuning. Traditional metrics such as PSNR/MSE which are differentiable remain important but do not capture perceptual visual criteria. In this paper we present a DNN which models a popular perceptual video metric VMAF. In so doing, we introduce a differentiable loss function that closely matches the behaviour of a perceptual metric. Employing degradation generated with H.265 compression, our model achieves a 4.41% RMSE in predicting VMAF. This can now be deployed as a video based loss function in video enhancement and compression tasks. Darren Ramsook, Anil C. Kokaram, Noel E. O'Connor, Neil Birkbeck, Yeping Su, Balu Adsumilli |
PCS | 2 |
| 2021 | Near Optimal Per-Clip Lagrangian Multiplier Prediction in HEVCabstractThe majority of internet traffic is video content. This drives the demand for video compression to deliver high quality video at low target bitrates. Optimising the parameters of a video codec for a specific video clip (per-clip optimisation) has been shown to yield significant bitrate savings. In previous work we have shown that per-clip optimisation of the Lagrangian multiplier leads to up to 24% BD-Rate improvement. A key component of these algorithms is modeling the R-D characteristic across the appropriate bitrate range. This is computationally heavy as it usually involves repeated video encodes of the high resolution material at different parameter settings. This work focuses on reducing this computational load by deploying a NN operating on lower bandwidth features. Our system achieves BD-Rate improvement in approximately 90% of a large corpus with comparable results to previous work in direct optimisation. Daniel Ringis, François Pitié, Anil C. Kokaram |
PCS | 3 |
| 2020 | A Bayesian View of Frame Interpolation and a Comparison with Existing Motion Picture Effects ToolsabstractFrame interpolation is the process of synthesising a new frame in-between existing frames in an image sequence. It has emerged as a key module in motion picture effects. Previous work either relies on two frame interpolation based entirely on optic flow, or recently DNNs. This paper presents a new algorithm based on multiframe motion interpolation motivated in a Bayesian sense. We also present the first comparison using industrial toolkits used in the post production industry today. We find that the latest Convolutional Neural Network approaches do not significantly outperform explicit motion based techniques. Anil C. Kokaram, Davinder Singh |
ICIP | 1 |
| 2020 | Motion-based frame interpolation for film and television effectsabstractFrame interpolation is the process of synthesising a new frame in‐between existing frames in an image sequence. It has emerged as a key algorithmic module in motion picture effects. In the context of this special issue, this study provides a review of the technology used to create in‐between frames and presents a Bayesian framework that generalises frame interpolation algorithms using the concept of motion interpolation. Unlike existing literature in this area, the authors also compare performance using the top industrial toolkits used in the post production industry. They find that all successful techniques employ motion‐based interpolation, and the commercial version of the Bayesian approach performs best. Another goal of this study is to compare the performance gains with recent convolutional neural network (CNN) algorithms against the traditional explicit model‐based approaches. They find that CNNs do not clearly outperform the explicit motion‐based techniques, and require significant compute resources, but provide complementary improvements in certain types of sequences. Anil C. Kokaram, Davinder Singh, Damien Kelly, Bill Collis, Kim Libreri |
IET Comput. Vis. | 1 |
| 2018 | Optimized Transcoding for Large Scale Adaptive Streaming Using Playback StatisticsabstractHTTP-based video streaming techniques have now been widely deployed to deliver video streams over communication networks. With these techniques, a video player can dynamically select a stream from a set of pre-encoded representations of the same source based on available bandwidth and viewport size. Since bitrate accounts for most of the cost incurred in these platforms, the problem is to minimize bitrate while maximizing quality. In this paper we use measurements on the actual usage of millions of video clips to create probability distributions of available bandwidth and viewport sizes. These probability distributions inform the optimization process and the resulting scheme demonstrates an overall bandwidth reduction of 9.7% compared with existing techniques without loss of delivered quality. Yao-Chung Lin, Steve Benting, Anil C. Kokaram |
ICIP | 4 |
| 2017 | A no-reference video quality predictor for compression and scaling artifactsabstractNo-Reference (NR) video quality assessment (VQA) models are gaining popularity as they offer scope for broader applicability to user-uploaded video-centric services such as YouTube and Facebook, where the pristine references are unavailable. However, there are few, well-performing NR-VQA models owing to the difficulty of the problem. We propose a novel NR video quality predictor that solely relies on the `quality-aware' natural statistical models in the space-time domain. The proposed quality predictor called Self-reference based LEarning-free Evaluator of Quality (SLEEQ) consists of three components: feature extraction in the spatial and temporal domains, motion-based feature fusion, and spatial-temporal feature pooling to derive a single quality score for a given video. SLEEQ achieves higher than 0.9 correlation with the subjective video quality scores on tested public databases and thus outperforms the existing NR VQA models. Deepti Ghadiyaram, Sasi Inguva, Anil C. Kokaram |
ICIP | 4 |
| 2017 | A software radio LTE network testbed for video quality of experience experimentationabstractThis paper presents a novel Long Term Evolution (LTE) mobile wireless testbed for streaming multimedia Quality of Experience (QoE) experimentation. The software-radio based design of the testbed provides an extremely flexible architecture which offers detailed insights into all elements of the entire LTE network protocol stack from physical to application layer. With the testbed, we aim to bridge the gap between video content providers such as YouTube and commercial LTE operators by identifying and testing viable approaches to enhance end-user quality of experience (QoE). In this work, we focus on the design and implementation of the automated and instrumented testbed, and share results from our initial experiments. Ismael Gómez Miguelez, Paul D. Sutton, Avishek Nag, Ahmed A. S. Seleim, Linda Doyle, Vivek Ramachandran, Anil C. Kokaram |
QoMEX | 7 |
| 2016 | A cloud-based large-scale distributed video analysis systemabstractDigital content consumption is exploding thanks to the advances of the distributed cloud-computing infrastructures and the consumer electronics. Further challenges have been posed to engineers and researchers to satisfy the ever increasing user needs not only for high quality video delivery, but also for richer experience. In order to support various video analysis tasks in addition to transcoding, a software platform is designed based on the Google cloud computing infrastructure with the features to be flexible, scalable, robust, and secure. In this paper, we discuss the scope, requirements, constraints, features of such a system, the problems we met and how they are resolved. Yongzhe Wang, Wei-Ta Chen, Huahui Wu, Anil C. Kokaram, Jaron Schaeffer |
ICIP | 4 |
| 2016 | A perceptual visibility metric for banding artifactsabstractBanding is a common video artifact caused by compressing low texture regions with coarse quantization. Relatively few previous attempts exist to address banding and none incorporate subjective testing for calibrating the measurement. In this paper, we propose a novel metric that incorporates both edge length and contrast across the edge to measure video banding. We further introduce both reference and non-reference metrics. Our results demonstrate that the new metrics have a very high correlation with subjective assessment and certainly outperforms PSNR, SSIM, and VQM. Yilin Wang 0001, Sang-Uok Kum, Anil C. Kokaram |
ICIP | 4 |
| 2016 | A Perceptual Quality Metric for Videos Distorted by Spatially Correlated NoiseabstractAssessing the perceptual quality of videos is critical for monitoring and optimizing video processing pipelines. In this paper, we focus on predicting the perceptual quality of videos distorted by noise. Existing video quality metrics are tuned for "white", i.e., spatially uncorrelated noise. However, white noise is very rare in real videos. Based on our analysis of the noise correlation patterns in a broad and comprehensive video set, we build a video database that simulates the commonly encountered noise characteristics. Using the database, we develop a perceptual quality assessment algorithm that explicitly incorporates the noise correlations. Experimental results show that, for videos with spatially correlated noises, the proposed algorithm presents high accuracy in predicting perceptual qualities. Mohammad Izadi, Anil C. Kokaram |
ACM Multimedia | 3 |
| 2016 | Geometry-driven quantization for omnidirectional image codingabstractIn this paper we propose a method to adapt the quantization tables of typical block-based transform codecs when the input to the encoder is a panoramic image resulting from equirectangular projection of a spherical image. When the visual content is projected from the panorama to the viewport, a frequency shift is occurring. The quantization can be adapted accordingly: the quantization step sizes that would be optimal to quantize the transform coefficients of the viewport image block, can be used to quantize the coefficients of the panoramic block. As a proof of concept, the proposed quantization strategy has been used in JPEG compression. Results show that a rate reduction up to 2.99% can be achieved for the same perceptual quality of the spherical signal with respect to a standard quantization. Francesca De Simone, Pascal Frossard, Paul Wilkins, Neil Birkbeck, Anil C. Kokaram |
PCS | 5 |
| 2016 | Bitrate classification of twice-encoded audio using objective quality featuresabstractWhen a user uploads audio files to a music streaming service, these files are subsequently re-encoded to lower bitrates to target different devices, e.g. low bitrate for mobile. To save time and bandwidth uploading files, some users encode their original files using a lossy codec. The metadata for these files cannot always be trusted as users might have encoded their files more than once. Determining the lowest bitrate of the files allows the streaming service to skip the process of encoding the files to bitrates higher than that of the uploaded files, saving on processing and storage space. This paper presents a model that uses quality predictions from ViSQOLAudio, a full reference objective audio quality metric, as features in combination with a multi-class support vector machine classifier. An experiment on twice-encoded files found that low bitrate codecs could be classified using audio quality features. The experiment also provides insights into the implications of multiple transcodes from a quality perspective. Colm Sloan, Naomi Harte, Damien Kelly, Anil C. Kokaram, Andrew Hines |
QoMEX | 4 |
| 2016 | Double-Tip Artifact Removal From Atomic Force Microscopy ImagesabstractThe Atomic Force Microscope (AFM) allows the measurement of interactions at interfaces with nanoscale resolution. Imperfections in the shape of the tip often lead to the presence of imaging artefacts such as the blurring and repetition of objects within images. Generally, these artefacts can only be avoided by discarding data and replacing the probe. Under certain circumstances (e.g., rare, high value samples, or extensive chemical/physical tip modification) such an approach is not feasible. Here, we apply a novel deblurring technique, using a Bayesian framework, to yield a reliable estimation of the real surface topography without any prior knowledge of the tip geometry (blind reconstruction). A key contribution is to leverage the significant recently successful body of work in natural image deblurring to solve this problem. We focus specifically on the 'double-tip' effect, where two asperities 1 are present on the tip, each contributing to the image formation mechanism. Finally, we demonstrate that the proposed technique successfully removes the 'double-tip' effect from high resolution AFM images which demonstrate this artefact whilst preserving feature resolution. Yun-feng Wang, Jason I. Kilpatrick, Suzanne P. Jarvis, Francis M. Boland, Anil C. Kokaram, David Corrigan |
IEEE Trans. Image Process. | 5 |
| 2015 | Multipass encoding for reducing pulsing artifacts in cloud based video transcodingabstractLow latency video transcoding is an important feature in video sharing platforms. Typically this is achieved by splitting a clip into short segments, followed by parallel encoding of the segments. However, this introduces quality artifacts due to transients in encoder rate control. We present a model for more effectively predicting the behavior of the encoder in each encoding pass, which minimizes this problem. We learn the model using measurements from more than 500 video clips, hence ensuring reliable estimation. Finally, we show that the proposed multipass strategy delivers more stable reconstructed quality for parallel video transcoding frameworks. Yao-Chung Lin, Hugh Denman, Anil C. Kokaram |
ICIP | 3 |
| 2014 | Temporal synchronization of multiple audio signalsabstractGiven the proliferation of consumer media recording devices, events often give rise to a large number of recordings. These recordings are taken from different spatial positions and do not have reliable timestamp information. In this paper, we present two robust graph-based approaches for synchronizing multiple audio signals. The graphs are constructed atop the over-determined system resulting from pairwise signal comparison using cross-correlation of audio features. The first approach uses a Minimum Spanning Tree (MST) technique, while the second uses Belief Propagation (BP) to solve the system. Both approaches can provide excellent solutions and robustness to pairwise outliers, however the MST approach is much less complex than BP. In addition, an experimental comparison of audio features-based synchronization shows that spectral flatness outperforms the zero-crossing rate and signal energy. Julius Kammerl, Neil Birkbeck, Sasi Inguva, Damien Kelly, Andrew J. Crawford, Hugh Denman, Anil C. Kokaram, Caroline Pantofaru |
ICASSP | 7 |
| 2014 | Automated registration of low and high resolution atomic force microscopy images using scale invariant featuresabstractThis paper introduces a method for registering scans acquired by Atomic Force Microscopy (AFM). Due to compromises between scan size, resolution, and scan rate, high resolution data is only attainable in a very limited field of view. The proposed method uses a sparse set of feature matches between the low and high resolution AFM scans and maps them onto a common coordinate system. This can provide a wider field of view of the sample and give context to the regions where high resolution AFM data has been obtained. The algorithm employs a robust approach overcoming complications due to temporal sample changes and sample drift of the AFM system which becomes significant at higher-resolutions. To our knowledge, this is the first approach for automatic high resolution AFM image registration. Experimental results show the correctness and robustness of our approach and shows that the estimated transforms can be used to deduce plausible measures of sample drift. Yun-feng Wang, Jason I. Kilpatrick, Suzanne P. Jarvis, Francis M. Boland, Anil C. Kokaram, David Corrigan |
ICIP | 5 |
| 2014 | Perceived Audio Quality for Streaming Stereo MusicabstractUsers of audio-visual streaming services expect an ever increasing quality of experience. Channel bandwidth remains a bottleneck commonly addressed with lossy compression schemes for both the video and audio streams. Anecdotal evidence suggests a strongly perceived link between bit rate and quality. This paper presents three audio quality listening experiments using the ITU MUSHRA methodology to assess a number of audio codecs typically used by streaming services. They were assessed for a range of bit rates using three presentation modes: consumer and studio quality headphones and loudspeakers. Our results indicate that with consumer quality headphones, listeners were not differentiating between codecs with bit rates greater than 48 kb/s (p>=0.228). For studio quality headphones and loudspeakers aac-lc at 128 kb/s and higher was differentiated over other codecs (p<=0.001). The results provide insights into quality of experience that will guide future development of objective audio quality metrics. Andrew Hines, Eoin Gillen, Damien Kelly, Jan Skoglund, Anil C. Kokaram, Naomi Harte |
ACM Multimedia | 5 |
| 2013 | A Non-parametric Framework for Document Bleed-through RemovalabstractThis paper presents recent work on a new framework for non-blind document bleed-through removal. The framework includes image preprocessing to remove local intensity variations, pixel region classification based on a segmentation of the joint recto-verso intensity histogram and connected component analysis on the subsequent image labelling. Finally restoration of the degraded regions is performed using exemplar-based image in painting. The proposed method is evaluated visually and numerically on a freely available database of 25 scanned manuscript image pairs with ground truth, and is shown to outperform recent non-blind bleed-through removal techniques. Róisín Rowley-Brooke, François Pitié, Anil C. Kokaram |
CVPR | 3 |
| 2013 | Robustness of speech quality metrics to background noise and network degradations: Comparing ViSQOL, PESQ and POLQAabstractThe Virtual Speech Quality Objective Listener (ViSQOL) is a new objective speech quality model. It is a signal based full reference metric that uses a spectro-temporal measure of similarity between a reference and a test speech signal. ViSQOL aims to predict the overall quality of experience for the end listener whether the cause of speech quality degradation is due to ambient noise, or transmission channel degradations. This paper describes the algorithm and tests the model using two speech corpora: NOIZEUS and E4. The NOIZEUS corpus contains speech under a variety of background noise types, speech enhancement methods, and SNR levels. The E4 corpus contains voice over IP degradations including packet loss, jitter and clock drift. The results are compared with the ITU-T objective models for speech quality: PESQ and POLQA. The behaviour of the metrics are also evaluated under simulated time warp conditions. The results show that for both datasets ViSQOL performed comparably with PESQ. POLQA was shown to have lower correlation with subjective scores than the other metrics for the NOIZEUS database. Andrew Hines, Jan Skoglund, Anil C. Kokaram, Naomi Harte |
ICASSP | 3 |
| 2013 | Monitoring the effects of temporal clipping on voIP speech qualityabstractThis paper presents work on a real-time temporal clipping monitoring tool for VoIP. Temporal clipping can occur as a result of voice activity detection (VAD) or echo cancellation where comfort noise in used in place of clipped speech segments. The algorithm presented will form part of a no-reference objective model for quantifying perceived speech quality in VoIP. The overall approach uses a modular design that will help pinpoint the reason for degradations in addition to quantifying their impact on speech quality. The new algorithm was tested for VAD compared over a range of thresholds and varied speech frame sizes. The results are compared to objective Mean Opinion Scores (MOS-LQO) from POLQA. The results show that the proposed algorithm can efficiently predict temporal clipping in speech and correlates well with the full reference quality predictions from POLQA. The model shows good potential for use in a real-time monitoring tool. Index Terms: temporal clipping, VAD, VoIP, POLQA 1. Andrew Hines, Jan Skoglund, Anil C. Kokaram, Naomi Harte |
INTERSPEECH | 3 |
| 2012 | A Ground Truth Bleed-Through Document Image Database
Róisín Rowley-Brooke, François Pitié, Anil C. Kokaram |
TPDL | 3 |
| 2012 | Measuring noise correlation for improved video denoisingabstractThe vast majority of previous work in noise reduction for visual media has assumed uncorrelated, white, noise sources. In practice this is almost always violated by real media. Film grain noise is never white, and this paper highlights that the same applies to almost all consumer video content. We therefore present an algorithm for measuring the spatial and temporal spectral density of noise in archived video content, be it consumer digital camera or film orginated. As an example of how this information can be used for video denoising, the spectral density is then used for spatio-temporal noise reduction in the Fourier frequency domain. Results show improved performance for noise reduction in an easily pipelined system. Anil C. Kokaram, Damien Kelly, Hugh Denman, Andrew Crawford |
ICIP | 1 |
| 2012 | Algorithms for the Digital Restoration of Torn FilmsabstractThis paper presents algorithms for the digital restoration of films damaged by tear. As well as causing local image data loss, a tear results in a noticeable relative shift in the frame between the regions at either side of the tear boundary. This paper describes a method for delineating the tear boundary and for correcting the displacement. This is achieved using a graph-cut segmentation framework that can be either automatic or interactive when automatic segmentation is not possible. Using temporal intensity differences to form the boundary conditions for the segmentation facilitates the robust division of the frame. The resulting segmentation map is used to calculate and correct the relative displacement using a global-motion estimation approach based on motion histograms. A high-quality restoration is obtained when a suitable missing-data treatment algorithm is used to recover any missing pixel intensities. David Corrigan, Anil C. Kokaram, Naomi Harte |
IEEE Trans. Image Process. | 2 |
| 2011 | Reflection detection in image sequencesabstractReflections in image sequences consist of several layers superimposed over each other. This phenomenon causes many image processing techniques to fail as they assume the presence of only one layer at each examined site e.g. motion estimation and object recognition. This work presents an automated technique for detecting reflections in image sequences by analyzing motion trajectories of feature points. It models reflection as regions containing two different layers moving over each other. We present a strong detector based on combining a set of weak detectors. We use novel priors, generate sparse and dense detection maps and our results show high detection rate with rejection to pathological motion and occlusion. Mohamed A. Elgharib, François Pitié, Anil C. Kokaram |
CVPR | 3 |
| 2011 | Voxel-based Viterbi Active Speaker Tracking (V-VAST) with best view selection for video lecture post-productionabstractAn automated system is presented for reducing a multi-view lecture recording into a single view video containing a best view summary of active speakers. The system uses skin color detection and voxel-based analysis in locating likely speaker locations. Using time-delay estimates from multiple micro phones, speech activity is analyzed for each speaker position. The Viterbi algorithm is then used to estimate a track of the active speaker which maximizes the observed speech activity. This novel approach is termed Voxel-based Viterbi Active Speaker Tracking (V-VAST) and is shown to track speakers with an accuracy of 0.23m. Using the tracking information, the system then extracts from the available camera views the most frontal face view of the active speaker to display. Damien Kelly, Anil C. Kokaram, Francis M. Boland |
ICASSP | 2 |
| 2011 | Cellsnake: A new active contour technique for cell/fibre segmentationabstractActive contours are a well known segmentation toolkit and widely adopted in various forms for biological image analysis. Most techniques are commonly based on object geometry but overlapping regions cause severe problems to contour propagation. In this paper, we propose a novel active contour technique (“cellsnake”) for solving this problem with an application to cell and fibre segmentation. Given that the transparency of overlapped objects is unavailable, we present a new set of contour forces derived from a-priori knowledge of cell geometry that allows the contour to deform correctly in those regions. We combine these terms with other existing forces and we show that cellsnake gives appropriate shape estimation of the objects especially in areas ahowing overlapping. Kangyu Pan, Anil C. Kokaram, Kerry Gilmore, Michael J. Higgins, Robert Kapsa, Gordon G. Wallace |
ICIP | 2 |
| 2010 | Semi-automatic motion based segmentation using long term motion trajectoriesabstractSemi-automated object segmentation is an important step in the cinema post-production workflow. We propose a dense motion based segmentation process that employs sparse feature based trajectories estimated across a long sequence of frames, articulated with a Bayesian framework. The algorithm first classifies the sparse trajectories into sparsely defined objects. Then the sparse object trajectories together with motion model side information are used to generate a dense object segmentation of each video frame. Unlike previous work, we do not use the sparse trajectories only to propose motion models, but instead use their position and motion throughout the sequence as part of the classification of pixels in the second step. Furthermore, we introduce novel colour and motion priors that employ the sparse trajectories to make explicit the spatiotemporal smoothness constraints important for long term motion segmentation. Gary Baugh, Anil C. Kokaram |
ICIP | 2 |
| 2010 | Gaussian mixture models for spots in microscopy using a new split/merge em algorithmabstractIn confocal microscopy imaging, target objects are labeled with fluorescent markers in the living specimen, and usually appear as spots in the observed images. Spot detection and analysis is therefore an important task but it is still heavily reliant on manual analysis. In this paper, a novel shape modeling algorithm is proposed for automating the detection and analysis of the spots of interest. The algorithm exploits a Gaussian mixture model to characterize the spatial intensity distribution of the spots, and estimates parameters using a novel split-and-merge expectation maximization (SMEM) algorithm. In previous work the split step is random which is an issue for biological analysis where repeatability is important. The new split/merge steps are deterministic, hence more useful, and further do not impact adversely on the optimality of the final result. Index Terms — Gaussian mixture model, split-and-merge EM algorithm, spot analysis, mRNA, shape modeling Kangyu Pan, Anil C. Kokaram, Jens Hillebrand, Mani Ramaswami |
ICIP | 2 |
| 2010 | Matting with a depth mapabstractDepth maps are becoming a readily available commodity of the stereo pipeline. We propose to make use of this new free information to improve a key step of postproduction that is matting. We extend the work of Levin et al on closed form matting to introduce two new depth-aware techniques. First we explore how depth can be used as an extra channel in the matting process. Then we see how depth can be used as a diffusion guide for matting. Our results show that both techniques can reduce the amount of time needed to pull a matte. François Pitié, Anil C. Kokaram |
ICIP | 2 |
| 2009 | Extraction of non-binary blotch mattesabstractAutomated blotch removal is important in film restoration and typically involves a detection/interpolation step. Current algorithms model the corruption as a binary mixture between the original, clean images and an opaque (dirt) field. This typically causes incomplete blotch removal that manifests as blotch haloes in reconstruction. This paper proposes a new approach by modeling the corruption as a continuous mixture between the two components and generating a solution using a Bayesian framework. We use novel priors, propose a computationally efficient scheme for implementation and our results show more complete blotch reconstruction. Mohamed A. Elgharib, François Pitié, Anil C. Kokaram |
ICIP | 3 |
| 2008 | Image inpainting with a wavelet domain Hidden Markov tree modelabstractWe present a novel technique for image inpainting, the problem of filling-in missing image parts. Image inpainting is ill-posed and we adopt a probabilistic model-based approach to regularize it. The main elements of our image model are, first, an over-complete complex-wavelet image representation, which ensures good shift invariance and directional selectivity and, second, a discrete-state/continuous-observation hidden Markov tree model for the wavelet coefficients, which captures key statistical properties of natural image wavelet responses, such as heavy-tailed histograms and persistence of large wavelet coefficients across scales. We show how these ideas can be integrated into a multi-scale generative process for natural images and present alternative deterministic and Markov chain Monte Carlo algorithms for image inpainting under this model. We demonstrate the effectiveness of the method in digitally restoring images of ancient wall-paintings. George Papandreou, Petros Maragos, Anil C. Kokaram |
ICASSP | 3 |
| 2008 | Feature-based object modelling for visual surveillanceabstractThis paper introduces a new feature-based technique for implicitly modelling objects in visual surveillance. Previous work has generally employed background subtraction and other image or motion based object segmentation schemes for the first step in identifying objects worthy of attention. Given that background subtraction is a notoriously noisy process, this paper investigates an alternative strategy by instead employing feature (SIFT [1]) clustering to characterise objects. The segmentation step is therefore performed on the sparse feature space instead of the image data itself. The paper also presents an application employing this idea for automatic detection of illegal dumping from CCTV footage. The Viterbi algorithm then allows robust tracking [2] of objects generated from the spatial clustering of these sparse foreground feature maps. Gary Baugh, Anil C. Kokaram |
ICIP | 2 |
| 2008 | Implicit spatial inference with sparse local featuresabstractThis paper introduces a novel way to leverage the implicit geometry of sparse local features (e.g. SIFT operator) for the purposes of object detection and segmentation. A two-class Bayesian scheme is used as a framework, and the likelihood is derived from the real-valued classification of machine learning algorithm Gentle AdaBoost, whose output is transformed to a probabilistic distribution using either of two models investigated; Log-Sigmoid or Bi-Gaussian. The main contribution is a novel scheme for the injection of prior contextual spatial information. This occurs on a uniquely designed Markov Random Field defined by Delaunay Tri- angulation of the feature points. Our experiments show that this framework is useful for object detection and segmentation, and we achieve good, mostly invariant results in these tasks. Deirdre O'Regan, Anil C. Kokaram |
ICIP | 2 |
| 2008 | The path assigned mean shift algorithm: A new fast mean shift implementation for colour image segmentationabstractThis paper presents a novel method for colour image segmentation derived from the mean shift theorem. When applied to colour image segmentation tasks, the path assigned mean shift algorithm performed 1.5 to 5 times faster than existing fast mean shift methods such as the hierarchical 'neighbourhood consistency' FMS Method proposed by Zhang with comparable results. The complexity of the new PAMS algorithm can be represented as O(Phi2) where Phi represents the total number of unassigned points per iteration of the algorithm. Akash Pooransingh, Cathy-Ann Radix, Anil C. Kokaram |
ICIP | 3 |
| 2007 | Automated Segmentation of Torn Frames using the Graph Cuts TechniqueabstractFilm Tear is a form of degradation in archived film and is the physical ripping of the film material. Tear causes displacement of a region of the degraded frame and the loss of image data along the boundary of tear. In [1], a restoration algorithm was proposed to correct the displacement in the frame introduced by the tear by estimating the global motion of the 2 regions either side of the tear. However, the algorithm depended on a user-defined segmentation to divide the frame. This paper presents a new fully-automated segmentation algorithm which divides affected frames along the tear. The algorithm employs the graph cuts optimisation technique and uses temporal intensity differences, rather than spatial gradient, to describe the boundary properties of the segmentation. Segmentations produced with the proposed algorithm agree well with the perceived correct segmentation. David Corrigan, Naomi Harte, Anil C. Kokaram |
ICIP (1) | 3 |
| 2007 | Multi-Scale Semi-Transparent Blotch Removal on Archived Photographs using Bayesian Matting Techniques and Visibility LawsabstractThis paper presents an automatic technique to remove semi-transparent blotches (due to moisture) from archived photographs and documents. Blotches are processed in the HSV space. While chroma components are processed using a simple texture synthesis method, the intensity component is split into an over-complete wavelet representation. In the approximation band, the blotch is modelled as an alpha matte which reduces the intensity of the image in a non-uniform yet smooth manner. The alpha matte is estimated using a Bayesian approach and its effect reversed. Wavelet details are left unchanged in the case of perfect semi-transparency or attenuated using visibility laws whenever dirt and dust cause spurious edges. Experimental results achieved on many historical photographs show the effectiveness of the proposed approach. Andrew J. Crawford, Vittoria Bruni, Anil C. Kokaram, Domenico Vitulano |
ICIP (1) | 3 |
| 2007 | Bayesian Example Based Segmentation using a Hybrid Energy ModelabstractThis paper describes a supervised segmentation algorithm which draws inspiration from recent advances in non-parametric texture synthesis. A set of example images which have been segmented a priori are used as a guide in the segmentation process. This new algorithm is built on the Bayesian framework and combines the strengths of both parametric and non-parametric modelling techniques. The suitability of the wavelet transform for texture modelling is highlighted and an outlier class condition is introduced as a means to increase the flexibility of the algorithm. Segmentation results demonstrate the potential of this new algorithm. Claire Gallagher, Anil C. Kokaram |
ICIP (2) | 2 |
| 2007 | Ten Years of Digital Visual Restoration SystemsabstractWith the growth in marketability of all forms of visual media comes the need to exploit archive material more effectively, and a requirement to guarantee picture quality regardless of the originating medium. Some kind of touch-up or restoration is rapidly becoming a "must have" in the visual production cycle. This paper reviews the development and impact of digital restoration systems in the broadcast television and film industry over the last ten years. It presents some unifying ideas as well as the relevance and challenges of defect treatment in such harsh picture environments. Anil C. Kokaram |
ICIP (4) | 1 |
| 2007 | Rotation Detection using the Curl EquationabstractRotational types of motion can often be seen in video sequences. However, not a lot of research has been done to investigate rotational motion models for use in video. Analysing this unique type of motion could be very useful. For example, if the the centre of rotation of a spinning object can be efficiently identified, extraction and tracking of it can be made easier by grouping points moving at the same radial speed. It could also improve compression by recording rotation variables. In this paper, we introduce a method for finding the centre of rotation of a rotating object and a basic approach for modelling the rotation for improved image quality. The method requires an initial block based translational motion field. Daire Lennon, Naomi Harte, Anil C. Kokaram |
ICIP (1) | 3 |
| 2007 | Online Parsing of Sports Coaching Video Through Intrinsic Motion AnalysisabstractAutomatic record and review of actions in sports training sessions is of great benefit to both coach and athlete. Many coaching sessions involve repetition of particular actions to hone technique, such as a swing from a tennis racket, golf club, or cricket bat. These actions can be defined by unique motion signatures. A method is proposed to parse the training video using motion into browseable actions. The method aims to avoid the intensive explicit computation of player silhouette and motion vector fields, allowing for a real-time, online application on standard hardware. Dan Ring, Anil C. Kokaram |
ICIP (4) | 2 |
| 2007 | Automated colour grading using colour distribution transfer
François Pitié, Anil C. Kokaram, Rozenn Dahyot |
Comput. Vis. Image Underst. | 2 |
| 2006 | Robust Global Motion Estimation From Mpeg Streamswith a Gradient Based RefinementabstractIn this paper a new robust algorithm for global motion estimation from an MPEG stream is proposed. The approach is to use a hybrid of two previous global motion estimation algorithms in order to improve the accuracy and robustness of the estimate. The first algorithm is a non-parametric technique which uses the block based motion vectors obtained from an MPEG2 stream and which also generates a coarse background segmentation of the frame. The second algorithm is a gradient based technique which estimates the parameters from the image data directly. In this work, an initial estimate for the global motion is obtained using the non-parametric approach. The parameters are then refined using the gradient based estimation and the segmentation field. This hybrid approach shows significant reduction in mean compensation error over the non-parametric approach. The algorithm is then applied to the problem of mosaicking in sports sequences in which the global motion parameters can be used to make a panorama of an entire shot David Corrigan, Anil C. Kokaram, Renan Coudray, Bernard Besserer |
ICASSP (2) | 2 |
| 2006 | Pathological Motion Detection for Robust Missing Data Treatment in Degraded Archived MediaabstractThis paper outlines an algorithm to improve the robustness of missing data treatment to pathological motion (PM). PM can cause misdiagnosis of clean image data as missing data. The proposed algorithm uses a probabilistic framework to jointly detect PM and missing data by exploiting more temporal information than is typically used for missing data detection and by exploiting the local smoothness assumption of motion fields. The results of the framework are compared to an equivalent missing data detector without PM detection and the framework is shown to prevent the misdiagnosis of missing data due to PM. David Corrigan, Naomi Harte, Anil C. Kokaram |
ICIP | 3 |
| 2006 | Discrete wavelet packet transform and ensembles of lazy and eager learners for music genre classification
Marco Grimaldi, Padraig Cunningham, Anil C. Kokaram |
Multim. Syst. | 3 |
| 2005 | N-Dimensional Probablility Density Function Transfer and its Application to Colour TransferabstractThis article proposes an original method to estimate a continuous transformation that maps one N-dimensional distribution to another. The method is iterative, non-linear, and is shown to converge. Only 1D marginal distribution is used in the estimation process, hence involving low computation costs. As an illustration this mapping is applied to color transfer between two images of different contents. The paper also serves as a central focal point for collecting together the research activity in this area and relating it to the important problem of automated color grading François Pitié, Anil C. Kokaram, Rozenn Dahyot |
ICCV | 2 |
| 2005 | Nonparametric wavelet based texture synthesisabstractThis paper presents a new algorithm for synthesising image texture. Texture synthesis is an important process in image post-production. Previous approaches can be classified as either parametric or nonparametric. Of these nonparametric approaches have achieved the most impressive results. Unfortunately, these methods generally suffer from high computational cost and difficulty in handling scale in the synthesis process. This paper introduces a new idea of using wavelet decomposition as a basis for nonparametric texture synthesis. The results show an order of magnitude improvement in computational speed and a better approximation of the dominant scale in the synthesised texture. Claire Gallagher, Anil C. Kokaram |
ICIP (2) | 2 |
| 2005 | Off-line multiple object tracking using candidate selection and the Viterbi algorithmabstractThis paper presents a probabilistic framework for off-line multiple object tracking. At each timestep, a small set of deterministic candidates is generated which is guaranteed to contain the correct solution. Tracking an object within video then becomes possible using the Viterbi algorithm. In contrast with particle filter methods where candidates are numerous and random, the proposed algorithm involves a few candidates and results in a deterministic solution. Moreover, we consider here off-line applications where past and future information is exploited. This paper shows that, although basic and very simple, this candidate selection allows the solution of many tracking problems in different real-world applications and offers a good alternative to particle filter methods for off-line applications. François Pitié, Sid-Ahmed Berrani, Anil C. Kokaram, Rozenn Dahyot |
ICIP (3) | 3 |
| 2005 | Classification and representation of semantic content in broadcast tennis videosabstractThis paper investigates the semantic analysis of broadcast tennis footage. We consider the spatio-temporal behaviour of an object in the footage as being the embodiment of a semantic event. This object is tracked using a colour based particle filter. The video syntax and audio features are used to help delineate the temporal boundaries of these events. For broadcast tennis footage, the system firstly parses the video sequence based on the geometry of the content in view and classifies the clip as a particular view type. The temporal behaviour of the serving player is modelled using a HMM. As a result, each model is representative of a particular semantic episode. Events are then summarised using a number of synthesised keyframes. Niall Rea, Rozenn Dahyot, Anil C. Kokaram |
ICIP (3) | 3 |
| 2004 | Modeling high level structure in sports with motion driven HMMsabstractIn this paper, we investigate the retrieval of dynamic events that occur in broadcast sports footage. Dynamic events in sports are important in so far as they are related to the game semantics. Thus far, the temporal interleaving of camera views has been used to infer these types of events. We propose the use of the spatio-temporal behaviour of an object in the footage as an embodiment of a semantic event. This is accomplished by modeling the evolution of the position of the object with a hidden Markov model (HMM). Snooker is used as an example for the purpose of this research. The system firstly parses the video sequence based on the geometry of the content in the camera view and classifies the footage as a particular view type. Secondly, we consider the relative position of the white ball on the snooker table over the duration of a clip to embody semantic events. A colour based particle filter is employed to robustly track the snooker balls. The temporal behaviour of the white ball is modeled using a HMM where each model is representative of a particular semantic episode. Upon collision of the white ball with another coloured ball, a separate track is instantiated. Niall Rea, Rozenn Dahyot, Anil C. Kokaram |
ICASSP (3) | 3 |
| 2004 | Automated treatment of film tear in degraded archived media
David Corrigan, Anil C. Kokaram |
ICIP | 2 |
| 2004 | Gradient based dominant motion estimation with integral projections for real time video stabilisationabstractThis paper presents a new expression of the relationship between integral projections and motion in an image pair. The resulting new multiresolution gradient based approach is used to estimate dominant motion in image sequences degraded by random shake. The paper also describes an implementation using the GPU as a coprocessor for the CPU that allows, for the first time, real time video stabilisation in software on broadcast standard definition television images. Andrew Crawford, Hugh Denman, Francis Kelly, François Pitié, Anil C. Kokaram |
ICIP | 5 |
| 2004 | Two layer segmentation for handling pathological motion in degraded post production mediaabstractThe paper presents a mechanism for dealing with incorrect motion estimation in degraded post production image sequences. This tends to be caused by pathological motion (a combination of motion blur, complex foreground motion, shadows, etc.). We describe a method for estimating where such regions are likely to occur by segmenting sequences into foreground and background motion. We show it is possible to produce a conservative matte of which regions in a sequence are foreground, and that blotch (dirt and sparkle defects) detection can be adapted accordingly in such areas. Benjamin Kent, Anil C. Kokaram, Bill Collis |
ICIP | 2 |
| 2004 | Inlier modeling for multimedia data analysisabstractThis paper presents a robust method to estimate the unknown standard deviation of a centred normal distribution from a mixture density. This method is applied to different signal processing problems. The first one concerns silence segmentation from audio data. The second application deals with colour class parameter extraction. In this later case, the mean is also estimated from the observations. Rozenn Dahyot, Niall Rea, Anil C. Kokaram, Nick G. Kingsbury |
MMSP | 3 |
| 2004 | A statistical framework for picture reconstruction using 2D AR models
Anil C. Kokaram |
Image Vis. Comput. | 1 |
| 2004 | On missing data treatment for degraded video and film archives: a survey and a new Bayesian approachabstractImage sequence restoration has been steadily gaining in importance with the increasing prevalence of visual digital media. The demand for content increases the pressure on archives to automate their restoration activities for preservation of the cultural heritage that they hold. There are many defects that affect archived visual material and one central issue is that of Dirt and Sparkle, or "Blotches." Research in archive restoration has been conducted for more than a decade and this paper places that material in context to highlight the advances made during that time. The paper also presents a new and simpler Bayesian framework that achieves joint processing of noise, missing data, and occlusion. Anil C. Kokaram |
IEEE Trans. Image Process. | 1 |
| 2003 | Joint audio visual retrieval for tennis broadcastsabstractIn recent years, there has been increasing work in the area of content retrieval for sports. The idea is generally to extract important events or create summaries to allow personalisation of the media stream. While previous work in sports analysis has employed either the audio or video stream to achieve some goal, there is little work that explores how much can be achieved by combining the two streams. This paper combines both audio and image features to identify the key episode in tennis broadcasts. The image feature is based on image moments and is able to capture the essence of scene geometry without recourse to 3D modelling. The audio feature uses PCA to identify the sound of the ball hitting the racket. The features are modelled as stochastic processes and the work combines the features using a likelihood approach. The results show that combining the features yields a much more robust system than using the features separately. Rozenn Dahyot, Anil C. Kokaram, Niall Rea, Hugh Denman |
ICASSP (3) | 2 |
| 2003 | A Bayesian framework for recursive object removal in movie post-productionabstractSome of the most convincing film and video effects are created in digital post-production by removing apparatus that supports or manipulates actors and objects. Wires and people, for instance, can be removed by digitally painting them out of the scene provided some 'clean plate' image is available for pasting in the missing regions. This paper addresses the problem when no such plate is available. Object removal requires the estimation of the motion of the hidden material and then the reconstruction of the missing image data. Using the notion of temporal motion smoothness, this paper articulates the two problems using a Bayesian framework and so develops a unique tool for automated object removal. The tool is currently being tested in the film effects industry and initial feedback is very positive. Anil C. Kokaram, Bill Collis |
ICIP (1) | 1 |
| 2003 | Sport video shot segmentation and classification
Rozenn Dahyot, Niall Rea, Anil C. Kokaram |
VCIP | 3 |
| 2003 | Content-based analysis for video from snooker broadcasts
Hugh Denman, Niall Rea, Anil C. Kokaram |
Comput. Vis. Image Underst. | 3 |
| 2002 | Parametric texture synthesis for filling holes in picturesabstractThis paper presents a framework for "filling in" missing gaps in images and particularly patches with texture. The underlying idea is to construct a parametric model of the p.d.f. of the texture to be re-synthesised and then draw samples from that p.d.f. to create the resulting reconstruction. A Bayesian approach is used to repose 2D autoregressive models as generative models for texture (using the Gibbs sampler) given surrounding boundary conditions. A fast implementation is presented that iterates between pixelwise updates and blockwise parametric model estimation. The novel ideas in this paper are joint parameter estimation and fast, efficient texture reconstruction using linear models. Anil C. Kokaram |
ICIP (1) | 1 |
| 2002 | Suppression of moire patterns via spectral analysis
Denis N. Sidorov, Anil C. Kokaram |
VCIP | 2 |
| 2002 | An improved error resilience scheme for highly compressed dataabstractThe coding of highly compressed data streams involves removing as much redundancy from the stream as possible. However, as redundancy is removed, so is the ability of the decoder to recover from error conditions caused by wireless channels, or other lossy communications links. Standard techniques for providing some protection to the stream against channel errors usually involve adding a controlled amount of redundancy back into the stream. Such redundancy might take the form of resynchronization markers, which enable the decoder to restart the decoding process from a known state, in the event of transmission errors. Transcoding schemes, i.e. algorithms which rearrange the data without altering the bitstream length, can be very successful too; the Error Resilient Entropy Coding (EREC) scheme is particularly applicable to compressed video streams. EREC can be quite easily incorporated into the stack architecture of common communications protocols. This paper presents a modification to EREC which greatly improves its ability to recover uncorrupted data in the event of errors. The new scheme, called Code-Aligned EREC (CA-EREC) achieves this by making minor, imperceptible adjustments to some of the data in the bitstream. This allows the codewords to align with the standard EREC packet boundaries. This paper demonstrates that a subtle change to the standard EREC scheme reduces code loss to an absolute minimum in the event of errors in the channel. Although the scheme presented here has a wider applicability, this paper focuses on video coding, and more specifically on the MPEG-4 video coding standard. Lorcán Mac Manus, Anil C. Kokaram |
WCNC | 2 |
| 2001 | Error-resilience in multimedia applications over ad-hoc networksabstractAd-hoc networking has been of increasing interest in recent years. It encapsulates the ultimate notion of ubiquitous communications with the absence of reliance on any existing network infrastructure. This paper presents a concept for robust operation of multimedia applications over such networks. Error resilient communication is achieved by using a new error detection and concealment technique that exploits information from the decoded image data itself as well as using information from the underlying network. This approach unifies information from both traditional computer science and signal processing domains. A layered architecture framework for the implementation of the proposed system is also described. Linda Doyle, Anil C. Kokaram, Donal O'Mahony |
ICASSP | 2 |
| 2001 | A new global motion estimation algorithm and its application to retrieval in sports eventsabstractWe propose a novel global motion estimation technique based on weighted gradient and displaced frame difference (DFD) associated with Wiener estimation. Then, we apply this technique to parse events of a high level of understanding in a cricket game. A user oriented analysis of the game then reveals a distinct connection between the global motion and specific events. By estimating global motion and analysing the temporal evolution of the estimated motion parameters, we present an effective process for the extraction of cricket events, leading to a success rate of 88.9%. Anil C. Kokaram, Perrine Delacourt |
MMSP | 1 |
| 1999 | Fast high quality interpolation of missing data in image sequences using a controlled pasting schemeabstractAn important topic in image restoration is interpolation of missing data in image sequences. Missing data is a result of dirt on film and of ageing processes where the film contents are replaced by data that bears little relationship with the original scene. We present a method for interpolating missing data with the aim of achieving higher fidelity and more consistency in the interpolated results than can be achieved by existing methods. This is done by combining autoregressive models and Markov-random field techniques. Experimental results confirm the superior performance of the proposed method over existing methods. Peter M. B. van Roosmalen, Anil C. Kokaram, Jan Biemond |
ICASSP | 2 |
| 1997 | Line registration of jittered videoabstractWith the imminent widespread availability of digital video broadcasts and the subsequent increase in the demand for broadcast material, image sequence restoration is an increasing source of concern for both archivists and broadcasters. This paper presents a two stage technique for registering the lines in video data digitized from a noisy source. In such situations the horizontal synchronization pulses may not have the correct amplitude causing the loss of 'lock' in the digitizing apparatus. The effect is that the image lines are randomly shifted horizontally with respect to their true locations. This manifests as jagged vertical edges in the observed sequence, an annoying artefact. The algorithm presented here relies on a two-dimensional autoregressive (2D AR) model of the image to measure the line displacements using a multiresolution scheme. Anil C. Kokaram, Peter J. W. Rayner, Peter M. B. van Roosmalen, Jan Biemond |
ICASSP | 1 |
| 1997 | A gradient based fast search algorithm for warping motion compensation schemesabstractThis paper describes a fast search algorithm for warping motion compensation schemes which is based on a first order approximation of a Taylor series. In comparison to a full search algorithm the technique significantly reduces the required computational load (by a factor of approximately fifty) whilst maintaining the performance in terms of prediction error. Further gains in prediction error performance are expected as the new algorithm is investigated further. David Benedict Bradshaw, Nick G. Kingsbury, Anil C. Kokaram |
ICIP (3) | 3 |
| 1997 | Joint Detection, Interpolation, Motion and Parameter Estimation forImage Sequences with Missing DataabstractThis paper presents methods for detection and reconstruction of 'missing' data in image sequences which can be modelled using 3-dimensional autoregressive (3D-AR) models. The interpolation of missing data is important in many areas of image processing, including the restoration of degraded motion pictures, reconstruction of drop-outs in digital video and automatic 're-touching' of old photographs. Here a probabilistic Bayesian framework is adopted. The method assumes no prior knowledge of the motion field or 3D-AR model parameters as these are estimated jointly with the missing image pixels. Incorporating a degradation model into the framework allows detection to proceed jointly with interpolation. Anil C. Kokaram, Simon J. Godsill |
ICIP (2) | 1 |
| 1997 | Optimal Schemes for Motion Estimation Using Colour Image SequencesabstractThis paper describes a method for incorporating the chrominance information when estimating the motion in a colour image sequence. It is based on a maximum likelihood formulation of the motion estimation problem which assumes homogeneous additive Gaussian noise in each colour component, with known inter-field correlation statistics. The formulation is applied to the complex-wavelet-domain matching algorithm of Magarey and Kingsbury (see Proc. IEEE Int. Conf. on Image Processing, p.969-72, 1996). We also define a noise-decorrelating colour space transform which provides a simple implementation of the ML formulation in the wavelet domain. Results for noisy synthesised colour sequences with known motion and noise statistics demonstrate the superiority of the exact ML formulation over straightforward, unweighted three-component estimation, most noticeably in high noise conditions. Julian Magarey, Anil C. Kokaram, Nick G. Kingsbury |
ICIP (2) | 2 |
| 1996 | A System for Reconstruction of Missing Data in Image Sequences Using Sampled 3D AR Models and MRF Motion Priors
Anil C. Kokaram, Simon J. Godsill |
ECCV (2) | 1 |
| 1996 | A sampling based approach to line scratch removal from motion picture framesabstractWe address the problem of detecting, and subsequently removing, 'line scratch' distortion in motion picture frames. A model for the lines' interaction with the image data is constructed. A sampling based algorithm based on the reversible jump Markov chain Monte Carlo framework is developed which enables automatic determination of both the unknown number of lines present, together with the lines' parameters. Previous work has not attempted to automatically determine the number of lines present. Our approach is widely applicable in many object recognition problems, where the number of objects is unknown. Robin D. Morris, William J. Fitzgerald 0001, Anil C. Kokaram |
ICIP (1) | 3 |
| 1995 | Detection of missing data in image sequencesabstractBright and dark flashes are typical artifacts in degraded motion picture material. The distortion is referred to as "dirt and sparkle" in the motion picture industry. This is caused either by dirt becoming attached to the frames of the film, or by the film material being abraded. The visual result is random patches of the frames having grey level values totally unrelated to the initial information at those sites. To restore the film without causing distortion to areas of the frames that are not affected, the locations of the blotches must be identified. Heuristic and model-based methods for the detection of these missing data regions are presented in this paper, and their action on simulated and real sequences is compared. Anil C. Kokaram, Robin D. Morris, William J. Fitzgerald 0001, Peter J. W. Rayner |
IEEE Trans. Image Process. | 1 |
| 1995 | Interpolation of missing data in image sequencesabstractThis paper presents a number of model based interpolation schemes tailored to the problem of interpolating missing regions in image sequences. These missing regions may be of arbitrary size and of random, but known, location. This problem occurs regularly with archived film material. The film is abraded or obscured in patches, giving rise to bright and dark flashes, known as "dirt and sparkle" in the motion picture industry. Both 3-D autoregressive models and 3-D Markov random fields are considered in the formulation of the different reconstruction processes. The models act along motion directions estimated using a multiresolution block matching scheme. It is possible to address this sort of impulsive noise suppression problem with median filters, and comparisons with earlier work using multilevel median filters are performed. These comparisons demonstrate the higher reconstruction fidelity of the new interpolators. Anil C. Kokaram, Robin D. Morris, William J. Fitzgerald 0001, Peter J. W. Rayner |
IEEE Trans. Image Process. | 1 |
| 1994 | Detection and Interpolation of Replacement Noise in Motion Picture Sequences Using 3D Autoregressive NodellingabstractPerhaps the most common distortion in degraded film or video material is the occurrence of blotches of varying colour randomly dispersed in each frame. The blotches are caused by the abrasion of the film or collection of dirt on the film as it passes through the projection apparatus. They manifest as flashes of bright and dark areas. The blotches represent regions of missing data (i.e. they replace the original data), and the paper discusses a model based technique for suppressing this artefact.> Anil C. Kokaram, Peter J. W. Rayner |
ISCAS | 1 |