EDBT 2026 Demo / reviewers in the wild / expert
David Bull 0001
dblp:00/202 · also Dave Bull 0001, David R. Bull
· DBLP profile ↗
281ranked-venue papers
5as first author
61since 2021 · last 2026
0000-0001-7634-190XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 230 · 2 first-author · 52 since 2021Artificial intelligence and machine learning · 26 · 11 since 2021Systems, architecture and hardware · 15 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Computer networks · 6Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Splatography: Sparse Multi-View Dynamic Gaussian Splatting for Filmmaking ChallengesabstractDeformable Gaussian Splatting (GS) accomplishes photorealistic dynamic 3-D reconstruction from dense multi-view video (MVV) by learning to deform a canonical GS representation. However, in filmmaking, tight budgets can result in sparse camera configurations, which limits state-of-the-art (SotA) methods when capturing complex dynamic features. To address this issue, we introduce an approach that splits the canonical Gaussians and deformation field into foreground and background components using a sparse set of masks for frames at$t=0$. Each representation is separately trained on different loss functions during canonical pre-training. Then, during dynamic training, different parameters are modeled for each deformation field following common filmmaking practices. The foreground stage contains diverse dynamic features so changes in color, position and rotation are learned. While, the background containing film-crew and equipment, is typically dimmer and less dynamic so only changes in point position are learned. Experiments on 3-D and 2.5-D entertainment datasets show that our method produces SotA qualitative and quantitative results; up to 3 PSNR higher with half the model size on 3-D scenes. Unlike the SotA and without the need for dense mask supervision, our method also produces segmented dynamic reconstructions including transparent and dynamic textures. Code and video comparisons are available online: http://bit.ly/4oqzZrO Adrian Azzarelli, Nantheera Anantrasirichai, David Bull 0001 |
3DV | 3 |
| 2026 | Bayesian Neural Networks for One-to-Many Mapping in Image EnhancementabstractIn image enhancement tasks, such as low-light and underwater image enhancement, a degraded image can correspond to multiple plausible target images due to dynamic photography conditions. This naturally results in a one-to-many mapping problem. To address this, we propose a Bayesian Enhancement Model (BEM) that incorporates Bayesian Neural Networks (BNNs) to capture data uncertainty and produce diverse outputs. To enable fast inference, we introduce a BNN-DNN framework: a BNN is first employed to model the one-to-many mapping in a low-dimensional space, followed by a Deterministic Neural Network (DNN) that refines fine-grained image details. Extensive experiments on multiple low-light and underwater image enhancement benchmarks demonstrate the effectiveness of our method. Guoxi Huang, Ruirui Lin, Zipeng Qi, David Bull 0001, Nantheera Anantrasirichai |
AAAI | 5 |
| 2026 | GFix: Perceptually Enhanced Gaussian Splatting Video Compressionabstract3D Gaussian Splatting (3DGS) enhances 3D scene reconstruction through explicit representation and fast rendering, demonstrating potential benefits for various low-level vision tasks, including video compression. However, existing 3DGS-based video codecs generally exhibit more noticeable visual artifacts and relatively low compression ratios. In this paper, we specifically target the perceptual enhancement of 3DGS-based video compression, based on the assumption that artifacts from 3DGS rendering and quantization resemble noisy latents sampled during diffusion training. Building on this premise, we propose a content-adaptive framework, GFix, comprising a streamlined, single-step diffusion model that serves as an off-the-shelf neural enhancer. Moreover, to increase compression efficiency, We propose a modulated LoRA scheme that freezes the low-rank decompositions and modulates the intermediate hidden states, thereby achieving efficient adaptation of the diffusion backbone with highly compressible updates. Experimental results show that GFix delivers strong perceptual quality enhancement, outperforming GSVC with up to 72.1% BD-rate savings in LPIPS and 21.4% in FID. Siyue Teng, Ge Gao 0005, Duolikun Danier, Yuxuan Jiang 0015, Fan Zhang 0017, Nantheera Anantrasirichai, Thomas Davis, Zoe Liu, David Bull 0001 |
ISCAS | 9 |
| 2026 | SAM3-LiteText: An Anatomical Study of the SAM3 Text Encoder for Efficient Vision-Language SegmentationabstractVision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we propose SAM3-LiteText, a lightweight text encoding framework that replaces the original SAM3 text encoder with a compact MobileCLIP student that is optimized by knowledge distillation. Extensive experiments on image and video segmentation benchmarks show that SAM3-LiteText reduces text encoder parameters by up to 88%, substantially reducing static memory footprint, while maintaining segmentation performance comparable to the original model. Code: https://github.com/SimonZeng7108/efficientsam3/tree/sam3_litetext. Chengxi Zeng, Yuxuan Jiang 0015, Ge Gao 0005, Shuai Wang 0054, Duolikun Danier, Bin Zhu 0006, Stevan Rudinac, David Bull 0001, Fan Zhang 0017 |
ICMR | 8 |
| 2026 | A Mamba-Based Perceptual Loss Function for Learning-Based UGC Transcoding
Zihao Qi, Chen Feng 0008, Fan Zhang 0017, Xiaozhong Xu, Shan Liu 0001, David Bull 0001 |
QoMEX | 6 |
| 2026 | FGSVQA: Frequency-Guided Short-Form Video Quality Assessment
Xinyi Wang 0011, Angeliki V. Katsenou, Junxiao Shen, David Bull 0001 |
QoMEX | 4 |
| 2026 | DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object DetectionabstractAtmospheric Turbulence (AT) degrades the clarity and accuracy of surveillance imagery, posing challenges not only for visualization quality but also for object classification and scene tracking. Deep learning-based methods have been proposed to improve visual quality, but spatio-temporal distortions remain a significant issue. Although deep learning-based object detection performs well under normal conditions, it struggles to operate effectively on sequences distorted by atmospheric turbulence. In this paper, we propose a novel framework that learns to compensate for distorted features while simultaneously improving visualization and object detection. This end-to-end training strategy leverages and exchanges knowledge of low-level distorted features in the AT mitigator with semantic features extracted in the object detector. Specifically, in the AT mitigator a 3D Mamba-based structure is used to handle the spatio-temporal displacements and blurring caused by turbulence. Optimization is achieved through back-propagation in both the AT mitigator and object detector. Our proposed DMAT outperforms state-of-the-art AT mitigation and object detection systems up to a 15% improvement on datasets corrupted by generated turbulence. The code is available at https://github.com/pui-nantheera/DMAT and datasets are available at https://zenodo.org/records/17673509. Paul R. Hill, Alin Achim, David Bull 0001, Nantheera Anantrasirichai |
WACV | 4 |
| 2026 | CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed VideoabstractThe prevalence of user-generated content (UGC) on platforms such as YouTube and TikTok has rendered no-reference (NR) perceptual video quality assessment (VQA) vital for optimizing video delivery. Nonetheless, the characteristics of non-professional acquisition and the subsequent transcoding of UGC video on sharing platforms present significant challenges for NR-VQA. Although NR-VQA models attempt to infer mean opinion scores (MOS), their modeling of subjective scores for compressed content remains limited due to the absence of fine-grained perceptual annotations of artifact types. To address these challenges, we propose CAMP-VQA, a novel NR-VQA framework that exploits the semantic understanding capabilities of large vision-language models. Our approach introduces a quality-aware prompting mechanism that integrates video metadata (e.g., resolution, frame rate, bitrate) with key fragments extracted from inter-frame variations to guide the BLIP-2 pretraining approach in generating fine-grained quality captions. A unified architecture has been designed to model perceptual quality across three dimensions: semantic alignment, temporal characteristics, and spatial characteristics. These multimodal features are extracted and fused, then regressed to video quality scores. Extensive experiments on a wide variety of UGC datasets demonstrate that our model consistently outperforms existing NR-VQA methods, achieving improved accuracy without the need for costly manual fine-grained annotations. Our method achieves the best performance in terms of average rank and linear correlation (SRCC: 0.928, PLCC: 0.938) compared to state-of-the-art methods. The source code and trained models, along with a user-friendly demo, are available at: https://github.com/xinyiW915/CAMP-VQA. Xinyi Wang 0011, Angeliki V. Katsenou, Junxiao Shen, David Bull 0001 |
WACV | 4 |
| 2026 | ViVo: A Dataset for Human Volumetric Video Reconstruction and Compression
Adrian Azzarelli, Ge Gao 0005, Ho Man Kwan, Fan Zhang 0017, Nantheera Anantrasirichai, Oliver Moolan-Feroze, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | FCVSR: A Frequency-Aware Method for Compressed Video Super-ResolutionabstractCompressed video super-resolution (SR) aims to generate high-resolution (HR) videos from the corresponding low-resolution (LR) compressed videos. Recently, some compressed video SR methods attempt to exploit the spatio-temporal information in the frequency domain, showing great promise in super-resolution performance. However, these methods do not differentiate various frequency subbands spatially or capture the temporal frequency dynamics, potentially leading to suboptimal results. In this paper, we propose a deep frequency-based compressed video SR model (FCVSR) consisting of a motion-guided adaptive alignment (MGAA) network and a multi-frequency feature refinement (MFFR) module. Additionally, a frequency-aware contrastive loss is proposed for training FCVSR, in order to reconstruct finer spatial details. The proposed model has been evaluated on three public compressed video super-resolution datasets, with results demonstrating its effectiveness when compared to existing works in terms of super-resolution performance and complexity. Fan Zhang 0017, Feiyu Chen 0001, Shuyuan Zhu, David Bull 0001, Bing Zeng 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | PNVC: Towards Practical INR-based Video CompressionabstractNeural video compression has recently demonstrated significant potential to compete with conventional video codecs in terms of rate-quality performance. These learned video codecs are however associated with various issues related to decoding complexity (for autoencoder-based methods) and/or system delays (for implicit neural representation (INR) based models), which currently prevent them from being deployed in practical applications. In this paper, targeting a practical neural video codec, we propose a novel INR-based coding framework, PNVC, which innovatively combines autoencoder-based and overfitted solutions. Our approach benefits from several design innovations, including a new structural reparameterization-based architecture, hierarchical quality control, modulation-based entropy modeling, and scale-aware positional embedding. Supporting both low delay (LD) and random access (RA) configurations, PNVC outperforms existing INR-based codecs, achieving nearly 35%+ BD-rate savings against HEVC HM 18.0 (LD) - almost 10% more compared to one of the state-of-the-art INR-based codecs, HiNeRV and 5% more over VTM 20.0 (LD), while maintaining 20+ FPS decoding speeds for 1080p content. This represents an important step forward for INR-based video coding, moving it towards practical deployment. Ge Gao 0005, Ho Man Kwan, Fan Zhang 0017, David Bull 0001 |
AAAI | 4 |
| 2025 | HIIF: Hierarchical Encoding based Implicit Image Function for Continuous Super-resolutionabstractRecent advances in implicit neural representations (INRs) have shown significant promise in modeling visual signals for various low-vision tasks including image super-resolution (ISR). INR-based ISR methods typically learn continuous representations, providing flexibility for generating high-resolution images at any desired scale from their low-resolution counterparts. However, existing INR-based ISR methods utilize multi-layer perceptrons for parameterization in the network; this does not take account of the hierarchical structure existing in local sampling points and hence constrains the representation capability. In this paper, we propose a new Hierarchical encoding based Implicit Image Function for continuous image super-resolution, HIIF, which leverages a novel hierarchical positional encoding that enhances the local implicit representation, enabling it to capture fine details at multiple scales. Our approach also embeds a multi-head linear attention mechanism within the implicit attention network by taking additional non-local information into account. Our experiments show that, when integrated with different backbone encoders, HIIF outperforms the state-of-the-art continuous image super-resolution methods by up to 0.17dB in PSNR. The source code of HIIF will be made publicly available at https://github.com/YuxuanJJ/HIIF. Yuxuan Jiang 0015, Ho Man Kwan, Tianhao Peng 0004, Ge Gao 0005, Fan Zhang 0017, Joel Sole, David Bull 0001 |
CVPR | 8 |
| 2025 | Multi-Scale Denoising in the Feature Space for Low-Light Instance SegmentationabstractInstance segmentation for low-light imagery remains largely unexplored due to the challenges imposed by such conditions, for example shot noise due to low photon count, color distortions and reduced contrast. In this paper, we propose an end-to-end solution to address this challenging task. Our proposed method implements weighted non-local blocks (wNLB) in the feature extractor. This integration enables an inherent denoising process at the feature level. As a result, our method eliminates the need for aligned ground truth images during training, thus supporting training on real-world low-light datasets. We introduce additional learnable weights at each layer in order to enhance the network’s adaptability to real-world noise characteristics, which affect different feature scales in different ways. Experimental results on several object detectors show that the proposed method outperforms the pre-trained networks with an Average Precision (AP) improvement of at least +7.6, with the introduction of wNLB further enhancing AP by upto +1.3. Joanne Lin, Nantheera Anantrasirichai, David Bull 0001 |
ICASSP | 3 |
| 2025 | GIViC: Generative Implicit Video Compression
Ge Gao 0005, Siyue Teng, Tianhao Peng 0004, Fan Zhang 0017, David Bull 0001 |
ICCV | 5 |
| 2025 | Blind Video Super-Resolution Based on Implicit KernelsabstractBlind video super-resolution (BVSR) is a low-level vision task which aims to generate high-resolution videos from low-resolution counterparts in unknown degradation scenarios. Existing approaches typically predict blur kernels that are spatially invariant in each video frame or even the entire video. These methods do not consider potential spatio-temporal varying degradations in videos, resulting in suboptimal BVSR performance. In this context, we propose a novel BVSR model based on Implicit Kernels, BVSR-IK, which constructs a multi-scale kernel dictionary parameterized by implicit neural representations. It also employs a newly designed recurrent Transformer to predict the coefficient weights for accurate filtering in both frame correction and feature alignment. Experimental results have demonstrated the effectiveness of the proposed BVSR-IK, when compared with four state-of-the-art BVSR models on three commonly used datasets, with BVSR-IK outperforming the second best approach, FMA-Net, by up to 0.59 dB in PSNR. Source code will be available at https://github.com/QZ1-boy/BVSR-IK. Yuxuan Jiang 0015, Shuyuan Zhu, Fan Zhang 0017, David Bull 0001, Bing Zeng 0001 |
ICCV | 5 |
| 2025 | DIVA-VQA: Detecting Inter-Frame Variations in UGC Video QualityabstractThe rapid growth of user-generated (video) content (UGC) has driven increased demand for research on no-reference (NR) perceptual video quality assessment (VQA). NR-VQA is a key component for large-scale video quality monitoring in social media and streaming applications where a pristine reference is not available. This paper proposes a novel NR-VQA model based on spatio-temporal fragmentation driven by inter-frame variations. By leveraging these inter-frame differences, the model progressively analyses quality-sensitive regions at multiple levels: frames, patches, and fragmented frames. It integrates frames, fragmented residuals, and fragmented frames aligned with residuals to effectively capture global and local information. The model extracts both 2D and 3D features in order to characterize these spatio-temporal variations. Experiments conducted on five UGC datasets and against state-of-the-art models ranked our proposed method among the top 2 in terms of average rank correlation (DIVA-VQA-L: 0.898 and DIVA-VQA-B: 0.886). The improved performance is offered at a low runtime complexity, with DIVA-VQA-B ranked top and DIVA-VQA-L third on average compared to the fastest existing NR-VQA method. Code and models are publicly available at: https://github.com/xinyiW915/DIVA-VQA. Xinyi Wang 0011, Angeliki V. Katsenou, David Bull 0001 |
ICIP | 3 |
| 2025 | Enhancing HDR Video Compression based on Deep Effective Bit Depth AdaptationabstractIt is well known that high dynamic range (HDR) videos enhance immersive visual experiences compared to conventional standard dynamic range content. However, HDR content is typically more challenging to encode due to the increased detail associated with the wider dynamic range. In this work, we improve HDR compression performance using an Effective Bit Depth Adaptation approach (EBDA), which reduces the effective bit depth of the original video content before encoding and reconstructs the full bit depth using a CNN-based up-sampling method at the decoder. The up-sampling deep network is based on a new version of Multi-frame MFRNet, MF-MFRNet. This approach has been integrated into the EBDA framework with two Versatile Video Coding (VVC) reference models: VTM 16.2 and the Fraunhofer Versatile Video Encoder (VVenC 1.4.0). The proposed approach has been evaluated under the JVET HDR Common Test Conditions using the Random Access configuration. The results show evident coding gains over both the original VTM 16.2 and VVenC 1.4.0 on all JVET HDR tested sequences, with average bitrate savings of 3.1% and 4.8% based on PSNR and 7.8% and 9.6% based on VMAF against VTM and VVenC respectively. The source code of multi-frame MFRNet has been released at https://github.com/fan-aaron-zhang/MF-MFRNet. Chen Feng 0008, Zihao Qi, Duolikun Danier, Fan Zhang 0017, Xiaozhong Xu, Shan Liu 0001, David Bull 0001 |
ISCAS | 7 |
| 2025 | BVI-CR: A Multi-View Human Dataset for Volumetric Video CompressionabstractThe advances in immersive technologies and 3D reconstruction have enabled the creation of digital replicas of real-world objects and environments with fine details. These processes generate vast amounts of 3D data, requiring more efficient compression methods to satisfy the memory and bandwidth constraints associated with data storage and transmission. However, the development and validation of efficient 3D data compression methods are constrained by the lack of comprehensive and high-quality volumetric video datasets, which typically require much more effort to acquire and consume increased resources compared to 2D image and video databases. To bridge this gap, we present an open multi-view volumetric human dataset, denoted BVI-CR, which contains 18 multi-view RGB-D captures and their corresponding textured polygonal meshes, depicting a range of diverse human actions. Each video sequence contains 10 views in 1080p resolution with durations between 10-15 seconds at 30FPS. Using BVI-CR, we benchmarked three conventional and neural coordinate-based multi-view video compression methods, following the MPEG MIV Common Test Conditions, and reported their rate quality performance based on various quality metrics. The results show the great potential of neural representation based methods in volumetric video compression compared to conventional video coding methods (with an up to 38% average coding gain in PSNR). This dataset provides a development and validation platform for a variety of tasks including volumetric reconstruction, compression, and quality assessment. The database will be shared publicly at https://github.com/fan-aaron-zhang/bvi-cr. Ge Gao 0005, Adrian Azzarelli, Ho Man Kwan, Nantheera Anantrasirichai, Fan Zhang 0017, Will Andrew, Oliver Moolan-Feroze, David Bull 0001 |
ISCAS | 8 |
| 2025 | RTSR: A Real-Time Super-Resolution Model for AV1 Compressed ContentabstractSuper-resolution (SR) is a key technique for improving the visual quality of video content by increasing its spatial resolution while reconstructing fine details. SR has been employed in many applications including video streaming, where compressed low-resolution content is typically transmitted to end users and then reconstructed with a higher resolution and enhanced quality. To support real-time playback, it is important to implement fast SR models while preserving reconstruction quality; however, most existing solutions, in particular those based on complex deep neural networks, fail to do so. To address this issue, this paper proposes a low-complexity SR method, RTSR, designed to enhance the visual quality of compressed video content, focusing on resolution up-scaling from a) 360p to 1080p and from b) 540p to 4K. The proposed approach utilizes a Convolutional Neural Network (CNN)-based network architecture, which was optimized for AOMedia Video 1 (AV1SVT)-encoded content at various quantization levels based on a dual-teacher knowledge distillation method. This method was submitted to the AIM 2024 Video Super-Resolution Challenge, specifically targeting the Efficient/Mobile Real-Time Video SuperResolution competition. It achieved the best trade-off between complexity and coding performance (measured in PSNR, SSIM and VMAF) among all six submissions. The code will be available at https://github.com/YuxuanJJ/RTSR. Yuxuan Jiang 0015, Jakub Nawala, Chen Feng 0008, Fan Zhang 0017, Joel Sole, David Bull 0001 |
ISCAS | 7 |
| 2025 | Ultra-lightweight Neural Video Representation Compression
Ho Man Kwan, Tianhao Peng 0004, Ge Gao 0005, Fan Zhang 0017, Mike Nilsson, Andrew Gower, David Bull 0001 |
PCS | 7 |
| 2025 | MVAD: A Multiple Visual Artifact Detector for Video StreamingabstractVisual artifacts are often introduced into streamed video content, due to prevailing conditions during content production and delivery. Since these can degrade the quality of the user's experience, it is important to automatically and accurately detect them in order to enable effective quality measurement and enhancement. Existing detection methods often focus on a single type of artifact and/or determine the presence of an artifact through thresholding objective quality indices. Such approaches have been reported to offer inconsistent prediction performance and are also impractical for real-world applications where multiple artifacts co-exist and interact. In this paper, we propose a Multiple Visual Artifact Detector, MVAD, for video streaming which, for the first time, is able to detect multiple artifacts using a single framework that is not reliant on video quality assessment models. Our approach employs a new Artifact-aware Dynamic Feature Extractor (ADFE) to obtain artifact-relevant spatial features within each frame for multiple artifact types. The extracted features are further processed by a Recurrent Memory Vision Transformer (RMViT) module, which captures both short-term and long-term temporal information within the input video. The proposed network architecture is optimized in an end-to-end manner based on a new, large and diverse training database that is generated by simulating the video streaming pipeline and based on Adversarial Data Augmentation. This model has been evaluated on two video artifact databases, Maxwell and BVI-Artifact, and achieves consistent and improved prediction results for ten target visual artifacts when compared to seven existing single and multiple artifact detectors. The source code and training database will be available at https://chenfeng-bristol.github.io/MVAD/. Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, Alex Mackin, Andrew Collins 0007, David Bull 0001 |
WACV | 6 |
| 2025 | UW-GS: Distractor-Aware 3D Gaussian Splatting for Enhanced Underwater Scene Reconstructionabstract3D Gaussian splatting (3DGS) offers the capability to achieve real-time high quality 3D scene rendering. However, 3DGS assumes that the scene is in a clear medium environment and struggles to generate satisfactory representations in underwater scenes, where light absorption and scattering are prevalent and moving objects are involved. To overcome these, we introduce a novel Gaussian Splatting-based method, UW-GS, designed specifically for underwater applications. It introduces a color appearance that models distance-dependent color variation, employs a new physics-based density control strategy to enhance clarity for distant objects, and uses a binary motion mask to handle dynamic content. Optimized with a well-designed loss function supporting for scattering media and strengthened by pseudo-depth maps, UW-GS outperforms existing methods with PSNR gains up to 1.26dB. To fully verify the effectiveness of the model, we also developed a new underwater dataset, S-UW, with dynamic object masks. The code of UW-GS and S-UW will be available at https://github.com/WangHaoran16/UW-GS. Nantheera Anantrasirichai, Fan Zhang 0017, David Bull 0001 |
WACV | 4 |
| 2024 | LDMVFI: Video Frame Interpolation with Latent Diffusion ModelsabstractExisting works on video frame interpolation (VFI) mostly employ deep neural networks that are trained by minimizing the L1, L2, or deep feature space distance (e.g. VGG loss) between their outputs and ground-truth frames. However, recent works have shown that these metrics are poor indicators of perceptual VFI quality. Towards developing perceptually-oriented VFI methods, in this work we propose latent diffusion model-based VFI, LDMVFI. This approaches the VFI problem from a generative perspective by formulating it as a conditional generation problem. As the first effort to address VFI using latent diffusion models, we rigorously benchmark our method on common test sets used in the existing VFI literature. Our quantitative experiments and user study indicate that LDMVFI is able to interpolate video content with favorable perceptual quality compared to the state of the art, even in the high-resolution regime. Our code is available at https://github.com/danier97/LDMVFI. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
AAAI | 3 |
| 2024 | MTKD: Multi-Teacher Knowledge Distillation for Image Super-Resolution
Yuxuan Jiang 0015, Chen Feng 0008, Fan Zhang 0017, David Bull 0001 |
ECCV (39) | 4 |
| 2024 | Rate-Quality or Energy-Quality Pareto Fronts for Adaptive Video Streaming?abstractAdaptive video streaming is a key enabler for optimising the delivery of offline encoded video content. The research focus to date has been on optimisation, based solely on rate-quality curves. This paper adds an additional dimension, the energy expenditure, and explores construction of bitrate ladders based on decoding energy-quality curves rather than the conventional rate-quality curves. Pareto fronts are extracted from the rate-quality and energy-quality spaces to select optimal points. Bitrate ladders are constructed from these points using conventional rate-based rules together with a novel qualitybased approach. Evaluation on a subset of YouTube-UGC videos encoded with x. 265 shows that the energy-quality ladders reduce energy requirements by $28-31 \%$ on average at the cost of slightly higher bitrates. The results indicate that optimising based on energy-quality curves rather than rate-quality curves and using quality levels to create the rungs could potentially improve energy efficiency for a comparable quality of experience. Angeliki V. Katsenou, Xinyi Wang 0011, Daniel Schien, David Bull 0001 |
ICIP | 4 |
| 2024 | A Spatio-Temporal Aligned SUNet Model For Low-Light Video EnhancementabstractDistortions caused by low-light conditions are not only visually unpleasant but also degrade the performance of computer vision tasks. The restoration and enhancement have proven to be highly beneficial. However, there are only a limited number of enhancement methods explicitly designed for videos acquired in low-light conditions. We propose a Spatio-Temporal Aligned SUNet (STA-SUNet) model using a Swin Transformer as a backbone to capture low light video features and exploit their spatio-temporal correlations. The STA-SUNet model is trained on a novel, fully registered dataset (BVI), which comprises dynamic scenes captured under varying light conditions. It is further analysed comparatively against various other models over three test datasets. The model demonstrates superior adaptivity across all datasets, obtaining the highest PSNR and SSIM values. It is particularly effective in extreme low-light conditions, yielding fairly good visualisation results. Ruirui Lin, Nantheera Anantrasirichai, Alexandra Malyugina, David Bull 0001 |
ICIP | 4 |
| 2024 | NVRC: Neural Video Representation CompressionabstractRecent advances in implicit neural representation (INR)-based video coding have
demonstrated its potential to compete with both conventional and other learning-
based approaches. With INR methods, a neural network is trained to overfit a
video sequence, with its parameters compressed to obtain a compact representation
of the video content. However, although promising results have been achieved,
the best INR-based methods are still out-performed by the latest standard codecs,
such as VVC VTM, partially due to the simple model compression techniques
employed. In this paper, rather than focusing on representation architectures, which
is a common focus in many existing works, we propose a novel INR-based video
compression framework, Neural Video Representation Compression (NVRC),
targeting compression of the representation. Based on its novel quantization and
entropy coding approaches, NVRC is the first framework capable of optimizing an
INR-based video representation in a fully end-to-end manner for the rate-distortion
trade-off. To further minimize the additional bitrate overhead introduced by the
entropy models, NVRC also compresses all the network, quantization and entropy
model parameters hierarchically. Our experiments show that NVRC outperforms
many conventional and learning-based benchmark codecs, with a 23% average
coding gain over VVC VTM (Random Access) on the UVG dataset, measured
in PSNR. As far as we are aware, this is the first time an INR-based video codec
achieving such performance. Ho Man Kwan, Ge Gao 0005, Fan Zhang 0017, Andrew Gower, David Bull 0001 |
NeurIPS | 5 |
| 2024 | Accelerating Learnt Video Codecs with Gradient Decay and Layer-Wise DistillationabstractIn recent years, end-to-end learnt video codecs have demonstrated their potential to compete with conventional coding algorithms in term of compression efficiency. However, most learning-based video compression models are associated with high computational complexity and latency, in particular at the decoder side, which limits their deployment in practical applications. In this paper, we present a novel model-agnostic pruning scheme based on gradient decay and adaptive layer-wise distillation. Gradient decay enhances parameter exploration during sparsification whilst preventing runaway sparsity and is superior to the standard Straight-Through Estimation. The adaptive layer-wise distillation regulates the sparse training in various stages based on the distortion of intermediate features. This stage-wise design efficiently updates parameters with minimal computational overhead. The proposed approach has been applied to three popular end-to-end learnt video codecs, FVC, DCVC, and DCVC-HEM. Results confirm that our method yields up to 65% reduction in MACs and 2× speedup with less than 0.3dB drop in BD-PSNR. Supporting code and supplementary material can be downloaded from: https://jasminepp.github.io/lightweighltdvc/. Tianhao Peng 0004, Ge Gao 0005, Heming Sun, Fan Zhang 0017, David Bull 0001 |
PCS | 5 |
| 2024 | RankDVQA-Mini: Knowledge Distillation-Driven Deep Video Quality AssessmentabstractDeep learning-based video quality assessment (deep VQA) has demonstrated significant potential in surpassing conventional metrics, with promising improvements in terms of correlation with human perception. However, the practical deployment of such deep VQA models is often limited due to their high computational complexity and large memory requirements. To address this issue, we aim to significantly reduce the model size and runtime of one of the state-of-the-art deep VQA methods, RankDVQA, by employing a two-phase workflow that integrates pruning-driven model compression with multilevel knowledge distillation. The resulting lightweight full reference quality metric, RankDVQA-mini, requires less than 10% of the model parameters compared to its full version (14% in terms of FLOPs), while still retaining a quality prediction performance that is superior to most existing deep VQA methods. The source code of the RankDVQA-mini has been released at https://chenfeng-bristolgithub.io/RankDVQA-mini/ for public evaluation. Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, Benoit Vallade, Alex Mackin, David Bull 0001 |
PCS | 7 |
| 2024 | BVI-Artefact: An Artefact Detection Benchmark Dataset for Streamed VideosabstractProfessionally generated content (PGC) streamed online can contain visual artefacts that degrade the quality of user experience. These artefacts arise from different stages of the streaming pipeline, including acquisition, post-production, compression, and transmission. To better guide streaming experience enhancement, it is important to detect specific artefacts at the user end in the absence of a pristine reference. In this work, we address the lack of a comprehensive benchmark for artefact detection within streamed PGC, via the creation and validation of a large database, BVI-Artefact. Considering the ten most relevant artefact types encountered in video streaming, we collected and generated 480 video sequences, each containing various artefacts with associated binary artefact labels. Based on this new database, existing artefact detection methods are benchmarked, with results showing the challenging nature of this tasks and indicating the requirement of more reliable artefact detection methods. To facilitate further research in this area, we have made BVI-Artifact publicly available at bttps://chenfeng-bristol.github.io/BVI=Artefact/ Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, Alex Mackin, Andy Collins, David Bull 0001 |
PCS | 6 |
| 2024 | Compressing Deep Image Super-Resolution ModelsabstractDeep learning techniques have been applied in the context of image super-resolution (SR), achieving remarkable advances in terms of reconstruction performance. Existing techniques typically employ highly complex model structures which result in large model sizes and slow inference speeds. This often leads to high energy consumption and restricts their adoption for practical applications. To address this issue, this work employs a three-stage workflow for compressing deep SR models which significantly reduces their memory requirement. Restoration performance has been maintained through teacher-student knowledge distillation using a newly designed distillation loss. We have applied this approach to two popular image super-resolution networks, SwinIR and EDSR, to demonstrate its effectiveness. The resulting compact models, SwinIRmini and EDSRmini, attain an 89% and 96% reduction in both model size and floating-point operations (FLOPs) respectively, compared to their original versions. They also retain competitive super-resolution performance compared to their original models and other commonly used SR approaches. The source code and pretrained models for these two lightweight SR approaches are released at https://pikapi22.github.io/CDISM/. Yuxuan Jiang 0015, Jakub Nawala, Fan Zhang 0017, David Bull 0001 |
PCS | 4 |
| 2024 | Comparative Study of Hardware and Software Power Measurements in Video CompressionabstractThe environmental impact of video streaming services has been discussed as part of the strategies towards sustainable information and communication technologies. A first step towards that is the energy profiling and assessment of energy consumption of existing video technologies. This paper presents a comprehensive study of power measurement techniques for video encoding and decoding that is comparing the use of hardware and software power meters. An experimental methodology to ensure reliability of measurements is introduced. Key findings demonstrate the high correlation of hardware and software based energy measurements for the case of two video codecs across different spatial and temporal resolutions at a lower computational overhead. Angeliki V. Katsenou, Xinyi Wang 0011, Daniel Schien, David Bull 0001 |
PCS | 4 |
| 2024 | Immersive Video Compression Using Implicit Neural RepresentationsabstractRecent work on implicit neural representations (INRs) has evidenced their potential for efficiently representing and encoding conventional video content. In this paper we, for the first time, extend their application to immersive (multi-view) videos, by proposing MV-HiNeRV, a new INR-based immersive video codec. MV-HiNeRV is an enhanced version of a state-of-the-art INR-based video codec, HiNeRV, which was developed for single-view video compression. We have modified the model to learn a different group of feature grids for each view, and share the learnt network parameters among all views. This enables the model to effectively exploit the spatio-temporal and the inter-view redundancy that exists within multi-view videos. The proposed codec was used to compress multi-view texture and depth video sequences in the MPEG Immersive Video (MIV) Common Test Conditions, and tested against the MIV Test model (TMIV) that uses the VVenC video codec. The results demonstrate the superior performance of MV-HiNeRV, with significant coding gains (up to 72.33%) over TMIV. The implementation of MV-HiNeRV is published for further development and evaluation11https://hmkx.github.io/mv-hinerv/. Ho Man Kwan, Fan Zhang 0017, Andrew Gower, David Bull 0001 |
PCS | 4 |
| 2024 | Full-Reference Video Quality Assessment for User Generated Content TranscodingabstractUnlike video coding for professional content, the delivery pipeline of User Generated Content (UGC) involves transcoding where unpristine reference content needs to be compressed repeatedly. In this work, we observe that existing full-/no-reference quality metrics fail to accurately predict the perceptual quality difference between transcoded UGC content and the corresponding unpristine references. Therefore, they are unsuited for guiding the rate-distortion optimisation process in the transcoding process. In this context, we propose a bespoke full-reference deep video quality metric for UGC transcoding. The proposed method features a transcoding-specific weakly supervised training strategy employing a quality ranking-based Siamese structure. The proposed method is evaluated on the YouTube-UGC VP9 subset and the LIVE-Wild database, demonstrating state-of-the-art performance compared to existing VQA methods. The source code of the developed quality metric and the associated training data are available from https://zihaoq1:github/io/FRUGC/. Zihao Qi, Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, Xiaozhong Xu, Shan Liu 0001, David Bull 0001 |
PCS | 7 |
| 2024 | Comparative Analysis of Subjective Evaluations for Traditional and Neural-Based Video Enhancement TechniquesabstractThis work evaluates the effectiveness of modern video restoration methods, contrasting neural network-based techniques with traditional statistical algorithms to improve perceived video quality. Our analysis focused on three distinct methods: VBM4D, CVEGAN, and Ramsook, assessing their performance using pairwise subjective assessments with a compressed baseline. Results indicate a significant disparity between objective and subjective evaluations, with traditional methods like VBM4D showing limited improvements in perceptual quality, as demonstrated by a statistically non-significant increase in Mean-Opinion-Score (MOS). In contrast, the neural-based methods, CVEGAN and Ramsook, showed statistically significant improvements in subjective video quality. The findings highlight the superior capability of neural approaches to enhance perceptual quality, suggesting that current objective metrics may not fully capture quality as perceived by human observers. This study also contributes the results of the comparative analysis and the dataset to the research community. Darren Ramsook, Vibhoothi, Anil C. Kokaram, Angeliki V. Katsenou, David Bull 0001 |
QoMEX | 5 |
| 2024 | BVI-AOM: A New Training Dataset for Deep Video Compression OptimizationabstractDeep learning is now playing an important role in enhancing the performance of conventional hybrid video codecs. These learning-based methods typically require diverse and representative training material for optimization in order to achieve model generalization and optimal coding performance. However, existing datasets either offer limited content variability or come with restricted licensing terms constraining their use to research purposes only. To address these issues, we propose a new training dataset, named BVI-AOM, which contains 956 uncompressed sequences at various resolutions from 270p to 2160p, covering a wide range of content and texture types. The dataset comes with more flexible licensing terms and offers competitive performance when used as a training set for optimizing deep video coding tools. The experimental results demonstrate that when used as a training set to optimize two popular network architectures for two different coding tools, the proposed dataset leads to additional bitrate savings of up to 0.29 and 2.98 percentage points in terms of PSNR-Y and VMAF, respectively, compared to an existing training dataset, BVI-DVC, which has been widely used for deep video coding. The BVI-AOM dataset is available at https://github.com/fan-aaron-zhang/bvi-aom. Jakub Nawala, Yuxuan Jiang 0015, Fan Zhang 0017, Joel Sole, David Bull 0001 |
VCIP | 6 |
| 2024 | Benchmarking Conventional and Learned Video Codecs with a Low-Delay ConfigurationabstractRecent advances in video compression have seen significant coding performance improvements with the development of new standards and learning-based video codecs. However, most of these works focus on application scenarios that allow a certain amount of system delay (e.g., Random Access mode in MPEG codecs), which is not always acceptable for live delivery. This paper conducts a comparative study of state-of-the-art conventional and learned video coding methods based on a low delay configuration. Specifically, this study includes two MPEG standard codecs (H.266/VVC VTM and JVET ECM), two AOM codecs (AV1 libaom and AVM), and two recent neural video coding models (DCVC-DC and DCVC-FM). To allow a fair and meaningful comparison, the evaluation was performed on test sequences defined in the AOM and MPEG common test conditions in the YCbCr 4:2:0 color space. The evaluation results show that the JVET ECM codecs offer the best overall coding performance among all codecs tested, with a 16.1% (based on PSNR) average BD-rate saving over AOM AVM, and 11.0% over DCVC-FM. We also observed inconsistent performance with the learned video codecs, DCVC-DC and DCVC-FM, for test content with large background motions. Siyue Teng, Yuxuan Jiang 0015, Ge Gao 0005, Fan Zhang 0017, Thomas Davis, Zoe Liu, David Bull 0001 |
VCIP | 7 |
| 2024 | RankDVQA: Deep VQA based on Ranking-inspired Hybrid TrainingabstractIn recent years, deep learning techniques have shown significant potential for improving video quality assessment (VQA), achieving higher correlation with subjective opinions compared to conventional approaches. However, the development of deep VQA methods has been constrained by the limited availability of large-scale training databases and ineffective training methodologies. As a result, it is difficult for deep VQA approaches to achieve consistently superior performance and model generalization. In this context, this paper proposes new VQA methods based on a two-stage training methodology which motivates us to develop a large-scale VQA training database without employing human subjects to provide ground truth labels. This method was used to train a new transformer-based network architecture, exploiting quality ranking of different distorted sequences rather than minimizing the difference from the ground-truth quality labels. The resulting deep VQA methods (for both full reference and no reference scenarios), FR- and NR-RankDVQA, exhibit consistently higher correlation with perceptual quality compared to the state-of-the-art conventional and deep VQA methods, with average SROCC values of 0.8972 (FR) and 0.7791 (NR) over eight test sets without performing cross-validation. The source code of the proposed quality metrics and the large training database are available at https://chenfeng-bristol.github.io/RankDVQA. Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
WACV | 4 |
| 2024 | CVEGAN: A perceptually-inspired GAN for Compressed Video EnhancementabstractWe propose a new Generative Adversarial Network for Compressed Video frame quality Enhancement (CVEGAN). The CVEGAN generator benefits from the use of a novel Mul2Res block (with multiple levels of residual learning branches), an enhanced residual non-local block (ERNB) and an enhanced convolutional block attention module (ECBAM). The ERNB has also been employed in the discriminator to improve the representational capability. The training strategy has also been re-designed specifically for video compression applications, to employ a relativistic sphere GAN (ReSphereGAN) training methodology together with new perceptual loss functions. The proposed network has been fully evaluated in the context of two typical video compression enhancement tools: post-processing (PP) and spatial resolution adaptation (SRA). CVEGAN has been fully integrated into the MPEG HEVC and VVC video coding test models (HM 16.20 and VTM 7.0) and experimental results demonstrate significant coding gains (up to 28% for PP and 38% for SRA compared to the anchor) over existing state-of-the-art architectures for both coding tools across multiple datasets based on the HM 16.20. The respective gains for VTM 7.0 are up to 8.0% for PP and up to 20.3% for SRA. Fan Zhang 0017, David Bull 0001 |
Signal Process. Image Commun. | 3 |
| 2023 | ST-MFNET Mini: Knowledge Distillation-Driven Frame InterpolationabstractCurrently, one of the major challenges in deep learning-based video frame interpolation (VFI) is the large model size and high computational complexity associated with many high performance VFI approaches. In this paper, we present a distillation-based two-stage workflow for obtaining compressed VFI models which perform competitively compared to the state of the art, but with significantly reduced model size and complexity. Specifically, an optimization-based network pruning method is applied to a state of the art frame interpolation model, ST-MFNet, which suffers from large model size. The resulting network architecture achieves a 91% reduction in parameter numbers and a 35% increase in speed. The performance of the new network is further enhanced through a teacher-student knowledge distillation training process using a Laplacian distillation loss. The final low complexity model, ST-MFNet Mini, achieves a comparable performance to most existing high-complexity VFI methods, only outperformed by the original ST-MFNet. Our source code is available at https://github.com/crispianm/ST-MFNet-Mini Crispian Morris, Duolikun Danier, Fan Zhang 0017, Nantheera Anantrasirichai, David Bull 0001 |
ICIP | 5 |
| 2023 | HiNeRV: Video Compression with Hierarchical Encoding-based Neural RepresentationabstractLearning-based video compression is currently a popular research topic, offering the potential to compete with conventional standard video codecs. In this context, Implicit Neural Representations (INRs) have previously been used to represent and compress image and video content, demonstrating relatively high decoding speed compared to other methods. However, existing INR-based methods have failed to deliver rate quality performance comparable with the state of the art in video compression. This is mainly due to the simplicity of the employed network architectures, which limit their representation capability. In this paper, we propose HiNeRV, an INR that combines light weight layers with novel hierarchical positional encodings. We employs depth-wise convolutional, MLP and interpolation layers to build the deep and wide network architecture with high capacity. HiNeRV is also a unified representation encoding videos in both frames and patches at the same time, which offers higher performance and flexibility than existing methods. We further build a video codec based on HiNeRV and a refined pipeline for training, pruning and quantization that can better preserve HiNeRV's performance during lossy model compression. The proposed method has been evaluated on both UVG and MCL-JCV datasets for video compression, demonstrating significant improvement over all existing INRs baselines and competitive performance when compared to learning-based codecs (72.3\% overall bit rate saving over HNeRV and 43.4\% over DCVC on the UVG dataset, measured in PSNR). Ho Man Kwan, Ge Gao 0005, Fan Zhang 0017, Andrew Gower, David Bull 0001 |
NeurIPS | 5 |
| 2023 | A multiple-UAV architecture for autonomous media production
Ioannis Mademlis, Arturo Torres-González, Jesús Capitán, Maurizio Montagnuolo, Alberto Messina, Fulvio Negro, Cédric Le Barz, Rita Cunha, Bruno J. Guerreiro, Fan Zhang 0017, Stephen Boyle, Gregoire Guerout, Anastasios Tefas, Nikos Nikolaidis 0001, David Bull 0001, Ioannis Pitas |
Multim. Tools Appl. | 16 |
| 2023 | A topological loss function for image Denoising on a new BVI-lowlight datasetabstractAlthough image denoising algorithms have attracted significant research attention, surprisingly few have been proposed for, or evaluated on, noise from imagery acquired under real low-light conditions. Moreover, noise characteristics are often assumed to be spatially invariant, leading to edges and textures being distorted after denoising. Here, we introduce a novel topological loss function which is based on persistent homology. The method performs in the space of image patches, where topological invariants are calculated and represented in persistent diagrams. The loss function is a combination of ℓ1 or ℓ2 losses with the new persistence-based topological loss. We compare its performance across popular denoising architectures and loss functions, training the networks on our new comprehensive dataset of natural images captured in low-light conditions – BVI-LOWLIGHT. Analysis reveals that this approach outperforms existing methods, adapting well to complex structures and suppressing common artifacts. Alexandra Malyugina, Nantheera Anantrasirichai, David Bull 0001 |
Signal Process. | 3 |
| 2023 | BVI-VFI: A Video Quality Database for Video Frame InterpolationabstractVideo frame interpolation (VFI) is a fundamental research topic in video processing, which is currently attracting increased attention across the research community. While the development of more advanced VFI algorithms has been extensively researched, there remains little understanding of how humans perceive the quality of interpolated content and how well existing objective quality assessment methods perform when measuring the perceived quality. In order to narrow this research gap, we have developed a new video quality database named BVI-VFI, which contains 540 distorted sequences generated by applying five commonly used VFI algorithms to 36 diverse source videos with various spatial resolutions and frame rates. We collected more than 10,800 quality ratings for these videos through a large scale subjective study involving 189 human subjects. Based on the collected subjective scores, we further analysed the influence of VFI algorithms and frame rates on the perceptual quality of interpolated videos. Moreover, we benchmarked the performance of 33 classic and state-of-the-art objective image/video quality metrics on the new database, and demonstrated the urgent requirement for more accurate bespoke quality assessment methods for VFI. To facilitate further research in this area, we have made BVI-VFI publicly available at https://github.com/danier97/BVI-VFI-database. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame InterpolationabstractVideo frame interpolation (VFI) is currently a very active research topic, with applications spanning computer vision, post production and video encoding. VFI can be extremely challenging, particularly in sequences containing large motions, occlusions or dynamic textures, where existing approaches fail to offer perceptually robust inter-polation performance. In this context, we present a novel deep learning based VFI method, ST-MFNet, based on a Spatio-Temporal Multi-Flow architecture. ST-MFNet employs a new multi-scale multi-flow predictor to estimate many-to-one intermediate flows, which are combined with conventional one-to-one optical flows to capture both large and complex motions. In order to enhance interpolation performance for various textures, a 3D CNN is also employed to model the content dynamics over an extended temporal window. Moreover, ST-MFNet has been trained within an ST-GAN framework, which was originally developedfor texture synthesis, with the aim of further improving perceptual interpolation quality. Our approach has been comprehensively evaluated - compared with fourteen state-of-the-art VFI algorithms - clearly demonstrating that ST-MFNet consistently outperforms these benchmarks on var-ied and representative test datasets, with significant gains up to 1.09dB in PSNR for cases including large motions and dynamic textures. Our source code is available at https://github.com/danielism97/ST-MFNet. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
CVPR | 3 |
| 2022 | A Subjective Quality Study for Video Frame InterpolationabstractVideo frame interpolation (VFI) is one of the fundamental research areas in video processing and there has been extensive research on novel and enhanced interpolation algorithms. The same is not true for quality assessment of the interpolated content. In this paper, we describe a subjective quality study for VFI based on a newly developed video database, BVI-VFI. BVI-VFI contains 36 reference sequences at three different frame rates and 180 distorted videos generated using five conventional and learning based VFI algorithms. Subjective opinion scores have been collected from 60 human participants, and then employed to evaluate eight popular quality metrics, including PSNR, SSIM and LPIPS which are all commonly used for assessing VFI methods. The results indicate that none of these metrics provide acceptable correlation with the perceived quality on interpolated content, with the best-performing metric, LPIPS, offering a SROCC value below 0.6. Our findings show that there is an urgent need to develop a bespoke perceptual quality metric for VFI. The BVI-VFI dataset is publicly available and can be accessed at https://danielism97.github.io/BVI-VFI/. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
ICIP | 3 |
| 2022 | Enhancing Deformable Convolution Based Video Frame Interpolation with Coarse-To-Fine 3d CnnabstractThis paper presents a new deformable convolution-based video frame interpolation (VFI) method, using a coarse to fine 3D CNN to enhance the multi-flow prediction. This model first extracts spatio-temporal features at multiple scales using a 3D CNN, and estimates multi-flows using these features in a coarse-to-fine manner. The estimated multi-flows are then used to warp the original input frames as well as context maps, and the warped results are fused by a synthesis network to produce the final output. This VFI approach has been fully evaluated against 12 state-of-the-art VFI methods on three commonly used test databases. The results evidently show the effectiveness of the proposed method, which offers superior interpolation performance over other state of the art algorithms, with PSNR gains up to 0.19dB. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
ICIP | 3 |
| 2022 | A CNN-Based Post-Processor for Perceptually-Optimized Immersive Media CompressionabstractIn recent years, resolution adaptation based on deep neural networks has enabled significant performance gains for conventional (2D) video codecs. This paper investigates the effectiveness of spatial resolution resampling in the context of immersive content. The proposed approach reduces the spatial resolution of input multi-view videos before encoding, and reconstructs their original resolution after decoding. During the up-sampling process, an advanced CNN model is used to reduce potential re-sampling, compression, and synthesis artifacts. This work has been fully tested with the TMIV coding standard using a Versatile Video Coding (VVC) codec. The results demonstrate that the proposed method achieves significant rate-quality performance improvement for the majority of the test sequences, with an average BD-VMAF improvement of 3.07 over all sequences. Angeliki V. Katsenou, Fan Zhang 0017, David Bull 0001 |
ICIP | 3 |
| 2022 | ViSTRA3: Video Coding with Deep Parameter Adaptation and Post ProcessingabstractThis paper presents a deep learning-based video compression framework (ViSTRA3), which has been employed to generate compression results for the ISCAS 2022 Grand Challenge on Neural Network-based Video Coding. The proposed framework intelligently adapts video format parameters of the input video before encoding, subsequently employing a CNN at the decoder to restore their original format and enhance reconstruction quality. ViSTRA3 has been integrated with the H.266/VVC Test Model VTM 14.0, and evaluated under the Joint Video Exploration Team Common Test Conditions. Bjønegaard Delta (BD) measurement results show that the proposed framework consistently outperforms the original VVC VTM, with average BD-rate savings of 1.8% and 3.7% based on the assessment of PSNR and VMAF. Chen Feng 0008, Duolikun Danier, Charlie Tan, Fan Zhang 0017, David Bull 0001 |
ISCAS | 5 |
| 2022 | FloLPIPS: A Bespoke Video Quality Metric for Frame InterpolationabstractVideo frame interpolation (VFI) serves as a useful tool for many video processing applications. Recently, it has also been applied in the video compression domain for enhancing both conventional video codecs and learning-based compression architectures. While there has been an increased focus on the development of enhanced frame interpolation algorithms in recent years, the perceptual quality assessment of interpolated content remains an open field of research. In this paper, we present a bespoke full reference video quality metric for VFI, FloLPIPS, that builds on the popular perceptual image quality metric, LPIPS, which captures the perceptual degradation in extracted image feature space. In order to enhance the performance of LPIPS for evaluating interpolated content, we re-designed its spatial feature aggregation step by using the temporal distortion (through comparing optical flows) to weight the feature difference maps. Evaluated on the BVI-VFI database, which contains 180 test sequences with various frame interpolation artefacts, FloLPIPS shows superior correlation performance (with statistical significance) with subjective ground truth over 12 popular quality assessors. To facilitate further research in VFI quality assessment, our code is publicly available at https://danielism97.github.io/F1oLPIPS. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
PCS | 3 |
| 2022 | Study of compression statistics and prediction of rate-distortion curves for video texture
Angeliki V. Katsenou, Mariana Afonso, David Bull 0001 |
Signal Process. Image Commun. | 3 |
| 2022 | On the Immersive Properties of High Dynamic Range VideoabstractThis paper presents the results from two studies which used a dual-task methodology to measure an audience's experience of immersion while watching video under typical television viewing conditions. Immersion was measured while participants watched either a high dynamic range, wide color gamut video or a standard dynamic range, standard color gamut video, in high definition or ultra-high definition. Other video parameters were carefully measured and controlled. The study found that high dynamic range, wide color gamut video is significantly more immersive than standard dynamic range, standard color gamut video in the chosen configuration. However, there was no evidence of significant differences in immersion between high-definition and ultra-high-definition resolutions. Stephen J. Hinde, Katy C. Noland, Graham A. Thomas, David Bull 0001, Iain D. Gilchrist |
ACM Trans. Appl. Percept. | 4 |
| 2022 | BVI-DVC: A Training Database for Deep Video CompressionabstractDeep learning methods are increasingly being applied in the optimisation of video compression algorithms and can achieve significantly enhanced coding gains, compared to conventional approaches. Such approaches often employ Convolutional Neural Networks (CNNs) which are trained on databases with relatively limited content coverage. In this paper, a new extensive and representative video database, BVI-DVC,is presented for training CNN-based video compression systems, with specific emphasis on machine learning tools that enhance conventional coding architectures, including spatial resolution and bit depth up-sampling, post-processing and in-loop filtering. BVI-DVC contains 800 sequences at various spatial resolutions from 270p to 2160p and has been evaluated on ten existing network architectures for four different coding tools. Experimental results show that this database produces significant improvements in terms of coding gains over five existing (commonly used) image/video training databases under the same training and evaluation configurations. The overall additional coding improvements by using the proposed database for all tested coding modules and CNN architectures are up to 10.3% based on the assessment of PSNR and 8.1% based on VMAF. Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Contextual Colorization and Denoising for Low-Light Ultra High Resolution SequencesabstractLow-light image sequences generally suffer from spatiotemporal incoherent noise, flicker and blurring of moving objects. These artefacts significantly reduce visual quality and, in most cases, post-processing is needed in order to generate acceptable quality. Most state-of-the-art enhancement methods based on machine learning require ground truth data but this is not usually available for naturally captured low light sequences. We tackle these problems with an unpaired learning method that offers simultaneous colorization and denoising. Our approach is an adaptation of the CycleGAN structure. To overcome the excessive memory limitations associated with ultra high resolution content, we propose a multiscale patch-based framework, capturing both local and contextual features. Additionally, an adaptive temporal smoothing technique is employed to remove flickering artefacts. Experimental results show that our method outperforms existing approaches in terms of subjective quality and that it is robust to variations in brightness levels and noise. Nantheera Anantrasirichai, David Bull 0001 |
ICIP | 2 |
| 2021 | Enhancing VMAF through New Feature Integration and Model CombinationabstractVMAF is a machine learning based video quality assessment method, originally designed for streaming applications, which combines multiple quality metrics and video features through SVM regression. It offers higher correlation with subjective opinions compared to many conventional quality assessment methods. In this paper we propose enhancements to VMAF through the integration of new video features and alternative quality metrics (selected from a diverse pool) alongside multiple model combination. The proposed combination approach enables training on multiple databases with varying content and distortion characteristics. Our enhanced VMAF method has been evaluated on eight HD video databases, and consistently outperforms the original VMAF model (0.6.1) and other benchmark quality metrics, exhibiting higher correlation with subjective ground truth data. Fan Zhang 0017, Angeliki V. Katsenou, Christos G. Bampis, Lukas Krasula, Zhi Li 0001, David Bull 0001 |
PCS | 6 |
| 2021 | Texture-aware Video Frame InterpolationabstractTemporal interpolation has the potential to be a powerful tool for video compression. Existing methods for frame interpolation do not discriminate between video textures and generally invoke a single general model capable of interpolating a wide range of video content. However, past work on video texture analysis and synthesis has shown that different textures exhibit vastly different motion characteristics and they can be divided into three classes (static, dynamic continuous and dynamic discrete). In this work, we study the impact of video textures on video frame interpolation, and propose a novel framework where, given an interpolation algorithm, separate models are trained on different textures. Our study shows that video texture has significant impact on the performance of frame interpolation models and it is beneficial to have separate models specifically adapted to these texture classes, instead of training a single model that tries to learn generic motion. Our results demonstrate that models fine-tuned using our framework achieve, on average, a 0.3dB gain in PSNR on the test set used. Duolikun Danier, David Bull 0001 |
PCS | 2 |
| 2021 | VMAF-based Bitrate Ladder Estimation for Adaptive StreamingabstractIn HTTP Adaptive Streaming, video content is conventionally encoded by adapting its spatial resolution and quantization level to best match the prevailing network state and display characteristics. It is well known that the traditional solution, of using a fixed bitrate ladder, does not result in the highest quality of experience for the user. Hence, in this paper, we introduce a content-driven approach for estimating the bitrate ladder, based on spatio-temporal features extracted from the uncompressed content. The method implements a content-driven interpolation. It uses the extracted features to train a machine learning model to infer the curvature points of the Rate-VMAF curves in order to guide a set of initial encodings. We employ the VMAF quality metric as a means of perceptually conditioning the estimation. When compared to the generation of a reference ladder using exhaustive encoding, 76.63% the estimated ladder's Rate-VMAF points are identical to those of the reference ladder. The proposed method benefits from a significant (77.4%) reduction in the number of encodes required with only a small (1.04%) average Bj⊘ntegaard Delta Rate increase. Angeliki V. Katsenou, Fan Zhang 0017, Kyle Swanson, Mariana Afonso, Joel Sole, David Bull 0001 |
PCS | 6 |
| 2021 | A Subjective Study on Videos at Various Bit DepthsabstractBit depth adaptation, where the bit depth of a video sequence is reduced before transmission and up-sampled during display, can potentially reduce data rates with limited impact on perceptual quality. In this context, we conducted a subjective study on a UHD video database, BVI-BD, to explore the relationship between bit depth and visual quality. In this paper, three bit depth adaptation methods are investigated, including linear scaling, error diffusion, and a novel adaptive Gaussian filtering approach for up-sampling. The results from a subjective experiment indicate that above a critical bit depth, bit depth adaptation has no significant impact on perceptual quality, while reducing the amount information that is required to be transmitted. Below the critical bit depth, the more ‘advanced’ adaptation methods can be used to retain ‘good’ visual quality down to around 2 bits per color channel for the experimental setup - far lower than the common 8 bits per color channel. A selection of image quality metrics were benchmarked on the subjective data, and analysis indicates that a bespoke quality metric may be required to enable accurate bit depth adaptation. Alex Mackin, Fan Zhang 0017, David Bull 0001 |
PCS | 4 |
| 2021 | ViSTRA2: Video coding using spatial resolution and effective bit depth adaptation
Fan Zhang 0017, Mariana Afonso, David Bull 0001 |
Signal Process. Image Commun. | 3 |
| 2021 | Detecting Ground Deformation in the Built Environment Using Sparse Satellite InSAR Data With a Convolutional Neural NetworkabstractThe large volumes of Sentinel-1 data produced over Europe are being used to develop pan-national ground motion services. However, simple analysis techniques like thresholding cannot detect and classify complex deformation signals reliably making providing usable information to a broad range of nonexpert stakeholders a challenge. Here, we explore the applicability of deep learning approaches by adapting a pretrained convolutional neural network (CNN) to detect deformation in a national-scale velocity field. For our proof-of-concept, we focus on the U.K. where previously identified deformation is associated with coal-mining, ground water withdrawal, landslides, and tunneling. The sparsity of measurement points and the presence of spike noise make this a challenging application for deep learning networks, which involve calculations of the spatial convolution between images. Moreover, insufficient ground truth data exist to construct a balanced training data set, and the deformation signals are slower and more localized than in previous applications. We propose three enhancement methods to tackle these problems: 1) spatial interpolation with modified matrix completion; 2) a synthetic training data set based on the characteristics of the real U.K. velocity map; and 3) enhanced overwrapping techniques. Using velocity maps spanning 2015-2019, our framework detects several areas of coal mining subsidence, uplift due to dewatering, slate quarries, landslides, and tunnel engineering works. The results demonstrate the potential applicability of the proposed framework to the development of automated ground motion analysis systems. Nantheera Anantrasirichai, Juliet Biggs, Krisztina Kelevitz, Zahra Sadeghi, Tim J. Wright, Alin Achim, David Bull 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2021 | BVI-SynTex: A Synthetic Video Texture Dataset for Video Compression and Quality AssessmentabstractHighly textured video content is challenging to compress since the bit-rate to video quality trade-off is high and complex perceptual masking influences performance. Test datasets that cover a wide range of texture types are thus important for codec evaluation, but few exist. In order to study the properties of video texture, this paper introduces a Synthetic video Texture dataset (BVI-SynTex) that was generated using a Computer-Generated Imagery (CGI) environment. It contains 196 sequences clustered in three different texture types and offers the capability of being able to generate many versions of the same scene with different video parameters. It therefore provides a flexible basis for studying the influence of texture type and parameters on video compression and perceived video quality. A thorough validation and comparison of BVI-SynTex with similarly textured natural video content is performed. The comparisons show that BVI-SynTex exhibits a comparable coverage over the spatial and temporal domain and that it produces similar encoding statistics to real video datasets. A subset of the BVI-SynTex dataset was selected to perform a subjective evaluation of compression using the MPEG HEVC codec.The results show the impact of the content parameters to both the compression efficiency and the perceived quality. The publicly available BVI-SynTex dataset contains all source sequences, the objective and subjective analysis results, providing a valuable resource for the research community. Angeliki V. Katsenou, Goce Dimitrov, David Bull 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Gan-Based Effective Bit Depth Adaptation for Perceptual Video CompressionabstractResolution and effective bit depth (EBD) adaptation have been recently utilised in video compression to improve coding efficiency. This type of approach dynamically reduces spatial/temporal resolutions and effective bit depth at the encoder and restores the original video formats during decoding. In this paper, a convolutional neural networks (CNN) based EBD adaptation method is presented for perceptual video compression, in which the employed CNN models are trained using a generative adversarial network (GAN), with perception-based loss functions. This method was integrated into the HEVC HM 16.20 reference software and fully evaluated on test sequences from the JVET Common Test Conditions using the Random Access configuration. The results show significant coding gains achieved on all test sequences with an overall bit rate saving of 24.8% (Bjøntegaard Delta measurement) based on a perceptual quality metric, VMAF. Fan Zhang 0017, David Bull 0001 |
ICME | 3 |
| 2020 | Enhancing VVC Through Cnn-Based Post-ProcessingabstractThis paper presents a new Convolutional Neural Network (CNN) based post-processing approach for video compression, which is applied at the decoder to improve the reconstruction quality. This method has been integrated with the Versatile Video Coding Test Model (VTM) 4.0.1, and evaluated using the Random Access (RA) configuration using the Joint Video Exploration Team (JVET) Common Test Conditions (CTC). The results show coding gains on all tested sequences at various spatial resolutions over different quantisation parameter ranges, with average bit rate savings (based on Bjøntegaard Delta measurements) of 3.90% and 4.13%, when PSNR and VMAF are used as quality metrics respectively. The computational complexities of different CNN architecture variants have also been investigated. Fan Zhang 0017, Chen Feng 0008, David Bull 0001 |
ICME | 3 |
| 2019 | Defectnet: Multi-Class Fault Detection on Highly-Imbalanced DatasetsabstractAs a data-driven method, the performance of deep convolutional neural networks (CNN) relies heavily on training data. The prediction results of traditional networks give a bias toward larger classes, which tend to be the background in the semantic segmentation task. This becomes a major problem for fault detection, where the targets appear very small on the images and vary in both types and sizes. In this paper we propose a new network architecture, DefectNet, that offers multi-class (including but not limited to) defect detection on highly-imbalanced datasets. DefectNet consists of two parallel paths, which are a fully convolutional network and a dilated convolutional network to detect large and small objects respectively. We propose a hybrid loss maximising the usefulness of a dice loss and a cross entropy loss, and we also employ the leaky rectified linear unit (ReLU) to deal with rare occurrence of some targets in training batches. The prediction results show that our DefectNet outperforms state-of-the-art networks for detecting multi-class defects with the average accuracy improvement of approximately 10% on a wind turbine. Nantheera Anantrasirichai, David Bull 0001 |
ICIP | 2 |
| 2019 | A Subjective Study of Viewing Experience for Drone Videos Using Simulated ConteabstractThis paper presents subjective evaluation results on the viewing experience of simulated aerial videos shot at different drone heights. A total of fifty video sequences were generated using a simulation engine, Unreal Engine 4, for two racing scenarios and five different shot types. Twenty human viewers were then employed to participate a subjective experiment, providing their preference opinions on viewing experience of these videos. Through the subjective test, optimal parameters of UAV height have been identified for the evaluated shot types and scenarios. These will provide recommendation of default shot parameters for drone operation in autonomous shooting and flight planning. Stephen Boyle, Fan Zhang 0017, David Bull 0001 |
ICIP | 3 |
| 2019 | A Subjective Comparison of AV1 and HEVC for Adaptive Video StreamingabstractIn this paper we compare the performance of two state-of-the-art competing codecs, AV1 and HEVC, in the context of adaptive streaming. We specifically consider a Dynamic Optimizer (DO) methodology that is content-aware and selects the resolution of the video sequence after constructing the convex hull of the Rate-Quality curves of all considered resolutions. We start with an objective evaluation of the Dynamic Optimizer, based on both PSNR and VMAF quality metrics. The Rate-VMAF curves show an average of 6.3% BD-Rate gain of AV1 over HEVC, while the Rate-PSNR curves an show an average BD-Rate loss of 1.8%. We then report subjective tests which evaluate the perceived quality of the selected bitstreams generated by the two codecs. In this case it was found that, for most rate points, the difference in the perceived quality between HEVC and AV1 is not significant. Angeliki V. Katsenou, Fan Zhang 0017, Mariana Afonso, David Bull 0001 |
ICIP | 4 |
| 2019 | A Synthetic Video Dataset for Video Compression EvaluationabstractIn this paper, a new Synthetic video Texture dataset (SynTex) is introduced. It was generated using a Computer Graphics Imagery (CGI) environment and offers the capability of being able to generate many versions of the same scenes with different video parameters. This will support research in video compression enabling researchers to understand and model the relationship between video content and its coding parameters. To validate that SynTex is suitable for this purpose, firstly, typical spatio-temporal descriptors were calculated and compared against existing real video datasets with similar parameters. Then, the encoding statistics of SynTex were extracted using the HEVC reference software and compared to natural video datasets. The comparisons show that SynTex exhibits a comparable coverage over the spatial and temporal domain and it has similar encoding statistics to real video datasets. Angeliki V. Katsenou, David Bull 0001 |
ICIP | 3 |
| 2019 | A Frame Rate Conversion Method Based on a Virtual Shutter AngleabstractIn this paper a new method for frame rate conversion is presented. Utilising motion compensated frame prediction, our method mitigates the spatial distortions associated with traditional methods, while facilitating the conversion between a wider range of frame rates - useful when converting to/between legacy video formats. Using the concept of a virtual shutter angle, our method provides content providers with greater flexibility over the motion characteristic of their video sequences. We also propose a Gaussian weighting scheme which attempts to emulate a video sequence as if it was captured natively at the converted frame rate. A subjective experiment which compared our method to averaging frames at a range of frame rates, has shown that our method results in significantly higher visual quality across a range of content - especially for high down-sample factors. Alex Mackin, Fan Zhang 0017, David Bull 0001 |
ICIP | 3 |
| 2019 | Enhanced Video Compression Based on Effective Bit Depth AdaptationabstractThis paper presents a novel Convolutional Neural Network (CNN) based effective bit depth adaptation approach (EBDA-CNN) for video compression. It applies effective bit depth down-sampling before encoding and reconstructs the original bit depth using a deep CNN based up-sampling method at the decoder. The proposed approach has been integrated with the High Efficiency Video Coding reference software HM 16.20, and evaluated under the Joint Video Exploration Team Common Test Conditions using the Random Access configuration. The results show consistent coding gains on all tested sequences, with an average bitrate saving of 6.4%, based on Bjøntegaard Delta measurements using PSNR. Fan Zhang 0017, Mariana Afonso, David Bull 0001 |
ICIP | 3 |
| 2019 | Content-gnostic Bitrate Ladder Prediction for Adaptive Video StreamingabstractA challenge that many video providers face is the heterogeneity of networks and display devices for streaming, as well as dealing with a wide variety of content with different encoding performance. In the past, a fixed bit rate ladder solution based on a "fitting all" approach has been employed. However, such a content-tailored solution is highly demanding; the computational and financial cost of constructing the convex hull per video by encoding at all resolutions and quantization levels is huge. In this paper, we propose a content-gnostic approach that exploits machine learning to predict the bit rate ranges for different resolutions. This has the advantage of significantly reducing the number of encodes required. The first results, based on over 100 HEVC-encoded sequences demonstrate the potential, showing an average Bjøntegaard Delta Rate (BDRate) loss of 0.51% and an average BDPSNR loss of 0.01 dB compared to the ground truth, while significantly reducing the number of pre-encodes required when compared to two other methods (by 81%-94%). Angeliki V. Katsenou, Joel Sole, David Bull 0001 |
PCS | 3 |
| 2019 | A multi-metric approach for block-level video quality assessmentabstractDeveloping an objective video quality metric that accurately estimates perceived video quality is challenging. Developing a metric that can additionally be embedded in the rate distortion optimization process of a video codec can be even harder given that decisions have to be made locally. In this paper, we present a method for combining a number of existing state of the art objective video quality metrics at the coding block level by employing a fusion of local content features for deciding how to best utilize the chosen metrics. Our results indicate promising performance in terms of the correlation of the developed locally-acting quality metric with the overall perceived quality of the video. Miltiadis Alexios Papadopoulos, Angeliki V. Katsenou, Dimitris Agrafiotis, David Bull 0001 |
Signal Process. Image Commun. | 4 |
| 2019 | Video Compression Based on Spatio-Temporal Resolution AdaptationabstractA video compression framework based on spatio-temporal resolution adaptation (ViSTRA) is proposed, which dynamically resamples the input video spatially and temporally during encoding, based on a quantisation-resolution decision, and reconstructs the full resolution video at the decoder. Temporal upsampling is performed using frame repetition, whereas a convolutional neural network super-resolution model is employed for spatial resolution upsampling. ViSTRA has been integrated into the high efficiency video coding reference software (HM 16.14). Experimental results verified via an international challenge show significant improvements, with BD-rate gains of 15% based on PSNR and an average MOS difference of 0.5 based on subjective visual quality tests. Mariana Afonso, Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Rate-Distortion Optimization Using Adaptive Lagrange MultipliersabstractIn current standardized hybrid video encoders, the Lagrange multiplier determination model is a key component in rate-distortion optimization. This originated some 20 years ago based on an entropy-constrained high-rate approximation and experimental results obtained using an H.263 reference encoder on limited test material. In this paper, we present a comprehensive analysis of the results of a Lagrange multiplier selection experiment conducted on various video content using H.264/AVC and HEVC reference encoders. These results show that the original Lagrange multiplier selection methods, employed in both video encoders, are able to achieve optimum rate-distortion performance for I and P frames, but fail to perform well for B frames. The relationship is identified between the optimum Lagrange multipliers for B frames and distortion information obtained from the experimental results, leading to a novel Lagrange multiplier determination approach. The proposed method adaptively predicts the optimum Lagrange multiplier for B frames based on the distortion statistics of recent reconstructed frames. After integration into both H.264/AVC and HEVC reference encoders, this approach was evaluated on 36 test sequences with various resolutions and differing content types. The results show consistent bitrate savings for various hierarchical B frame configurations with minimal additional complexity. BD savings average approximately 3% when constant quantization parameter (QP) values are used for all frames, and 0.5% when non-zero QP offset values are employed for different B frame hierarchical levels. Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | A Study of High Frame Rate Video FormatsabstractHigh frame rates are acknowledged to increase the perceived quality of certain video content. However, the lack of high frame rate test content has previously restricted the scope of research in this area-especially in the context of immersive video formats. This problem has been addressed through the publication of a high frame rate video database BVI-HFR, which was captured natively at 120 fps. BVI-HFR spans a variety of scenes, motions, and colors, and is shown to be representative of BBC broadcast content. In this paper, temporal down-sampling is utilized to enable both subjective and objective comparisons across a range frame rates. A large-scale subjective experiment has demonstrated that high frame rates lead to increases in perceived quality, and that a degree of content dependence exists-notably related to camera motion. Various image and video quality metrics have been benchmarked on these subjective evaluations, and analysis shows that those which explicitly account for temporal distortions (e.g., FRQM) provide improved correlation with subjective opinions compared to generic quality metrics such as PSNR. Alex Mackin, Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Image Fusion Using Belief PropagationabstractThis paper describes the application of belief propagation methods to image fusion within a complex wavelet decomposition (the Dual Tree Complex Wavelet Transform: DT-CWT). Belief propagation within each transform subband iterates through a lattice based Bayesian belief network. This leads to precisely controlled spatial coherence of subband coefficient fusion through the definition of belief graph probabilities. This results in a significant improvement in quantitatively measured fusion performance for a large database of over 160 fusion image pairs from a range of fusion applications including remote sensing, multi-focus and multi-modal sources. Improvements in qualitative image fusion performance is also demonstrated. Paul R. Hill, David Bull 0001 |
ICASSP | 2 |
| 2018 | Atmospheric Turbulence Mitigation for Sequences with Moving Objects Using Recursive Image FusionabstractThis paper describes a new method for mitigating the effects of atmospheric distortion on observed sequences that include large moving objects. In order to provide accurate detail from objects behind the distorting layer, we solve the space-variant distortion problem using recursive image fusion based on the Dual Tree Complex Wavelet Transform (DT-CWT). The moving objects are detected and tracked using the improved Gaussian mixture models (GMM) and Kalman filtering. New fusion rules are introduced which work on the magnitudes and angles of the DT-CWT coefficients independently to achieve a sharp image and to reduce atmospheric distortion, respectively. The subjective results show that the proposed method achieves better video quality than other existing methods with competitive speed. Nantheera Anantrasirichai, Alin Achim, David Bull 0001 |
ICIP | 3 |
| 2018 | A Study of Subjective Video Quality at Various Spatial ResolutionsabstractIn this paper we present the BVI-SR video database, which contains 24 unique video sequences at a range of spatial resolutions up to UHD-1 (3840p). These sequences were used as the basis for a large-scale subjective experiment exploring the relationship between visual quality and spatial resolution when using three distinct spatial adaptation filters (including a CNN-based super-resolution method). The results demonstrate that while spatial resolution has a significant impact on mean opinion scores (MOS), no significant reduction in visual quality between UHD-1 and HD resolutions for the super-resolution method is reported. A selection of image quality metrics were benchmarked on the subjective evaluations, and analysis indicates that VIF offers the best performance. Alex Mackin, Mariana Afonso, Fan Zhang 0017, David Bull 0001 |
ICIP | 4 |
| 2018 | Perceptually-Aligned Frame Rate Selection Using Spatio-Temporal FeaturesabstractDuring recent years, the standardisation committees on video compression and broadcast formats have worked on extending practical video frame rates up to 120 frames per second. Generally, increased video frame rates have been shown to improve immersion, but at the cost of higher bit rates. Taking into consideration that the benefits of high frame rates are content dependent, a decision mechanism that recommends the appropriate frame rate for the specific content would provide benefits prior to compression and transmission. Furthermore, this decision mechanism must take account of the perceived video quality. The proposed method extracts and selects suitable spatio-temporal features and uses a supervised machine learning technique to build a model that is able to predict, with high accuracy, the lowest frame rate for which the perceived video quality is indistinguishable from that of video at the acquisition frame rate. The results show that it is a promising tool for prior to compression and delivery processing of videos, such as content-aware frame rate adaptation. Angeliki V. Katsenou, David Bull 0001 |
PCS | 3 |
| 2018 | SRQM: A Video Quality Metric for Spatial Resolution AdaptationabstractThis paper presents a full reference objective video quality metric (SRQM), which characterises the relationship between variations in spatial resolution and visual quality in the context of adaptive video formats. SRQM uses wavelet decomposition, subband combination with perceptually inspired weights, and spatial pooling, to estimate the relative quality between the frames of a high resolution reference video, and one that has been spatially adapted through a combination of down and upsampling. The BVI-SR video database is used to benchmark SRQM against five commonly-used quality metrics. The database contains 24 diverse video sequences that span a range of spatial resolutions up to UHD-1 (3840×2160). An indepth analysis demonstrates that SRQM is statistically superior to the other quality metrics for all tested adaptation filters, and all with relatively low computational complexity. Alex Mackin, Mariana Afonso, Fan Zhang 0017, David Bull 0001 |
PCS | 4 |
| 2018 | Superpixel-Level CFAR Detectors for Ship Detection in SAR ImageryabstractSynthetic aperture radar (SAR) is one of the most widely employed remote sensing modalities for large-scale monitoring of maritime activity. Ship detection in SAR images is a challenging task due to inherent speckle, discernible sea clutter, and the little exploitable shape information the targets present. Constant false alarm rate (CFAR) detectors, utilizing various sea clutter statistical models and thresholding schemes, are near ubiquitous in the literature. Very few of the proposed CFAR variants deviate from the classical CFAR topology; this letter proposes a modified topology, utilizing superpixels (SPs) in lieu of rectangular sliding windows to define CFAR guardbands and background. The aim is to achieve better target exclusion from the background band and reduced false detections. The performance of this modified SP-CFAR algorithm is demonstrated on TerraSAR-X and SENTINEL-1 images, achieving superior results in comparison to classical CFAR for various background distributions. Odysseas A. Pappas, Alin Achim, David Bull 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Fixation Prediction and Visual Priority Maps for Biped LocomotionabstractThis paper presents an analysis of the low-level features and key spatial points used by humans during locomotion over diverse types of terrain. Although, a number of methods for creating saliency maps and task-dependent approaches have been proposed to estimate the areas of an image that attract human attention, none of these can straightforwardly be applied to sequences captured during locomotion, which contain dynamic content derived from a moving viewpoint. We used a novel learning-based method for creating a visual priority map informed by human eye tracking data. Our proposed priority map is created based on two fixation types: first exploiting the observation that humans search for safe foot placement and second that they observe the edges of a path as a guide to safe traversal of the terrain. Texture features and the difference between them, observed at the region around an eye position, are employed within a support vector machine to create a visual priority map for biped locomotion. The results show that our proposed method outperforms the state-of-the-art, particularly for more complex terrains, where achieving smooth locomotion needs more attention on the traversing path. Nantheera Anantrasirichai, Katherine A. J. Daniels, Jeremy F. Burn, Iain D. Gilchrist, David Bull 0001 |
IEEE Trans. Cybern. | 5 |
| 2018 | BVI-HD: A Video Quality Database for HEVC Compressed and Texture Synthesized ContentabstractThis paper introduces a new high-definition video quality database, referred to as BVI-HD, which contains 32 reference and 384 distorted video sequences plus subjective scores. The reference material in this database was carefully selected to optimize the coverage range and distribution uniformity of five low-level video features, while the included 12 distortions, using both original high efficiency video coding (HEVC) and HEVC with synthesis mode, represent state-of-the-art approaches to compression. The range of quantization parameters included in the database for HEVC compression was determined by a subjective study, the results of which indicate that a wider range of QP values should be used than the current recommendation. The subjective opinion scores for all 384 distorted videos were collected from a total of 86 subjects, using a double stimulus test methodology. Based on these results, we compare the subjective quality between HEVC and synthesised content, and evaluate the performance of nine state-of-the-art, full-reference objective quality metrics. This database has now been made available online, representing a valuable resource to those concerned with compression performance evaluation and objective video quality assessment. Fan Zhang 0017, Felix Mercer Moss, Roland Baddeley, David Bull 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | Inpainting-based error concealment for low-delay video communicationabstractError concealment (EC) is one of the target applications of inpainting techniques. Some methods combine the estimated lost motion vectors (MVs) with the exemplar-based inpainting technique to recover the lost regions. Due to the erroneous motion vectors that might indicate a moving object as background object and vice versa, these methods are still showing visual artifacts in the recovered regions. In this paper, a concept of motion map that can be easily generated in the decoder side is introduced and it is combined with the exemplar-based inpainting technique. The proposed method introduces an adaptive search window size that trades-off the quality and complexity. Moreover, an optional blending technique is proposed to limit the spatio-temporal artifacts. Experiments show that the proposed method improves the visual quality with 5dB on average relative to the state-of-the-art inpainting-based EC method. Ahmed Aldahdooh, Marcus Barkowsky, David Bull 0001, Patrick Le Callet |
ICASSP | 3 |
| 2017 | Line detection in speckle images using Radon transform and ℓ1 regularizationabstractBoundaries and lines in medical images are important structures as they can delineate between tissue types, organs, and membranes. Although, a number of image enhancement and segmentation methods have been proposed to detect lines, none of these have considered line artefacts, which are more difficult to visualise as they are not physical structures, yet are still meaningful for clinical interpretation. This paper presents a novel method to restore lines, including line artefacts, in speckle images. We address this as a sparse estimation problem using a convex optimisation technique based on a Radon transform and sparsity regularisation (ℓ1norm). This problem divides into subproblems which are solved using the alternating direction method of multipliers, thereby achieving line detection and deconvolution simultaneously. The results for both simulated and in vivo ultrasound images show that the proposed method outperforms existing methods, in particular for detecting B-lines in lung ultrasound images, where the performance can be improved by up to 30 %. Nantheera Anantrasirichai, Marco Allinovi, Wesley Hayes, David Bull 0001, Alin Achim |
ICASSP | 4 |
| 2017 | Superpixel-guided CFAR detection of ships at sea in SAR imageryabstractSynthetic Aperture Radar (SAR) has over the years evolved to be one of the most promising remote sensing modalities for large-scale monitoring of the ocean and maritime activity. The detection of ships at sea in SAR imagery is a challenging task, as it requires the detection of small targets with little exploitable spatial information within a high resolution image. We present a novel method for the detection of ships based on superpixel segmentation and subsequent statistical characterisation, with no prior land masking. Our method acts as a bound to a CFAR detector, greatly reducing false positives. We present results on SENTINEL-1 imagery, demonstrating the detection performance of our algorithm. Odysseas A. Pappas, Alin Achim, David Bull 0001 |
ICASSP | 3 |
| 2017 | Low complexity video coding based on spatial resolution adaptationabstractIn this paper, a novel spatial resolution adaptation approach for video compression is proposed. Its ability to dynamically apply downsampling to frames exhibiting low spatial detail delivers improved rate distortion performance, together with a reduction in computational complexity of the encoding process. This method is based on an experimental investigation of the dependence between the QP threshold, which determines when to encode lower resolution frames, and the distortion obtained after downsampling/upsampling. The proposed approach is integrated with the High Efficiency Video Coding (HEVC) reference codec for intra coding, and evaluated on 15 high-resolution test sequences with varying levels of spatial detail. The results show a promising average bitrate savings of approximately 4% (B-D measurements), and significant complexity reduction (29% on average). Mariana Afonso, Fan Zhang 0017, Angeliki V. Katsenou, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 5 |
| 2017 | Blind high dynamic range image quality assessment using deep learningabstractIn this paper we propose a No-Reference Image Quality Assessment (NR-IQA) method on High Dynamic Range (HDR) images by combining deep Convolutional Neural Networks (CNNs) with saliency maps. The proposed method utilises the power of deep CNN architectures to extract quality features which can be applied cross HDR and Standard Dynamic Range (SDR) domains. To introduce human visual system to CNNs, a saliency map algorithm is used to select a subset of salient image patches to evaluate on. Our CNN-based method delivers a state-of-the-art performance in HDR NR-IQA experiment, competitive with full reference IQA methods. Sen Jia 0002, Yang Zhang 0003, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 4 |
| 2017 | Investigating the impact of high frame rates on video compressionabstractIn this paper we investigate the impact of frame rate variation on HEVC video compression, and demonstrate that high frame rates (60+ fps) can lead to increased perceptual quality, notably in high bitrate environments. In order to quantify content dependence, a novel way of partitioning video sequences into categories is proposed. Results show that rate-quality performance is improved at higher frame rates for video sequences with camera motion, whereas lower frame rates are favorable in sequences with complex motion (e.g. dynamic textures). We calculate that 60 fps and 120 fps are optimal choices of frame rate at bitrates of 3 Mbps and 7 Mbps respectively, demonstrating that increased frame rates are both feasible and desirable, given current broadcast data rates. Alex Mackin, Fan Zhang 0017, Miltiadis Alexios Papadopoulos, David Bull 0001 |
ICIP | 4 |
| 2017 | Video quality enhancement via QP adaptation based on perceptual coding mapsabstractThis paper introduces a method for adapting block quantisation parameter values in HEVC video compression based on perceptual coding maps. These maps are computed per block taking into account masking effects. Masking levels are calculated using spatial, temporal and foveation features that are extracted from the video and are stored in a perceptual coding map. The produced map drives a QP adaptation process that aims to redistribute coding bits in the frame so that the perceived quality is improved, especially at those bitrates where coding artifacts become visible (mid to high QP values). The subjective performance evaluation that was conducted showed that the proposed method can offer a measurable improvement in perceived quality relative to a constant QP approach, with Bjontegaard mean opinion scores (MOS) gains reaching almost 9% for the test sequences used. The paper additionally highlights the need for further work in order to increase gains in perceived quality and optimise parameter selection. Miltiadis Alexios Papadopoulos, Yashas Rai, Angeliki V. Katsenou, Dimitris Agrafiotis, Patrick Le Callet, David Bull 0001 |
ICIP | 6 |
| 2017 | Synthesis of fine details in B picture for dynamic texturesabstractDynamic textures are characterized with irregular motions that are often challenging for motion compensation as applied in the state of the art video codecs. Due to rapid and randomly evolving nature of such a signal, it is accompanied with very high energy in the residual. As a result, B-pictures as used in HEVC layer are relative expensive to code. This leads to an overall increase in the bitrate. Further, increasing QPoffsetworsens the quality by forcing lower rate to these B-pictures, leading to strong blurring and blocking artefacts. In this paper, we exploit Steerable Pyramid (SP) for coding pictures with tid> 2 in a downsampled format. At the decoder side, details are synthesized for these low resolution pictures by adding back the high frequencies using motion compensation from the nearest key picture followed by an inverse SP transform. The paper synthesizes details for the dynamic textures that are expensive to code. Our investigation shows up to 31% saving in bitrate, while visual quality is kept acceptable. Uday Singh Thakur, Madhukar Bhat, Max Bläser, Mathias Wien, David Bull 0001, Jens-Rainer Ohm |
ICIP | 5 |
| 2017 | A frame rate dependent video quality metric based on temporal wavelet decomposition and spatiotemporal poolingabstractThis paper presents an objective quality metric (FRQM), which characterises the relationship between variations in frame rate and perceptual video quality. The proposed method estimates the relative quality of a low frame rate video with respect to its higher frame rate counterpart, through temporal wavelet decomposition, subband combination and spatiotemporal pooling. FRQM was tested alongside six commonly used quality metrics (two of which explicitly relate frame rate variation to perceptual quality), on the publicly available BVI-HFR video database, that spans a diverse range of scenes and frame rates, up to 120fps. Results show that FRQM offers significant improvement over all other tested quality assessment methods with relatively low complexity. Fan Zhang 0017, Alex Mackin, David Bull 0001 |
ICIP | 3 |
| 2017 | Understanding video texture - A basis for video compressionabstractEncoding spatio-temporally varying textures is challenging for standardised video encoders, with significantly more bits required for textured blocks compared to non-textured blocks. It is therefore beneficial to understand video textures in terms of both their spatio-temporal characteristics and their encoding statistics in order to optimize coding modes and performance. To this end, we examine the classification of video texture based on encoder performance. For this purpose, we employ spatio-temporal features and follow a two-step feature selection process by employing unsupervised machine learning approaches across the selected feature space. Finally, supervised machine learning approaches are applied on the set of the selected features that support classification prior to encoding with up to 95.1% accuracy. The results of this study offer the potential to underpin a new informed approach to a new informed approach to codec configuration and mode selection. Angeliki V. Katsenou, Thomas Ntasios, Mariana Afonso, Dimitris Agrafiotis, David Bull 0001 |
MMSP | 5 |
| 2017 | Perceptual Image Fusion Using WaveletsabstractA perceptual image fusion method is proposed that employs explicit luminance and contrast masking models. These models are combined to give the perceptual importance of each coefficient produced by the dual-tree complex wavelet transform of each input image. This combined model of perceptual importance is used to select which coefficients are retained and furthermore to determine how to present the retained information in the most effective way. This paper is the first to give a principled approach to image fusion from a perceptual perspective. Furthermore, the proposed method is shown to give improved quantitative and qualitative results compared with previously developed methods. Paul R. Hill, Mohammed E. Al-Mualla, David Bull 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Line Detection as an Inverse Problem: Application to Lung Ultrasound ImagingabstractThis paper presents a novel method for line restoration in speckle images. We address this as a sparse estimation problem using both convex and non-convex optimization techniques based on the Radon transform and sparsity regularization. This breaks into subproblems, which are solved using the alternating direction method of multipliers, thereby achieving line detection and deconvolution simultaneously. We include an additional deblurring step in the Radon domain via a total variation blind deconvolution to enhance line visualization and to improve line recognition. We evaluate our approach on a real clinical application: the identification of B-lines in lung ultrasound images. Thus, an automatic B-line identification method is proposed, using a simple local maxima technique in the Radon transform domain, associated with known clinical definitions of line artefacts. Using all initially detected lines as a starting point, our approach then differentiates between B-lines and other lines of no clinical significance, including Z-lines and A-lines. We evaluated our techniques using as ground truth lines identified visually by clinical experts. The proposed approach achieves the best B-line detection performance as measured by the F score when a non-convex [Formula: see text] regularization is employed for both line detection and deconvolution. The F scores as well as the receiver operating characteristic (ROC) curves show that the proposed approach outperforms the state-of-the-art methods with improvements in B-line detection performance of 54%, 40%, and 33% for [Formula: see text], [Formula: see text], and [Formula: see text], respectively, and of 24% based on ROC curve evaluations. Nantheera Anantrasirichai, Wesley Hayes, Marco Allinovi, David Bull 0001, Alin Achim |
IEEE Trans. Medical Imaging | 4 |
| 2016 | Improved illumination invariant homomorphic filtering using the dual tree complex wavelet transformabstractA novel adaptation of the two dimensional Homomorphic filter is introduced using the Dual Tree Complex Wavelet Transform for improved illumination invariant processing. The Homomorphic filter is conventionally implemented within the log-Fourier domain using an isotropic high-pass filter based on the assumption that the illumination signal occupies low spatial frequencies. In this case however, low frequency structural reflectance content will be incorrectly attenuated. Our method implements the Homomorphic filter using the DT-CWT and exploits the property of cross scale persistence (for structural content) to generate a filter that retains cross scale content and therefore reduces incorrect attenuation of structural reflectance content. Paul R. Hill, Harish Bhaskar, Mohammed E. Al-Mualla, David Bull 0001 |
ICASSP | 4 |
| 2016 | An adaptive resolution rate control method for intra coding in HEVCabstractPrevious work has shown that spatial resampling can improve rate-distortion performance by providing a higher and more consistent level of video quality at low bitrates. Rate control aims to regulate the video bitrate in accordance to the bit budget. While this is a well studied problem in the single resolution case, very little progress has been made on the adaptive resolution case. In this paper we present an enhanced method of rate control for intra coding that allows the algorithm to learn from previously coded frames and make more accurate predictions, resulting in a lower average mismatch ratio. Our main contribution, however, lies in our adaptive resolution approach where the best scale factor is selected after prediction of the best Quantisation Parameter (QP). We show that our method closely conforms to the bit budget and provides a more stable bitrate. The likelihood of frame skipping is therefore reduced and a more consistent level of video quality is provided compared to standard single resolution methods. Brett Hosking, Dimitris Agrafiotis, David Bull 0001, Nick Eastern |
ICASSP | 3 |
| 2016 | Visual salience and priority estimation for locomotion using a deep convolutional neural networkabstractThis paper presents a novel method of salience and priority estimation for the human visual system during locomotion. This visual information contains dynamic content derived from a moving viewpoint. The priority map, ranking key areas on the image, is created from probabilities of gaze fixations, merged from bottom-up features and top-down control on the locomotion. Two deep convolutional neural networks (CNNs), inspired by models of the primate visual system, are employed to capture local salience features and compute probabilities. The first network operates through the foveal and peripheral areas around the eye positions. The second network obtains the importance of fixated points that have long durations or multiple visits, of which such areas need more times to process or to recheck to ensure smooth locomotion. The results show that our proposed method outperforms the state-of-the-art by up to 30 %, computed from average of four well known metrics for saliency estimation. Nantheera Anantrasirichai, Iain D. Gilchrist, David Bull 0001 |
ICIP | 3 |
| 2016 | Fixation identification for low-sample-rate mobile eye trackersabstractThis paper presents a novel method of fixation identification for mobile eye trackers. The most significant benefit of our method over the state-of-the-art is that it achieves high accuracy for low-sample-rate devices worn during locomotion. This in turn delivers higher quality datasets for further use in human behaviour research, robotics and the development of guidance aids for the visually impaired. The proposed method employs temporal characteristics of the eye positions combined with statistical visual features extracted using a deep convolutional neural network, inspired by models of the primate visual system, through the fovea and peripheral areas around the eye positions. The results show that the proposed method outperforms existing methods by up to 16 % in terms of classification accuracy. Nantheera Anantrasirichai, Iain D. Gilchrist, David Bull 0001 |
ICIP | 3 |
| 2016 | Compressive imaging using approximate message passing and a Cauchy prior in the wavelet domainabstractApproximate Message Passing (AMP) is an iterative reconstruction algorithm that performs signal denoising within a compressive sensing framework. We propose the use of heavy tailed distribution based image denoising, specifically using a Cauchy prior based Maximum A-Posteriori (MAP) estimate within a wavelet based AMP compressive sensing structure. The use of this MAP denoising algorithm provides extremely fast convergence for image based compressive sensing. The proposed method converges approximately twice as fast as the compared AMP methods whilst providing superior final MSE results over a range of measurement rates. Paul R. Hill, Adrian Basarab, Denis Kouame, David Bull 0001, Alin Achim |
ICIP | 5 |
| 2016 | The visibility of motion artifacts and their effect on motion qualityabstractThe visibility of motion artifacts in a video sequence e.g. motion blur and temporal aliasing, affects perceived motion quality. The frame rate required to render these motion artifacts imperceptible is far higher than is currently feasible or specified in current video formats. This paper investigates the perception of temporal aliasing and its associated artifacts below this frame rate, along with their influence on motion quality, with the aim of making suitable frame rate recommendations for future formats. Results show impairment in motion quality due to temporal aliasing can be tolerated to a degree, and that it may be acceptable to sample at frame rates 50% lower than those needed to eliminate perceptible temporal aliasing. Alex Mackin, Katy C. Noland, David Bull 0001 |
ICIP | 3 |
| 2016 | What's on TV: A large scale quantitative characterisation of modern broadcast video contentabstractVideo databases, used for benchmarking and evaluating the performance of new video technologies, should represent the full breadth of consumer video content. The parameterisation of video databases using low-level features has proven to be an effective way of quantifying the diversity within a database. However, without a comprehensive understanding of the importance and relative frequency and of these features in the content people actually consume, the utility of such information is limited. Here, we present a large-scale analysis of programming on BBC One and CBeebies, the most popular television channels in the United Kingdom for adults and children, respectively. Twenty video features are extracted from almost three thousand television programmes shown throughout 2015 before principal components analysis is used to identify just five factors representing the most variation. The meaning and relative significance of these five factors together with the shape of their frequency distributions represent highly valuable information for researchers wanting to model the diversity of modern consumer content in representative video databases. Felix Mercer Moss, Fan Zhang 0017, Roland Baddeley, David Bull 0001 |
ICIP | 4 |
| 2016 | An adaptive QP offset determination method for HEVCabstractThis paper investigates the effect that the QP offset value has on the coding performance of HEVC. We relate QP offset to the type of texture content present in the sequence. These then are used to develop a low-complexity adaptive QP offset selection method. This enables in-loop configuration of the QP offset parameter in a way that is content dependent and utilizes available encoding statistics. The proposed adaptive method is found to offer average bitrate reductions ranging from 1.38% for dynamic texture sequences up to 1.59% for static texture sequences relative to the QP offset used in the JCT-VC common test conditions. Miltiadis Alexios Papadopoulos, Fan Zhang 0017, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 4 |
| 2016 | HEVC enhancement using content-based local QP selectionabstractInspired by recent advances in objective video quality assessment, this paper proposes a novel, local quantisation parameter (QP) determination approach for perceptual video compression, based on the experimental results of a QP selection test. This method has been fully integrated into the High Efficiency Video Coding (HEVC) reference codec for intra coding, which predicts coding tree unit (CTU) level QPs to achieve optimised rate quality performance. The proposed approach consistently shows bitrate savings based on perceptual quality metrics and Bjontegaard delta measurements, with minimal complexity increase over the original codec. Fan Zhang 0017, David Bull 0001 |
ICIP | 2 |
| 2016 | Video texture analysis based on HEVC encoding statisticsabstractIn this paper, an extensive study of different video texture properties based on encoding statistics extracted from the HEVC HM reference software is presented. Mode selection, partitioning, motion vectors and bitrate allocation are among the statistics obtained from the encoder. For this study, a new dataset of homogeneous static and dynamic video textures, HomTex, is proposed. A comprehensive investigation of the results reveals a significant variability of coding statistics within dynamic textures, suggesting that this category should be further split into two relevant subcategories, continuous dynamic textures and discrete dynamic textures. This case is supported by an unsupervised learning approach on the statistics extracted. Finally, following the results obtained, some suggestions of improvements in video texture coding are presented. Mariana Afonso, Angeliki V. Katsenou, Fan Zhang 0017, Dimitris Agrafiotis, David Bull 0001 |
PCS | 5 |
| 2016 | Enhancement of intra-coded pictures for greater coding efficiencyabstractAllocating low bit-budgets to intra-coded pictures can often lead to high quantisation and loss of important high frequency information. Coding at lower resolutions can alleviate some of this distortion when utilising effective resampling techniques and filters that minimise the additional distortion introduced as a result of spatial resampling. However, it is not always possible to reconstruct each portion of a picture to the same or similar level of quality; within natural scenes frequency content tends to vary spatially and therefore a fixed scale factor may either cause loss of high frequencies or oversampling of low frequency content. SHVC provides many of the tools required for enhancing rate-distortion performance of intra-coded pictures but it is not used for this purpose. In this paper we analyse the performance of a mixed resolution coding approach by using inter-layer prediction within SHVC to construct intra-coded pictures as a combination of different spatial resolutions. We evaluate the performance of coding with and without an enhancement layer for varying bit-budgets and compare these results to standard HEVC to provide insight regarding the optimisation of intra-coded pictures. Brett Hosking, Dimitris Agrafiotis, David Bull 0001, Nick Easton |
PCS | 3 |
| 2016 | Predicting video rate-distortion curves using textural featuresabstractThis work addresses the problem of predicting the compression efficiency of a video codec solely from features extracted from uncompressed content. Towards this goal, we have used a database of videos of homogeneous texture and extracted both spatial and frequency domain features. The videos are encoded using High Efficiency Video Coding (HEVC) reference codec at different quantization scales and their Rate-Distortion (RD) curves are modelled using linear regression. Using the extracted features and the fitted parameters of the RD model, a Support Vector Regression Model (SVRM) is trained to learn the relationship of the textural features with the RD curves. The SVRM is tested using iterative five-fold cross-validation. The presented experimental results demonstrate that RD curve characteristics can be predicted based on the textural features of the uncompressed videos, which offers potential benefits for encoder optimization. Angeliki V. Katsenou, Mariana Afonso, Dimitris Agrafiotis, David Bull 0001 |
PCS | 4 |
| 2016 | Support for reduced presentation durations in subjective video quality assessment
Felix Mercer Moss, Chun-Ting Yeh, Fan Zhang 0017, Roland Baddeley, David Bull 0001 |
Signal Process. Image Commun. | 5 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | On the Optimal Presentation Duration for Subjective Video Quality AssessmentabstractSubjective quality assessment is an essential component of modern image and video processing for both the validation of objective metrics and the comparison of coding methods. However, the standard procedures used to collect data can be prohibitively time consuming. One way of increasing the efficiency of data collection is to reduce the duration of test sequences from a 10-s length currently used in most subjective video quality assessment (VQA) experiments. Here, we explore the impact of reducing sequence length upon perceptual accuracy when identifying compression artifacts. A group of four reference sequences, together with five levels of distortion, are used to compare the subjective ratings of viewers watching videos between 1.5 and 10 s long. We identify a smooth function indicating that accuracy increases linearly as the length of the sequences increases from 1.5 to 7 s. The accuracy of observers viewing 1.5-s sequences was significantly inferior to those viewing sequences of 5, 7, and 10 s. We argue that sequences between 5 and 10 s produce satisfactory levels of accuracy but the practical benefits of acquiring more data lead us to recommend the use of 5-s sequences for future VQA studies that use the double stimulus continuous quality scale methodology. Felix Mercer Moss, Fan Zhang 0017, Roland Baddeley, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | A Perception-Based Hybrid Model for Video Quality AssessmentabstractIt is known that the human visual system (HVS) employs independent processes (distortion detection and artifact perception-also often referred to as near-threshold and suprathreshold distortion perception) to assess video quality for various distortion levels. Visual masking effects also play an important role in video distortion perception, especially within spatial and temporal textures. In this paper, a novel perception-based hybrid model for video quality assessment is presented. This simulates the HVS perception process by adaptively combining noticeable distortion and blurring artifacts using an enhanced nonlinear model. Noticeable distortion is defined by thresholding absolute differences using spatial and temporal tolerance maps that characterize texture masking effects, and this makes a significant contribution to quality assessment when the quality of the distorted video is similar to that of the original video. Characterization of blurring artifacts, estimated by computing high frequency energy variations and weighted with motion speed, is found to further improve metric performance. This is especially true for low quality cases. All stages of our model exploit the orientation selectivity and shift invariance properties of the dual-tree complex wavelet transform. This not only helps to improve the performance but also offers the potential for new low complexity in-loop application. Our approach is evaluated on both the Video Quality Experts Group (VQEG) full reference television Phase I and the Laboratory for Image and Video Engineering (LIVE) video databases. The resulting overall performance is superior to the existing metrics, exhibiting statistically better or equivalent performance with significantly lower complexity. Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | High Dynamic Range Video Compression Exploiting Luminance MaskingabstractThe human visual system (HVS) exhibits nonlinear sensitivity to the distortions introduced by lossy image and video coding. This effect is due to the luminance masking, contrast masking, and spatial and temporal frequency masking characteristics of the HVS. This paper proposes a novel perception-based quantization to remove nonvisible information in high dynamic range (HDR) color pixels by exploiting luminance masking so that the performance of the High Efficiency Video Coding (HEVC) standard is improved for HDR content. A profile scaling based on a tone-mapping curve computed for each HDR frame is introduced. The quantization step is then perceptually tuned on a transform unit basis. The proposed method has been integrated into the HEVC reference model for the HEVC range extensions (HM-RExt), and its performance was assessed by measuring the bitrate reduction against the HM-RExt. The results indicate that the proposed method achieves significant bitrate savings, up to 42.2%, with an average of 12.8%, compared with HEVC at the same quality (based on HDR-visible difference predictor-2 and subjective evaluations). Yang Zhang 0003, Matteo Naccari, Dimitris Agrafiotis, Marta Mrak, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Contrast Sensitivity of the Wavelet, Dual Tree Complex Wavelet, Curvelet, and Steerable Pyramid TransformsabstractAccurate estimation of the contrast sensitivity of the human visual system is crucial for perceptually based image processing in applications such as compression, fusion and denoising. Conventional contrast sensitivity functions (CSFs) have been obtained using fixed-sized Gabor functions. However, the basis functions of multiresolution decompositions such as wavelets often resemble Gabor functions but are of variable size and shape. Therefore to use the conventional CSFs in such cases is not appropriate. We have therefore conducted a set of psychophysical tests in order to obtain the CSF for a range of multiresolution transforms: the discrete wavelet transform, the steerable pyramid, the dual-tree complex wavelet transform, and the curvelet transform. These measures were obtained using contrast variation of each transforms' basis functions in a 2AFC experiment combined with an adapted version of the QUEST psychometric function method. The results enable future image processing applications that exploit these transforms such as signal fusion, superresolution processing, denoising and motion estimation, to be perceptually optimized in a principled fashion. The results are compared with an existing vision model (HDR-VDP2) and are used to show quantitative improvements within a denoising application compared with using conventional CSF values. Paul R. Hill, Alin Achim, Mohammed E. Al-Mualla, David Bull 0001 |
IEEE Trans. Image Process. | 4 |
| 2015 | Robust texture features based on undecimated dual-tree complex wavelets and local magnitude binary patternsabstractImage degradation due to illumination change, blur and noise can have a significant influence on classification performance, and yet no descriptors that perform well under these conditions exist. We propose a novel method for obtaining texture features, robust to these distortions, based on an undecimated dual-tree complex wavelet transform (UDT-CWT)1. As the UDT-CWT provides a local spatial relationship between scales, we can straightforwardly create bit-planes of the images representing local phases of wavelet coefficients. Magnitudes of the UDT-CWT are captured via a local binary pattern (LBP), after discarding some of the finest scales that are most affected by the blur and noise. A histogram of the resulting binary code words then forms the features used in texture classification. Results show that our approach outperforms existing methods, that claim to be invariant to feature degradations. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
ICIP | 3 |
| 2015 | Parameter optimisation for vision guided terrestrial locomotion: Multi-frameabstractDuring locomotion, Autonomous Ground Vehicles (AGV) need to map the surface of the ground ahead in order to plan a safe route of passage. Algorithms that achieve this are all derived from a basic requirement to track pixels between successive frames. This paper establishes the effect of tracking error on depth estimation with different image-plane geometries and presents an argument for tight control of the camera orientation. It is found that the effect of tracking error is invariant with camera angle for a concave image-plane but varies non-linearly with angle for both flat and inverse-flat image-planes. Pairs of frames are produced by each camera from a simulated scene and a selection of algorithms are used to compute the optical flow. It is shown that current state-of-the-art optical flow algorithms can successfully track pixels across concave and inverse-flat image-planes. This results in comparable depth estimation accuracy to a flat image-plane but allows for improved object tracking. G. P. Daniels, David Bull 0001, Jeremy F. Burn |
ICIP | 2 |
| 2015 | A study of subjective video quality at various frame ratesabstractThis paper presents a new video database (BVI-HFR), which contains content with a variety of frame rates from 15Hz to 120Hz, that can be used to demonstrate the benefits and limitations of higher frame rates, as well as investigating the role that frame rates play from capture to delivery. A characterization of the video database using low-level descriptors is also provided, which establishes that it successfully spans a variety of scene types and motions, and compares well to existing video databases. Subjective evaluations performed on the video database, have demonstrated a significant relationship between frame rates and perceived quality, up to 120Hz. They also confirm that the relationship between frame rate and perceived quality is content dependent. Alex Mackin, Fan Zhang 0017, David Bull 0001 |
ICIP | 3 |
| 2015 | A video texture database for perceptual compression and quality assessmentabstractThis paper presents a new publicly available video texture database (BVI Texture) that contains test sequences and subjective opinion scores. The database exhibits a wide range of static and dynamic textures together with some mixed content. Each sequence is indexed using various video feature descriptors that characterize its spatial activity, temporal activity, static texture content and dynamic texture content. Moreover, rate/distortion results for the new dataset are presented after compression using HEVC, alongside subjective quality evaluation data. The BVI texture database will provide utility in testing quality assessment metrics and emerging video compression methods, particularly those based on texture analysis and synthesis. Miltiadis Alexios Papadopoulos, Fan Zhang 0017, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 4 |
| 2015 | Superpixel-based statistical anomaly detection for sense and avoidabstractThis paper presents a novel preprocessing method for detecting small objects of interest within a high-resolution image, applied to the problem of visually detecting possible aircraft collisions (Sense and Avoid) for UAV platforms. The method is based on superpixel image segmentation combined with subsequent statistical analysis and anomaly detection. The existence of a possible target within a superpixel is described in terms of how it affects the local superpixel statistics and this signature statistical profile is consequently used to identify regions of interest throughout the image. The approach eliminates upwards of 90% of the total image area, significantly reducing the workload of further processing stages. Odysseas A. Pappas, Alin Achim, David Bull 0001 |
ICIP | 3 |
| 2015 | High dynamic range content calibration for accurate acquisition and displayabstractThis paper presents an end-to-end workflow for high dynamic range (HDR) content acquisition and HDR representation that aims to reproduce in a perceptually realistic manner the appearance of the captured scene. The proposed workflow includes a camera-independent colour calibration method for accurate HDR content acquisition and a perceptually optimized HDR signal representation method that takes into account display and viewing conditions. Subjective evaluation of the proposed and other methods indicate that the proposed workflow offers statistically significant improvements in the quality of the HDR image/video displayed to the viewers relative to the anchor methods. Yang Zhang 0003, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 3 |
| 2015 | An adaptive Lagrange multiplier determination method for rate-distortion optimisation in hybrid video codecsabstractThis paper describes an adaptive Lagrange multiplier determination method for rate-quality optimisation in video compression. Inspired by the experimental results of a Lagrange multiplier selection test, the presented approach adaptively estimates the optimum Lagrange multiplier for different video content, based on distortion statistics of recently encoded frames. The proposed algorithm has been fully integrated into both the H.264 and HEVC reference codecs, and is used in rate-distortion optimisation for encoding B frames. The results show promising (up to 11% on the sequences tested) overall bitrate savings, for a minimal increase in complexity, on various types of test content based on Bjontegaard delta measurements. Fan Zhang 0017, David Bull 0001 |
ICIP | 2 |
| 2015 | Hierarchical video summarization with loitering indicationabstractIn this paper, a hierarchical and informative summarization framework is proposed, which facilitates rapid video browsing. Moreover, a method for loitering detection is exploited to indicate potential abnormal behaviors. The hierarchical framework includes two levels: a holistic-level and an object-level. The holistic-level summarization provides viewers with a comprehensive and compact representation of the original video, while the object-level summarization extracts the narrative information of each object, including trajectory, direction, time, changes of appearance and indication of the loitering behavior. The two summarizations are formulated as two different energy minimization problems, which are solved by the proposed heuristic algorithms. Our framework is evaluated on two publicly available datasets. Experimental results demonstrate that the proposed method performs favourably in providing holistic- and object-level information, fast browsing, and loitering detection. Ruipeng Lu, Hua Yang 0001, Ji Zhu 0002, Shuang Wu 0001, Jia Wang 0004, David Bull 0001 |
VCIP | 6 |
| 2015 | Spectrum efficient cross-layer adaptation of Raptor codes for video multicasting over mobile broadband networks
Victoria Sgardoni, David Bull 0001, Andrew R. Nix |
Pervasive Mob. Comput. | 2 |
| 2015 | Undecimated Dual-Tree Complex Wavelet Transforms
Paul R. Hill, Nantheera Anantrasirichai, Alin Achim, Mohammed E. Al-Mualla, David Bull 0001 |
Signal Process. Image Commun. | 5 |
| 2015 | Terrain Classification From Body-Mounted Cameras During Human LocomotionabstractThis paper presents a novel algorithm for terrain type classification based on monocular video captured from the viewpoint of human locomotion. A texture-based algorithm is developed to classify the path ahead into multiple groups that can be used to support terrain classification. Gait is taken into account in two ways. Firstly, for key frame selection, when regions with homogeneous texture characteristics are updated, the frequency variations of the textured surface are analyzed and used to adaptively define filter coefficients. Secondly, it is incorporated in the parameter estimation process where probabilities of path consistency are employed to improve terrain-type estimation. When tested with multiple classes that directly affect mobility-a hard surface, a soft surface, and an unwalkable area-our proposed method outperforms existing methods by up to 16%, and also provides improved robustness. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
IEEE Trans. Cybern. | 3 |
| 2014 | Orientation estimation for planar textured surfaces based on complex waveletsabstractThe gradient of a road or terrain influences the appropriate speed and power of a vehicle traversing it. Therefore, gradient prediction is necessary if autonomous vehicles are to optimise their locomotion. This paper presents a novel texture-based method for estimating the orientation of planar surfaces under the basic assumption of homogeneity. Based on a backprojection technique, we propose a new method for measuring texture consistency in the Dual-Tree Complex Wavelet Transform (DT-CWT) domain. Texture histograms computed with a new anti-aliasing compensation approach are employed to find the rotation of a planar surface. The proposed method performs well with various types of textured surfaces and outperforms other existing methods with significantly reduced computational complexity up to 35%. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
ICIP | 3 |
| 2014 | Robust texture features for blurred images using Undecimated Dual-Tree Complex WaveletsabstractThis paper presents a new descriptor for texture classification. The descriptor is rotationally invariant and blur insensitive, which provides great benefits for various applications that suffer from out-of-focus content or involve fast moving or shaking cameras. We employ an Undecimated Dual-Tree Complex Wavelet Transform (UDT-CWT) [1] to extract texture features. As the UDT-CWT fully provides local spatial relationship between scales and subband orientations, we can straightforwardly create bit-planes of the images representing local phases of wavelet coefficients. We also discard some of the finest decomposition levels where are most affected by the blur. A histogram of the resulting code words is created and used as features in texture classification. Experimental results show that our approach outperforms existing methods by up to 40% for synthetic blurs and up to 30% for natural video content due to camera motion when walking. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
ICIP | 3 |
| 2014 | Feature-based registration for correlative light and electron microscopy imagesabstractIn this paper we present a feature-based registration algorithm for largely misaligned bright-field light microscopy images and transmission electron microscopy images. We first detect cell centroids, using a gradient-based single-pass voting algorithm. Images are then aligned by finding the flip, translation and rotation parameters, which maximizes the overlap between pseudo-cell-centers. We demonstrate the effectiveness of our method, by comparing it to manually aligned images. Combining registered light and electron microscopy images together can reveal details about cellular structure with spatial and high-resolution information. David Nam, Judith Mantell, Lorna Hodgson, David Bull 0001, Paul Verkade, Alin Achim |
ICIP | 4 |
| 2014 | A Novel Framework for Segmentation of Secretory Granules in Electron Micrographs
David Nam, Judith Mantell, David Bull 0001, Paul Verkade, Alin Achim |
Medical Image Anal. | 3 |
| 2014 | Dual-tree complex wavelet coefficient magnitude modelling using the bivariate Cauchy-Rayleigh distribution for image denoising
Paul R. Hill, Alin Achim, David Bull 0001, Mohammed E. Al-Mualla |
Signal Process. | 3 |
| 2013 | Curvelet fusion of panchromatic and SAR satellite imagery using fractional lower order momentsabstractThis paper presents a novel fusion method aimed at combining panchromatic and synthetic aperture radar(SAR) satellite imagery. The presented method seeks to combine the advantages of the two modalities while simultaneously minimizing the effect of artifacts inherent in SAR images. The alpha-stable distribution is used to model the curvelet decomposition coefficients of the image, as it caters for their heavy-tailed nature. Coefficients are fused using a weighted average rule with the saliency and match measures derived from the fractional lower-order moments of the alpha-stable distribution [3]. Experimental results show this method to provide high-quality results that improve the perceptive quality of the image without introducing any additional artifacts. Odysseas A. Pappas, Alin Achim, David Bull 0001 |
AVSS | 3 |
| 2013 | Insulin Granule Segmentation in 3-D TEM Beta Cell TomogramsabstractDavid Nam1 [email protected] Judith Mantell2,3 [email protected] David Bull1 [email protected] Paul Verkade2,3,4/shared last author [email protected] Alin Achim1 [email protected] 1 Visual Information Laboratory University of Bristol, UK 2 Wolfson Bioimaging Facility University of Bristol, UK 3 School of Biochemistry University of Bristol, UK 4 School of Physiology and Pharmacology University of Bristol, UK David Nam, Judith Mantell, David Bull 0001, Paul Verkade, Alin Achim |
BMVC | 3 |
| 2013 | Projective image restoration using sparsity regularizationabstractThis paper presents a method of image restoration for projective ground images which lie on a projection orthogonal to the camera axis. The ground images are initially transformed using homography, and then the proposed image restoration is applied. The process is performed in the dual-tree complex wavelet transform domain in conjunction with L0 reweighting and L2 minimisation (L0RL2) employed to solve this ill-posed problem. We also propose instant estimation of a blur kernel arising from the projective transform and the subsequent interpolation of sparse data. Subjective results show significant improvement of image quality. Furthermore, classification of surface type at various distances (evaluated using a support vector machine classifier) is also improved for the images restored using our proposed algorithm. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
ICIP | 3 |
| 2013 | Gaze location prediction for broadcast football video using Bayesian integration of low level features and top-down cuesabstractAccurate prediction of the viewer's gaze location has the potential to improve bit allocation, rate control, error resilience and quality evaluation in video compression. With complex contexts, such as that of broadcast football video, the potential reward is even higher given that compression and transmission of this type of content is challenging. In this paper we propose a gaze location prediction system for high definition broadcast football video. The proposed system employs Bayesian integration of bottom-up features and context specific top-down cues. Our results show that the proposed model has better gaze prediction performance than other top-down models that we adapted to this context. Qin Cheng, Dimitris Agrafiotis, Alin Achim, David Bull 0001 |
ICIP | 4 |
| 2013 | Scalable video fusionabstractA novel system is introduced that is able to fuse two or more sets of multimodal videos in the transform domain. This is achieved without drift and produces an embedded bitstream that offers fine grain scalability. Previous attempts to fuse in the transform domain have not been possible for video compression systems due to the complications of predictive loops within conventional video encoding. The compression system is based on an optimised spatiotemporal codec using the 3D Discrete Dual-tree Wavelet Transform (DDWT) together with a bit plane encoding method (SPIHT) and a coefficient sparsification process (noise shaping). Together, these methods can efficiently encode a video sequence without the need for motion compensation due to the directional (in space and time) selectivity of the transform. This system offers extremely flexible video fusion in dynamic bandwidth environments where there are variable client receiving capabilities. Paul R. Hill, Alin Achim, David Bull 0001 |
ICIP | 3 |
| 2013 | Image denoising using dual tree statistical models for complex wavelet transform coefficient magnitudesabstractWavelet shrinkage is a standard technique for denoising natural images. Originally proposed for univariate shrinkage in the Discrete Wavelet Transform (DWT) domain, it has since been optimised through the exploitation of translationally invariant wavelet decompositions such as the Dual-Tree Complex Wavelet Transform (DT-CWT) alongside bivariate analysis techniques that condition the shrinkage on spatially related coefficients across neighbouring scales. These more recent techniques have denoised the real and imaginary components of the DT-CWT coefficients separately. Processing real and imaginary components separately has been found to lead to an increase in the phase noise of the transform which in turn affects denoising performance. On this basis, the work presented in this paper offers improved denoising performance through modelling the bivariate distribution of the coefficient magnitudes. The results were compared to the current state of the art non-local means denoising technique BM3D, showing clear subjective improvements, through the retention of high frequency structural and textural information. The paper also compares objective measures, using both PSNR and the more perceptually valid structural similarity measure (SSIM). Whereas PSNR results were slightly below those for BM3D, those for SSIM showed closer correlation with subjective assessment, indicating improvements over BM3D for most noise levels on the images tested. Paul R. Hill, Alin Achim, David Bull 0001, Mohammed E. Al-Mualla |
ICIP | 3 |
| 2013 | Context-based video codingabstractWe present a video CODEC framework which exploits extrinsic scene knowledge to condition a perspective motion model. An approximate textural-geometric model of the scene is prepared prior to coding. During coding, the locations of planar surfaces in the scene are tracked, facilitating the computation of accurate perspective motion warp parameters. These algorithms are integrated with H.264 into a hybrid CODEC framework, achieving savings of up to 48% for equivalent visual quality. Richard George Vigars, Andrew Calway, David Bull 0001 |
ICIP | 3 |
| 2013 | Visual masking phenomena with high dynamic range contentabstractHigh Dynamic Range (HDR) technology (capture and display), can offer high levels of immersion through a dynamic range that meets and exceeds that of the Human Visual System (HVS). This increase in immersion comes at the cost of higher bitrate requirements, which necessitate the development of efficient HDR-relevant coding solutions. Efficient perception-based compression of HDR imagery requires models that capture accurately the various masking effects experienced by the HVS under HDR conditions, so that bits are not wasted coding redundant imperceptible information. In this paper we present two psychovisual experiments that we carried out with the aid of a high dynamic range display, in order to determine potential differences between Standard Dynamic Range (SDR) and HDR edge masking (EM) and luminance masking (LM) effects. The EM experimental results indicate that the visibility threshold is higher for the case of HDR content than SDR, especially on the dark background side of an edge. The LM experimental results suggest that the HDR visibility threshold is higher compared to SDR for both dark and bright luminance backgrounds. Yang Zhang 0003, Dimitris Agrafiotis, Matteo Naccari, Marta Mrak, David Bull 0001 |
ICIP | 5 |
| 2013 | Quality assessment methods for perceptual video compressionabstractThis paper describes a quality assessment model for perceptual video compression applications (PVM), which stimulates visual masking and distortion-artefact perception using an adaptive combination of noticeable distortions and blurring artefacts. The method shows significant improvement over existing quality metrics based on the VQEG database, and provides compatibility with in-loop rate-quality optimisation for next generation video codecs due to its latency and complexity attributes. Performance comparison are validated against a range of different distortion types. Fan Zhang 0017, David Bull 0001 |
ICIP | 2 |
| 2013 | High dynamic range video compression by intensity dependent spatial quantization in HEVCabstractThe Human Visual System (HVS) shows non-linear sensitivity to the distortion introduced by lossy image and video coding. This non-linear sensitivity is due to luminance masking, contrast masking and the spatial and temporal frequency masking phenomena of the HVS. This paper proposes a perception-based quantization method that exploits luminance masking in the HVS in order to enhance the performance of the High Efficiency Video Coding (HEVC) standard for the case of High Dynamic Range (HDR) video content. A profile scaling based on a tone-mapping curve computed for each HDR frame is introduced. The quantization step is then perceptually tuned on a Transform Unit (TU) basis. The proposed method has been integrated into the reference codec considered for the HEVC range extensions and its performance was assessed by measuring the bitrate reduction against the codec without perceptual quantization. The HDR-VDP-2 image quality metric was employed to measure the compressed picture quality. For the same quality level, an average bitrate reduction of 9% is achieved across all tested HDR sequences. Yang Zhang 0003, Matteo Naccari, Dimitris Agrafiotis, Marta Mrak, David Bull 0001 |
PCS | 5 |
| 2013 | Atmospheric Turbulence Mitigation Using Complex Wavelet-Based FusionabstractRestoring a scene distorted by atmospheric turbulence is a challenging problem in video surveillance. The effect, caused by random, spatially varying, perturbations, makes a model-based solution difficult and in most cases, impractical. In this paper, we propose a novel method for mitigating the effects of atmospheric distortion on observed images, particularly airborne turbulence which can severely degrade a region of interest (ROI). In order to extract accurate detail about objects behind the distorting layer, a simple and efficient frame selection method is proposed to select informative ROIs only from good-quality frames. The ROIs in each frame are then registered to further reduce offsets and distortions. We solve the space-varying distortion problem using region-level fusion based on the dual tree complex wavelet transform. Finally, contrast enhancement is applied. We further propose a learning-based metric specifically for image quality assessment in the presence of atmospheric distortion. This is capable of estimating quality in both full- and no-reference scenarios. The proposed method is shown to significantly outperform existing methods, providing enhanced situational awareness in a range of surveillance scenarios. Nantheera Anantrasirichai, Alin Achim, Nick G. Kingsbury, David Bull 0001 |
IEEE Trans. Image Process. | 4 |
| 2013 | Gaze Location Prediction for Broadcast Football VideoabstractThe sensitivity of the human visual system decreases dramatically with increasing distance from the fixation location in a video frame. Accurate prediction of a viewer's gaze location has the potential to improve bit allocation, rate control, error resilience, and quality evaluation in video compression. Commercially, delivery of football video content is of great interest because of the very high number of consumers. In this paper, we propose a gaze location prediction system for high definition broadcast football video. The proposed system uses knowledge about the context, extracted through analysis of a gaze tracking study that we performed, to build a suitable prior map. We further classify the complex context into different categories through shot classification thus allowing our model to prelearn the task pertinence of each object category and build the prior map automatically. We thus avoid the limitation of assigning the viewers a specific task, allowing our gaze prediction system to work under free-viewing conditions. Bayesian integration of bottom-up features and top-down priors is finally applied to predict the gaze locations. Results show that the prediction performance of the proposed model is better than that of other top-down models that we adapted to this context. Qin Cheng, Dimitris Agrafiotis, Alin Achim, David Bull 0001 |
IEEE Trans. Image Process. | 4 |
| 2012 | Mitigating the effects of atmospheric distortion using DT-CWT fusionabstractThis paper describes a new method for mitigating the effects of atmospheric distortion on observed images, particularly airborne turbulence which degrades a region of interest (ROI). In order to provide accurate detail from objects behind the distorting layer, a simple and efficient frame selection method is proposed to pick informative ROIs from only good-quality frames. We solve the space-variant distortion problem using region-based fusion based on the Dual Tree Complex Wavelet Transform (DT-CWT). We also propose an object alignment method for pre-processing the ROI since this can exhibit significant offsets and distortions between frames. Simple haze removal is used as the final step. The proposed method performs very well with atmospherically distorted videos and outperforms other existing methods. Nantheera Anantrasirichai, Alin Achim, David Bull 0001, Nick G. Kingsbury |
ICIP | 3 |
| 2012 | The Undecimated Dual Tree Complex Wavelet Transform and its application to bivariate image denoising using a Cauchy modelabstractThe Undecimated Dual Tree Complex Wavelet Transform (UDTCWT) is introduced together with its application to image denoising. The UDT-CWT extends the traditional DT-CWT using the methods of filter upsampling and the removal of downsampling developed for the Undecimated Discrete Wavelet Transform (UDWT). The UDTCWT results in a one-to-one relationship between co-located complex coefficients in all subbands and offers improved lower scale subband localisation together with improved directional selectivity (compared to the UDWT). These properties of the UDT-CWT have been exploited in the presented bivariate shrinkage denoising algorithm and gives quantitative improvements in the application to the denoising of images. Paul R. Hill, Alin Achim, David Bull 0001 |
ICIP | 3 |
| 2012 | Perceptually lossless High Dynamic Range image compression with JPEG 2000abstractHigh Dynamic Range (HDR) technology offers high levels of immersion with a dynamic range meeting and exceeding that of the Human Visual System (HVS). A primary drawback of HDR images and video is that memory and bandwidth requirements are significantly higher than for conventional images and video. Many bits can be wasted coding redundant imperceptible information. The challenge is therefore to develop means for efficiently compressing HDR imagery to a manageable bit rate without compromising perceptual quality. In this paper, an HDR image compression method, based on an HVS optimized wavelet subband weighting method is proposed. The method has been fully integrated into a JPEG 2000 codec. Experimental results indicate that the proposed method outperforms previous approaches and operates in accordance with characteristics of the HVS, tested objectively using a HDR Visible Difference Predictor (VDP). Yang Zhang 0003, Erik Reinhard, David Bull 0001 |
ICIP | 3 |
| 2012 | Unsupervised video anomaly detection using feature clusteringabstractThis study addresses the problem of automatic anomaly detection for surveillance applications. A general framework for anomalous event detection in uncrowded scenes has been developed which consists of the following key components: (i) an efficient foreground detection model based on a Gaussian mixture model (GMM), which can selectively update pixel information in each image region; (ii) an adaptive foreground object tracker that combines the merits of Kalman, mean-shift and particle filtering; (iii) a feature clustering algorithm, which can automatically choose the optimal number of clusters in the training data for scene pattern modelling; (iv) a statistical scene modeller based on Bayesian theory and GMM, which combines trajectory-based and region-based information for enhanced anomaly detection. The resulting approach achieves fully unsupervised anomaly detection in surveillance video. The experimental results show improved detection performance compared with the state-of-the-art methods. Alin Achim, David Bull 0001 |
IET Signal Process. | 3 |
| 2012 | Perception-oriented video coding based on image analysis and completion: A review
Patrick Ndjiki-Nya, Dimitar Doshkov, Hagen Kaprykowsky, Fan Zhang 0017, David Bull 0001, Thomas Wiegand 0001 |
Signal Process. Image Commun. | 5 |
| 2012 | Rate-Distortion-Optimized Video Transmission Using Pyramid Vector QuantizationabstractConventional video compression relies on interframe prediction (motion estimation), intra frame prediction and variable-length entropy encoding to achieve high compression ratios but, as a consequence, produces an encoded bitstream that is inherently sensitive to channel errors. In order to ensure reliable delivery over lossy channels, it is necessary to invoke various additional error detection and correction methods. In contrast, techniques such as Pyramid Vector Quantisation have the ability to prevent error propagation through the use of fixed length codewords. This paper introduces an efficient rate distortion optimisation algorithm for intra-mode PVQ which offers similar compression performance to intra H.264/AVC and Motion JPEG 2000 while offering inherent error resilience. The performance of our enhanced codec is evaluated for HD content in the context of a realistic (IEEE 802.11n) wireless environment. We show that PVQ provides high tolerance to corrupted data compared to the state of the art while obviating the need for complex encoding tools. Syed Mohsin Matloob Bokhari, Andrew R. Nix, David Bull 0001 |
IEEE Trans. Image Process. | 3 |
| 2011 | Robust video transmission using Pyramid Vector QuantisationabstractVideo encoders such as H.264 and VC-1 rely on variable length coding, motion estimation and intra prediction to achieve high compression ratios but, as a consequence, produce an encoded bitstream that is extremely sensitive to channel errors. Pyramid Vector Quantisation, which generates fixed length codes, has been shown to provide excellent error resilience for wireless transmission of HD video. This paper presents an efficient rate distortion optimisation algorithm for PVQ which yields impressive intra coding performance with low complexity. The compression performance of PVQ is comparable with intra-H.264 and intra-VC-1. The error resilience of PVQ is also compared with H.264 and VC-1 in a representative MIMO Wireless LAN environment, showing the superior performance of PVQ. Syed Mohsin Matloob Bokhari, David Bull 0001, Andrew R. Nix |
ICIP | 2 |
| 2011 | Perception-based high dynamic range video compression with optimal bit-depth transformationabstractHigh Dynamic Range (HDR) technology is able to offer high levels of immersion with a dynamic range comparable to the Human Visual System (HVS). A primary drawback of HDR is that its memory and bandwidth requirements are significantly higher than for conventional video. The challenge is thus to develop means for efficiently compressing the video to a manageable bitrate without compromising perceptual quality. In this paper, we propose an HDR compression method based on an optimized bit-depth transformation, and HVS model based wavelet transform denoising. Experimental results indicate that the proposed method outperforms previous approaches and operates in accordance with characteristics of the HVS, tested objectively using a Visible Difference Predictor (VDP). Yang Zhang 0003, Erik Reinhard, David Bull 0001 |
ICIP | 3 |
| 2011 | Mobile WiMAX video quality and transmission efficiencyabstractThis paper explores rules and defines cross layer parameters that enhance the received video quality per unit bandwidth while observing strict video QoS for unicast video transmission over mobile WiMAX. Channel capacity and goodput are estimated using a detailed MAC/PHY simulator. The simulator supports link adaptation of the PHY layer Modulation and Coding Scheme (MCS) alongside MAC layer Automatic Repeat Request (ARQ). A new performance indicator, PSNR per Hz, is defined to capture the video PSNR attained for a given effective bandwidth. For the 3GPP urban micro channel, MCS and MAC layer block lifetime are jointly determined as a function of average channel SNR in order to enhance video PSNR and goodput. At low SNR higher level QAM is inefficient (as a result of excessive MAC layer ARQ or high packet error rate) for video transfer. Results show that video efficiency improves for block lifetimes up to 100ms, since the greater use of ARQ enables more spectrally efficient MCS modes to be applied. Victoria Sgardoni, David Halls, Syed Mohsin Matloob Bokhari, David Bull 0001, Andrew R. Nix |
PIMRC | 4 |
| 2011 | Colour volumetric compression for realistic view synthesis applications
Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, Akbar Sheikh Akbari, David Bull 0001 |
Multim. Tools Appl. | 5 |
| 2011 | Localization of Mobile Nodes in Wireless Networks with Correlated in Time Measurement NoiseabstractWireless sensor networks are an inherent part of decision making, object tracking, and location awareness systems. This work is focused on simultaneous localization of mobile nodes based on received signal strength indicators (RSSIs) with correlated in time measurement noises. Two approaches to deal with the correlated measurement noises are proposed in the framework of auxiliary particle filtering: with a noise augmented state vector and the second approach implements noise decorrelation. The performance of the two proposed multimodel auxiliary particle filters (MM AUX-PFs) is validated over simulated and real RSSIs and high localization accuracy is demonstrated. Lyudmila Mihaylova, Donka S. Angelova, David Bull 0001, Cedric Nishan Canagarajah |
IEEE Trans. Mob. Comput. | 3 |
| 2010 | High definition wireless video transmission using Pyramid Vector QuantisationabstractVideo Source encoders such as H.264 and VC-1 are in common use today and provide impressive compression performance. These use variable length coding and predictive coding to achieve high compression ratios but as a consequence render the encoded bitstream extremely sensitive to channel errors. In this paper we revisit an alternative technique, Pyramid Vector Quantisation, which generates fixed length codes. Intra mode PVQ is presented as a means of robustly compressing and wirelessly transmitting HD Video. The paper shows that PVQ is an effective codec for indoor wireless transmission of HD video. Comparing the required channel SNR for an H.264 based system with a PVQ based system shows that PVQ can operate at a much lower SNR (with gains of up to 11 dB) while still providing excellent video quality and compression performance. Syed Mohsin Matloob Bokhari, David Bull 0001, Andrew R. Nix |
ICIP | 2 |
| 2010 | Automatic multi-camera placement and optimisation using ray tracingabstractAn automatic method for rapid and optimised surveillance camera deployment is proposed for outdoor settings. Given a target scene. potential camera locations are generated taking account of real-world camera parameters and environmental constraints. For example, each point on a target must be viewed by at least two cameras. We employ ray tracing to determine the visibility of any point from a given camera and camera-to-point viewing angle is used to measure visibility quality. Gradient ascent is used to optimise camera orientation for maximum quality coverage. Camera network optimisation was also addressed, and a new combinatorial algorithm is proposed. A comparative study showed that our approach significantly outperforms competing techniques providing solutions which better cover the search space, while offering multiple alternative solutions. Samira Bouyagoub, David Bull 0001, Cedric Nishan Canagarajah, Andrew R. Nix |
ICIP | 2 |
| 2010 | Automatic contrast enhancement of low-light images based on local statistics of wavelet coefficientsabstractThis paper describes a new method for contrast enhancement in images of low-light or unevenly illuminated scenes based on statistical modelling of wavelet coefficients of the image. A non-linear enhancement function has been designed based on the local dispersion of the wavelet coefficients modelled as a bivariate Cauchy distribution. Within the same statistical framework, a simultaneous noise reduction in the image is performed by means of a shrinkage function, thus preventing noise amplification. The proposed enhancement method has been shown to perform very well with insufficiently illuminated and noisy images, outperforming other conventional methods, in terms of contrast enhancement and noise reduction in the output image. Artur Loza, David Bull 0001, Alin Achim |
ICIP | 2 |
| 2010 | Region-based texture modelling for next generation video codecsabstractThis paper describes a texture model application for future video compression algorithms. This employs region-based texture analysis techniques and advanced motion models to synthesise and warp video frames rather than encode the whole image or the prediction residual after traditional motion compensation. The proposed texture warping method and the dynamic texture model are integrated into an H.264 framework together with video quality assessment modules to prevent video artefacts. The results show impressive bitrate savings, up to 47%, over H.264 with similar visual quality. Fan Zhang 0017, David Bull 0001, Cedric Nishan Canagarajah |
ICIP | 2 |
| 2010 | Enhanced video compression with region-based texture modelsabstractThis paper presents a region-based video compression algorithm based on texture warping and synthesis. Instead of encoding whole images or prediction residuals after translational motion estimation, this algorithm employs a perspective motion model to warp static textures and uses a texture synthesis approach to synthesise dynamic textures. Spatial and temporal artefacts are prevented by an in-loop video quality assessment module. The proposed method has been integrated into an H.264 video coding framework. The results show significant bitrate savings, up to 55%, compared with H.264, for similar visual quality. Fan Zhang 0017, David Bull 0001 |
PCS | 2 |
| 2010 | Non-Gaussian model-based fusion of noisy images in the wavelet domain
Artur Loza, David Bull 0001, Cedric Nishan Canagarajah, Alin Achim |
Comput. Vis. Image Underst. | 2 |
| 2010 | Sequential Monte Carlo methods for contour tracking of contaminant clouds
Mohamed Hisham Jaward, David Bull 0001, Cedric Nishan Canagarajah |
Signal Process. | 2 |
| 2010 | A video error resilience redundant slices algorithm and its performance relative to other fixed redundancy schemes
Pierre Ferré, Dimitris Agrafiotis, David Bull 0001 |
Signal Process. Image Commun. | 3 |
| 2010 | Sub-pixel motion estimation using kernel methods
Paul R. Hill, David Bull 0001 |
Signal Process. Image Commun. | 2 |
| 2010 | Secure transcoders for single layer video data
Nithin M. Thomas, David W. Redmill, David Bull 0001 |
Signal Process. Image Commun. | 3 |
| 2010 | In-Band Disparity Compensation for Multiview Image Compression and View SynthesisabstractThis paper presents a novel framework to achieve scalable multiview image compression and view synthesis. The open-loop wavelet-lifting scheme for geometric filtering has been exploited to achieve signal-to-noise ratio scalability and view-type scalability (mono, stereo, or multiview). Spatial scalability is achieved by employing in-band prediction which removes correlations among subbands (level-by-level) via shift-invariant references obtained by overcomplete discrete wavelet transforms. We propose a novel in-band disparity compensated view filtering approach, akin to motion compensated temporal filtering, for achieving a scalable multiview codec. In our codec, hybrid prediction is proposed to deal with occlusions, and a novel cost function in dynamic programming (DP) for disparity estimation is introduced to improve view synthesis quality. Experiments show comparable results at full resolution and significant improvements at coarser resolutions, compared to a conventional spatial prediction scheme. View synthesis efficiency is extensively improved by utilizing disparity estimation from the proposed DP approach. Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Enhanced spatially interleaved DVC using diversity and selective feedbackabstractSystems with cheap/simple/power efficient encoders but complex decoders make applications such as low cost, low power remote sensors practical. Bandwidth considerations however are still an issue and compression efficiency has to remain high. In this paper, we present a distributed video codec (DVC) that we are developing with the aim of achieving such a low power paradigm at the cost of only a small compression performance deficit relative to the current state of the art, H.264. The proposed system employs spatial interleaving of KEY and Wyner-Ziv data which allows efficient side information (SI) generation through block-based error concealment, a Gray code that increases the accuracy of bit probability estimation, and a diversity scheme that produces more reliable results by exploiting multiple SI generated data. Simulation results show an improvement of the proposed scheme over H.264 intra coding of up to 1.5 dB. We additionally propose two mechanisms for selective parity bit feedback requests that can further reduce the WZ bitrate by up to 15%. Nantheera Anantrasirichai, Dimitris Agrafiotis, David Bull 0001 |
ICASSP | 3 |
| 2009 | Kernel based sub-pixel motion estimationabstractModern video codecs such as MPEG2, MPEG4-ASP and H.264 depend on computationally complex sub-pixel motion estimation (ME) to optimise rate-distortion efficiency. Sub-pixel ME is implemented within these standards using interpolated values at 1/2 or 1/4 pixel accuracy. This paper presents a novel scheme for sub-pixel ME using a model free kernel method utilising the result of the whole-pixel SAD distribution applied to an H.264 encoder. While giving an equivalent rate distortion performance, this approach approximately halves the number of quarter-pixel search positions giving an overall speed up of approximately 10% compared to the EPZS quarter-pixel method (the state of the art H.264 optimised sub-pixel motion estimator). Paul R. Hill, David Bull 0001 |
ICIP | 2 |
| 2009 | Unsupervised image compression using graphcut texture synthesisabstractAn unsupervised image compression-by-synthesis system is proposed utilising wavelet based image segmentation and analysis combined with patch based texture synthesis. High perceptual quality is ensured using an artefact detection algorithm in the encoder loop. EBCOT is used to transform code texture samples and residual image data. Resulting bitrate savings of up to approximately 17% over JPEG2000 for little change in perceptual quality have been shown. Stephen Ierodiaconou, James Byrne, David Bull 0001, David W. Redmill, Paul R. Hill |
ICIP | 3 |
| 2009 | GMM-based efficient foreground detection with adaptive region updateabstractThe accurate detection of moving objects is an important step in the process of tracking and recognition in many real-time video surveillance applications. In this paper, we propose a combination of block-based detection and a pixel-based Gaussian Mixture Model (GMM) for moving object detection. Compared with traditional pixel-based algorithms which update all pixels for every frame, our algorithm has the ability to selectively update region information within each frame, while offering the capability to refine the silhouette of a foreground object. The algorithm offers an efficient trade-off between complexity and detection performance. The results show improved detection in the presence of high camera noise, high level compression artefacts, camera movements and dynamic background conditions. Alin Achim, David Bull 0001 |
ICIP | 3 |
| 2009 | Perception-oriented video coding based on texture analysis and synthesisabstractPerception-oriented video coding based on texture analysis and synthesis has gained importance over the past few years. Hence, the present paper overviews a selection of related approaches that have been proposed over the past decades. They are also referred to as content-based video coding (CBVC) methods in this paper. For better insight into CBVC, an overview on texture analysis and synthesis is also given. Furthermore, the principles common to a careful selection of CBVC methods are depicted and the requirements of each of the fundamental modules are extensively discussed in the context of the limitations of state-of-the-art hybrid video codecs like H.264/AVC. Patrick Ndjiki-Nya, David Bull 0001, Thomas Wiegand 0001 |
ICIP | 2 |
| 2009 | A novel H.264 SVC encryption scheme for secure bit-rate transcodingabstractThis paper presents a novel architecture for the secure delivery of encrypted H.264 SVC bitstreams. It relies on a block cipher and stream cipher used in a novel way that would allow an intermediary transcoder to truncate the bitstream to the appropriate bit-rate without decrypting the data. The system, called SVC-sec, is compared to other architectures presented in the literature and it is shown that SVC-sec offers many benefits, particularly when used with FGS streams. Nithin M. Thomas, David Bull 0001, David W. Redmill |
PCS | 2 |
| 2009 | Structural similarity-based object tracking in multimodality surveillance videos
Artur Loza, Lyudmila Mihaylova, David Bull 0001, Cedric Nishan Canagarajah |
Mach. Vis. Appl. | 3 |
| 2009 | A Multicue Bayesian State Estimator for Gaze Prediction in Open Signed VideoabstractWe propose a multicue gaze prediction framework for open signed video content, the benefits of which include coding gains without loss of perceived quality. We investigate which cues are relevant for gaze prediction and find that shot changes, facial orientation of the signer and face locations are the most useful. We then design a face orientation tracker based upon grid-based likelihood ratio trackers, using profile and frontal face detections. These cues are combined using a grid-based Bayesian state estimation algorithm to form a probability surface for each frame. We find that this gaze predictor outperforms a static gaze prediction and one based on face locations within the frame. Sam J. C. Davies, Dimitris Agrafiotis, Cedric Nishan Canagarajah, David Bull 0001 |
IEEE Trans. Multim. | 4 |
| 2008 | Contour tracking of contaminant clouds with sequential Monte Carlo methodsabstractContour tracking for a single source emission is addressed in this paper. This problem is solved by estimating the contour boundary positions using a set of particle filters. The use of Sequential Monte Carlo techniques enables the tracking to performed when the measurements are noisy and the tracking results also includes the estimation uncertainty. The proposed technique is illustrated for a SCIPUFF generated single emission scenario and simulation experiments showed the successful tracking throughout the tracking period. Mohamed Hisham Jaward, David Bull 0001, Cedric Nishan Canagarajah |
ICASSP | 2 |
| 2008 | A concealment based approach to distributed video codingabstractThis paper presents a concealment based approach to distributed video coding that uses hybrid key/WZ frames via an FMO type interleaving of macroblocks. Our motivation stems from a previous work of ours that showed promising results relative to the more common approach of splitting the sequence in key and WZ frames. In this paper, we extend our previous scheme to the case of I-B-P frame structures and transform domain DVC. We additionally introduce a number of enhancements at the decoder including use of spatio-temporal concealment for generating the side information on a MB basis, mode selection for switching between the two concealment approaches and for deciding how the correlation noise is estimated, local (MB wise) correlation noise estimation and modified B frame quantisation. The results presented indicate considerable improvement (up to 30%) compared to corresponding frame extrapolation and frame interpolation schemes. Nantheera Anantrasirichai, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 3 |
| 2008 | Unsupervised image compression-by-synthesis within a JPEG frameworkabstractAn image compression scheme is proposed, utilising wavelet- based image segmentation and texture analysis, and patch- based texture synthesis. This has been incorporated into a JPEG framework. Homogeneous textured regions are identified and removed prior to transform coding. These regions are then replaced at the decoder by synthesis from marked samples, and colour matched to ensure similarity to the original. Experimental results on natural images show bitrate savings of over 18% compared with JPEG for little change in measured visual quality. James Byrne, Stephen Ierodiaconou, David Bull 0001, David W. Redmill, Paul R. Hill |
ICIP | 3 |
| 2008 | A gaze prediction technique for open signed video content using a track before detect algorithmabstractThis paper proposes a gaze prediction model for open signed video content. A face detection algorithm is used to locate faces across each frame in both profile and frontal orientations. A grid-based likelihood ratio track before detect routine is used to predict the orientation of the signer's head, which allows the gaze location to be localised to either the signer or the inset. The face detections are then used to narrow down the gaze prediction further. The gaze predictor is able to predict the results of an eye tracking study with up to 95% accuracy, and an average accuracy of over 80%. Sam J. C. Davies, Dimitris Agrafiotis, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP | 4 |
| 2008 | A framework for dense optical flow from multiple sparse hypothesesabstractOptical flow forms an important initial processing stage for many machine vision tasks. A framework is presented for the recovery of dense optical flows from image sequences containing large motions. Sparse feature correspondences are used to assign multiple optical flow hypotheses to each image pixel which are then independently refined to produce a further set of refined hypotheses. One final flow is selected for each pixel from these refined flows by seeking to minimize the local matching error. Dense optical flows from image sequences with small motions are successfully recovered. In image sequences with very large motions, a clear increase in optical flow accuracy is observed when compared to a hierarchical approach to optical flow estimation. Timothy M. A. Smith, David W. Redmill, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP | 4 |
| 2008 | Enhancing video quality for multiple-description MIMO transmission through unequal power allocation between eigen-modesabstractPromising performance of multiple-description coding (MDC) as a suitable video decomposition for MIMO (multiple-input-multiple-output) wireless transmission, reported previously by a number of authors, is further enhanced in this paper by investigating power allocation in MIMO systems that employ singular value decomposition (SVD). Inspired by the water-filling power allocation strategy, results demonstrate that further improvements over conventional, single-description coding (SDC) are possible if, for low SNRs, stronger sub-channel is boosted at the expense of the weaker channel. For high SNR values, the converse is true. A simple yet effective power allocation strategy whereby the bulk of the transmit power - the highest value considered was nine tenths - was allocated to the stronger channel (weaker channel) for very low (very high) SNRs yielded improvements of up to 4 dB (1 dB) in average MDC PSNR over the nominal - equal power allocation - case. The PSNR gains for low SNR values even exceeded the gains arising from water-filling. Milos Tesanovic, David Bull 0001, Angela Doufexi |
ICIP | 2 |
| 2008 | Enhanced MIMO wireless video communication using multiple-description coding
Milos Tesanovic, David Bull 0001, Angela Doufexi, Andrew R. Nix |
Signal Process. Image Commun. | 2 |
| 2007 | Time Varying Volumetric Scene Reconstruction Using Scene FlowabstractTraditional volumetric scene reconstruction algorithms involve the evaluation of many millions of voxels which is highly time consuming. This paper presents an efficient algorithm based of future frame prediction that can dramatically reduce the number of voxels to be evaluated in time varying scenes. The new prediction method, combining scene flow and morphological dilations, is evaluated against a simple model dilation method. Results show the proposed method outperforms a simple dilation method and has the potential to improve the efficiency of volumetric scene reconstruction algorithms while retaining quality given accurate optical flows. 1 Timothy M. A. Smith, David W. Redmill, Cedric Nishan Canagarajah, David Bull 0001 |
BMVC | 4 |
| 2007 | The Effect of Pixel-Level Fusion on Object Tracking in Multi-Sensor Surveillance VideoabstractThis paper investigates the impact of pixel-level fusion of videos from visible (VIZ) and infrared (IR) surveillance cameras on object tracking performance, as compared to tracking in single modality videos. Tracking has been accomplished by means of a particle filter which fuses a colour cue and the structural similarity measure (SSIM). The highest tracking accuracy has been obtained in IR sequences, whereas the VIZ video showed the worst tracking performance due to higher levels of clutter. However, metrics for fusion assessment clearly point towards the supremacy of the multiresolutional methods, especially Dual Tree-Complex Wavelet Transform method. Thus, a new, tracking-oriented metric is needed that is able to accurately assess how fusion affects the performance of the tracker. Nedeljko Cvejic, Stavri G. Nikolov, Henry D. Knowles, Artur Loza, Alin Achim, David Bull 0001, Cedric Nishan Canagarajah |
CVPR | 6 |
| 2007 | Scanpath assessment of visible and infrared side-by-side and fused video displaysabstractAdvances in fusion of multi-sensor inputs have necessitated the creation of more sophisticated fused image assessment techniques. The current work extends previous studies investigating participant accuracy in tracking individuals in a video sequence. Participants were shown visible and IR videos individually and the two video inputs side-by-side, as well as averaged, discrete wavelet transform, and dual- tree complex wavelet transform fused videos. Two scenarios were shown to participants: one featured a camouflaged man walking down a pathway through foliage and across a clearing; the other featured several individuals moving around the clearing. The side-by-side scanpath data were analysed by studying how often participants looked at the visible and infrared sides, and analysing how accurately participants tracked the given target, and compared with previously analysed data. The results of this study are discussed in the context of wider applications to image assessment, and the potential for modelling human scanpath performance. Timothy D. Dixon, Jian Li 0020, Jan Noyes, Tom Troscianko, Stavri G. Nikolov, John Joseph Lewis, Eduardo Fernández Canga, David Bull 0001, Cedric Nishan Canagarajah |
FUSION | 8 |
| 2007 | Statistical Model-based fusion of noisy multi-band images in the wavelet domainabstractA new method for multimodal image fusion, based on statistical modelling of wavelet coefficients, is proposed in this paper. The algorithm draws from the Weighted Average scheme, but incorporates Laplacian bivariate parent-child statistical dependencies. The interscale dependency is brought in the form of shrinkage functions. The proposed method has been shown to perform very well with noisy datasets, outperforming other conventional methods in terms of fusion quality and noise reduction in the fused output. Artur Loza, Alin Achim, David Bull 0001, Cedric Nishan Canagarajah |
FUSION | 3 |
| 2007 | Colour Volumetric Compression for Realistic View Synthesis ApplicationsabstractThe colour volumetric data which is constructed from a set of multi-view images is capable of providing realistic immersive experience. However it is not widely applicable due to its manifold increase in bandwidth. This paper presents a novel framework to achieve scalable volumetric compression. Based on wavelet transformation, data rearrangement algorithm is proposed to compact volumetric data leading to high efficiency of transformation. The colour data is also rearranged by using the characteristics of human eye sensitivity. Moreover, the pre-processing for adaptive resolution is proposed in this paper. The low resolution overcomes the limitation of the data transmission at low bit rate, whilst the fine resolution improves the synthesised images' quality. The results of our proposed schemes show significant improvement of the compression performance over the traditional 3D coding. Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001 |
ICASSP (1) | 4 |
| 2007 | Fusion Methods for Side Information Generation in Multi-View Distributed Video Coding SystemsabstractThe generation of the side information is an important part in the design of a distributed video coding (DVC) system as it directly relates to the system's rate distortion performance. In multi-view systems spatial/view inter-camera correlations can be exploited alongside temporal/motion intra camera ones for generating the side information as accurately as possible. In this paper, algorithms for fusing multiple such side information estimates, generated with temporal and view interpolation methods, are proposed. Two algorithms are suggested along with weighted fusion schemes for combining multiple methods. The results presented indicate performance improvements of up to 2 dB compared to exiting fusion approaches. Pierre Ferré, Dimitris Agrafiotis, David Bull 0001 |
ICIP (6) | 3 |
| 2007 | Analysis of IEEE 802.11N-Like Transmission Techniques with and Without Prior CSI for Video ApplicationsabstractPrevious research into MIMO systems has focused little on optimisation for video transmission. In terms of multimedia transmission, spatial multiplexing (SM) is commonly proposed as the most suitable MIMO technique. Most SM-based video transport schemes look to exploit the multiplexing gain, which comes at the expense of a relatively high SNR, unless the channel state is known at the transmitter. Space-time block coding (STBC) is an attractive technique that does not provide the mapping flexibility of SM techniques, but does dramatically reduce the packet-error rate. This paper compares these techniques in terms of decoded video quality through simulations based on practical transmission scenarios. The use of multiple-description coding (MDC) is further proposed to provide a new class of wireless video transmission algorithms. Milos Tesanovic, David Bull 0001, Angela Doufexi, Andrew R. Nix |
ICIP (6) | 2 |
| 2007 | A Novel Secure H.264 Transcoder using Selective EncryptionabstractIn digital broadcast TV systems, video data is normally encrypted before transmission. For in-home redistribution, it is often necessary to transcode the bitstream to achieve optimum utilization of available bandwidth. If a signal is decrypted before transcoding and re-encrypted, this may lead to a security loophole. This paper presents a solution in the form of a novel H.264 selective encryption algorithm that encrypts sign bits of transform coefficients and motion vectors to allow secure transcoding without decryption. The performance of this system is compared with I-frame encryption. The results show that sign encryption is more secure than I-frame encryption and has a lower complexity. A hybrid system using a modified transcoder and sign encryption is found to give an optimal compromise between security and transcoding performance. Nithin M. Thomas, Damien Lefol, David Bull 0001, David W. Redmill |
ICIP (4) | 3 |
| 2007 | Frame Delay and Loss Analysis for Video Transmission over time-correlated 802.11A/G channelsabstractThis paper presents simulation results for the transmission of unicast MAC frames over 802.11a/g. Fading channel models at various Doppler frequencies are developed to generate time- correlated SNR waveforms. These are then used together with a bit accurate MAC/PHY simulator to estimate the frame loss rate, the transmission delay, and the jitter for a steady flow of transmit frames. Time correlated channels are required to correctly simulate the bursty nature of packet loss in a wireless channel. The Doppler spread is shown to have a strong effect on the performance of the ARQ mechanism in the MAC layer. Delay is computed as the sum of the transmission delay and the accumulated queuing delay in the MAC buffer. Delay and frame loss are compared for time correlated and time uncorrelated fading channels. Compared to the slow fading case, in a fast fading channel fewer retransmissions are required and the end-to-end delay is significantly reduced. When channel conditions are poor the simulated delay and frame loss rate are seriously underestimated when time uncorrelated fading is assumed. To analyze the performance of video codecs, we show that a time correlated channel model must be combined with a dedicated 802.11a/g MAC/PHY simulation. Victoria Sgardoni, Pierre Ferré, Angela Doufexi, Andrew R. Nix, David Bull 0001 |
PIMRC | 5 |
| 2007 | Hybrid key/Wyner-Ziv frames with flexible macroblock ordering for improved low delay distributed video codingabstractThis paper proposes a concealment based approach to generating the side information and estimating the correlation noise for low-delay, pixel-based, distributed video coding. The proposed method employs a macroblock pattern similar to the one used in the dispersed type FMO of H.264 for grouping the macroblocks of each frame into intra coded (key) and Wyner-Ziv groups. Temporal concealment is then used at the decoder for "concealing" the missing macroblocks (estimating the side information-predicting the Wyner-Ziv macroblocks). The actual intra coded/decoded macroblocks are used for estimating the correlation noise. The results indicate significant performance improvements relative to existing motion extrapolation based approaches (up to 25% bit rate reduction). Dimitris Agrafiotis, Pierre Ferré, David Bull 0001 |
VCIP | 3 |
| 2007 | Comparison of standard-based H.264 error-resilience techniques and multiple-description coding for robust MIMO-enabled video transmissionabstractMIMO (multiple-input-multiple-output) systems offer potential for throughput increase and enhanced quality of service for multimedia transmission. The underlying multipath environment requires new error-resilience techniques if the obtained benefits are to be fully exploited. Different MIMO architectures produce error-patterns of somewhat diverse characteristics. This paper proposes the use of multiple-description coding (MDC) as an approach that outperforms the standard-based error-resilience techniques in the majority of these cases. Results obtained from the random packet-error generator are furthered through the use of realistic MIMO channel scenarios and argue in favour of the deployment of an MDC-based video transmission system. Singular value decomposition (SVD) is used to create orthogonal sub-channels within a MIMO system which provide, depending on their respective gains and fading characteristics, an efficient means of mapping video content. Results indicate improvements in average PSNR of decoded test-sequences of up to 3 dB (5dB in the region of high PERs) compared to standard, single-description video transmission. This is also supported by significant subjective quality enhancements. Milos Tesanovic, David Bull 0001, Dimitris Agrafiotis, Angela Doufexi |
VCIP | 2 |
| 2007 | Applied Multi-Dimensional FusionabstractThe purpose of the Applied Multi-dimensional Fusion Project is to investigate the benefits that data fusion and related techniques may bring to future military Intelligence Surveillance Target Acquisition and Reconnaissance systems. In the course of this work, it is intended to show the practical application of some of the best multi-dimensional fusion research in the UK. This paper highlights the work done in the area of multi-spectral synthetic data generation, super-resolution, joint fusion and blind image restoration, multi-resolution target detection and identification and assessment measures for fusion. The paper also delves into the future aspirations of the work to look further at the use of hyper-spectral data and hyper-spectral fusion. The paper presents a wide work base in multi-dimensional fusion that is brought together through the use of common synthetic data, posing real-life problems faced in the theatre of war. Work done to date has produced practical pertinent research products with direct applicability to the problems posed. Asher Mahmood, Philip M. Tudor, William Oxford, Robert Hansford, James D. B. Nelson, Nick G. Kingsbury, Antonis Katartzis, Maria Petrou, Nikolaos Mitianoudis, Tania Stathaki, Alin Achim, David Bull 0001, Cedric Nishan Canagarajah, Stavri G. Nikolov, Artur Loza, Nedeljko Cvejic |
Comput. J. | 12 |
| 2007 | Sequential Monte Carlo tracking by fusing multiple cues in video sequences
Paul Brasnett, Lyudmila Mihaylova, David Bull 0001, Cedric Nishan Canagarajah |
Image Vis. Comput. | 3 |
| 2007 | An efficient complexity-scalable video transcoder with mode refinement
Damien Lefol, David Bull 0001, Cedric Nishan Canagarajah, David W. Redmill |
Signal Process. Image Commun. | 2 |
| 2007 | Towards efficient context-specific video coding based on gaze-tracking analysisabstractThis article discusses a framework for model-based, context-dependent video coding based on exploitation of characteristics of the human visual system. The system utilizes variable-quality coding based on priority maps which are created using mostly context-dependent rules. The technique is demonstrated through two case studies of specific video context, namely open signed content and football sequences. Eye-tracking analysis is employed for identifying the characteristics of each context, which are subsequently exploited for coding purposes, either directly or through a gaze prediction model. The framework is shown to achieve a considerable improvement in coding efficiency. Dimitris Agrafiotis, Sam J. C. Davies, Cedric Nishan Canagarajah, David Bull 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2007 | Mobility Tracking in Cellular Networks Using Particle FilteringabstractMobility tracking based on data from wireless cellular networks is a key challenge that has been recently investigated both from a theoretical and practical point of view. This paper proposes Monte Carlo techniques for mobility tracking in wireless communication networks by means of received signal strength indications. These techniques allow for accurate estimation of mobile station's (MS) position and speed. The command process of the MS is represented by a first-order Markov model which can take values from a finite set of acceleration levels. The wide range of acceleration changes is covered by a set of preliminary determined acceleration values. A particle filter and a Rao-Blackwellised particle filter are proposed and their performance is evaluated both over synthetic and real data. A comparison with an extended Kalman filter (EKF) is performed with respect to accuracy and computational complexity. With a small number of particles the RBPF gives more accurate results than the PF and the EKF. A posterior Cramer Rao lower bound (PCRLB) is calculated and it is compared with the filters' root- mean-square error performance. Lyudmila Mihaylova, Donka S. Angelova, S. Honary, David Bull 0001, Cedric Nishan Canagarajah, Branko Ristic 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2006 | Adaptive Region-Based Multimodal Image Fusion Using ICA BasesabstractIn this paper, we present a novel multimodal image fusion algorithm in ICA domain. It uses segmentation to determine the most important regions in the input images and consequently fuses the ICA coefficients from given regions using the Piella fusion metric to maximise the quality of the fused image. The proposed method exhibits significantly higher performance than the basic ICA algorithm and improvement over other state-of-the-art algorithms Nedeljko Cvejic, David Bull 0001, Cedric Nishan Canagarajah |
FUSION | 3 |
| 2006 | Scanpath Analysis of Fused Multi-Sensor Images with Luminance Change: A Pilot StudyabstractImage fusion is the process of combining images of differing modalities, such as visible and infrared (IR) images. Significant work has recently been carried out comparing methods of fused image assessment, with findings strongly suggesting that a task-centred approach would be beneficial to the assessment process. The current paper reports a pilot study analysing eye movements of participants involved in four tasks. The first and second tasks involved tracking a human figure wearing camouflage clothing walking through thick undergrowth at light and dark luminance levels, whilst the third and fourth task required tracking an individual in a crowd, again at two luminance levels. Participants were shown the original visible and IR images individually, pixel-averaged, contrast pyramid, and dual-tree complex wavelet fused video sequences. They viewed each display and sequence three times to compare inter-subject scanpath variability. This paper describes the initial analysis of the eye-tracking data gathered from the pilot study. These were also compared with computational metric assessment of the image sequences Timothy D. Dixon, Jian Li 0020, Jan Noyes, Tom Troscianko, Stavri G. Nikolov, John Joseph Lewis, Eduardo Fernández Canga, David Bull 0001, Cedric Nishan Canagarajah |
FUSION | 8 |
| 2006 | Uni-Modal Versus Joint Segmentation for Region-Based Image FusionabstractA number of segmentation techniques are compared with regard to their usefulness for region-based image and video fusion. In order to achieve this, a new multi-sensor data set is introduced containing a variety of infra-red, visible and pixel fused images together with manually produced "ground truth" segmentations. This enables the objective comparison of joint and unimodal segmentation techniques. A clear advantage to using joint segmentation over unimodal segmentation, when dealing with sets of multi-modal images, is shown. The relevance of these results to region-based image fusion is confirmed with task-based analysis and a quantitative comparison of the fused images produced using the various segmentation algorithms John Joseph Lewis, Stavri G. Nikolov, Cedric Nishan Canagarajah, David Bull 0001, Alexander Toet |
FUSION | 4 |
| 2006 | Structural Similarity-Based Object Tracking in Video SequencesabstractThis paper addresses the problem of object tracking in video sequences. The use of a structural similarity measure for tracking is proposed. The measure reflects the distance between two images by comparing their structural and spatial characteristics and has shown to be robust to illumination and contrast changes. As a result it guarantees robustness of the tracking process under changes in the environment. The previously used Bhattacharyya distance is not robust to such changes. Additionally, when a tracker is run with the Bhattacharyya distance, histograms should be calculated in order to find the likelihood function of the measurements. With the new function there is no need to calculate histograms. A particle filter (PF) is implemented where this measure is used for computing the distance between the reference and current frame. The algorithm performance has been tested and evaluated over real-world video sequences, and has been shown to outperform methods based on colour and edge histograms Artur Loza, Lyudmila Mihaylova, Cedric Nishan Canagarajah, David Bull 0001 |
FUSION | 4 |
| 2006 | Algorithms for Mobile Nodes Self-Localisation in Wireless Ad Hoc NetworksabstractThis paper addresses the problem of position localisation of mobile nodes in ad hoc wireless networks based on received signal strength indicator measurements. Node mobility is modelled as a linear system driven by a discrete command Markov process. Self-localisation of mobile nodes is performed via an interacting multiple model filter consisting of a bank of unscented Kalman filters (IMM-UKF). Estimation of the mobility state, which comprises the position, speed and acceleration of the mobile nodes is accomplished. The performance of the IMM- UKF filter is investigated and compared to a multiple model particle filter (MM PF) by Monte Carlo simulation Lyudmila Mihaylova, Donka S. Angelova, Cedric Nishan Canagarajah, David Bull 0001 |
FUSION | 4 |
| 2006 | Improving Performance of ICA Domain Multimodal Image Fusion in Presence of NoiseabstractIn this paper we present a novel fusion method in ICA domain. The method uses nonlinear shrinkage of the coefficients in ICA domain to determine the ICA coefficients to be used in the fused image reconstruction, so that the noise transferred from input images into the fused output is minimized. Experimental results have shown that it has improved performance in the presence of noise, compared to the basic ICA algorithm. Compared to standard multiresolutional fusion methods, the noise in the fused image is visually less annoying, while the important detail is still adequately represented. The proposed method exhibits considerably higher fusion performance, measured by Piella and Petrovic fusion metric, than the basic ICA algorithm and improvement over other state-of-the-art algorithms. Nedeljko Cvejic, David Bull 0001, Cedric Nishan Canagarajah |
FUZZ-IEEE | 2 |
| 2006 | Dynamic Programming for Multi-View Disparity/Depth EstimationabstractA novel algorithm for disparity/depth estimation from multi-view images is presented. A dynamic programming approach with window-based correlation and a novel cost function is proposed. The smoothness of disparity/depth map is embedded in dynamic programming approach, whilst the window-based correlation increases reliability. The enhancement methods are included, i.e. adaptive window size and shiftable window are used to increase reliability in homogenous areas and to increase sharpness at object boundaries. First, the algorithms estimate depth maps along a single camera axis. The algorithms exploits then combines the depth estimates from different axis to derive a suitable depth map for multi-view images. The proposed scheme outperforms existing approaches in parallel and in the non-parallel camera configurations Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001 |
ICASSP (2) | 4 |
| 2006 | Multiple Priority Region of Interest Coding with H.264abstractThis paper describes a modified rate control algorithm for H.264 that can accommodate multiple priority levels given a region of interest (Rol). The modified method allows better control of the quality of the RoI and gradual variation of the quality in the rest of the video frame through a bit redistribution process that is based on a number of parameters, including characteristics of the Rol, user input and perceptual factors. Dimitris Agrafiotis, David Bull 0001, Cedric Nishan Canagarajah, Nawat Kamnoonwatana |
ICIP | 2 |
| 2006 | Volumetric Representation for Sparse Multi-ViewsabstractIn this paper, we propose a novel volumetric representation for a sparse set of calibrated multi-view images of a non-Lambertian scene. The depth map of each reference view is registered into a volume and a simple algorithm to shape the volume is introduced. Particular colours are defined for each voxel to render a smooth and realistic image. Synthesized results demonstrate the good performance with the proposed scheme, both for parallel-camera and non-parallel-camera geometries. Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001 |
ICIP | 4 |
| 2006 | Region-Based Multimodal Image Fusion using ICA BasesabstractIn this paper, we present a novel region-based multimodal image fusion algorithm in the ICA domain. It uses segmentation to determine the most important regions in the input images and consequently fuses the ICA coefficients from the given regions. The proposed method exhibits considerably higher performance than the basic ICA algorithm and shows improvement over other state-of-the-art algorithms Nedeljko Cvejic, John Joseph Lewis, David Bull 0001, Cedric Nishan Canagarajah |
ICIP | 3 |
| 2006 | Macroblock-Level Mode Based Adaptive in-Band Motion Compensated Temporal FilteringabstractThis paper presents an adaptive in-band motion compensated temporal filtering (MCTF) scheme for 3D wavelet based scalable video coding. The proposed scheme solves the motion mismatch problem when motion vectors from the LL subband are inaccurately applied to the highpass subbands in decoding high spatial resolution video. Specifically, we compare the macroblock residue energy in the highpass frames obtained by using motion vectors from both the LL and highpass subbands, and then adaptively transmit different sets of motion vectors based on whether mismatch has occurred in the highpass subbands. Macroblocks in the higher temporal levels favour the selection of highpass subbands' motion vectors because the motion estimation process becomes less accurate as temporal level increases. The modes information, which specifies whether the LL subband motion vectors or the highpass subbands' motion vectors are used by the current macroblock, is coded by run-length coding. Experimental results show that the proposed scheme improves both the visual quality and PSNR for high resolution decoding with comparison to other in-band MCTF schemes. Furthermore, our scheme requires only modifications when performing MCTF in the highpass subbands, thus, the original strength of in-band MCTF for decoding low spatial resolution video is well preserved. Anyu Gao, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP | 3 |
| 2006 | Interpolation Free Sub-Pixel Motion Estimation for H.264abstractSub-pixel motion compensation plays an important role in compression efficiency within modern video codecs such as MPEG2, MPEG4 and H.264. Sub-pixel motion compensation is implemented within these standards using interpolated pixel values at 1/2 or 1/4 pixel accuracy. Such interpolation gives a good reduction in residual energy for each predicted macroblock and therefore improves compression. However, such interpolation is very computationally complex for the encoder. This is especially true for H.264 where the cost of an exhaustive set of macroblock segmentations need to be estimated in order to obtain an optimal mode for prediction. This paper presents a novel interpolation-free scheme for sub-pixel motion compensation using the result of the full pixel SAD distribution of each motion compensated block applied to an H.264 encoder. This system produces reduced complexity motion compensation with a controllable trade-off between compression performance and encoder speed. These methods facilitate the generation of a real time software H.264 encode. Paul R. Hill, Tuan-Kiang Chiew, David Bull 0001 |
ICIP | 3 |
| 2006 | Mode Refinement Algorithm for H.264 Inter Frame RequantizationabstractThe latest video coding standard H.264 has been recently approved and has already been adopted for numerous applications including HD-DVD and satellite broadcast. To allow interconnectivity between different applications using H.264, transcoding will be a key factor. When requantizing a bitstream the incoming coding decisions are usually kept unchanged to reduce the complexity, but it can have a major impact on the coding efficiency. This paper proposes a novel algorithm for mode refinement of inter prediction in the case of requantization of H.264 bitstreams. The proposed approach gives a comparable quality to a full search for a fraction of its complexity by exploiting the statistical properties of the mode distribution and motion vector refinement. Damien Lefol, David Bull 0001 |
ICIP | 2 |
| 2006 | Enhanced spatial error concealment with directional entropy based interpolation switchingabstractThis paper describes a spatial error concealment method that uses edge related information for concealing missing macroblocks in a way that not only preserves existing edges but also avoids introducing new strong ones. The method relies on a novel switching algorithm which uses the directional entropy of neighboring edges for choosing between two interpolation methods, a directional along detected edges or a bilinear using the nearest neighboring pixels. Results show that the performance of the proposed method is subjectively and objectively (PSNR wise) better compared to both 'single interpolation' and to edge strength based switching methods Dimitris Agrafiotis, David Bull 0001, Cedric Nishan Canagarajah |
ISCAS | 2 |
| 2006 | Mode refinement algorithm for H.264 intra frame requantizationabstractThe latest video coding standard H.264 has been recently approved and has already been adopted for numerous applications including HD-DVD and satellite broadcast. To allow interconnectivity between different applications using H.264, transcoding will be a key factor. When requantizing a bit stream the incoming coding decisions are usually kept unchanged to reduce the complexity, but it can have a major impact on the coding efficiency. This paper proposes a novel algorithm for mode refinement of intra prediction in the case of requantization of H.264 intra frames. The proposed approach gives a comparable quality to a full search for a fraction of its complexity by exploiting the statistical properties of the mode distribution Damien Lefol, David Bull 0001, Cedric Nishan Canagarajah |
ISCAS | 2 |
| 2006 | Enhanced Error-Resilient Video Transport Over MIMO Systems Using Multiple DescriptionsabstractMuch of the work on wireless transmission over the past several years has focused on simulation and deployment of multiple-input-multiple-output (MIMO) systems. These systems provide benefits of improved robustness and enhanced throughput at relatively low cost. Despite the increased understanding of the performance of MIMO systems, little is known about which combination of channel and source coding yields the best results for video transport. It is clear that new ways of providing error-resilience that emerge from MIMO architectures need to be developed which can cope with the particularities of video content. This paper proposes a new scheme for video transmission using multiple-description coding (MDC). Two complementary MIMO techniques, space-time block coding (STBC) and spatial multiplexing (SM), are employed. The quality of the reconstructed video, already enhanced by the inherent MIMO systems' properties, is further improved through the use of MDC. Milos Tesanovic, David Bull 0001, Angela Doufexi |
VTC Fall | 2 |
| 2006 | Spatial error concealment with edge related perceptual considerations
Dimitris Agrafiotis, David Bull 0001, Cedric Nishan Canagarajah |
Signal Process. Image Commun. | 2 |
| 2006 | A perceptually optimised video coding system for sign language communication at low bit rates
Dimitris Agrafiotis, Cedric Nishan Canagarajah, David Bull 0001, Jim Kyle, Helen Seers, Matt Dye |
Signal Process. Image Commun. | 3 |
| 2006 | Methods for the assessment of fused imagesabstractThe prevalence of image fusion---the combination of images of different modalities, such as visible and infrared radiation---has increased the demand for accurate methods of image-quality assessment. The current study used a signal-detection paradigm, identifying the presence or absence of a target in briefly presented images followed by an energy mask, which was compared with computational metric and subjective quality assessment results. In Study 1, 18 participants were presented with fused infrared-visible light images, with a soldier either present or not. Two independent variables, image-fusion method (averaging, contrast pyramid, dual-tree complex wavelet transform) and JPEG compression (no compression, low and high compression), were used in a repeated-measures design. Participants were presented with images and asked to state whether or not they detected the target. In addition, subjective ratings and metric results were obtained. This process was repeated in Study 2, using JPEG2000 compression. The results showed a significant effect for fusion but not compression in JPEG2000 images, while JPEG images showed significant effects for both fusion and compression. Subjective ratings differed, especially for JPEG2000 images, while metric results for both JPEG and JPEG2000 showed similar trends. These results indicate that objective and subjective ratings can differ significantly, and subjective ratings should, therefore, be used with care. Timothy D. Dixon, Eduardo Fernández Canga, Jan Noyes, Tom Troscianko, Stavri G. Nikolov, David Bull 0001, Cedric Nishan Canagarajah |
ACM Trans. Appl. Percept. | 6 |
| 2006 | Enhanced Error Concealment With Mode SelectionabstractDelay sensitive video transmission over error prone networks can suffer from packet erasures when channel conditions are not favorable. Use of error concealment (EC) at the video decoder is necessary in such cases to prevent error induced artefacts making the affected video frames visibly intolerable. This paper proposes an EC method that incorporates enhanced temporal and spatial concealment elements, the use of which is controlled by a mode selection (MS) algorithm well matched to the characteristics of the temporal concealment approach. The performance of the individual enhancements and of the MS algorithm are compared with the respective features of the method employed in the H.264 joint model (JM) decoder and with other state of the art methods. The overall performance of the proposed method is shown to offer significant gains (up to 9 dB) compared to that of the JM decoder for a wide range of natural and animation image sequences without any considerable increase in complexity Dimitris Agrafiotis, David Bull 0001, Cedric Nishan Canagarajah |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Interpolation Free Subpixel Accuracy Motion EstimationabstractSubpixel motion estimation plays an important role in compression efficiency within modern video codecs such as MPEG2, MPEG4, and H.264. Subpixel motion estimation is implemented within these standards using interpolated values at 1/2 or 1/4 pixel accuracy. Such interpolation gives a good reduction in residual energy for each predicted macroblock and, therefore, improves compression. However, this leads to a significant increase in computational complexity at the encoder. This is especially true for H.264 where the cost of an exhaustive set of macroblock segmentations need to be estimated in order to obtain an optimal mode for prediction. This paper presents a novel interpolation-free scheme for subpixel motion estimation using the result of the full pixel sum of absolute difference distribution of each motion compensated block applied to an H.264 encoder. This system produces reduced complexity motion estimation with a controllable tradeoff between compression performance and encoder speed. These methods facilitate the generation of a real time software H.264 encoder Paul R. Hill, Tuan-Kiang Chiew, David Bull 0001, Cedric Nishan Canagarajah |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | On the performance of fast motion estimation for multiple-reference motion compensated predictionabstractMultiple reference motion compensated prediction achieves significant prediction gains, but at the expense of a significant increase in computational complexity. This paper investigates and compares the performance of a number of fast motion estimation techniques when extended to work for multiple reference prediction. In general, algorithms that are designed to exploit the properties of the long-term block motion field, were found to provide superior performance and better compromise between prediction quality and computational complexity. Mohammed E. Al-Mualla, David Bull 0001 |
ICIP (1) | 2 |
| 2005 | Multi-view image coding with wavelet lifting and in-band disparity compensationabstractThis paper presents a novel framework to achieve scalable multi-view image coding. As open loop operation, the wavelet lifting scheme for geometric filtering has been exploited to overcome the limitation of SNR scalability and to attain view scalability. The essential key for achieving the spatial scalability is the in-band prediction. It removes correlations among subbands level-by-level via shift-invariant references obtained by overcomplete discrete wavelet transforms (ODWT). Additionally, the proposed disparity compensated view filtering is allowed to exploit the different filters and estimation parameters for each resolution level. The experiments show comparable results at full resolution and the significant improvement at coarser resolution over the conventional spatial prediction scheme. Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP (3) | 3 |
| 2005 | Image fusion using a new framework for complex wavelet transformsabstractImage fusion is the process of extracting meaningful visual information from two or more images and combining them to form one fused image. Image fusion is important within many different image processing fields from remote sensing to medical applications. Previously, real valued wavelet transforms have been used for image fusion. Although this technique has provided improvements over more naive methods, this transform suffers from the shift variance and lack of directionality associated with its wavelet bases. These problems have been overcome by the use of a reversible and discrete complex wavelet transform (the dual tree complex wavelet transform DT-CWT). However, the existing structure of this complex wavelet decomposition enforces a very strict choice of filters in order to achieve a necessary quarter shift in coefficient output. This paper therefore introduces an alternative structure to the DT-CWT that is more flexible in its potential choice of filters and can be implemented by the combination of four normally structured wavelet transforms. The use of these more common wavelet transforms enables this method to make use of existing optimised wavelet decomposition and recomposition methods, code and filter choice. Paul R. Hill, David Bull 0001, Cedric Nishan Canagarajah |
ICIP (2) | 2 |
| 2005 | Error concealment for slice group based multiple description video codingabstractThis paper develops error concealment methods for multiple description video coding (MDC) in order to adapt to error prone packet networks. The three-loop slice group MDC approach of D. Wang et al. (2005) is used. MDC is very suitable for multiple channel environments, and especially able to maintain acceptable quality when some of these channels fail completely, i.e. in an on-off MDC environment, without experiencing any drifting problem. Our MDC scheme coupled with the proposed concealment approaches proved to be suitable not only for the on-off MDC environment case (data from one channel fully lost), but also for the case where only some packets are lost from one or both channels. Copying video and using motion vectors from correct descriptions are combined together for concealment prior to applying traditional methods. Results are compared to the traditional error concealment method proposed in the H.264 reference software, showing significant improvements for both the balanced and unbalanced channel cases. Cedric Nishan Canagarajah, Dimitris Agrafiotis, David Bull 0001 |
ICIP (1) | 4 |
| 2005 | Combined morphological-spectral unsupervised image segmentationabstractThe goal of segmentation is to partition an image into disjoint regions, in a manner consistent with human perception of the content. For unsupervised segmentation of general images, however, there is the competing requirement not to make prior assumptions about the scene. Here, a two-stage method for general image segmentation is proposed, which is capable of processing both textured and nontextured objects in a meaningful fashion. The first stage extracts texture features from the subbands of the dual-tree complex wavelet transform. Oriented median filtering is employed, to circumvent the problem of texture feature response at step edges in the image. From the processed feature images, a perceptual gradient function is synthesised, whose watershed transform provides an initial segmentation. The second stage of the algorithm groups together these primitive regions into meaningful objects. To achieve this, a novel spectral clustering technique is proposed, which introduces the weighted mean cut cost function for graph partitioning. The ability of the proposed algorithm to generalize across a variety of image types is demonstrated. Robert J. O'Callaghan, David Bull 0001 |
IEEE Trans. Image Process. | 2 |
| 2004 | Gaze-contingent display using texture mapping and OpenGL: system and applicationsabstractThis paper describes a novel gaze-contingent display (GCD) using texture mapping and OpenGL. This new system has a number of key features: (a) it is platform independent, i.e. it runs on different computers and under different operating systems; (b) it is eyetracker independent, since it provides an interactive focus+context display that can be easily integrated with any eye-tracker that provides real-time 2-D gaze estimation; (c) it is flexible in that it provides for straightforward modification of the main GCD parameters, including size and shape of the window and its border; and (d) through the use of OpenGL extensions it can perform local real-time image analysis within the GCD window. The new GCD system implementation is described in detail and some performance figures are given. Several applications of this system are studied, including gaze-contingent multi-resolution displays, gaze-contingent multi-modality displays, and gaze-contingent image analysis. Stavri G. Nikolov, Timothy D. Newman, David Bull 0001, Cedric Nishan Canagarajah, Michael G. Jones, Iain D. Gilchrist |
ETRA | 3 |
| 2004 | A video coding system for sign language communication at low bit ratesabstractThe ability to communicate remotely through the use of video as promised by wireless networks and already practised over fixed networks, is for deaf people as important as voice telephony is for hearing people. Sign languages are visual-spatial languages and as such demand good image quality for interaction and understanding. In this paper, based on analysis of the viewer's perceptual behavior and the video content involved we propose a sign language video coding system using foveated processing, which can lead to bit rate savings without compromising the comprehension of the coded sequence. We support this claim with the results of an initial comprehension assessment trial of such coded sequences by deaf users. Dimitris Agrafiotis, Cedric Nishan Canagarajah, David Bull 0001, Jim Kyle, Helen Seers, Matt Dye |
ICIP | 3 |
| 2004 | Multi-resolution parametric region tracking for 2D object replacement in videoabstractThis paper develops an efficient parametric multiresolution region tracking algorithm which is applied to the task of object replacement in video sequences. The tracker relies on gradient-based techniques to provide efficient estimation of the target location and pose in each frame. The tracking algorithm improves upon similar efficient parametric algorithms by increasing the distance an object can move from frame to frame. Experimental results provide clear evidence of the improved performance over existing region tracking. The parameters that estimate the pose and location of the target are used to initialise the object replacement algorithm. The object replacement uses the pose estimates with texture mapping techniques to perform image warping of an arbitrary sized replacement image. The location estimates are used to accurately insert the replacement image into the current frame. Paul Brasnett, David Bull 0001, Cedric Nishan Canagarajah |
ICIP | 2 |
| 2004 | Slice group based multiple description video coding using motion vector estimationabstractThis paper proposes a novel scheme for multiple description video coding approach using slice group coding tool proposed in H.264. Independent motion compensation loops are maintained for two descriptions in the encoder. In each description, one slice group is encoded as main information, while the other slice group is encoded very coarsely, as redundancy, to keep basic information. This coarse slice group can be encoded using normal encoding with coarse quantizer, or using predicted motion vectors by the main slice group. Mode decision is made to select best encoding method. Results show that our scheme works very well and keeps subjective quality very good for middle and higher bitrate. Redundancy and side quality can be controlled by changing parameters of coarsely coded slice group, unlike some other MDC methods with fixed redundancy. Cedric Nishan Canagarajah, David Bull 0001 |
ICIP | 3 |
| 2004 | Queue-based block matching algorithm for video compression and motion segmentationabstractThis paper addresses two issues related to motion estimation using the block matching algorithms (BMA): (1) determining the reliability of the motion vectors of each block, and (2) imposing smoothness constraint to the motion vector field. We introduce a new robust reliability measure to represent the confidence level of the motion vector from the cost function distribution and propose a novel algorithm that incorporates smoothness constraint into the motion vector field evaluation by implementing a priority queue structure based on the reliability measure. In this framework, a smooth motion vector field is evaluated in a single pass without going through iterations typical of many existing optical flow estimation algorithms. Hence it is fast and can easily be incorporated into real-time applications for video compression as well as image segmentation. Tuan-Kiang Chiew, James T. Chung-How, David Bull 0001, Cedric Nishan Canagarajah |
VCIP | 3 |
| 2004 | Interpolation-free subpixel refinement for block-based motion estimationabstractThis paper proposes a low-complexity sub-pixel refinement to motion estimation based on full-search block matching algorithm (BMA) at integer-pixel accuracy. This algorithm eliminates the need to produce interpolated reference frames, which is may be too memory- and processor- intensive, for some real-time mobile applications. The algorithm assumes the BMA is done at pixel resolution and the (sum-of-absolute-differences) SADs of the candidate motion vector and its neighbouring vectors are available for each block. The proposed method than models the SAD distribution around the candidate motion vector and its neighbouring points. Actual minimum point at sub-pixel resolution is then computed according to the model used. 3 variations of the parabolic model are considered and simulations using the H.263 standard encoder on several test sequences reveal an improvement of 1.0 dB over integer-accuracy motion estimation. Albeit its simplicity, some test cases come close to the results obtained by actual interpolated reference frames. Tuan-Kiang Chiew, James T. Chung-How, David Bull 0001, Cedric Nishan Canagarajah |
VCIP | 3 |
| 2004 | Throughput analysis of IEEE 802.11 and IEEE 802.11e MACabstractThis work focuses on the throughput analysis of the IEEE 802.11 medium access control (MAC). In the IEEE 802.11 standard, the main access scheme is called the distributed coordination function (DCF) and is the basis for the other access schemes, such as the point coordination function (PCF) and the enhanced DCF (EDCF) of the new IEEE 802.11e MAC. Since the IEEE 802.11 MAC does not support quality of service (QoS), the IEEE is currently working on the final draft of an enhanced version known as the IEEE 802.11e. In this enhanced standard, QoS and service differentiations are supported. The throughput and delay performance of the DCF/PCF of the IEEE 802.11 MAC are presented for different packet lengths and different numbers of users. Throughput performances are also detailed for the EDCF. Pierre Ferré, Angela Doufexi, Andrew R. Nix, David Bull 0001 |
WCNC | 4 |
| 2004 | Range and throughput enhancement of wireless local area networks using smart sectorised antennasabstractAt present, wireless local area networks (WLANs) such as HiperLAN2 and 802.11a are being developed and deployed around the world. In this letter, the use of sectorized antennas is considered as a means to improve the physical layer performance of WLANs. Results demonstrate that throughput and range can be enhanced and/or the transmit power can be reduced. However, these benefits are achieved with a small increase in multiple access protocol overhead. Simulations are performed using measured wideband channels in the 5-GHz band. In cases where the channel exhibits strong Rician characteristics, gains as high as 13 dB are demonstrated. These benefits substantially outweigh the associated medium access control overhead. Angela Doufexi, Simon Armour, Andrew R. Nix, Peter Karlsson, David Bull 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2003 | A simple scheme for enhanced reassignment of the smoothed pseudo Wigner-Ville representation of noisy signalsabstractThe reassignment procedure has often been employed to improve readability of some time-frequency representations (TFRs). When processing noisy signals, the problem of sensitivity of the technique to noise is encountered. A simple modification of the reassignment method is proposed, based on a thresholding operation. Specifically, by preventing the reassignment of the distribution coefficients below the noise dependent threshold and replacing them with zeros, the enhanced signatures on the time-frequency plane are obtained. This method is compared with other techniques, such as the reassigned spectrogram (RSP) and the supervised reassigned spectrogram (SRSP). An experimental test of these algorithms as the instantaneous frequency (IF) estimators for a chirp signal have shown that our method improves the accuracy of the estimation for heavy noise. Artur Loza, Cedric Nishan Canagarajah, David Bull 0001 |
ICASSP (6) | 3 |
| 2003 | An evaluation of the performance of IEEE 802.11a and 802.11g wireless local area networks in a corporate office environmentabstractIn recent years there has been considerable interest in the development of standards for wireless local area networks. In particular, IEEE's 802.11a and 802.11g both employ coded orthogonal frequency division multiplexing (COFDM) but operate in different frequency bands. In this paper, the performance and relative merits of 802.11a and 802.11g are compared for the scenario of a corporate office wireless LAN application. It is shown that for comparable scenarios 802.11g achieves superior range but that 802.11a achieves higher data rates. Thus the two standards are found to have complimentary strengths and weaknesses. Angela Doufexi, Simon Armour, Beng-Sin Lee, Andrew R. Nix, David Bull 0001 |
ICC | 5 |
| 2003 | Region of interest coding of volumetric medical imagesabstractThree-dimensional wavelet coding of volumetric medical images provides better coding performance compared to corresponding 2D methods by exploiting the inter-slice correlation that exists in such data. It introduces however latencies when it comes to transmitting specific parts of the volume. This paper presents an extension to 3D-SPlHT which allows 3D region of interest (ROI) coding. ROl coding enables faster reconstruction of diagnostically useful regions in volumetric datasets by assigning higher priority to them in the bitstream. It also introduces the possibility for increased compression performance, by allowing certain parts of the volume to be coded in a lossy manner while others are coded losslessly. The necessary modifications to 3D-SPIHT for ROI coding are described and methods for specifying a 3D ROI without adding a significant overhead are suggested. Results are presented highlighting the benefits of the ROl extension. Dimitris Agrafiotis, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP (3) | 3 |
| 2003 | Optimized sign language video coding based on eye-tracking analysis
Dimitris Agrafiotis, Cedric Nishan Canagarajah, David Bull 0001, Matt Dye, Helen Twyford, Jim Kyle, James T. Chung-How |
VCIP | 3 |
| 2003 | Bayesian approach to attack characterization using robust watermarks
Henry D. Knowles, Dominique A. Winne, Cedric Nishan Canagarajah, David Bull 0001 |
VCIP | 4 |
| 2003 | Efficient watermarking system with increased reliability for video authentication
Dominique A. Winne, Henry D. Knowles, David Bull 0001, Cedric Nishan Canagarajah |
VCIP | 3 |
| 2003 | Image segmentation using a texture gradient based watershed transformabstractThe segmentation of images into meaningful and homogenous regions is a key method for image analysis within applications such as content based retrieval. The watershed transform is a well established tool for the segmentation of images. However, watershed segmentation is often not effective for textured image regions that are perceptually homogeneous. In order to segment such regions properly, the concept of the "texture gradient" is introduced. Texture information and its gradient are extracted using a novel nondecimated form of a complex wavelet transform. A novel marker location algorithm is subsequently used to locate significant homogeneous textured or non textured regions. A marker driven watershed transform is then used to segment the identified regions properly. The combined algorithm produces effective texture and intensity based segmentation for application to content based image retrieval. Paul R. Hill, Cedric Nishan Canagarajah, David Bull 0001 |
IEEE Trans. Image Process. | 3 |
| 2002 | Image Fusion Using Complex WaveletsabstractThe fusion of images is the process of combining two or more images into a single image retaining important features from each. Fusion is an important technique within many disparate fields such as remote sensing, robotics and medical applications. Wavelet based fusion techniques have been reasonably effective in combining perceptually important image features. Shift invariance of the wavelet transform is important in ensuring robust subband fusion. Therefore, the novel application of the shift invariant and directionally selective Dual Tree Complex Wavelet Transform (DT-CWT) to image fusion is now introduced. This novel technique provides improved qualitative and quantitative results compared to previous wavelet fusion methods. 1 Paul R. Hill, Cedric Nishan Canagarajah, David Bull 0001 |
BMVC | 3 |
| 2002 | Packet loss resilient videoconferencing using H.263+abstractReal-time video transmission over the Internet is becoming increasingly desirable for videoconferencing and other interactive applications. The reliable transport protocols used by the Internet were designed mainly for non-real time data, and cannot be used for delay critical applications, which must therefore be able to cope with packet loss. Motion compensated video coding is especially sensitive to loss because of temporal error propagation. In this paper, the effect of packet loss on H.263+ video transmitted using the real-time transport protocol (RIP) is assessed. Various algorithms are described to minimise or prevent temporal error propagation. These algorithms do not rely on retransmissions and do not introduce more than one frame delay, and are therefore suitable for real-time and multicast applications. It is shown that the robustness of H.263+ video can be greatly improved using these techniques. Such techniques can also be applied to other motion compensated video coding standards such as MPEG-4 and H.26L. David Bull 0001, James T. Chung-How |
ICASSP | 1 |
| 2002 | Texture gradient based watershed segmentationabstractThe watershed transform is a well established tool for the segmentation of images. However, watershed segmentation is often not effective for textured regions that are perceptually homogeneous. Such regions are usually inaccurately over-segmented with no reference to any texture changes. We now introduce a novel concept of “texture gradient” implemented using a non-decimated complex wavelet transform. A novel marker location algorithm is subsequently used to locate significant homogeneous textured or non textured regions. A marker driven watershed transform is then used to properly segment the identified regions. The combined algorithm produces effective texture and intensity based segmentation for the application to content based retrieval of images. Paul R. Hill, Cedric Nishan Canagarajah, David Bull 0001 |
ICASSP | 3 |
| 2002 | Improved illumination-invariant descriptors for robust colour object recognitionabstractColour object recognition is heavily influenced by variation in the scene illumination conditions. This paper proposes a set of illumination-invariant descriptors of image content. The descriptors are based on a moment-based approach to histogram comparison and, in the case of an object imaged under two different lighting conditions, permit straightforward recovery of the illumination change involved. The efficacy of the descriptors is compared experimentally with a variety of existing techniques, using an established methodology and an existing purpose-built dataset. The evidence suggests that the new descriptors outperform existing techniques in the area of colour object recognition. Robert J. O'Callaghan, David Bull 0001 |
ICASSP | 2 |
| 2002 | Spatial digital watermark for MPEG-2 video authentication and tamper detectionabstractThe widespread adoption of digital video techniques has generated a requirement for authenticity verification in applications such as criminal evidence, insurance claims and commercial databases. This work addresses problems that arise from a spatial digital watermarking technique developed to detect frame reordering and dropping scenarios. It discusses the differences between mutual frame types at different bit-rates. Many papers consider detection after MPEG-2 decoding as a naïve approach. However, this approach does offer significant advantages for a slight increase of computation load. This paper also establishes a link between the detector performance and the sequence content. The uniqueness of this work is the comparison of the test results using 18 different standard MPEG test-sequences. The functionality of the algorithm is demonstrated with a simulated attack. Dominique A. Winne, Henry D. Knowles, David Bull 0001, Cedric Nishan Canagarajah |
ICASSP | 3 |
| 2002 | A scale invariant distance measure for texture retrievalabstractWe propose a similarity measure between two textures based on moments of the Fourier magnitude spectrum. The resulting distance is robust to changes in scale as well as to spatial shifts and grey-scale transforms of the texture samples. This type of invariant distance has applications to content-based image retrieval and classification tasks. We test the performance of the algorithm in a retrieval scenario using texture patches from the Brodatz album. The results indicate that the distance measure emulates human similarity perception in comparing textures. Robert J. O'Callaghan, David Bull 0001 |
ICIP (1) | 2 |
| 2002 | Region-based H.263+ coding for real-time video communication
Tuan-Kiang Chiew, James T. Chung-How, David Bull 0001, Cedric Nishan Canagarajah |
VCIP | 3 |
| 2002 | Efficient methodology for hand-coding video algorithms for VLIW-type processors
Oliver P. Sohm, David Bull 0001, Cedric Nishan Canagarajah |
Signal Process. Image Commun. | 2 |
| 2001 | Virtual liver biopsy: image processing and 3D visualizationabstractThis paper presents results in the image processing and visualization aspects of a virtual liver biopsy system (a system for simulating the medical procedure of liver biopsy). The creation of 3D models from 2D images of the organs involved is described, and segmentation requirements of this process are discussed. Endoscopic images of the liver that simulate the needle's point of view are created by means of combined volume and surface rendering. For this purpose ray casting is used with the ray start and end points being constrained within a surface rendered environment. Visualization of the needle insertion process from an exterior point of view is presented. A real-time sectional imaging component is also used in which the displayed 2D-image section of the 3D volume tracks the tip of the needle. Dimitris Agrafiotis, Michael G. Jones, Stavri G. Nikolov, M. Halliwell, David Bull 0001, Cedric Nishan Canagarajah |
ICIP (2) | 5 |
| 2001 | Statistical wavelet subband modelling for texture classificationabstractSimple wavelet and wavelet packet transforms have often been used for texture characterisation through the analysis of spatial-frequency content. However, most previous methods make no use of any statistical analysis of the transforms' subbands. A novel method is now presented for modelling the multivariate distributions of subband coefficients by considering spatially related coefficients. The Bhattacharya and divergence metrics are then used to produce an improved texture classification method for the application to content based image retrieval. Paul R. Hill, David Bull 0001, Cedric Nishan Canagarajah |
ICIP (1) | 2 |
| 2001 | Rotationally invariant texture based featuresabstractContent-based retrieval is ultimately dependent on the features used for the annotation of data and its efficiency is dependent on the invariance and robust properties of these features. For texture-based features an important form of invariance is rotational invariance. Novel rotationally invariant texture-based features are introduced that are extracted from a polar Fourier transform (PFT). The PFT is similar to the discrete Fourier transform in two dimensions but uses transform parameters radius and angle rather than the Cartesian co-ordinates. The PFT is discretised appropriately across the angular and radial frequency space with the transform magnitudes forming the rotationally invariant features. These features although rotationally invariant, capture the angular distribution together with the radial distribution of frequency within texture. Preliminary results show the method to give better results than rotationally variant and invariant Gabor filter schemes. Paul R. Hill, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP (2) | 3 |
| 2001 | Genetic stereo matching using complex conjugate wavelet pyramidsabstractA new genetic algorithm-based optimisation technique for stereo matching using complex conjugate wavelet pyramids is proposed. Reliable disparity fields are estimated in the wavelet domain with low computational cost. The new cost function is composed of the differences in wavelet coefficient values, plus vertical discontinuity and ordering constraints. Within homogenous regions, smoothness constraints on the disparity field are also employed. A genetic algorithm is used, where previously estimated vectors at the former image hierarchy are used to predict the corresponding search space of chromosomes, and to correct each newly calculated set of disparity vectors. This significantly reduces computational complexity compared to other methods, whilst maintaining robust performance. L. J. Luo, D. R. Clewer, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP (2) | 4 |
| 2001 | Video object tracking using region split and merge and a Kalman filter tracking algorithmabstractThis paper proposes a reliable method for tracking the trajectory of video objects using the vector Kalman predictor. Video objects, within the scope of this paper, are defined as groups of image pixels coherent spatially as well as in their values of luminance. The extent to which the quality of the unsupervised region split and merge segmentation affects the accuracy of the tracker is discussed, alongside the improvements made in the segmentation process as a result of the feedback from the tracking algorithm. The overall low complexity of the system, and the time savings made in using a spiral search algorithm, provide this method with prospects of being implemented in real-time. S. A. Vigus, David Bull 0001, Cedric Nishan Canagarajah |
ICIP (1) | 2 |
| 2001 | Viterbi decoding strategies for 5 GHz wireless LAN systemsabstractStandards for the operation of wireless local area network (WLAN) technology in the 5 GHz band have been developed in Europe, North America and Japan. These systems employ orthogonal frequency division multiplexing (OFDM) technology, and utilize forward error correction (FEC) coding, based on a 1/2-rate convolutional encoder, to combat frequency selective fading caused by multipath channels. This paper examines methods, based on the popular Viterbi algorithm, for maximum-likelihood (ML) decoding of this convolutional code within the WLAN baseband receiver. Although the paper focuses on the European HIPERLAN/2 standard, because of international physical (PHY) layer harmonization, the work is equally applicable to the North American and Japanese WLAN systems. Received error rate results, based on a software simulation of the HIPERLAN/2 PHY layer, are presented for both hard and soft decision decoding strategies. In addition, the effect of quantization on soft decision decoding is investigated. The results highlight the importance of the decoder design on the performance of the WLAN. Michael R. G. Butler, Simon Armour, Paul N. Fletcher, Andrew R. Nix, David Bull 0001 |
VTC Fall | 5 |
| 2001 | Loss resilient H.263+video over the Internet
James T. Chung-How, David Bull 0001 |
Signal Process. Image Commun. | 2 |
| 2001 | Simplex minimization for single- and multiple-reference motion estimationabstractBlock-matching motion estimation (BMME) can be formulated as a 2-D constrained minimization problem. This problem can, therefore, be solved with reduced complexity using optimization techniques. This paper proposes a novel fast BMME algorithm called the simplex minimization search (SMS). The algorithm is based on the simplex minimization (SM) optimization method. The initialization procedure, termination criterion, and constraints on the independent variables of the search are designed to take advantage of the characteristics of the BMME problem and the properties of the block motion fields of typical video sequences. Simulation results show that the proposed algorithm outperforms other fast BMME algorithms, providing better prediction quality, a smoother motion field, and higher speed-up ratio. This paper also investigates the properties of the multiple-reference (MR) block motion field. Guided by the results of this investigation, the paper extends the SMS algorithm to the MR case. Three MR SMS algorithms are proposed, providing different degrees of compromise between prediction quality and computational complexity. Simulation results using 50 reference frames indicate that the proposed MR algorithms have a computational complexity comparable to that of single-reference full-search while still maintaining the prediction gain of MR motion estimation. Mohammed E. Al-Mualla, Cedric Nishan Canagarajah, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2000 | Robust matching pursuits video transmission using the HIPERLAN/1 air interface standardabstractMatching pursuits over a basis of separable Gabor functions has been demonstrated to outperform DCT methods for low bit rate video coding. This paper introduces an error resilient implementation of the matching pursuits algorithm, based on the error resilient positional code. Coded video is transmitted using the simulated HIPERLAN/1 air interface standard, which recommends ARQ as a means of overcoming channel errors. This may be unsuitable for real time and broadcast applications. Therefore, a modified HIPERLAN/1 receiver is proposed in this paper, which does not use ARQ to retransmit erroneous packets but instead passes them to the video decoder to exploit error resilience. This strategy provides an acceptable reconstruction quality for average bit error rates equal to 1 in 1000 and is superior to a standard compliant system in the absence of ARQ. This confirms that wireless LAN standards should support a transparent mode for video applications. Przemyslaw Czerepinski, M. Fahim Tariq, David Bull 0001, Cedric Nishan Canagarajah, Andrew R. Nix |
ICASSP | 3 |
| 2000 | On the Performance of Temporal Error Concealment for Long-Term Motion-Compensated PredictionabstractThis paper investigates the performance of different temporal concealment techniques when incorporated within a long-term motion compensated video codec. In particular, the paper finds a combination of techniques that best recovers the spatial-temporal components of a damaged long-term motion vector. It is demonstrated also that accurate recovery of the spatial components of the damaged motion vector is, in general, more important and more complex than the recovery of the temporal component. Such findings are explained in view of the properties of the long-term block motion field. The paper also, proposes a novel long-term multihypothesis temporal concealment technique where the damaged block is concealed using the average of a number of candidates from the several reference frames available at the decoder. The superior performance of this technique is demonstrated over a range of block error rates. Mohammed E. Al-Mualla, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP | 3 |
| 2000 | Robust H.263+ Video for Real-Time Internet ApplicationsabstractAny real-time interactive video coding algorithm used over the Internet needs to be able to cope with packet loss, since the existing error recovery mechanisms are not suitable for real-time data. In this paper, a robust H.263+ video codec suitable for real-time interactive and multicast Internet applications is proposed. Initially, the robustness to packet loss of H.263 video packetised according to the RTP-H.263+ payload format specifications is assessed. Two techniques are proposed to minimise temporal propagation-selective FEC of the motion information and the use of periodic reference frames. It is shown that when these two techniques are combined, the robustness to loss of H.263+ video is greatly improved. James T. Chung-How, David Bull 0001 |
ICIP | 2 |
| 2000 | Wipe Production in MPEG-2 Compressed VideoabstractWith the increased role of technology in video production, several types of complex video special effects editing have begun to appear. We consider wiping special effects editing in MPEG-2 compressed video without full frame decompression and motion estimation. We estimated the DCT coefficients and use these coefficients together with the existing motion vectors to produce these special effects editing in the compressed domain. The results show that both the objective and subjective quality of the edited video in the compressed domain closely follows the quality of the video edited in the uncompressed domain at the same bit rate. Warnakulasuriya Anil Chandana Fernando, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP | 3 |
| 2000 | Statistical Feature Extraction from Compressed Video SequencesabstractTo maximise the benefits from data compression, it would be advantageous to develop algorithms that do not require decompression to extract relevant information for post processing. In this paper, a novel technique for extracting variance is proposed for MPEG-2 compressed video using Parseval's theorem. Results show that the estimated variance closely matches with the actual variance. Furthermore, this technique is applied to identify scene changes in the compressed domain. Warnakulasuriya Anil Chandana Fernando, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP | 3 |
| 2000 | Rotationally Invariant Texture Features Using the Dual-Tree Complex Wavelet TransformabstractNew rotationally invariant texture feature extraction methods are introduced that utilise the dual-tree complex wavelet transform (DT-CWT). The complex wavelet transform is a new technique that uses a dual tree of wavelet filters to obtain the real and imaginary parts of complex wavelet coefficients. When applied in two dimensions the DT-CWT produces shift invariant orientated subbands. Both isotropic and anisotropic rotationally invariant features can be extracted from the energies of these subbands. Using simple minimum distance classifiers, the classification performance of the proposed feature extraction methods were tested with rotated sample textures. The anisotropic features gave the best classification results for the rotated texture tests, outperforming a similar method using a real wavelet decomposition. Paul R. Hill, David Bull 0001, Cedric Nishan Canagarajah |
ICIP | 2 |
| 2000 | A Hierarchical Genetic Disparity Estimation Algorithm for Multiview Image SynthesisabstractA hierarchical genetic algorithm for disparity estimation is presented. The goal, to estimate reliable disparity fields with low computational cost, is reached using a hierarchical genetic matching procedure. Firstly, each hierarchical image of the stereo pair is divided into sets of feature points and non-feature points. The image morphological gradient for feature points and the disparity Laplacian function for non-feature points are incorporated into the matching function to serve as an adaptive smoothness term. Meanwhile, the vertical-discontinuity constraint and the ordering constraint are also proposed to smooth out vertical disparity discontinuities and to obtain a more reliable disparity estimation. In the hierarchical genetic matching procedure, previously estimated vectors at the former image hierarchy are used to predict the corresponding searching space of chromosomes, and to correct each newly calculated set of disparity vectors. This significantly increases the accuracy of disparity estimation. L. J. Luo, D. R. Clewer, David Bull 0001, Cedric Nishan Canagarajah |
ICIP | 3 |
| 2000 | Fusion of 2-D Images Using Their Multiscale EdgesabstractA new framework for fusion of 2D images based on their multiscale edges is described in this paper. The new method uses the multiscale edge representation of images proposed by Mallat and Hwang (1992). The input images are fused using their multiscale edges only. Two different algorithms for fusing the point representations and the chain representations of the multiscale edges (wavelet transform modulus maxima) are given. The chain representation has been found to provide numerous new alternatives for image fusion, since edge graph fusion techniques can be employed to combine the images. The new framework encompasses different levels, i.e. pixel and feature levels, of image fusion in the wavelet domain. Stavri G. Nikolov, David Bull 0001, Cedric Nishan Canagarajah, M. Halliwell, P. N. T. Wells |
ICPR | 2 |
| 2000 | Simplex minimisation for multiple-reference motion estimationabstractThis paper investigates the properties of the multiple-reference block motion field. Guided by the results of this investigation, the paper proposes three fast multiple-reference block matching motion estimation algorithms. The proposed algorithms are extensions of the single-reference simplex minimisation search (SMS) algorithm. The algorithms provide different degrees of compromise between prediction quality and computational complexity. Simulation results using a multi-frame memory of 50 frames indicate that the proposed multiple-reference SMS algorithms have a computational complexity comparable to that of single-reference full-search while still maintaining the prediction gain of multiple-reference motion estimation. Mohammed E. Al-Mualla, Cedric Nishan Canagarajah, David Bull 0001 |
ISCAS | 3 |
| 2000 | Video special effects editing in MPEG-2 compressed videoabstractWith the increase of digital technology in video production, several types of complex video special effects editing have begun to appear in video clips. In this paper we consider fade-out and fade-in special effects editing in MPEG-2 compressed video without full frame decompression and motion estimation. We estimated the DCT coefficients and use these coefficients together with the existing motion vectors to produce these special effects editing in compressed domain. Results show that both objective and subjective quality of the edited video in compressed domain closely follows the quality of the edited video in uncompressed video at the same bit rate. Warnakulasuriya Anil Chandana Fernando, Cedric Nishan Canagarajah, David Bull 0001 |
ISCAS | 3 |
| 2000 | Fade-in and fade-out detection in video sequences using histogramsabstractThere is an increased need to extract key information automatically from video for the purposes of indexing, fast retrieval, and scene analysis. To support this vision, reliable scene change detection algorithms must be developed. Several algorithms have been proposed for both sudden and gradual scene change detection in uncompressed and compressed video. In this paper we present an algorithm for fade-in and fade-out scene change detection in both uncompressed and compressed video sequences using histograms. We use the properties of the fading operation and extract these features in the luminance histogram. Results show that the proposed algorithm can be used in both uncompressed and compressed video to detect fade regions with a high reliability and less computations. Warnakulasuriya A. C. Fernando, Cedric Nishan Canagarajah, David Bull 0001 |
ISCAS | 3 |
| 2000 | Use of linear transverse equalisers and channel state information in combined OFDM-equalisationabstractThe efficiency of a coded orthogonal frequency division multiplexing (COFDM) receiver can be improved by use of a pre-FFT equaliser (PFE). This technique is known as combined OFDM-equalisation. This paper considers the noise amplification effect of the PFE and its effect on the performance of a combined COFDM-equalisation modem and proposes a method to improve this performance. It first reviews and compares the conventional COFDM and combined COFDM-equalisation techniques. The PFE is then described and the conditions required to allow it to take the form of a decision feedback equaliser (DFE) are discussed. It is shown that these conditions often cannot be met. In these cases the PFE must take the form of a linear transversal equaliser (LTE). The performance of the LTE is inferior to that of the DFE due to significant noise amplification. Hence, if the performance of combined COFDM-equalisation is to match that of conventional COFDM, a method for mitigating the noise amplifying effect of an LTE type PFE is required. One such method is proposed. This technique exploits channel state information (CSI) available in a COFDM receiver in order to enhance the performance of a Viterbi convolutional decoder. This 'CSI modified' Viterbi algorithm is capable of mitigating the noise amplification occurring in the equaliser. To demonstrate this, the performance of conventional COFDM and combined COFDM-equalisation are compared by means of software simulation using both standard and CSI modified Viterbi decoding. The results for the standard Viterbi decoding illustrate the performance penalty due to noise amplification. The results for the CSI modified Viterbi algorithm demonstrate the effective mitigation of noise amplification and the comparable performance of conventional COFDM and combined COFDM-equalisation. Simon Armour, Andrew R. Nix, David Bull 0001 |
PIMRC | 3 |
| 2000 | Implementations of error-resilient transcoders for MPEG-2 video over HIPERLAN
Greg J. Cain, David W. Redmill, David Bull 0001 |
VCIP | 3 |
| 2000 | Matching pursuits video coding: Dictionaries and fast implementationabstractMatching pursuits over a basis of separable Gabor functions has been demonstrated to outperform DCT methods for displaced frame difference coding for video compression. Unfortunately, apart from very low bit-rate applications, the algorithm involves an extremely high computational load. This paper contains an original contribution to the issues of dictionary selection and fast implementation for matching pursuits video coding. First, it is shown that the PSNR performance of existing matching pursuits codecs can be improved and the implementation cost reduced by a better selection of dictionary functions. Secondly, dictionary factorization is put forward to further reduce implementation costs. A reduction of the computational load by a factor of 20 is achieved compared to implementations reported to date. For a majority of test conditions, this reduction is supplemented by an improvement in reconstruction quality. Finally, a pruned full-search algorithm is introduced, which offers significant quality gains compared to the better-known heuristic fast-search algorithm, while keeping the computational cost low. Przemyslaw Czerepinski, Colin Davies, Cedric Nishan Canagarajah, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 1999 | Wipe Scene Change Detection in Video SequencesabstractThis paper presents a novel algorithm for wipe scene change detection in video sequences. In the proposed scheme, each image in the sequence is mapped to a reduced image. Then we use statistical features and structural properties of the images to identify wipe transition region. Finally, Hough transform is used to analyse the wiping pattern and the direction of wiping. Results show that the algorithm is capable of detecting all wipe regions accurately even when the video sequence contains other special effects. Warnakulasuriya Anil Chandana Fernando, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP (3) | 3 |
| 1999 | Fade and Dissolve Detection in Uncompressed and Compressed Video SequencesabstractAutomatic identification of special effects is a prerequisite for video indexing and intelligent video encoding. In this paper we present an algorithm for fade and dissolve scene change detection in video sequences. We use statistical features of the images to identify these special effects in uncompressed video. DC-estimation is used to evaluate statistical features both in H.263 and MPEG-2 compressed video. Results show that these special effects can be identified accurately with the proposed scheme. Warnakulasuriya Anil Chandana Fernando, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP (3) | 3 |
| 1999 | Sudden scene change detection in MPEG-2 video sequencesabstractContent based indexing and retrieval of video is becoming increasingly important in many applications. Identifying scene changes and special effects in a video scene is an essential pre-requisite for automatic indexing. This paper presents a real time algorithm, which can detect abrupt scene changes in the compressed domain. It is based on the number of interpolated macroblocks (MBs) for a given B-frame as the main feature since it expresses a measure of how strong the previous and future I or P (I/P) frames are correlated. Experimental results show that this algorithm can detect most abrupt scene changes in MPEG-2 compressed video. Warnakulasuriya Anil Chandana Fernando, Cedric Nishan Canagarajah, David Bull 0001 |
MMSP | 3 |
| 1999 | Fast 2D-DCT implementations for VLIW processorsabstractThis paper analyzes various fast 2D-DCT algorithms regarding their suitability for VLIW processors. Operations for truncation or rounding which are usually neglected in proposals for fast algorithms have also been taken into consideration. Loeffler's algorithm with parallel multiplications was found to be most suitable due to its parallel structure. Oliver P. Sohm, David Bull 0001, Cedric Nishan Canagarajah |
MMSP | 2 |
| 1998 | Error Concealment using Motion Field Interpolation
Mohammed E. Al-Mualla, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP (3) | 3 |
| 1998 | Video Coding using a Fast Non-Separable Matching Pursuits AlgorithmabstractThis paper presents a fast matching pursuits algorithm. In addition to a fast implementation, the algorithm also allows the efficient use of non-separable dictionary basis functions. The use of non-separable components allows basis functions which provide a better match to diagonally orientated image features. In order to demonstrate the proposed method, a simple dictionary is developed and used. Even without any sophisticated entropy coding, the proposed system gives performance exceeding that of H.263. David W. Redmill, David Bull 0001, Przemyslaw Czerepinski |
ICIP (1) | 2 |
| 1997 | Artificial neural network analysis of noisy visual field data in glaucoma
D. B. Henson, Susan E. George, David Bull 0001 |
Artif. Intell. Medicine | 3 |
| 1996 | Non-linear perfect reconstruction filter banks for image codingabstractThis paper presents a new architecture for designing non-linear, critically decimated, perfect reconstruction filter banks. It is believed that such filters have a wide range of applications including that of still image and image sequence compression. Example filters are presented, for still image compression using a quincunx system. These filters are shown to out-perform other linear and non-linear methods, both in terms of subjective and objective image quality. David W. Redmill, David Bull 0001 |
ICIP (1) | 2 |
| 1996 | Error resilient arithmetic coding of still imagesabstractThis paper examines the use of arithmetic coding in conjunction with the error resilient entropy code (EREC). The constraints on the coding model are discussed and simulation results are presented and compared to those obtained using Huffman coding. These results show that without the EREC, arithmetic coding is less resilient than Huffman coding, while with the EREC both systems yield comparable results. David W. Redmill, David Bull 0001 |
ICIP (2) | 2 |
| 1995 | Algorithms for Flexible Equalisation in Wireless Communications
R. Perry, David Bull 0001, Andrew R. Nix |
ISCAS | 2 |
| 1994 | The Optimisation of Multiplier-Free Directed Graphs: an Approach using Genetic AlgorithmsabstractThis paper considers the problem of realising directed graphs using evolutionary optimisation methods. Graphs are constrained to have edge gains equal to powers of two and signal values at internal vertices are required to be weighted by elements of a given coefficient vector. The objective is to synthesise a graph with minimum complexity. The method is developed for the case of a single multiplicative coefficient using vertex cardinality as a measure of solution fitness and extended to the more general case of a multi-element coefficient vector with additional optimisation constraints. The potential of the approach is demonstrated using examples based on FIR digital filters.> David Bull 0001, Alexis Aladjidi |
ISCAS | 1 |
| 1994 | Gate Level Optimisation of Primitive Operator Digital Filters using a Carry Save DecompositionabstractThis paper introduces a method for optimising digital filter realisations at the gate level. The method is based on a derivative of the primitive operator approach of Bull and Horrocks which is extended using a carry-save decomposition of the primitive operator graph. This facilitates the generation of a set of Boolean expressions for the multiply-accumulate section of the filter which can be minimised using standard sum of products or Reed Muller techniques. The technique is fully described and results are presented for a representative range of FIR filters. Savings of up to 83% are obtained for sum-of-products minimisation when compared to a CSD coded hard-wired multiplier solution. Initial results suggest further improvements in excess of 20% for the Reed Muller case.> David Bull 0001, Graham Wacey |
ISCAS | 1 |
| 1993 | A compound primitive operator approach to the realisation of video sub-band filter banks
David Bull 0001, Graham Wacey, John J. Stone, Jon M. Soloff |
ICASSP (1) | 1 |
| 1993 | Realisation techniques for primitive operator infinite impulse response digital filters
David Bull 0001, David H. Horrocks |
ISCAS | 1 |
| 1993 | POFGEN: A design automation system for VLSI digital filters with invariant transfer function
Graham Wacey, David Bull 0001 |
ISCAS | 2 |
| 1993 | A Sequential Niche Technique for Multimodal Function OptimizationabstractA technique is described that allows unimodal function optimization methods to be extended to locate all optima of multimodal problems efficiently. We describe an algorithm based on a traditional genetic algorithm (GA). This technique involves iterating the GA but uses knowledge gained during one iteration to avoid re-searching, on subsequent iterations, regions of problem space where solutions have already been found. This gain is achieved by applying a fitness derating function to the raw fitness function, so that fitness values are depressed in the regions of the problem space where solutions have already been found. Consequently, the likelihood of discovering a new solution on each iteration is dramatically increased. The technique may be used with various styles of GAs or with other optimization methods, such as simulated annealing. The effectiveness of the algorithm is demonstrated on a number of multimodal test functions. The technique is at least as fast as fitness sharing methods. It provides an acceleration of between 1 and l0p on a problem with p optima, depending on the value of p and the convergence time complexity. David Beasley, David Bull 0001, Ralph R. Martin |
Evol. Comput. | 2 |
| 1989 | Response error estimates for FIR digital filters with floating point coefficientsabstractResponse error estimates are derived for direct-form FIR (finite impulse response) digital filters having low-pass, high-pass, and bandpass characteristics. Simple-to-apply formulas are obtained on the basis of two observations: first, that the discontinuous relationships between a floating can be replaced by a continuous approximation; and secondly, that in the classes of filter considered, the impulse response possesses a closed form. Formulas that show the effect of wordlength, filter order, and filter bandwidth are derived. Simulation studies that confirm these estimates are reported.> David H. Horrocks, David Bull 0001 |
ICASSP | 2 |