EDBT 2026 Demo / reviewers in the wild / expert
Fan Zhang 0017
dblp:21/3626-17
· DBLP profile ↗
79ranked-venue papers
13as first author
49since 2021 · last 2026
0000-0001-6623-9936ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 68 · 13 first-author · 38 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Systems, architecture and hardware · 7 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GFix: Perceptually Enhanced Gaussian Splatting Video Compressionabstract3D Gaussian Splatting (3DGS) enhances 3D scene reconstruction through explicit representation and fast rendering, demonstrating potential benefits for various low-level vision tasks, including video compression. However, existing 3DGS-based video codecs generally exhibit more noticeable visual artifacts and relatively low compression ratios. In this paper, we specifically target the perceptual enhancement of 3DGS-based video compression, based on the assumption that artifacts from 3DGS rendering and quantization resemble noisy latents sampled during diffusion training. Building on this premise, we propose a content-adaptive framework, GFix, comprising a streamlined, single-step diffusion model that serves as an off-the-shelf neural enhancer. Moreover, to increase compression efficiency, We propose a modulated LoRA scheme that freezes the low-rank decompositions and modulates the intermediate hidden states, thereby achieving efficient adaptation of the diffusion backbone with highly compressible updates. Experimental results show that GFix delivers strong perceptual quality enhancement, outperforming GSVC with up to 72.1% BD-rate savings in LPIPS and 21.4% in FID. Siyue Teng, Ge Gao 0005, Duolikun Danier, Yuxuan Jiang 0015, Fan Zhang 0017, Nantheera Anantrasirichai, Thomas Davis, Zoe Liu, David Bull 0001 |
ISCAS | 5 |
| 2026 | SAM3-LiteText: An Anatomical Study of the SAM3 Text Encoder for Efficient Vision-Language SegmentationabstractVision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we propose SAM3-LiteText, a lightweight text encoding framework that replaces the original SAM3 text encoder with a compact MobileCLIP student that is optimized by knowledge distillation. Extensive experiments on image and video segmentation benchmarks show that SAM3-LiteText reduces text encoder parameters by up to 88%, substantially reducing static memory footprint, while maintaining segmentation performance comparable to the original model. Code: https://github.com/SimonZeng7108/efficientsam3/tree/sam3_litetext. Chengxi Zeng, Yuxuan Jiang 0015, Ge Gao 0005, Shuai Wang 0054, Duolikun Danier, Bin Zhu 0006, Stevan Rudinac, David Bull 0001, Fan Zhang 0017 |
ICMR | 9 |
| 2026 | A Mamba-Based Perceptual Loss Function for Learning-Based UGC Transcoding
Zihao Qi, Chen Feng 0008, Fan Zhang 0017, Xiaozhong Xu, Shan Liu 0001, David Bull 0001 |
QoMEX | 3 |
| 2026 | Inter predictive coding for point cloud attributes with online coordinate alignment and multi-scale latent prediction
Yu Liu 0091, Shuyuan Zhu, Zeliang Li, Jeff Siu-Kei Au-Yeung, Fan Zhang 0017, Bing Zeng 0001 |
Neurocomputing | 5 |
| 2026 | Human-Inspired Perspectives: A Survey on AI Long-Term MemoryabstractWith the rapid advancement of AI systems, their abilities to store, retrieve, and utilize information over the long term - referred to as long-term memory - have become increasingly significant. These capabilities are crucial for enhancing the performance of AI systems across a wide range of tasks. However, there is currently no comprehensive survey that systematically investigates AI's long-term memory capabilities, formulates a theoretical framework, and inspires the development of next-generation AI long-term memory systems. This paper begins by introducing the mechanisms of human long-term memory, then explores AI long-term memory mechanisms, establishing a mapping between the two. Based on the mapping relationships identified, we extend the current cognitive architectures and propose the Cognitive Architecture of Self-Adaptive Long-term Memory (SALM). SALM provides a theoretical framework for the practice of AI long-term memory and holds potential for guiding the creation of next-generation long-term memory driven AI systems. Finally, we delve into the future directions and application prospects of AI long-term memory. Zihong He, Weizhe Lin, Fan Zhang 0017, Matt W. Jones, Laurence Aitchison, Xuhai Xu, Miao Liu 0007, Hai-Ning Liang, Per Ola Kristensson, Junxiao Shen |
Proc. IEEE | 4 |
| 2026 | ViVo: A Dataset for Human Volumetric Video Reconstruction and Compression
Adrian Azzarelli, Ge Gao 0005, Ho Man Kwan, Fan Zhang 0017, Nantheera Anantrasirichai, Oliver Moolan-Feroze, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | FCVSR: A Frequency-Aware Method for Compressed Video Super-ResolutionabstractCompressed video super-resolution (SR) aims to generate high-resolution (HR) videos from the corresponding low-resolution (LR) compressed videos. Recently, some compressed video SR methods attempt to exploit the spatio-temporal information in the frequency domain, showing great promise in super-resolution performance. However, these methods do not differentiate various frequency subbands spatially or capture the temporal frequency dynamics, potentially leading to suboptimal results. In this paper, we propose a deep frequency-based compressed video SR model (FCVSR) consisting of a motion-guided adaptive alignment (MGAA) network and a multi-frequency feature refinement (MFFR) module. Additionally, a frequency-aware contrastive loss is proposed for training FCVSR, in order to reconstruct finer spatial details. The proposed model has been evaluated on three public compressed video super-resolution datasets, with results demonstrating its effectiveness when compared to existing works in terms of super-resolution performance and complexity. Fan Zhang 0017, Feiyu Chen 0001, Shuyuan Zhu, David Bull 0001, Bing Zeng 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | PNVC: Towards Practical INR-based Video CompressionabstractNeural video compression has recently demonstrated significant potential to compete with conventional video codecs in terms of rate-quality performance. These learned video codecs are however associated with various issues related to decoding complexity (for autoencoder-based methods) and/or system delays (for implicit neural representation (INR) based models), which currently prevent them from being deployed in practical applications. In this paper, targeting a practical neural video codec, we propose a novel INR-based coding framework, PNVC, which innovatively combines autoencoder-based and overfitted solutions. Our approach benefits from several design innovations, including a new structural reparameterization-based architecture, hierarchical quality control, modulation-based entropy modeling, and scale-aware positional embedding. Supporting both low delay (LD) and random access (RA) configurations, PNVC outperforms existing INR-based codecs, achieving nearly 35%+ BD-rate savings against HEVC HM 18.0 (LD) - almost 10% more compared to one of the state-of-the-art INR-based codecs, HiNeRV and 5% more over VTM 20.0 (LD), while maintaining 20+ FPS decoding speeds for 1080p content. This represents an important step forward for INR-based video coding, moving it towards practical deployment. Ge Gao 0005, Ho Man Kwan, Fan Zhang 0017, David Bull 0001 |
AAAI | 3 |
| 2025 | HIIF: Hierarchical Encoding based Implicit Image Function for Continuous Super-resolutionabstractRecent advances in implicit neural representations (INRs) have shown significant promise in modeling visual signals for various low-vision tasks including image super-resolution (ISR). INR-based ISR methods typically learn continuous representations, providing flexibility for generating high-resolution images at any desired scale from their low-resolution counterparts. However, existing INR-based ISR methods utilize multi-layer perceptrons for parameterization in the network; this does not take account of the hierarchical structure existing in local sampling points and hence constrains the representation capability. In this paper, we propose a new Hierarchical encoding based Implicit Image Function for continuous image super-resolution, HIIF, which leverages a novel hierarchical positional encoding that enhances the local implicit representation, enabling it to capture fine details at multiple scales. Our approach also embeds a multi-head linear attention mechanism within the implicit attention network by taking additional non-local information into account. Our experiments show that, when integrated with different backbone encoders, HIIF outperforms the state-of-the-art continuous image super-resolution methods by up to 0.17dB in PSNR. The source code of HIIF will be made publicly available at https://github.com/YuxuanJJ/HIIF. Yuxuan Jiang 0015, Ho Man Kwan, Tianhao Peng 0004, Ge Gao 0005, Fan Zhang 0017, Joel Sole, David Bull 0001 |
CVPR | 5 |
| 2025 | GIViC: Generative Implicit Video Compression
Ge Gao 0005, Siyue Teng, Tianhao Peng 0004, Fan Zhang 0017, David Bull 0001 |
ICCV | 4 |
| 2025 | Blind Video Super-Resolution Based on Implicit KernelsabstractBlind video super-resolution (BVSR) is a low-level vision task which aims to generate high-resolution videos from low-resolution counterparts in unknown degradation scenarios. Existing approaches typically predict blur kernels that are spatially invariant in each video frame or even the entire video. These methods do not consider potential spatio-temporal varying degradations in videos, resulting in suboptimal BVSR performance. In this context, we propose a novel BVSR model based on Implicit Kernels, BVSR-IK, which constructs a multi-scale kernel dictionary parameterized by implicit neural representations. It also employs a newly designed recurrent Transformer to predict the coefficient weights for accurate filtering in both frame correction and feature alignment. Experimental results have demonstrated the effectiveness of the proposed BVSR-IK, when compared with four state-of-the-art BVSR models on three commonly used datasets, with BVSR-IK outperforming the second best approach, FMA-Net, by up to 0.59 dB in PSNR. Source code will be available at https://github.com/QZ1-boy/BVSR-IK. Yuxuan Jiang 0015, Shuyuan Zhu, Fan Zhang 0017, David Bull 0001, Bing Zeng 0001 |
ICCV | 4 |
| 2025 | Enhancing HDR Video Compression based on Deep Effective Bit Depth AdaptationabstractIt is well known that high dynamic range (HDR) videos enhance immersive visual experiences compared to conventional standard dynamic range content. However, HDR content is typically more challenging to encode due to the increased detail associated with the wider dynamic range. In this work, we improve HDR compression performance using an Effective Bit Depth Adaptation approach (EBDA), which reduces the effective bit depth of the original video content before encoding and reconstructs the full bit depth using a CNN-based up-sampling method at the decoder. The up-sampling deep network is based on a new version of Multi-frame MFRNet, MF-MFRNet. This approach has been integrated into the EBDA framework with two Versatile Video Coding (VVC) reference models: VTM 16.2 and the Fraunhofer Versatile Video Encoder (VVenC 1.4.0). The proposed approach has been evaluated under the JVET HDR Common Test Conditions using the Random Access configuration. The results show evident coding gains over both the original VTM 16.2 and VVenC 1.4.0 on all JVET HDR tested sequences, with average bitrate savings of 3.1% and 4.8% based on PSNR and 7.8% and 9.6% based on VMAF against VTM and VVenC respectively. The source code of multi-frame MFRNet has been released at https://github.com/fan-aaron-zhang/MF-MFRNet. Chen Feng 0008, Zihao Qi, Duolikun Danier, Fan Zhang 0017, Xiaozhong Xu, Shan Liu 0001, David Bull 0001 |
ISCAS | 4 |
| 2025 | BVI-CR: A Multi-View Human Dataset for Volumetric Video CompressionabstractThe advances in immersive technologies and 3D reconstruction have enabled the creation of digital replicas of real-world objects and environments with fine details. These processes generate vast amounts of 3D data, requiring more efficient compression methods to satisfy the memory and bandwidth constraints associated with data storage and transmission. However, the development and validation of efficient 3D data compression methods are constrained by the lack of comprehensive and high-quality volumetric video datasets, which typically require much more effort to acquire and consume increased resources compared to 2D image and video databases. To bridge this gap, we present an open multi-view volumetric human dataset, denoted BVI-CR, which contains 18 multi-view RGB-D captures and their corresponding textured polygonal meshes, depicting a range of diverse human actions. Each video sequence contains 10 views in 1080p resolution with durations between 10-15 seconds at 30FPS. Using BVI-CR, we benchmarked three conventional and neural coordinate-based multi-view video compression methods, following the MPEG MIV Common Test Conditions, and reported their rate quality performance based on various quality metrics. The results show the great potential of neural representation based methods in volumetric video compression compared to conventional video coding methods (with an up to 38% average coding gain in PSNR). This dataset provides a development and validation platform for a variety of tasks including volumetric reconstruction, compression, and quality assessment. The database will be shared publicly at https://github.com/fan-aaron-zhang/bvi-cr. Ge Gao 0005, Adrian Azzarelli, Ho Man Kwan, Nantheera Anantrasirichai, Fan Zhang 0017, Will Andrew, Oliver Moolan-Feroze, David Bull 0001 |
ISCAS | 5 |
| 2025 | AquaNeRF: Neural Radiance Fields in Underwater Media with Distractor RemovalabstractNeural radiance field (NeRF) research has made significant progress in modeling static video content captured in the wild. However, current models and rendering processes rarely consider scenes captured underwater, which are useful for studying and filming ocean life. They fail to address visual artifacts unique to underwater scenes, such as moving fish and suspended particles. This paper introduces a novel NeRF renderer and optimization scheme for an implicit MLP-based NeRF model. Our renderer reduces the influence of floaters and moving objects that interfere with static objects of interest by estimating a single surface per ray. We use a Gaussian weight function with a small offset to ensure that the transmittance of the surrounding media remains constant. Additionally, we enhance our model with a depth-based scaling function to upscale gradients for near-camera volumes. Overall, our method outperforms the baseline Nerfacto by approximately 7.5% and SeaThru-NeRF by 6.2% in terms of PSNR. Subjective evaluation also shows a significant reduction of artifacts while preserving details of static targets and background compared to the state of the arts. Luca Gough, Adrian Azzarelli, Fan Zhang 0017, Nantheera Anantrasirichai |
ISCAS | 3 |
| 2025 | RTSR: A Real-Time Super-Resolution Model for AV1 Compressed ContentabstractSuper-resolution (SR) is a key technique for improving the visual quality of video content by increasing its spatial resolution while reconstructing fine details. SR has been employed in many applications including video streaming, where compressed low-resolution content is typically transmitted to end users and then reconstructed with a higher resolution and enhanced quality. To support real-time playback, it is important to implement fast SR models while preserving reconstruction quality; however, most existing solutions, in particular those based on complex deep neural networks, fail to do so. To address this issue, this paper proposes a low-complexity SR method, RTSR, designed to enhance the visual quality of compressed video content, focusing on resolution up-scaling from a) 360p to 1080p and from b) 540p to 4K. The proposed approach utilizes a Convolutional Neural Network (CNN)-based network architecture, which was optimized for AOMedia Video 1 (AV1SVT)-encoded content at various quantization levels based on a dual-teacher knowledge distillation method. This method was submitted to the AIM 2024 Video Super-Resolution Challenge, specifically targeting the Efficient/Mobile Real-Time Video SuperResolution competition. It achieved the best trade-off between complexity and coding performance (measured in PSNR, SSIM and VMAF) among all six submissions. The code will be available at https://github.com/YuxuanJJ/RTSR. Yuxuan Jiang 0015, Jakub Nawala, Chen Feng 0008, Fan Zhang 0017, Joel Sole, David Bull 0001 |
ISCAS | 4 |
| 2025 | Cross-Space Alignment-based Attribute Artifact Removal for V-PCCabstractIn this paper, we propose a cross-space alignment-based attribute artifact removal method for V-PCC. Firstly, we construct a cross-space frame alignment module to align the adjacent projected 2D frames based on the temporal continuity between the 3D point clouds. Then, we design a multi-frame fusion network that is developed based on U-Net to fuse the current frame with the aligned frames to remove the artifact that occurred in the attribute. Experimental results demonstrate that our proposed method achieves superior performance for the enhancement of point clouds. Yu Liu 0091, Mingfei Hu, Zeliang Li, Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Lihuo He, Fan Zhang 0017 |
ISCAS | 8 |
| 2025 | Ultra-lightweight Neural Video Representation Compression
Ho Man Kwan, Tianhao Peng 0004, Ge Gao 0005, Fan Zhang 0017, Mike Nilsson, Andrew Gower, David Bull 0001 |
PCS | 4 |
| 2025 | MVAD: A Multiple Visual Artifact Detector for Video StreamingabstractVisual artifacts are often introduced into streamed video content, due to prevailing conditions during content production and delivery. Since these can degrade the quality of the user's experience, it is important to automatically and accurately detect them in order to enable effective quality measurement and enhancement. Existing detection methods often focus on a single type of artifact and/or determine the presence of an artifact through thresholding objective quality indices. Such approaches have been reported to offer inconsistent prediction performance and are also impractical for real-world applications where multiple artifacts co-exist and interact. In this paper, we propose a Multiple Visual Artifact Detector, MVAD, for video streaming which, for the first time, is able to detect multiple artifacts using a single framework that is not reliant on video quality assessment models. Our approach employs a new Artifact-aware Dynamic Feature Extractor (ADFE) to obtain artifact-relevant spatial features within each frame for multiple artifact types. The extracted features are further processed by a Recurrent Memory Vision Transformer (RMViT) module, which captures both short-term and long-term temporal information within the input video. The proposed network architecture is optimized in an end-to-end manner based on a new, large and diverse training database that is generated by simulating the video streaming pipeline and based on Adversarial Data Augmentation. This model has been evaluated on two video artifact databases, Maxwell and BVI-Artifact, and achieves consistent and improved prediction results for ten target visual artifacts when compared to seven existing single and multiple artifact detectors. The source code and training database will be available at https://chenfeng-bristol.github.io/MVAD/. Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, Alex Mackin, Andrew Collins 0007, David Bull 0001 |
WACV | 3 |
| 2025 | UW-GS: Distractor-Aware 3D Gaussian Splatting for Enhanced Underwater Scene Reconstructionabstract3D Gaussian splatting (3DGS) offers the capability to achieve real-time high quality 3D scene rendering. However, 3DGS assumes that the scene is in a clear medium environment and struggles to generate satisfactory representations in underwater scenes, where light absorption and scattering are prevalent and moving objects are involved. To overcome these, we introduce a novel Gaussian Splatting-based method, UW-GS, designed specifically for underwater applications. It introduces a color appearance that models distance-dependent color variation, employs a new physics-based density control strategy to enhance clarity for distant objects, and uses a binary motion mask to handle dynamic content. Optimized with a well-designed loss function supporting for scattering media and strengthened by pseudo-depth maps, UW-GS outperforms existing methods with PSNR gains up to 1.26dB. To fully verify the effectiveness of the model, we also developed a new underwater dataset, S-UW, with dynamic object masks. The code of UW-GS and S-UW will be available at https://github.com/WangHaoran16/UW-GS. Nantheera Anantrasirichai, Fan Zhang 0017, David Bull 0001 |
WACV | 3 |
| 2025 | Distortion-Induced Saliency Shifts in VideoabstractVisual saliency modelling is of fundamental importance in modern video processing and its applications. Our previous eye-tracking study revealed that signal distortions caused by editing, compression, or transmission alter gaze patterns and consequently induce saliency shifts in both spatial and temporal domains. Saliency shifts provide crucial insights into viewers’ behavioural responses to video distortions, facilitating the perception-based optimisation of video algorithms. However, the spatio-temporal saliency shifts and their measurable effects on perception related applications remain largely unexplored. In this paper, we first investigate the measurement of distortion-induced saliency shifts (DSS) in videos and analyse DSS behaviours as functions of video content, time order and critical distortion disruption. Second, based on our findings, we construct three vision models to quantitatively simulate distinct DSS behaviours and integrate them into a comprehensive DSS behaviour model. Finally, we demonstrate that the computational DSS model can enhance emerging video technologies. Xinbo Wu, Jianxun Lou, Zhengyan Dong, Fan Zhang 0017, Paul L. Rosin, Hantao Liu |
IEEE Trans. Multim. | 4 |
| 2024 | LDMVFI: Video Frame Interpolation with Latent Diffusion ModelsabstractExisting works on video frame interpolation (VFI) mostly employ deep neural networks that are trained by minimizing the L1, L2, or deep feature space distance (e.g. VGG loss) between their outputs and ground-truth frames. However, recent works have shown that these metrics are poor indicators of perceptual VFI quality. Towards developing perceptually-oriented VFI methods, in this work we propose latent diffusion model-based VFI, LDMVFI. This approaches the VFI problem from a generative perspective by formulating it as a conditional generation problem. As the first effort to address VFI using latent diffusion models, we rigorously benchmark our method on common test sets used in the existing VFI literature. Our quantitative experiments and user study indicate that LDMVFI is able to interpolate video content with favorable perceptual quality compared to the state of the art, even in the high-resolution regime. Our code is available at https://github.com/danier97/LDMVFI. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
AAAI | 2 |
| 2024 | MTKD: Multi-Teacher Knowledge Distillation for Image Super-Resolution
Yuxuan Jiang 0015, Chen Feng 0008, Fan Zhang 0017, David Bull 0001 |
ECCV (39) | 3 |
| 2024 | NVRC: Neural Video Representation CompressionabstractRecent advances in implicit neural representation (INR)-based video coding have
demonstrated its potential to compete with both conventional and other learning-
based approaches. With INR methods, a neural network is trained to overfit a
video sequence, with its parameters compressed to obtain a compact representation
of the video content. However, although promising results have been achieved,
the best INR-based methods are still out-performed by the latest standard codecs,
such as VVC VTM, partially due to the simple model compression techniques
employed. In this paper, rather than focusing on representation architectures, which
is a common focus in many existing works, we propose a novel INR-based video
compression framework, Neural Video Representation Compression (NVRC),
targeting compression of the representation. Based on its novel quantization and
entropy coding approaches, NVRC is the first framework capable of optimizing an
INR-based video representation in a fully end-to-end manner for the rate-distortion
trade-off. To further minimize the additional bitrate overhead introduced by the
entropy models, NVRC also compresses all the network, quantization and entropy
model parameters hierarchically. Our experiments show that NVRC outperforms
many conventional and learning-based benchmark codecs, with a 23% average
coding gain over VVC VTM (Random Access) on the UVG dataset, measured
in PSNR. As far as we are aware, this is the first time an INR-based video codec
achieving such performance. Ho Man Kwan, Ge Gao 0005, Fan Zhang 0017, Andrew Gower, David Bull 0001 |
NeurIPS | 3 |
| 2024 | Accelerating Learnt Video Codecs with Gradient Decay and Layer-Wise DistillationabstractIn recent years, end-to-end learnt video codecs have demonstrated their potential to compete with conventional coding algorithms in term of compression efficiency. However, most learning-based video compression models are associated with high computational complexity and latency, in particular at the decoder side, which limits their deployment in practical applications. In this paper, we present a novel model-agnostic pruning scheme based on gradient decay and adaptive layer-wise distillation. Gradient decay enhances parameter exploration during sparsification whilst preventing runaway sparsity and is superior to the standard Straight-Through Estimation. The adaptive layer-wise distillation regulates the sparse training in various stages based on the distortion of intermediate features. This stage-wise design efficiently updates parameters with minimal computational overhead. The proposed approach has been applied to three popular end-to-end learnt video codecs, FVC, DCVC, and DCVC-HEM. Results confirm that our method yields up to 65% reduction in MACs and 2× speedup with less than 0.3dB drop in BD-PSNR. Supporting code and supplementary material can be downloaded from: https://jasminepp.github.io/lightweighltdvc/. Tianhao Peng 0004, Ge Gao 0005, Heming Sun, Fan Zhang 0017, David Bull 0001 |
PCS | 4 |
| 2024 | RankDVQA-Mini: Knowledge Distillation-Driven Deep Video Quality AssessmentabstractDeep learning-based video quality assessment (deep VQA) has demonstrated significant potential in surpassing conventional metrics, with promising improvements in terms of correlation with human perception. However, the practical deployment of such deep VQA models is often limited due to their high computational complexity and large memory requirements. To address this issue, we aim to significantly reduce the model size and runtime of one of the state-of-the-art deep VQA methods, RankDVQA, by employing a two-phase workflow that integrates pruning-driven model compression with multilevel knowledge distillation. The resulting lightweight full reference quality metric, RankDVQA-mini, requires less than 10% of the model parameters compared to its full version (14% in terms of FLOPs), while still retaining a quality prediction performance that is superior to most existing deep VQA methods. The source code of the RankDVQA-mini has been released at https://chenfeng-bristolgithub.io/RankDVQA-mini/ for public evaluation. Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, Benoit Vallade, Alex Mackin, David Bull 0001 |
PCS | 4 |
| 2024 | BVI-Artefact: An Artefact Detection Benchmark Dataset for Streamed VideosabstractProfessionally generated content (PGC) streamed online can contain visual artefacts that degrade the quality of user experience. These artefacts arise from different stages of the streaming pipeline, including acquisition, post-production, compression, and transmission. To better guide streaming experience enhancement, it is important to detect specific artefacts at the user end in the absence of a pristine reference. In this work, we address the lack of a comprehensive benchmark for artefact detection within streamed PGC, via the creation and validation of a large database, BVI-Artefact. Considering the ten most relevant artefact types encountered in video streaming, we collected and generated 480 video sequences, each containing various artefacts with associated binary artefact labels. Based on this new database, existing artefact detection methods are benchmarked, with results showing the challenging nature of this tasks and indicating the requirement of more reliable artefact detection methods. To facilitate further research in this area, we have made BVI-Artifact publicly available at bttps://chenfeng-bristol.github.io/BVI=Artefact/ Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, Alex Mackin, Andy Collins, David Bull 0001 |
PCS | 3 |
| 2024 | Compressing Deep Image Super-Resolution ModelsabstractDeep learning techniques have been applied in the context of image super-resolution (SR), achieving remarkable advances in terms of reconstruction performance. Existing techniques typically employ highly complex model structures which result in large model sizes and slow inference speeds. This often leads to high energy consumption and restricts their adoption for practical applications. To address this issue, this work employs a three-stage workflow for compressing deep SR models which significantly reduces their memory requirement. Restoration performance has been maintained through teacher-student knowledge distillation using a newly designed distillation loss. We have applied this approach to two popular image super-resolution networks, SwinIR and EDSR, to demonstrate its effectiveness. The resulting compact models, SwinIRmini and EDSRmini, attain an 89% and 96% reduction in both model size and floating-point operations (FLOPs) respectively, compared to their original versions. They also retain competitive super-resolution performance compared to their original models and other commonly used SR approaches. The source code and pretrained models for these two lightweight SR approaches are released at https://pikapi22.github.io/CDISM/. Yuxuan Jiang 0015, Jakub Nawala, Fan Zhang 0017, David Bull 0001 |
PCS | 3 |
| 2024 | Immersive Video Compression Using Implicit Neural RepresentationsabstractRecent work on implicit neural representations (INRs) has evidenced their potential for efficiently representing and encoding conventional video content. In this paper we, for the first time, extend their application to immersive (multi-view) videos, by proposing MV-HiNeRV, a new INR-based immersive video codec. MV-HiNeRV is an enhanced version of a state-of-the-art INR-based video codec, HiNeRV, which was developed for single-view video compression. We have modified the model to learn a different group of feature grids for each view, and share the learnt network parameters among all views. This enables the model to effectively exploit the spatio-temporal and the inter-view redundancy that exists within multi-view videos. The proposed codec was used to compress multi-view texture and depth video sequences in the MPEG Immersive Video (MIV) Common Test Conditions, and tested against the MIV Test model (TMIV) that uses the VVenC video codec. The results demonstrate the superior performance of MV-HiNeRV, with significant coding gains (up to 72.33%) over TMIV. The implementation of MV-HiNeRV is published for further development and evaluation11https://hmkx.github.io/mv-hinerv/. Ho Man Kwan, Fan Zhang 0017, Andrew Gower, David Bull 0001 |
PCS | 2 |
| 2024 | Full-Reference Video Quality Assessment for User Generated Content TranscodingabstractUnlike video coding for professional content, the delivery pipeline of User Generated Content (UGC) involves transcoding where unpristine reference content needs to be compressed repeatedly. In this work, we observe that existing full-/no-reference quality metrics fail to accurately predict the perceptual quality difference between transcoded UGC content and the corresponding unpristine references. Therefore, they are unsuited for guiding the rate-distortion optimisation process in the transcoding process. In this context, we propose a bespoke full-reference deep video quality metric for UGC transcoding. The proposed method features a transcoding-specific weakly supervised training strategy employing a quality ranking-based Siamese structure. The proposed method is evaluated on the YouTube-UGC VP9 subset and the LIVE-Wild database, demonstrating state-of-the-art performance compared to existing VQA methods. The source code of the developed quality metric and the associated training data are available from https://zihaoq1:github/io/FRUGC/. Zihao Qi, Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, Xiaozhong Xu, Shan Liu 0001, David Bull 0001 |
PCS | 4 |
| 2024 | BVI-AOM: A New Training Dataset for Deep Video Compression OptimizationabstractDeep learning is now playing an important role in enhancing the performance of conventional hybrid video codecs. These learning-based methods typically require diverse and representative training material for optimization in order to achieve model generalization and optimal coding performance. However, existing datasets either offer limited content variability or come with restricted licensing terms constraining their use to research purposes only. To address these issues, we propose a new training dataset, named BVI-AOM, which contains 956 uncompressed sequences at various resolutions from 270p to 2160p, covering a wide range of content and texture types. The dataset comes with more flexible licensing terms and offers competitive performance when used as a training set for optimizing deep video coding tools. The experimental results demonstrate that when used as a training set to optimize two popular network architectures for two different coding tools, the proposed dataset leads to additional bitrate savings of up to 0.29 and 2.98 percentage points in terms of PSNR-Y and VMAF, respectively, compared to an existing training dataset, BVI-DVC, which has been widely used for deep video coding. The BVI-AOM dataset is available at https://github.com/fan-aaron-zhang/bvi-aom. Jakub Nawala, Yuxuan Jiang 0015, Fan Zhang 0017, Joel Sole, David Bull 0001 |
VCIP | 3 |
| 2024 | Benchmarking Conventional and Learned Video Codecs with a Low-Delay ConfigurationabstractRecent advances in video compression have seen significant coding performance improvements with the development of new standards and learning-based video codecs. However, most of these works focus on application scenarios that allow a certain amount of system delay (e.g., Random Access mode in MPEG codecs), which is not always acceptable for live delivery. This paper conducts a comparative study of state-of-the-art conventional and learned video coding methods based on a low delay configuration. Specifically, this study includes two MPEG standard codecs (H.266/VVC VTM and JVET ECM), two AOM codecs (AV1 libaom and AVM), and two recent neural video coding models (DCVC-DC and DCVC-FM). To allow a fair and meaningful comparison, the evaluation was performed on test sequences defined in the AOM and MPEG common test conditions in the YCbCr 4:2:0 color space. The evaluation results show that the JVET ECM codecs offer the best overall coding performance among all codecs tested, with a 16.1% (based on PSNR) average BD-rate saving over AOM AVM, and 11.0% over DCVC-FM. We also observed inconsistent performance with the learned video codecs, DCVC-DC and DCVC-FM, for test content with large background motions. Siyue Teng, Yuxuan Jiang 0015, Ge Gao 0005, Fan Zhang 0017, Thomas Davis, Zoe Liu, David Bull 0001 |
VCIP | 4 |
| 2024 | RankDVQA: Deep VQA based on Ranking-inspired Hybrid TrainingabstractIn recent years, deep learning techniques have shown significant potential for improving video quality assessment (VQA), achieving higher correlation with subjective opinions compared to conventional approaches. However, the development of deep VQA methods has been constrained by the limited availability of large-scale training databases and ineffective training methodologies. As a result, it is difficult for deep VQA approaches to achieve consistently superior performance and model generalization. In this context, this paper proposes new VQA methods based on a two-stage training methodology which motivates us to develop a large-scale VQA training database without employing human subjects to provide ground truth labels. This method was used to train a new transformer-based network architecture, exploiting quality ranking of different distorted sequences rather than minimizing the difference from the ground-truth quality labels. The resulting deep VQA methods (for both full reference and no reference scenarios), FR- and NR-RankDVQA, exhibit consistently higher correlation with perceptual quality compared to the state-of-the-art conventional and deep VQA methods, with average SROCC values of 0.8972 (FR) and 0.7791 (NR) over eight test sets without performing cross-validation. The source code of the proposed quality metrics and the large training database are available at https://chenfeng-bristol.github.io/RankDVQA. Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
WACV | 3 |
| 2024 | CVEGAN: A perceptually-inspired GAN for Compressed Video EnhancementabstractWe propose a new Generative Adversarial Network for Compressed Video frame quality Enhancement (CVEGAN). The CVEGAN generator benefits from the use of a novel Mul2Res block (with multiple levels of residual learning branches), an enhanced residual non-local block (ERNB) and an enhanced convolutional block attention module (ECBAM). The ERNB has also been employed in the discriminator to improve the representational capability. The training strategy has also been re-designed specifically for video compression applications, to employ a relativistic sphere GAN (ReSphereGAN) training methodology together with new perceptual loss functions. The proposed network has been fully evaluated in the context of two typical video compression enhancement tools: post-processing (PP) and spatial resolution adaptation (SRA). CVEGAN has been fully integrated into the MPEG HEVC and VVC video coding test models (HM 16.20 and VTM 7.0) and experimental results demonstrate significant coding gains (up to 28% for PP and 38% for SRA compared to the anchor) over existing state-of-the-art architectures for both coding tools across multiple datasets based on the HM 16.20. The respective gains for VTM 7.0 are up to 8.0% for PP and up to 20.3% for SRA. Fan Zhang 0017, David Bull 0001 |
Signal Process. Image Commun. | 2 |
| 2023 | ST-MFNET Mini: Knowledge Distillation-Driven Frame InterpolationabstractCurrently, one of the major challenges in deep learning-based video frame interpolation (VFI) is the large model size and high computational complexity associated with many high performance VFI approaches. In this paper, we present a distillation-based two-stage workflow for obtaining compressed VFI models which perform competitively compared to the state of the art, but with significantly reduced model size and complexity. Specifically, an optimization-based network pruning method is applied to a state of the art frame interpolation model, ST-MFNet, which suffers from large model size. The resulting network architecture achieves a 91% reduction in parameter numbers and a 35% increase in speed. The performance of the new network is further enhanced through a teacher-student knowledge distillation training process using a Laplacian distillation loss. The final low complexity model, ST-MFNet Mini, achieves a comparable performance to most existing high-complexity VFI methods, only outperformed by the original ST-MFNet. Our source code is available at https://github.com/crispianm/ST-MFNet-Mini Crispian Morris, Duolikun Danier, Fan Zhang 0017, Nantheera Anantrasirichai, David Bull 0001 |
ICIP | 3 |
| 2023 | HiNeRV: Video Compression with Hierarchical Encoding-based Neural RepresentationabstractLearning-based video compression is currently a popular research topic, offering the potential to compete with conventional standard video codecs. In this context, Implicit Neural Representations (INRs) have previously been used to represent and compress image and video content, demonstrating relatively high decoding speed compared to other methods. However, existing INR-based methods have failed to deliver rate quality performance comparable with the state of the art in video compression. This is mainly due to the simplicity of the employed network architectures, which limit their representation capability. In this paper, we propose HiNeRV, an INR that combines light weight layers with novel hierarchical positional encodings. We employs depth-wise convolutional, MLP and interpolation layers to build the deep and wide network architecture with high capacity. HiNeRV is also a unified representation encoding videos in both frames and patches at the same time, which offers higher performance and flexibility than existing methods. We further build a video codec based on HiNeRV and a refined pipeline for training, pruning and quantization that can better preserve HiNeRV's performance during lossy model compression. The proposed method has been evaluated on both UVG and MCL-JCV datasets for video compression, demonstrating significant improvement over all existing INRs baselines and competitive performance when compared to learning-based codecs (72.3\% overall bit rate saving over HNeRV and 43.4\% over DCVC on the UVG dataset, measured in PSNR). Ho Man Kwan, Ge Gao 0005, Fan Zhang 0017, Andrew Gower, David Bull 0001 |
NeurIPS | 3 |
| 2023 | A multiple-UAV architecture for autonomous media production
Ioannis Mademlis, Arturo Torres-González, Jesús Capitán, Maurizio Montagnuolo, Alberto Messina, Fulvio Negro, Cédric Le Barz, Rita Cunha, Bruno J. Guerreiro, Fan Zhang 0017, Stephen Boyle, Gregoire Guerout, Anastasios Tefas, Nikos Nikolaidis 0001, David Bull 0001, Ioannis Pitas |
Multim. Tools Appl. | 11 |
| 2023 | BVI-VFI: A Video Quality Database for Video Frame InterpolationabstractVideo frame interpolation (VFI) is a fundamental research topic in video processing, which is currently attracting increased attention across the research community. While the development of more advanced VFI algorithms has been extensively researched, there remains little understanding of how humans perceive the quality of interpolated content and how well existing objective quality assessment methods perform when measuring the perceived quality. In order to narrow this research gap, we have developed a new video quality database named BVI-VFI, which contains 540 distorted sequences generated by applying five commonly used VFI algorithms to 36 diverse source videos with various spatial resolutions and frame rates. We collected more than 10,800 quality ratings for these videos through a large scale subjective study involving 189 human subjects. Based on the collected subjective scores, we further analysed the influence of VFI algorithms and frame rates on the perceptual quality of interpolated videos. Moreover, we benchmarked the performance of 33 classic and state-of-the-art objective image/video quality metrics on the new database, and demonstrated the urgent requirement for more accurate bespoke quality assessment methods for VFI. To facilitate further research in this area, we have made BVI-VFI publicly available at https://github.com/danier97/BVI-VFI-database. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame InterpolationabstractVideo frame interpolation (VFI) is currently a very active research topic, with applications spanning computer vision, post production and video encoding. VFI can be extremely challenging, particularly in sequences containing large motions, occlusions or dynamic textures, where existing approaches fail to offer perceptually robust inter-polation performance. In this context, we present a novel deep learning based VFI method, ST-MFNet, based on a Spatio-Temporal Multi-Flow architecture. ST-MFNet employs a new multi-scale multi-flow predictor to estimate many-to-one intermediate flows, which are combined with conventional one-to-one optical flows to capture both large and complex motions. In order to enhance interpolation performance for various textures, a 3D CNN is also employed to model the content dynamics over an extended temporal window. Moreover, ST-MFNet has been trained within an ST-GAN framework, which was originally developedfor texture synthesis, with the aim of further improving perceptual interpolation quality. Our approach has been comprehensively evaluated - compared with fourteen state-of-the-art VFI algorithms - clearly demonstrating that ST-MFNet consistently outperforms these benchmarks on var-ied and representative test datasets, with significant gains up to 1.09dB in PSNR for cases including large motions and dynamic textures. Our source code is available at https://github.com/danielism97/ST-MFNet. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
CVPR | 2 |
| 2022 | A Subjective Quality Study for Video Frame InterpolationabstractVideo frame interpolation (VFI) is one of the fundamental research areas in video processing and there has been extensive research on novel and enhanced interpolation algorithms. The same is not true for quality assessment of the interpolated content. In this paper, we describe a subjective quality study for VFI based on a newly developed video database, BVI-VFI. BVI-VFI contains 36 reference sequences at three different frame rates and 180 distorted videos generated using five conventional and learning based VFI algorithms. Subjective opinion scores have been collected from 60 human participants, and then employed to evaluate eight popular quality metrics, including PSNR, SSIM and LPIPS which are all commonly used for assessing VFI methods. The results indicate that none of these metrics provide acceptable correlation with the perceived quality on interpolated content, with the best-performing metric, LPIPS, offering a SROCC value below 0.6. Our findings show that there is an urgent need to develop a bespoke perceptual quality metric for VFI. The BVI-VFI dataset is publicly available and can be accessed at https://danielism97.github.io/BVI-VFI/. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
ICIP | 2 |
| 2022 | Enhancing Deformable Convolution Based Video Frame Interpolation with Coarse-To-Fine 3d CnnabstractThis paper presents a new deformable convolution-based video frame interpolation (VFI) method, using a coarse to fine 3D CNN to enhance the multi-flow prediction. This model first extracts spatio-temporal features at multiple scales using a 3D CNN, and estimates multi-flows using these features in a coarse-to-fine manner. The estimated multi-flows are then used to warp the original input frames as well as context maps, and the warped results are fused by a synthesis network to produce the final output. This VFI approach has been fully evaluated against 12 state-of-the-art VFI methods on three commonly used test databases. The results evidently show the effectiveness of the proposed method, which offers superior interpolation performance over other state of the art algorithms, with PSNR gains up to 0.19dB. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
ICIP | 2 |
| 2022 | A CNN-Based Post-Processor for Perceptually-Optimized Immersive Media CompressionabstractIn recent years, resolution adaptation based on deep neural networks has enabled significant performance gains for conventional (2D) video codecs. This paper investigates the effectiveness of spatial resolution resampling in the context of immersive content. The proposed approach reduces the spatial resolution of input multi-view videos before encoding, and reconstructs their original resolution after decoding. During the up-sampling process, an advanced CNN model is used to reduce potential re-sampling, compression, and synthesis artifacts. This work has been fully tested with the TMIV coding standard using a Versatile Video Coding (VVC) codec. The results demonstrate that the proposed method achieves significant rate-quality performance improvement for the majority of the test sequences, with an average BD-VMAF improvement of 3.07 over all sequences. Angeliki V. Katsenou, Fan Zhang 0017, David Bull 0001 |
ICIP | 2 |
| 2022 | Analysis of Video Quality Induced Spatio-Temporal Saliency ShiftsabstractHuman viewers’ eye movements reflect their perceptual responses to visual signals. Previous research has shown that distortions in videos cause spatio-temporal gaze shifts, which means gaze behaviour is related to video quality perception. It would be highly beneficial to understand gaze behaviour of viewing videos of varying perceived quality. However, little is known about the interactions between gaze, video content and distortions. In this paper, based on our eye-tracking database for video quality (SVQ160), we perform systematic analyses to reveal the impact of video content (VC) and time order (TO) on gaze shifts. Findings and quantitative methods for gaze behaviour can be used to develop advanced video quality metrics and video processing algorithms. Xinbo Wu, Zhengyan Dong, Fan Zhang 0017, Paul L. Rosin, Hantao Liu |
ICIP | 3 |
| 2022 | ViSTRA3: Video Coding with Deep Parameter Adaptation and Post ProcessingabstractThis paper presents a deep learning-based video compression framework (ViSTRA3), which has been employed to generate compression results for the ISCAS 2022 Grand Challenge on Neural Network-based Video Coding. The proposed framework intelligently adapts video format parameters of the input video before encoding, subsequently employing a CNN at the decoder to restore their original format and enhance reconstruction quality. ViSTRA3 has been integrated with the H.266/VVC Test Model VTM 14.0, and evaluated under the Joint Video Exploration Team Common Test Conditions. Bjønegaard Delta (BD) measurement results show that the proposed framework consistently outperforms the original VVC VTM, with average BD-rate savings of 1.8% and 3.7% based on the assessment of PSNR and VMAF. Chen Feng 0008, Duolikun Danier, Charlie Tan, Fan Zhang 0017, David Bull 0001 |
ISCAS | 4 |
| 2022 | FloLPIPS: A Bespoke Video Quality Metric for Frame InterpolationabstractVideo frame interpolation (VFI) serves as a useful tool for many video processing applications. Recently, it has also been applied in the video compression domain for enhancing both conventional video codecs and learning-based compression architectures. While there has been an increased focus on the development of enhanced frame interpolation algorithms in recent years, the perceptual quality assessment of interpolated content remains an open field of research. In this paper, we present a bespoke full reference video quality metric for VFI, FloLPIPS, that builds on the popular perceptual image quality metric, LPIPS, which captures the perceptual degradation in extracted image feature space. In order to enhance the performance of LPIPS for evaluating interpolated content, we re-designed its spatial feature aggregation step by using the temporal distortion (through comparing optical flows) to weight the feature difference maps. Evaluated on the BVI-VFI database, which contains 180 test sequences with various frame interpolation artefacts, FloLPIPS shows superior correlation performance (with statistical significance) with subjective ground truth over 12 popular quality assessors. To facilitate further research in VFI quality assessment, our code is publicly available at https://danielism97.github.io/F1oLPIPS. Duolikun Danier, Fan Zhang 0017, David Bull 0001 |
PCS | 2 |
| 2022 | BVI-DVC: A Training Database for Deep Video CompressionabstractDeep learning methods are increasingly being applied in the optimisation of video compression algorithms and can achieve significantly enhanced coding gains, compared to conventional approaches. Such approaches often employ Convolutional Neural Networks (CNNs) which are trained on databases with relatively limited content coverage. In this paper, a new extensive and representative video database, BVI-DVC,is presented for training CNN-based video compression systems, with specific emphasis on machine learning tools that enhance conventional coding architectures, including spatial resolution and bit depth up-sampling, post-processing and in-loop filtering. BVI-DVC contains 800 sequences at various spatial resolutions from 270p to 2160p and has been evaluated on ten existing network architectures for four different coding tools. Experimental results show that this database produces significant improvements in terms of coding gains over five existing (commonly used) image/video training databases under the same training and evaluation configurations. The overall additional coding improvements by using the proposed database for all tested coding modules and CNN architectures are up to 10.3% based on the assessment of PSNR and 8.1% based on VMAF. Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Enhancing VMAF through New Feature Integration and Model CombinationabstractVMAF is a machine learning based video quality assessment method, originally designed for streaming applications, which combines multiple quality metrics and video features through SVM regression. It offers higher correlation with subjective opinions compared to many conventional quality assessment methods. In this paper we propose enhancements to VMAF through the integration of new video features and alternative quality metrics (selected from a diverse pool) alongside multiple model combination. The proposed combination approach enables training on multiple databases with varying content and distortion characteristics. Our enhanced VMAF method has been evaluated on eight HD video databases, and consistently outperforms the original VMAF model (0.6.1) and other benchmark quality metrics, exhibiting higher correlation with subjective ground truth data. Fan Zhang 0017, Angeliki V. Katsenou, Christos G. Bampis, Lukas Krasula, Zhi Li 0001, David Bull 0001 |
PCS | 1 |
| 2021 | VMAF-based Bitrate Ladder Estimation for Adaptive StreamingabstractIn HTTP Adaptive Streaming, video content is conventionally encoded by adapting its spatial resolution and quantization level to best match the prevailing network state and display characteristics. It is well known that the traditional solution, of using a fixed bitrate ladder, does not result in the highest quality of experience for the user. Hence, in this paper, we introduce a content-driven approach for estimating the bitrate ladder, based on spatio-temporal features extracted from the uncompressed content. The method implements a content-driven interpolation. It uses the extracted features to train a machine learning model to infer the curvature points of the Rate-VMAF curves in order to guide a set of initial encodings. We employ the VMAF quality metric as a means of perceptually conditioning the estimation. When compared to the generation of a reference ladder using exhaustive encoding, 76.63% the estimated ladder's Rate-VMAF points are identical to those of the reference ladder. The proposed method benefits from a significant (77.4%) reduction in the number of encodes required with only a small (1.04%) average Bj⊘ntegaard Delta Rate increase. Angeliki V. Katsenou, Fan Zhang 0017, Kyle Swanson, Mariana Afonso, Joel Sole, David Bull 0001 |
PCS | 2 |
| 2021 | A Subjective Study on Videos at Various Bit DepthsabstractBit depth adaptation, where the bit depth of a video sequence is reduced before transmission and up-sampled during display, can potentially reduce data rates with limited impact on perceptual quality. In this context, we conducted a subjective study on a UHD video database, BVI-BD, to explore the relationship between bit depth and visual quality. In this paper, three bit depth adaptation methods are investigated, including linear scaling, error diffusion, and a novel adaptive Gaussian filtering approach for up-sampling. The results from a subjective experiment indicate that above a critical bit depth, bit depth adaptation has no significant impact on perceptual quality, while reducing the amount information that is required to be transmitted. Below the critical bit depth, the more ‘advanced’ adaptation methods can be used to retain ‘good’ visual quality down to around 2 bits per color channel for the experimental setup - far lower than the common 8 bits per color channel. A selection of image quality metrics were benchmarked on the subjective data, and analysis indicates that a bespoke quality metric may be required to enable accurate bit depth adaptation. Alex Mackin, Fan Zhang 0017, David Bull 0001 |
PCS | 3 |
| 2021 | ViSTRA2: Video coding using spatial resolution and effective bit depth adaptation
Fan Zhang 0017, Mariana Afonso, David Bull 0001 |
Signal Process. Image Commun. | 1 |
| 2020 | Gan-Based Effective Bit Depth Adaptation for Perceptual Video CompressionabstractResolution and effective bit depth (EBD) adaptation have been recently utilised in video compression to improve coding efficiency. This type of approach dynamically reduces spatial/temporal resolutions and effective bit depth at the encoder and restores the original video formats during decoding. In this paper, a convolutional neural networks (CNN) based EBD adaptation method is presented for perceptual video compression, in which the employed CNN models are trained using a generative adversarial network (GAN), with perception-based loss functions. This method was integrated into the HEVC HM 16.20 reference software and fully evaluated on test sequences from the JVET Common Test Conditions using the Random Access configuration. The results show significant coding gains achieved on all test sequences with an overall bit rate saving of 24.8% (Bjøntegaard Delta measurement) based on a perceptual quality metric, VMAF. Fan Zhang 0017, David Bull 0001 |
ICME | 2 |
| 2020 | Enhancing VVC Through Cnn-Based Post-ProcessingabstractThis paper presents a new Convolutional Neural Network (CNN) based post-processing approach for video compression, which is applied at the decoder to improve the reconstruction quality. This method has been integrated with the Versatile Video Coding Test Model (VTM) 4.0.1, and evaluated using the Random Access (RA) configuration using the Joint Video Exploration Team (JVET) Common Test Conditions (CTC). The results show coding gains on all tested sequences at various spatial resolutions over different quantisation parameter ranges, with average bit rate savings (based on Bjøntegaard Delta measurements) of 3.90% and 4.13%, when PSNR and VMAF are used as quality metrics respectively. The computational complexities of different CNN architecture variants have also been investigated. Fan Zhang 0017, Chen Feng 0008, David Bull 0001 |
ICME | 1 |
| 2019 | A Subjective Study of Viewing Experience for Drone Videos Using Simulated ConteabstractThis paper presents subjective evaluation results on the viewing experience of simulated aerial videos shot at different drone heights. A total of fifty video sequences were generated using a simulation engine, Unreal Engine 4, for two racing scenarios and five different shot types. Twenty human viewers were then employed to participate a subjective experiment, providing their preference opinions on viewing experience of these videos. Through the subjective test, optimal parameters of UAV height have been identified for the evaluated shot types and scenarios. These will provide recommendation of default shot parameters for drone operation in autonomous shooting and flight planning. Stephen Boyle, Fan Zhang 0017, David Bull 0001 |
ICIP | 2 |
| 2019 | A Subjective Comparison of AV1 and HEVC for Adaptive Video StreamingabstractIn this paper we compare the performance of two state-of-the-art competing codecs, AV1 and HEVC, in the context of adaptive streaming. We specifically consider a Dynamic Optimizer (DO) methodology that is content-aware and selects the resolution of the video sequence after constructing the convex hull of the Rate-Quality curves of all considered resolutions. We start with an objective evaluation of the Dynamic Optimizer, based on both PSNR and VMAF quality metrics. The Rate-VMAF curves show an average of 6.3% BD-Rate gain of AV1 over HEVC, while the Rate-PSNR curves an show an average BD-Rate loss of 1.8%. We then report subjective tests which evaluate the perceived quality of the selected bitstreams generated by the two codecs. In this case it was found that, for most rate points, the difference in the perceived quality between HEVC and AV1 is not significant. Angeliki V. Katsenou, Fan Zhang 0017, Mariana Afonso, David Bull 0001 |
ICIP | 2 |
| 2019 | A Frame Rate Conversion Method Based on a Virtual Shutter AngleabstractIn this paper a new method for frame rate conversion is presented. Utilising motion compensated frame prediction, our method mitigates the spatial distortions associated with traditional methods, while facilitating the conversion between a wider range of frame rates - useful when converting to/between legacy video formats. Using the concept of a virtual shutter angle, our method provides content providers with greater flexibility over the motion characteristic of their video sequences. We also propose a Gaussian weighting scheme which attempts to emulate a video sequence as if it was captured natively at the converted frame rate. A subjective experiment which compared our method to averaging frames at a range of frame rates, has shown that our method results in significantly higher visual quality across a range of content - especially for high down-sample factors. Alex Mackin, Fan Zhang 0017, David Bull 0001 |
ICIP | 2 |
| 2019 | Enhanced Video Compression Based on Effective Bit Depth AdaptationabstractThis paper presents a novel Convolutional Neural Network (CNN) based effective bit depth adaptation approach (EBDA-CNN) for video compression. It applies effective bit depth down-sampling before encoding and reconstructs the original bit depth using a deep CNN based up-sampling method at the decoder. The proposed approach has been integrated with the High Efficiency Video Coding reference software HM 16.20, and evaluated under the Joint Video Exploration Team Common Test Conditions using the Random Access configuration. The results show consistent coding gains on all tested sequences, with an average bitrate saving of 6.4%, based on Bjøntegaard Delta measurements using PSNR. Fan Zhang 0017, Mariana Afonso, David Bull 0001 |
ICIP | 1 |
| 2019 | Video Compression Based on Spatio-Temporal Resolution AdaptationabstractA video compression framework based on spatio-temporal resolution adaptation (ViSTRA) is proposed, which dynamically resamples the input video spatially and temporally during encoding, based on a quantisation-resolution decision, and reconstructs the full resolution video at the decoder. Temporal upsampling is performed using frame repetition, whereas a convolutional neural network super-resolution model is employed for spatial resolution upsampling. ViSTRA has been integrated into the high efficiency video coding reference software (HM 16.14). Experimental results verified via an international challenge show significant improvements, with BD-rate gains of 15% based on PSNR and an average MOS difference of 0.5 based on subjective visual quality tests. Mariana Afonso, Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Rate-Distortion Optimization Using Adaptive Lagrange MultipliersabstractIn current standardized hybrid video encoders, the Lagrange multiplier determination model is a key component in rate-distortion optimization. This originated some 20 years ago based on an entropy-constrained high-rate approximation and experimental results obtained using an H.263 reference encoder on limited test material. In this paper, we present a comprehensive analysis of the results of a Lagrange multiplier selection experiment conducted on various video content using H.264/AVC and HEVC reference encoders. These results show that the original Lagrange multiplier selection methods, employed in both video encoders, are able to achieve optimum rate-distortion performance for I and P frames, but fail to perform well for B frames. The relationship is identified between the optimum Lagrange multipliers for B frames and distortion information obtained from the experimental results, leading to a novel Lagrange multiplier determination approach. The proposed method adaptively predicts the optimum Lagrange multiplier for B frames based on the distortion statistics of recent reconstructed frames. After integration into both H.264/AVC and HEVC reference encoders, this approach was evaluated on 36 test sequences with various resolutions and differing content types. The results show consistent bitrate savings for various hierarchical B frame configurations with minimal additional complexity. BD savings average approximately 3% when constant quantization parameter (QP) values are used for all frames, and 0.5% when non-zero QP offset values are employed for different B frame hierarchical levels. Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | A Study of High Frame Rate Video FormatsabstractHigh frame rates are acknowledged to increase the perceived quality of certain video content. However, the lack of high frame rate test content has previously restricted the scope of research in this area-especially in the context of immersive video formats. This problem has been addressed through the publication of a high frame rate video database BVI-HFR, which was captured natively at 120 fps. BVI-HFR spans a variety of scenes, motions, and colors, and is shown to be representative of BBC broadcast content. In this paper, temporal down-sampling is utilized to enable both subjective and objective comparisons across a range frame rates. A large-scale subjective experiment has demonstrated that high frame rates lead to increases in perceived quality, and that a degree of content dependence exists-notably related to camera motion. Various image and video quality metrics have been benchmarked on these subjective evaluations, and analysis shows that those which explicitly account for temporal distortions (e.g., FRQM) provide improved correlation with subjective opinions compared to generic quality metrics such as PSNR. Alex Mackin, Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | A Study of Subjective Video Quality at Various Spatial ResolutionsabstractIn this paper we present the BVI-SR video database, which contains 24 unique video sequences at a range of spatial resolutions up to UHD-1 (3840p). These sequences were used as the basis for a large-scale subjective experiment exploring the relationship between visual quality and spatial resolution when using three distinct spatial adaptation filters (including a CNN-based super-resolution method). The results demonstrate that while spatial resolution has a significant impact on mean opinion scores (MOS), no significant reduction in visual quality between UHD-1 and HD resolutions for the super-resolution method is reported. A selection of image quality metrics were benchmarked on the subjective evaluations, and analysis indicates that VIF offers the best performance. Alex Mackin, Mariana Afonso, Fan Zhang 0017, David Bull 0001 |
ICIP | 3 |
| 2018 | SRQM: A Video Quality Metric for Spatial Resolution AdaptationabstractThis paper presents a full reference objective video quality metric (SRQM), which characterises the relationship between variations in spatial resolution and visual quality in the context of adaptive video formats. SRQM uses wavelet decomposition, subband combination with perceptually inspired weights, and spatial pooling, to estimate the relative quality between the frames of a high resolution reference video, and one that has been spatially adapted through a combination of down and upsampling. The BVI-SR video database is used to benchmark SRQM against five commonly-used quality metrics. The database contains 24 diverse video sequences that span a range of spatial resolutions up to UHD-1 (3840×2160). An indepth analysis demonstrates that SRQM is statistically superior to the other quality metrics for all tested adaptation filters, and all with relatively low computational complexity. Alex Mackin, Mariana Afonso, Fan Zhang 0017, David Bull 0001 |
PCS | 3 |
| 2018 | BVI-HD: A Video Quality Database for HEVC Compressed and Texture Synthesized ContentabstractThis paper introduces a new high-definition video quality database, referred to as BVI-HD, which contains 32 reference and 384 distorted video sequences plus subjective scores. The reference material in this database was carefully selected to optimize the coverage range and distribution uniformity of five low-level video features, while the included 12 distortions, using both original high efficiency video coding (HEVC) and HEVC with synthesis mode, represent state-of-the-art approaches to compression. The range of quantization parameters included in the database for HEVC compression was determined by a subjective study, the results of which indicate that a wider range of QP values should be used than the current recommendation. The subjective opinion scores for all 384 distorted videos were collected from a total of 86 subjects, using a double stimulus test methodology. Based on these results, we compare the subjective quality between HEVC and synthesised content, and evaluate the performance of nine state-of-the-art, full-reference objective quality metrics. This database has now been made available online, representing a valuable resource to those concerned with compression performance evaluation and objective video quality assessment. Fan Zhang 0017, Felix Mercer Moss, Roland Baddeley, David Bull 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | Low complexity video coding based on spatial resolution adaptationabstractIn this paper, a novel spatial resolution adaptation approach for video compression is proposed. Its ability to dynamically apply downsampling to frames exhibiting low spatial detail delivers improved rate distortion performance, together with a reduction in computational complexity of the encoding process. This method is based on an experimental investigation of the dependence between the QP threshold, which determines when to encode lower resolution frames, and the distortion obtained after downsampling/upsampling. The proposed approach is integrated with the High Efficiency Video Coding (HEVC) reference codec for intra coding, and evaluated on 15 high-resolution test sequences with varying levels of spatial detail. The results show a promising average bitrate savings of approximately 4% (B-D measurements), and significant complexity reduction (29% on average). Mariana Afonso, Fan Zhang 0017, Angeliki V. Katsenou, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 2 |
| 2017 | Investigating the impact of high frame rates on video compressionabstractIn this paper we investigate the impact of frame rate variation on HEVC video compression, and demonstrate that high frame rates (60+ fps) can lead to increased perceptual quality, notably in high bitrate environments. In order to quantify content dependence, a novel way of partitioning video sequences into categories is proposed. Results show that rate-quality performance is improved at higher frame rates for video sequences with camera motion, whereas lower frame rates are favorable in sequences with complex motion (e.g. dynamic textures). We calculate that 60 fps and 120 fps are optimal choices of frame rate at bitrates of 3 Mbps and 7 Mbps respectively, demonstrating that increased frame rates are both feasible and desirable, given current broadcast data rates. Alex Mackin, Fan Zhang 0017, Miltiadis Alexios Papadopoulos, David Bull 0001 |
ICIP | 2 |
| 2017 | A frame rate dependent video quality metric based on temporal wavelet decomposition and spatiotemporal poolingabstractThis paper presents an objective quality metric (FRQM), which characterises the relationship between variations in frame rate and perceptual video quality. The proposed method estimates the relative quality of a low frame rate video with respect to its higher frame rate counterpart, through temporal wavelet decomposition, subband combination and spatiotemporal pooling. FRQM was tested alongside six commonly used quality metrics (two of which explicitly relate frame rate variation to perceptual quality), on the publicly available BVI-HFR video database, that spans a diverse range of scenes and frame rates, up to 120fps. Results show that FRQM offers significant improvement over all other tested quality assessment methods with relatively low complexity. Fan Zhang 0017, Alex Mackin, David Bull 0001 |
ICIP | 1 |
| 2016 | What's on TV: A large scale quantitative characterisation of modern broadcast video contentabstractVideo databases, used for benchmarking and evaluating the performance of new video technologies, should represent the full breadth of consumer video content. The parameterisation of video databases using low-level features has proven to be an effective way of quantifying the diversity within a database. However, without a comprehensive understanding of the importance and relative frequency and of these features in the content people actually consume, the utility of such information is limited. Here, we present a large-scale analysis of programming on BBC One and CBeebies, the most popular television channels in the United Kingdom for adults and children, respectively. Twenty video features are extracted from almost three thousand television programmes shown throughout 2015 before principal components analysis is used to identify just five factors representing the most variation. The meaning and relative significance of these five factors together with the shape of their frequency distributions represent highly valuable information for researchers wanting to model the diversity of modern consumer content in representative video databases. Felix Mercer Moss, Fan Zhang 0017, Roland Baddeley, David Bull 0001 |
ICIP | 2 |
| 2016 | An adaptive QP offset determination method for HEVCabstractThis paper investigates the effect that the QP offset value has on the coding performance of HEVC. We relate QP offset to the type of texture content present in the sequence. These then are used to develop a low-complexity adaptive QP offset selection method. This enables in-loop configuration of the QP offset parameter in a way that is content dependent and utilizes available encoding statistics. The proposed adaptive method is found to offer average bitrate reductions ranging from 1.38% for dynamic texture sequences up to 1.59% for static texture sequences relative to the QP offset used in the JCT-VC common test conditions. Miltiadis Alexios Papadopoulos, Fan Zhang 0017, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 2 |
| 2016 | HEVC enhancement using content-based local QP selectionabstractInspired by recent advances in objective video quality assessment, this paper proposes a novel, local quantisation parameter (QP) determination approach for perceptual video compression, based on the experimental results of a QP selection test. This method has been fully integrated into the High Efficiency Video Coding (HEVC) reference codec for intra coding, which predicts coding tree unit (CTU) level QPs to achieve optimised rate quality performance. The proposed approach consistently shows bitrate savings based on perceptual quality metrics and Bjontegaard delta measurements, with minimal complexity increase over the original codec. Fan Zhang 0017, David Bull 0001 |
ICIP | 1 |
| 2016 | Video texture analysis based on HEVC encoding statisticsabstractIn this paper, an extensive study of different video texture properties based on encoding statistics extracted from the HEVC HM reference software is presented. Mode selection, partitioning, motion vectors and bitrate allocation are among the statistics obtained from the encoder. For this study, a new dataset of homogeneous static and dynamic video textures, HomTex, is proposed. A comprehensive investigation of the results reveals a significant variability of coding statistics within dynamic textures, suggesting that this category should be further split into two relevant subcategories, continuous dynamic textures and discrete dynamic textures. This case is supported by an unsupervised learning approach on the statistics extracted. Finally, following the results obtained, some suggestions of improvements in video texture coding are presented. Mariana Afonso, Angeliki V. Katsenou, Fan Zhang 0017, Dimitris Agrafiotis, David Bull 0001 |
PCS | 3 |
| 2016 | Support for reduced presentation durations in subjective video quality assessment
Felix Mercer Moss, Chun-Ting Yeh, Fan Zhang 0017, Roland Baddeley, David Bull 0001 |
Signal Process. Image Commun. | 3 |
| 2016 | On the Optimal Presentation Duration for Subjective Video Quality AssessmentabstractSubjective quality assessment is an essential component of modern image and video processing for both the validation of objective metrics and the comparison of coding methods. However, the standard procedures used to collect data can be prohibitively time consuming. One way of increasing the efficiency of data collection is to reduce the duration of test sequences from a 10-s length currently used in most subjective video quality assessment (VQA) experiments. Here, we explore the impact of reducing sequence length upon perceptual accuracy when identifying compression artifacts. A group of four reference sequences, together with five levels of distortion, are used to compare the subjective ratings of viewers watching videos between 1.5 and 10 s long. We identify a smooth function indicating that accuracy increases linearly as the length of the sequences increases from 1.5 to 7 s. The accuracy of observers viewing 1.5-s sequences was significantly inferior to those viewing sequences of 5, 7, and 10 s. We argue that sequences between 5 and 10 s produce satisfactory levels of accuracy but the practical benefits of acquiring more data lead us to recommend the use of 5-s sequences for future VQA studies that use the double stimulus continuous quality scale methodology. Felix Mercer Moss, Fan Zhang 0017, Roland Baddeley, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | A Perception-Based Hybrid Model for Video Quality AssessmentabstractIt is known that the human visual system (HVS) employs independent processes (distortion detection and artifact perception-also often referred to as near-threshold and suprathreshold distortion perception) to assess video quality for various distortion levels. Visual masking effects also play an important role in video distortion perception, especially within spatial and temporal textures. In this paper, a novel perception-based hybrid model for video quality assessment is presented. This simulates the HVS perception process by adaptively combining noticeable distortion and blurring artifacts using an enhanced nonlinear model. Noticeable distortion is defined by thresholding absolute differences using spatial and temporal tolerance maps that characterize texture masking effects, and this makes a significant contribution to quality assessment when the quality of the distorted video is similar to that of the original video. Characterization of blurring artifacts, estimated by computing high frequency energy variations and weighted with motion speed, is found to further improve metric performance. This is especially true for low quality cases. All stages of our model exploit the orientation selectivity and shift invariance properties of the dual-tree complex wavelet transform. This not only helps to improve the performance but also offers the potential for new low complexity in-loop application. Our approach is evaluated on both the Video Quality Experts Group (VQEG) full reference television Phase I and the Laboratory for Image and Video Engineering (LIVE) video databases. The resulting overall performance is superior to the existing metrics, exhibiting statistically better or equivalent performance with significantly lower complexity. Fan Zhang 0017, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | A study of subjective video quality at various frame ratesabstractThis paper presents a new video database (BVI-HFR), which contains content with a variety of frame rates from 15Hz to 120Hz, that can be used to demonstrate the benefits and limitations of higher frame rates, as well as investigating the role that frame rates play from capture to delivery. A characterization of the video database using low-level descriptors is also provided, which establishes that it successfully spans a variety of scene types and motions, and compares well to existing video databases. Subjective evaluations performed on the video database, have demonstrated a significant relationship between frame rates and perceived quality, up to 120Hz. They also confirm that the relationship between frame rate and perceived quality is content dependent. Alex Mackin, Fan Zhang 0017, David Bull 0001 |
ICIP | 2 |
| 2015 | A video texture database for perceptual compression and quality assessmentabstractThis paper presents a new publicly available video texture database (BVI Texture) that contains test sequences and subjective opinion scores. The database exhibits a wide range of static and dynamic textures together with some mixed content. Each sequence is indexed using various video feature descriptors that characterize its spatial activity, temporal activity, static texture content and dynamic texture content. Moreover, rate/distortion results for the new dataset are presented after compression using HEVC, alongside subjective quality evaluation data. The BVI texture database will provide utility in testing quality assessment metrics and emerging video compression methods, particularly those based on texture analysis and synthesis. Miltiadis Alexios Papadopoulos, Fan Zhang 0017, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 2 |
| 2015 | A very low complexity reduced reference video quality metric based on spatio-temporal information selectionabstractThis paper presents a reduced reference video quality metric that exploits contrast and motion sensitivity characteristics of the HVS to perform a spatio-temporal selection of reference data. Spatio-temporal selection is realised through mapping of wavelet subbands to contrast sensitivity and through motion analysis. The proposed method is integrated with a modified SSIM-based framework to produce STIS-SSIM, a very low complexity reduced reference metric. The metric is shown to offer significant performance improvement over many existing full reference and reduced reference video quality metrics when tested on the LIVE video database. Fan Zhang 0017, Dimitris Agrafiotis |
ICIP | 2 |
| 2015 | An adaptive Lagrange multiplier determination method for rate-distortion optimisation in hybrid video codecsabstractThis paper describes an adaptive Lagrange multiplier determination method for rate-quality optimisation in video compression. Inspired by the experimental results of a Lagrange multiplier selection test, the presented approach adaptively estimates the optimum Lagrange multiplier for different video content, based on distortion statistics of recently encoded frames. The proposed algorithm has been fully integrated into both the H.264 and HEVC reference codecs, and is used in rate-distortion optimisation for encoding B frames. The results show promising (up to 11% on the sequences tested) overall bitrate savings, for a minimal increase in complexity, on various types of test content based on Bjontegaard delta measurements. Fan Zhang 0017, David Bull 0001 |
ICIP | 1 |
| 2013 | Quality assessment methods for perceptual video compressionabstractThis paper describes a quality assessment model for perceptual video compression applications (PVM), which stimulates visual masking and distortion-artefact perception using an adaptive combination of noticeable distortions and blurring artefacts. The method shows significant improvement over existing quality metrics based on the VQEG database, and provides compatibility with in-loop rate-quality optimisation for next generation video codecs due to its latency and complexity attributes. Performance comparison are validated against a range of different distortion types. Fan Zhang 0017, David Bull 0001 |
ICIP | 1 |
| 2012 | Perception-oriented video coding based on image analysis and completion: A review
Patrick Ndjiki-Nya, Dimitar Doshkov, Hagen Kaprykowsky, Fan Zhang 0017, David Bull 0001, Thomas Wiegand 0001 |
Signal Process. Image Commun. | 4 |
| 2010 | Region-based texture modelling for next generation video codecsabstractThis paper describes a texture model application for future video compression algorithms. This employs region-based texture analysis techniques and advanced motion models to synthesise and warp video frames rather than encode the whole image or the prediction residual after traditional motion compensation. The proposed texture warping method and the dynamic texture model are integrated into an H.264 framework together with video quality assessment modules to prevent video artefacts. The results show impressive bitrate savings, up to 47%, over H.264 with similar visual quality. Fan Zhang 0017, David Bull 0001, Cedric Nishan Canagarajah |
ICIP | 1 |
| 2010 | Enhanced video compression with region-based texture modelsabstractThis paper presents a region-based video compression algorithm based on texture warping and synthesis. Instead of encoding whole images or prediction residuals after translational motion estimation, this algorithm employs a perspective motion model to warp static textures and uses a texture synthesis approach to synthesise dynamic textures. Spatial and temporal artefacts are prevented by an in-loop video quality assessment module. The proposed method has been integrated into an H.264 video coding framework. The results show significant bitrate savings, up to 55%, compared with H.264, for similar visual quality. Fan Zhang 0017, David Bull 0001 |
PCS | 1 |