EDBT 2026 Demo / reviewers in the wild / expert
Yang Zhang 0003
dblp:06/6785-3
· DBLP profile ↗
19ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-2381-6067ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image CompressionabstractGenerative neural image compression supports data representation at extremely low bitrate, synthesizing details at the client and consistently producing highly realistic images. By leveraging the similarities between quantization error and additive noise, diffusion-based generative image compression codecs can be built using a latent diffusion model to "denoise" the artifacts introduced by quantization. However, we identify three critical gaps in previous approaches following this paradigm (namely, the noise level, noise type, and discretization gaps) that result in the quantized data falling out of the data distribution known by the diffusion model. In this work, we propose a novel quantization-based forward diffusion process with theoretical foundations that tackles all three aforementioned gaps. We achieve this through universal quantization with a carefully tailored quantization schedule and a diffusion model trained with uniform noise. Compared to previous work, our proposal produces consistently realistic and detailed reconstructions, even at very low bitrates. In such a regime, we achieve the best rate-distortion-realism performance, outperforming previous related works. Lucas Relic, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Christopher Schroers |
CVPR | 3 |
| 2025 | LDIP: Long Distance Information Propagation for Video Super-ResolutionabstractVideo super-resolution (VSR) methods typically exploit information across multiple frames to achieve high quality upscaling, with recent approaches demonstrating impressive performance. Nevertheless, challenges remain, particularly in effectively leveraging information over long distances. To address this limitation in VSR, we propose a strategy for long distance information propagation with a flexible fusion module that can optionally also assimilate information from additional high resolution reference images. We design our overall approach such that it can leverage existing pre-trained VSR backbones and adapt the feature upscaling module to support arbitrary scaling factors. Our experiments demonstrate that we can achieve state-of-theart results on perceptual metrics and deliver more visually pleasing results compared to existing solutions. Michael Bernasconi, Abdelaziz Djelouah, Yang Zhang 0003, Markus Gross 0001, Christopher Schroers |
ICCV | 3 |
| 2025 | Spatiotemporal Diffusion Priors for Extreme Video CompressionabstractDiffusion models have recently demonstrated impressive results in image compression, where the strong spatial prior enables the synthesis of fine details rather than allocating bits to transmit them. In this work, we propose to extend this paradigm to video compression by utilizing a generative spatiotemporal prior and present the first codec based on a video diffusion model. Our method operates by performing longcontext interpolation guided by sparse inter-frame predictions, thus requiring minimal motion information. To this end, we develop a sparse, bidirectional optical flow which serves as a bitrate-efficient motion conditioning in the diffusion decoding process. The resulting codec can compress videos to extremely low rates (as low as 0.01 bits per pixel) while maintaining realistic textures and motion, and outperforms both neural and traditional baselines on several benchmark datasets. Our method shows state-of-the art performance in perceptually-oriented distortion metrics, and, when considering rate-realism, we achieve an improvement in FID score of up to 73.3 at the same bitrate compared to the leading traditional video codec, VTM. Overall, we present an important first work examining spatiotemporal diffusion priors for video compression. Lucas Relic, André Emmenegger, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Christopher Schroers |
PCS | 4 |
| 2025 | CLIP-Fusion: A Spatio-Temporal Quality Metric for Frame Interpolation
Göksel Mert Çökmez, Yang Zhang 0003, Christopher Schroers, Tunç Ozan Aydin |
WACV | 2 |
| 2024 | Revitalizing Legacy Video Content: Deinterlacing with Bidirectional Information Propagation
Zhaowei Gao, Christopher Schroers, Yang Zhang 0003 |
BMVC | 4 |
| 2023 | Neural Video Compression with Spatio-Temporal Cross-Covariance TransformersabstractAlthough existing neural video compression~(NVC) methods have achieved significant success, most of them focus on improving either temporal or spatial information separately. They generally use simple operations such as concatenation or subtraction to utilize this information, while such operations only partially exploit spatio-temporal redundancies. This work aims to effectively and jointly leverage robust temporal and spatial information by proposing a new 3D-based transformer module: Spatio-Temporal Cross-Covariance Transformer (ST-XCT). The ST-XCT module combines two individual extracted features into a joint spatio-temporal feature, followed by 3D convolutional operations and a novel spatio-temporal-aware cross-covariance attention mechanism. Unlike conventional transformers, the cross-covariance attention mechanism is applied across the feature channels without breaking down the spatio-temporal features into local tokens. Such design allows for modeling global cross-channel correlations of the spatio-temporal context while lowering the computational requirement. Based on ST-XCT, we introduce a novel transformer-based end-to-end optimized NVC framework. ST-XCT-based modules are integrated into various key coding components of NVC, such as feature extraction, frame reconstruction, and entropy modeling, demonstrating its generalizability. Extensive experiments show that our ST-XCT-based NVC proposal achieves state-of-the-art compression performances on various standard video benchmark datasets. Lucas Relic, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Dong Xu 0001, Luping Zhou, Christopher Schroers |
ACM Multimedia | 4 |
| 2023 | Large-Scale Multi-Site Subjective Assessment on Image Banding ArtifactsabstractBanding largely occurs due to the finite bit depth representation of digital media and displays. It is a unique type of artifact that is both challenging to accurately measure and difficult to properly mitigate. These artifacts can also be observed across a wide spectrum of streaming services, no matter if professional studio productions or user-generated content. Inspired by the need to monitor and measure these banding artifacts, we present a large-scale multi-site subjective study on image banding artifacts. We designed a novel two-question-based subjective evaluation protocol that goes beyond giving ratings at the frame level, but also collects subjective opinions at the sub-frame level. The high-quality subjective data are collected from a cohort of 56 participants across 3 physical sites with calibrated environments, covering diverse backgrounds and experiences. Finally, we demonstrate that none of the common banding metrics are highly correlated with the subjective data. The paper calls for a need of continued efforts in modeling the perceptual effect of banding artifacts. Yuanyi Xue, Roberto Azevedo, Xuchang Huangfu, Yang Zhang 0003, Christopher Schroers, Scott Labrozzi |
QoMEX | 4 |
| 2022 | TempFormer: Temporally Consistent Transformer for Video Denoising
Yang Zhang 0003, Tunç Ozan Aydin |
ECCV (19) | 2 |
| 2021 | Deep HDR estimation with generative detail reconstructionabstractAbstract We study the problem of High Dynamic Range (HDR) image reconstruction from a Standard Dynamic Range (SDR) input with potential clipping artifacts. Instead of building a direct model that maps from SDR to HDR images as in previous work, we decompose an input SDR image into a base (low frequency) and detail layer (high frequency), and treat reconstructing these two layers as two separate problems. We propose a novel architecture that comprises individual components specially designed to handle both tasks. Specifically, our base layer reconstruction component recovers low frequency content and remaps the color gamut of the input SDR, whereas our detail layer reconstruction component, which builds upon prior work on image inpainting, hallucinates missing texture information. The output HDR prediction is produced by a final refinement stage. We present qualitative and quantitative comparisons with existing techniques where our method achieves state‐of‐the‐art performance. Yang Zhang 0003, Tunç Ozan Aydin |
Comput. Graph. Forum | 1 |
| 2018 | The Effect of Foveation on High Dynamic Range Video PerceptionabstractWhen watching a video, the viewer's eyes will fixate on a certain point within each frame. Areas far from the viewers gaze location are perceived with much lower visual acuity than those around the fixation point. This effect is known as foveation. In this paper, the effect of foveation on High Dynamic Range (HDR) video perception is investigated. Using eye tracking data recorded from six different HDR sequences, the bit depth of individual frames are variably encoded, with the pixels with the highest bit depth corresponding to areas around the most likely fixation point for the frame. The bit depth of pixels within the modified frame will then gradually reduce, dependent on how far the pixel is located from the fixation point. To lower the bit depth of the HDR content, a tone mapping operator (TMO) is used. The particular TMO that is used generates an optimal tone mapping curve for every frame, which is used for both tone mapping to reduce the bit depth, and for inverse tone mapping for display purposes. However, this procedure can often cause large amounts of flickering, as well as banding artefacts, which reduce the perceptual quality of the video. Methods to mitigate these effects are proposed and implemented in this paper. Subjective performance evaluations were carried out involving 17 participants in order to evaluate the proposed methodology. Results show that when the lowest bit depth is 8 bits, the modified video is indistinguishable from the original. However, when 6 bit regions are introduced, a significant difference is noticed. Dithering and increasing the foveation region significantly improves the perceptual quality of the modified sequence. Joshua Sowerby, Yang Zhang 0003, Dimitris Agrafiotis |
ACM Multimedia | 2 |
| 2018 | Saliency-based deep convolutional neural network for no-reference image quality assessmentabstractIn this paper, we proposed a novel method for No-Reference Image Quality Assessment (NR-IQA) by combining deep Convolutional Neural Network (CNN) with saliency map. We first investigate the effect of depth of CNNs for NR-IQA by comparing our proposed ten-layer Deep CNN (DCNN) for NR-IQA with the state-of-the-art CNN architecture proposed by Kang et al. ( 2014 ). Our results show that the DCNN architecture can deliver a higher accuracy on the LIVE dataset. To mimic human vision, we introduce saliency maps combining with CNN to propose a Saliency-based DCNN (SDCNN) framework for NR-IQA. We compute a saliency map for each image and both the map and the image are split into small patches. Each image patch is assigned with a patch importance value based on its saliency patch. A set of Salient Image Patches (SIPs) are selected according to their saliency and we only apply the model on those SIPs to predict the quality score for the whole image. Our experimental results show that the SDCNN framework is superior to other state-of-the-art approaches on the widely used LIVE dataset. The TID2008 and the CISQ image quality datasets are utilised to report cross-dataset results. The results indicate that our proposed SDCNN can generalise well on other datasets. Sen Jia 0002, Yang Zhang 0003 |
Multim. Tools Appl. | 2 |
| 2017 | Blind high dynamic range image quality assessment using deep learningabstractIn this paper we propose a No-Reference Image Quality Assessment (NR-IQA) method on High Dynamic Range (HDR) images by combining deep Convolutional Neural Networks (CNNs) with saliency maps. The proposed method utilises the power of deep CNN architectures to extract quality features which can be applied cross HDR and Standard Dynamic Range (SDR) domains. To introduce human visual system to CNNs, a saliency map algorithm is used to select a subset of salient image patches to evaluate on. Our CNN-based method delivers a state-of-the-art performance in HDR NR-IQA experiment, competitive with full reference IQA methods. Sen Jia 0002, Yang Zhang 0003, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 2 |
| 2017 | Performance evaluation of reverse tone mapping operators for dynamic range expansion of SDR video contentabstractWhen displaying Standard Dynamic Range (SDR) video on High Dynamic Range (HDR) displays a reverse tone mapping operation can be employed to expand the dynamic range of the SDR video to that offered by the display. This paper presents a subjective performance evaluation of existing reverse Tone Mapping Operators (rTMOs). The presented study evaluates the performance of rTMOs acting on both well exposed SDR video and video that exhibits exposure variation, as is often the case with user generated content (e.g. captured on mobile phones). The paper highlights flickering artefacts that arise in such cases and evaluates the effect on perceived quality of an adapted de-flickering method. Hu Hao, Yang Zhang 0003, Dimitris Agrafiotis, Matteo Naccari, Marta Mrak |
MMSP | 2 |
| 2016 | High Dynamic Range Video Compression Exploiting Luminance MaskingabstractThe human visual system (HVS) exhibits nonlinear sensitivity to the distortions introduced by lossy image and video coding. This effect is due to the luminance masking, contrast masking, and spatial and temporal frequency masking characteristics of the HVS. This paper proposes a novel perception-based quantization to remove nonvisible information in high dynamic range (HDR) color pixels by exploiting luminance masking so that the performance of the High Efficiency Video Coding (HEVC) standard is improved for HDR content. A profile scaling based on a tone-mapping curve computed for each HDR frame is introduced. The quantization step is then perceptually tuned on a transform unit basis. The proposed method has been integrated into the HEVC reference model for the HEVC range extensions (HM-RExt), and its performance was assessed by measuring the bitrate reduction against the HM-RExt. The results indicate that the proposed method achieves significant bitrate savings, up to 42.2%, with an average of 12.8%, compared with HEVC at the same quality (based on HDR-visible difference predictor-2 and subjective evaluations). Yang Zhang 0003, Matteo Naccari, Dimitris Agrafiotis, Marta Mrak, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | High dynamic range content calibration for accurate acquisition and displayabstractThis paper presents an end-to-end workflow for high dynamic range (HDR) content acquisition and HDR representation that aims to reproduce in a perceptually realistic manner the appearance of the captured scene. The proposed workflow includes a camera-independent colour calibration method for accurate HDR content acquisition and a perceptually optimized HDR signal representation method that takes into account display and viewing conditions. Subjective evaluation of the proposed and other methods indicate that the proposed workflow offers statistically significant improvements in the quality of the HDR image/video displayed to the viewers relative to the anchor methods. Yang Zhang 0003, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 1 |
| 2013 | Visual masking phenomena with high dynamic range contentabstractHigh Dynamic Range (HDR) technology (capture and display), can offer high levels of immersion through a dynamic range that meets and exceeds that of the Human Visual System (HVS). This increase in immersion comes at the cost of higher bitrate requirements, which necessitate the development of efficient HDR-relevant coding solutions. Efficient perception-based compression of HDR imagery requires models that capture accurately the various masking effects experienced by the HVS under HDR conditions, so that bits are not wasted coding redundant imperceptible information. In this paper we present two psychovisual experiments that we carried out with the aid of a high dynamic range display, in order to determine potential differences between Standard Dynamic Range (SDR) and HDR edge masking (EM) and luminance masking (LM) effects. The EM experimental results indicate that the visibility threshold is higher for the case of HDR content than SDR, especially on the dark background side of an edge. The LM experimental results suggest that the HDR visibility threshold is higher compared to SDR for both dark and bright luminance backgrounds. Yang Zhang 0003, Dimitris Agrafiotis, Matteo Naccari, Marta Mrak, David Bull 0001 |
ICIP | 1 |
| 2013 | High dynamic range video compression by intensity dependent spatial quantization in HEVCabstractThe Human Visual System (HVS) shows non-linear sensitivity to the distortion introduced by lossy image and video coding. This non-linear sensitivity is due to luminance masking, contrast masking and the spatial and temporal frequency masking phenomena of the HVS. This paper proposes a perception-based quantization method that exploits luminance masking in the HVS in order to enhance the performance of the High Efficiency Video Coding (HEVC) standard for the case of High Dynamic Range (HDR) video content. A profile scaling based on a tone-mapping curve computed for each HDR frame is introduced. The quantization step is then perceptually tuned on a Transform Unit (TU) basis. The proposed method has been integrated into the reference codec considered for the HEVC range extensions and its performance was assessed by measuring the bitrate reduction against the codec without perceptual quantization. The HDR-VDP-2 image quality metric was employed to measure the compressed picture quality. For the same quality level, an average bitrate reduction of 9% is achieved across all tested HDR sequences. Yang Zhang 0003, Matteo Naccari, Dimitris Agrafiotis, Marta Mrak, David Bull 0001 |
PCS | 1 |
| 2012 | Perceptually lossless High Dynamic Range image compression with JPEG 2000abstractHigh Dynamic Range (HDR) technology offers high levels of immersion with a dynamic range meeting and exceeding that of the Human Visual System (HVS). A primary drawback of HDR images and video is that memory and bandwidth requirements are significantly higher than for conventional images and video. Many bits can be wasted coding redundant imperceptible information. The challenge is therefore to develop means for efficiently compressing HDR imagery to a manageable bit rate without compromising perceptual quality. In this paper, an HDR image compression method, based on an HVS optimized wavelet subband weighting method is proposed. The method has been fully integrated into a JPEG 2000 codec. Experimental results indicate that the proposed method outperforms previous approaches and operates in accordance with characteristics of the HVS, tested objectively using a HDR Visible Difference Predictor (VDP). Yang Zhang 0003, Erik Reinhard, David Bull 0001 |
ICIP | 1 |
| 2011 | Perception-based high dynamic range video compression with optimal bit-depth transformationabstractHigh Dynamic Range (HDR) technology is able to offer high levels of immersion with a dynamic range comparable to the Human Visual System (HVS). A primary drawback of HDR is that its memory and bandwidth requirements are significantly higher than for conventional video. The challenge is thus to develop means for efficiently compressing the video to a manageable bitrate without compromising perceptual quality. In this paper, we propose an HDR compression method based on an optimized bit-depth transformation, and HVS model based wavelet transform denoising. Experimental results indicate that the proposed method outperforms previous approaches and operates in accordance with characteristics of the HVS, tested objectively using a Visible Difference Predictor (VDP). Yang Zhang 0003, Erik Reinhard, David Bull 0001 |
ICIP | 1 |