VLDB 2026 Research / reviewers in the wild / expert
Andrey Norkin
dblp:48/693
· DBLP profile ↗
24ranked-venue papers
12as first author
9since 2021 · last 2026
0000-0002-2417-1635ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Film Grain Synthesis with Debanding FeatureabstractFilm grain synthesis is a powerful tool that can significantly reduce the bitrate of a grainy video. It is typically used with noise removal before the compression, which can make banding more pronounced in the compressed video. When the synthesized grain is added, the banding can still be visible, even at mid QPs. This article describes three algorithms that can be used with the AV1/AV2 film grain synthesis to reduce visibility of underlying bands in the re-noised video. These changes to the film grain synthesis algorithm are computationally inexpensive and improve the perceptual video quality when banding is present. Andrey Norkin |
DCC | 1 |
| 2025 | Banding Prevention for AVM Video CodecabstractThis paper discusses sources of banding artifacts present in video codecs using an example of the AVM video codec and proposes solutions that help to significantly reduce these artifacts. The proposed approach shows a reduction of banding observed by visual inspection and a decrease in banding according to an objective banding metric. The approach does not add new tools to the video codec. There is a penalty of 1.32% in the PSNR-YUV BD-rate observed on a range of test sequences. The proposed solutions can also be applied to other hybrid video codecs. Andrey Norkin |
DCC | 1 |
| 2024 | "Discriminability-Experimental Cost" Tradeoff in Subjective Video Quality Assessment of Codec: DCR with EVP Rating Scale Versus ACR-HRabstractThis work uses naive observers to compare two subjective studies conducted in a controlled laboratory environment on SDR HD, UHD, and HDR UHD contents. These tests aim to compare the precision and accuracy of a modified Degradation Category Rating (DCR) and Absolute Category Rating with Hidden Reference (ACR-HR) subjective methods for video quality assessment. The modified version of the DCR method includes a repetition of both reference and distorted stimuli; and utilizes an 11-grade rating scale from Expert Viewing Protocol (EVP) of ITU-R BT.500-15 standards. In the second subjective protocol, ACR-HR operates without repetition and with the 5-grade quality scale from ITU standards. We extensively analyze the scale usage and compare Mean Opinion Score (MOS) discriminability in both subjective studies. We show that both methods can retrieve accurate MOS. However, the ACR-HR method achieves better discriminability among MOS than DCR with the EVP rating scale while reducing the experimental effort by a factor of two, i.e., the cost of the experiment. The findings of this work give new insight into how to perform cost-efficient subjective tests for video quality estimation with naive observers and how to retrieve good MOS estimates. Andreas Pastor, Ioannis Katsavounidis, Lukas Krasula, Andrey Norkin, Hassene Tmar, Patrick Le Callet |
PCS | 5 |
| 2024 | Convex Hull Prediction for Adaptive Video Streaming by Recurrent LearningabstractAdaptive video streaming relies on the construction of efficient bitrate ladders to deliver the best possible visual quality to viewers under bandwidth constraints. The traditional method of content dependent bitrate ladder selection requires a video shot to be pre-encoded with multiple encoding parameters to find the optimal operating points given by the convex hull of the resulting rate-quality curves. However, this pre-encoding step is equivalent to an exhaustive search process over the space of possible encoding parameters, which causes significant overhead in terms of both computation and time expenditure. To reduce this overhead, we propose a deep learning based method of content aware convex hull prediction. We employ a recurrent convolutional network (RCN) to implicitly analyze the spatiotemporal complexity of video shots in order to predict their convex hulls. A two-step transfer learning scheme is adopted to train our proposed RCN-Hull model, which ensures sufficient content diversity to analyze scene complexity, while also making it possible to capture the scene statistics of pristine source videos. Our experimental results reveal that our proposed model yields better approximations of the optimal convex hulls, and offers competitive time savings as compared to existing approaches. On average, the pre-encoding time was reduced by 53.8% by our method, while the average Bjøntegaard delta bitrate (BD-rate) of the predicted convex hulls against ground truth was 0.26%, and the mean absolute deviation of the BD-rate distribution was 0.57%. Somdyuti Paul, Andrey Norkin, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2023 | Self-Supervised Learning of Perceptually Optimized Block Motion Estimates for Video CompressionabstractBlock based motion estimation is integral to inter prediction processes performed in hybrid video codecs. Prevalent block matching based methods that are used to compute block motion vectors (MVs) rely on computationally intensive search procedures. They also suffer from the aperture problem, which tends to worsen as the block size is reduced. Moreover, the block matching criteria used in typical codecs do not account for the resulting levels of perceptual quality of the motion compensated pictures that are created upon decoding. Towards achieving the elusive goal of perceptually optimized motion estimation, we propose a search-free block motion estimation framework using a multi-stage convolutional neural network, which is able to conduct motion estimation on multiple block sizes simultaneously, using a triplet of frames as input. This composite block translation network (CBT-Net) is trained in a self-supervised manner on a large database that we created from publicly available uncompressed video content. We deploy the multi-scale structural similarity (MS-SSIM) loss function to optimize the perceptual quality of the motion compensated predicted frames. Our experimental results highlight the computational efficiency of our proposed model relative to conventional block matching based motion estimation algorithms, for comparable prediction errors. Further, when used to perform inter prediction in AV1, the MV predictions of the perceptually optimized model result in average Bjøntegaard-delta rate (BD-rate) improvements of -1.73% and -1.31% with respect to the MS-SSIM and Video Multi-Method Assessment Fusion (VMAF) quality metrics, respectively, as compared to the block matching based motion estimation system employed in the SVT-AV1 encoder. Somdyuti Paul, Andrey Norkin, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2022 | Generalized deblocking filter for AVMabstractThe AV1 deblocking filter did not sufficiently attenuate visual quality artifacts when used in the AOMedia video model (AVM), especially at lower bitrates. A generalized deblocking filter described in this paper uses one equation for any filter length. The filter brings PSNR-YUV BD-rate $-\mathbf{0. 1 7 \%},\ -0.90\%,-1.07\%$, and $-0.92\%$ on All Intra, Random Access, LowDelay, and Adaptive Streaming configurations in the AOMedia common test conditions. Improvement of visual quality is observed on a number of sequences, while the computational complexity is close to that of the AVM deblocking. Andrey Norkin |
PCS | 1 |
| 2022 | An Open Video Dataset For Screen Content CodingabstractIn recent years, screen content video is becoming increasingly popular in several major video applications, such as video recording and video conferencing. Due to the unique features of screen content videos that are not captured by camera sensors but produced artificially, dedicated coding tools have been developed for achieving significant compression efficiency gain. In recognition of the popularity of screen content applications, an open video dataset for screen content is proposed in this paper for the development of screen content coding technologies. The proposed video dataset consists of 12 typical screen content type video clips that are publicly available. In addition, to better understand the characteristics of the proposed video dataset, several major screen content coding tools in AOMedia Video 1 (AV1) have been evaluated on this dataset and analyzed in this paper. Yingbin Wang, Xin Zhao 0003, Xiaozhong Xu, Shan Liu 0001, Zhijun Lei, Mariana Afonso, Andrey Norkin, Thomas Daede |
PCS | 7 |
| 2021 | On visual masking estimation for adaptive quantization using steerable filters
Somdyuti Paul, Andrey Norkin, Alan C. Bovik |
Signal Process. Image Commun. | 2 |
| 2021 | ProxIQA: A Proxy Approach to Perceptual Optimization of Learned Image Compressionabstract(p = 1,2) norms has largely dominated the measurement of loss in neural networks due to their simplicity and analytical properties. However, when used to assess the loss of visual information, these simple norms are not very consistent with human perception. Here, we describe a different "proximal" approach to optimize image analysis networks against quantitative perceptual models. Specifically, we construct a proxy network, broadly termed ProxIQA, which mimics the perceptual model while serving as a loss layer of the network. We experimentally demonstrate how this optimization framework can be applied to train an end-to-end optimized image compression network. By building on top of an existing deep image compression model, we are able to demonstrate a bitrate reduction of as much as 31% over MSE optimization, given a specified perceptual quality (VMAF) level. Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Andrey Norkin, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2020 | Speeding Up VP9 Intra Encoder With Hierarchical Deep Learning-Based Partition PredictionabstractIn VP9 video codec, the sizes of blocks are decided during encoding by recursively partitioning 64×64 superblocks using rate-distortion optimization (RDO). This process is computationally intensive because of the combinatorial search space of possible partitions of a superblock. Here, we propose a deep learning based alternative framework to predict the intra-mode superblock partitions in the form of a four-level partition tree, using a hierarchical fully convolutional network (H-FCN). We created a large database of VP9 superblocks and the corresponding partitions to train an H-FCN model, which was subsequently integrated with the VP9 encoder to reduce the intra-mode encoding time. The experimental results establish that our approach speeds up intra-mode encoding by 69.7% on average, at the expense of a 1.71% increase in the Bjøntegaard-Delta bitrate (BD-rate). While VP9 provides several built-in speed levels which are designed to provide faster encoding at the expense of decreased rate-distortion performance, we find that our model is able to outperform the fastest recommended speed level of the reference VP9 encoder for the good quality intra encoding configuration, in terms of both speedup and BD-rate. Somdyuti Paul, Andrey Norkin, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2018 | Film Grain Synthesis for AV1 Video CodecabstractFilm grain is abundant in TV and movie content. It is often part of the creative intent and needs to be preserved while encoding. However, the random nature of film grain is difficult to compress using traditional coding tools. This paper describes a film grain modeling and synthesis algorithm proposed for the AV1 video codec. At the encoder, an autoregressive model of film grain is transmitted relative to a denoised signal, and the film grain strength is modeled as a function of intensity. The corresponding renoising at the decoder is implemented using an efficient block-based approach suitable for use in consumer electronic devices. Preliminary results indicate that the approach can give significant bitrate savings (up to 50%) on sequences with heavy film grain. Andrey Norkin, Neil Birkbeck |
DCC | 1 |
| 2018 | An Overview of Core Coding Tools in the AV1 Video CodecabstractAV1 is an emerging open-source and royalty-free video compression format, which is jointly developed and finalized in early 2018 by the Alliance for Open Media (AOMedia) industry consortium. The main goal of AV1 development is to achieve substantial compression gain over state-of-the-art codecs while maintaining practical decoding complexity and hardware feasibility. This paper provides a brief technical overview of key coding techniques in AV1 along with preliminary compression performance comparison against VP9 and HEVC. Yue Chen 0040, Debargha Mukherjee, Jingning Han, Adrian Grange, Yaowu Xu, Zoe Liu, Sarah Parker, Hui Su, Urvang Joshi, Ching-Han Chiang, Yunqing Wang, Paul Wilkins, Jim Bankoski, Luc N. Trudeau, Nathan E. Egge, Jean-Marc Valin, Thomas Davies 0002, Steinar Midtskogen, Andrey Norkin, Peter De Rivaz |
PCS | 20 |
| 2016 | Fast Algorithm for HDR Color ConversionabstractThe paper addresses a problem of perceptual artifacts that appear in Y'CbCr non-linear luminance 4:2:0 HDR video. A computationally inexpensive method is proposed for converting the 4:4:4 HDR video to Y'CbCr 4:2:0 nonconstant luminance format. The method removes artifacts in areas with saturated colors. The approach obtains results in one step, improving the average linear light PSNR by 2.16 dB and tPSNR metric by 1.99 dB on the investigated videos. Andrey Norkin |
DCC | 1 |
| 2016 | Fast algorithm for HDR video pre-processingabstractThe paper addresses a problem of perceptual artifacts that appear in saturated colors of Y'CbCr non-constant luminance 4:2:0 HDR video. A computationally inexpensive method is proposed for converting 4:4:4 HDR video to Y'CbCr 4:2:0 non-constant luminance format. The method shows similar objective performance to an iterative algorithm and outperforms a previously published explicit solution. The method removes artifacts in areas with saturated colors, obtaining results in one step. Compared to the iterative algorithm, the proposed solution significantly decreases the worst-case and average complexity. Andrey Norkin |
PCS | 1 |
| 2014 | HEVC-based deblocking filter with ramp preservation propertiesabstractThe paper presents an HEVC-based deblocking filter that improves perceptual quality of reconstructed video on the content that exhibits a lot of chaotic motion, such as water, fire or smoke while providing similar quality on “normal” video content, such as the content with linear motion. The filter is also capable of efficiently suppressing block artifacts in smooth areas with slowly changing samples intensity. The objective performance of the proposed deblocking filter is on average similar to the HEVC deblocking. Andrey Norkin |
ICIP | 1 |
| 2013 | Two HEVC encoder methods for block artifact reductionabstractThe HEVC deblocking filter significantly improves the subjective quality of coded video sequences at lower bitrates. During the final phase of HEVC standardization, it was shown that the reference software encoder may produce visible block artifacts on some sequences with content that shows chaotic motion, such as water or fire. The paper analyses the reasons for blocking artifacts in such sequences and describes two simple encoder-side methods that improve the subjective quality on these sequences without degrading the quality on other content and without significant bitrate increase. The effect on subjective quality has been evaluated by a formal subjective test. Andrey Norkin, Kenneth Andersson, Valentin Kulyk |
VCIP | 1 |
| 2012 | 3DTV: One stream for different screens: Keeping perceived scene proportions by adjusting camera parametersabstractStereo and 3D video material is usually optimized during the production phase for a particular display size and viewing distance. When the content is shown on a display of different size and/or from different viewing distance, the perceived proportions of objects, e.g. object's depth relative to object's size, will be distorted compared to the original viewing conditions. This can make the scene look unnatural and even lead to eye strain and fatigue when observing the content. This paper proposes equations for adapting rendering parameters to new viewing conditions so that the perceived proportions of the objects in 3D scene are exactly the same as for the reference viewing conditions. Andrey Norkin, Ivana Girdzijauskas |
PCS | 1 |
| 2012 | HEVC Deblocking FilterabstractThis paper describes the in-loop deblocking filter used in the upcoming High Efficiency Video Coding (HEVC) standard to reduce visible artifacts at block boundaries. The deblocking filter performs detection of the artifacts at the coded block boundaries and attenuates them by applying a selected filter. Compared to the H.264/AVC deblocking filter, the HEVC deblocking filter has lower computational complexity and better parallel processing capabilities while still achieving significant reduction of the visual artifacts. Andrey Norkin, Gisle Bjøntegaard, Arild Fuldseth, Matthias Narroschke, Masaru Ikeda, Kenneth Andersson, Minhua Zhou, Geert Van der Auwera |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Low complexity video coding and the emerging HEVC standardabstractThis paper describes a low complexity video codec with high coding efficiency. It was proposed to the High Efficiency Video Coding (HEVC) standardization effort of MPEG and VCEG, and has been partially adopted into the initial HEVC Test Model under Consideration design. The proposal utilizes a quad-tree structure with a support of large macroblocks of size 64×64 and 32×32, in addition to macroblocks of size 16×16. The entropy coding is done using a low complexity variable length coding based scheme with improved context adaptation over the H.264/AVC design. In addition, the proposal includes improved interpolation and deblocking filters, giving better coding efficiency while having low complexity. Finally, an improved intra coding method is presented. The subjective quality of the proposal is evaluated extensively and the results show that the proposed method achieves similar visual quality as H.264/AVC High Profile anchors with around 50% and 35% bit rate reduction for low delay and random-access experiments respectively at high definition sequences. This is achieved with less complexity than H.264/AVC Baseline Profile, making the proposal especially suitable for resource constrained environments. Kemal Ugur, Kenneth Andersson, Arild Fuldseth, Gisle Bjøntegaard, Lars Petter Endresen, Jani Lainema, Antti Hallapuro, Justin Ridge, Dmytro Rusanovskyy, Cixun Zhang, Andrey Norkin, Clinton Priddle, Thomas Rusert, Jonatan Samuelsson, Rickard Sjöberg, Zhuangfei Wu |
PCS | 11 |
| 2010 | High Performance, Low Complexity Video Coding and the Emerging HEVC StandardabstractThis paper describes a low complexity video codec with high coding efficiency. It was proposed to the high efficiency video coding (HEVC) standardization effort of moving picture experts group and video coding experts group, and has been partially adopted into the initial HEVC test model under consideration design. The proposal utilizes a quadtree-based coding structure with support for macroblocks of size 64$\,\times\,$64, 32$\,\times\,$32, and 16$\,\times\,$16 pixels. Entropy coding is performed using a low complexity variable length coding scheme with improved context adaptation compared to the context adaptive variable length coding design in H.264/AVC. The proposal's interpolation and deblocking filter designs improve coding efficiency, yet have low complexity. Finally, intra-picture coding methods have been improved to provide better subjective quality than H.264/AVC. The subjective quality of the proposed codec has been evaluated extensively within the HEVC project, with results indicating that similar visual quality to H.264/AVC High Profile anchors is achieved, measured by mean opinion score, using significantly fewer bits. Coding efficiency improvements are achieved with lower complexity than the H.264/AVC Baseline Profile, particularly suiting the proposal for high resolution, high quality applications in resource-constrained environments. Kemal Ugur, Kenneth Andersson, Arild Fuldseth, Gisle Bjøntegaard, Lars Petter Endresen, Jani Lainema, Antti Hallapuro, Justin Ridge, Dmytro Rusanovskyy, Cixun Zhang, Andrey Norkin, Clinton Priddle, Thomas Rusert, Jonatan Samuelsson, Rickard Sjöberg, Zhuangfei Wu |
IEEE Trans. Circuits Syst. Video Technol. | 11 |
| 2008 | Principal Component Analysis in Multiple Description Coding of Spectral ImagesabstractCommunications in general require protection due to error-prone channels. In geoscience and remote sensing, especially coded or compressed data, and results from classifications are vulnerable to transmissions errors. Multiple descriptions of data are one way for protection of communications over unreliable channels. This study concentrates on multiple description of spectral images as a way for providing scalable coding. The principal component analysis outputs the common content, the redundant part, for the two descriptions and then the integer wavelet transform selects different contents for those descriptions. In the experiments, the goal was to find good parameterization in the transmitter for generating the two descriptions which will allow perfect reconstruction if both of them are available at the receiver. The reconstruction quality for various number of principal components is demonstrated. Integer wavelet filter 5/3 showed the best performance among the implemented filters. Arto Kaarna, Andrey Norkin, Jaakko Astola |
IGARSS (3) | 2 |
| 2007 | Packet Loss Resilient Transmission of 3D ModelsabstractThis paper presents an efficient joint source-channel coding scheme based on forward error correction (FEC) for three dimensional (3D) models. The system employs a wavelet based zero-tree 3D mesh coder based on Progressive Geometry Compression (PGC). Reed-Solomon (RS) codes are applied to the embedded output bitstream to add resiliency to packet losses. Two-state Markovian channel model is employed to model packet losses. The proposed method applies approximately optimal and unequal FEC across packets. Therefore the scheme is scalable to varying network bandwidth and packet loss rates (PLR). In addition, Distortion-Rate (D-R) curve is modeled to decrease the computational complexity. Experimental results show that the proposed method achieves considerably better expected quality compared to previous packet-loss resilient schemes. M. Oguz Bici, Andrey Norkin, Gozde Bozdagi Akar |
ICIP (5) | 2 |
| 2007 | Wavelet-based multiple description coding of 3-D geometryabstractIn this work, we present a multiple description coding (MDC) scheme for reliable transmission of compressed three dimensional (3-D) meshes. It trades off reconstruction quality for error resilience to provide the best expected reconstruction of 3-D mesh at the decoder side. The proposed scheme is based on multiresolution geometry compression achieved by using wavelet transform and modified SPIHT algorithm. The trees of wavelet coefficients are divided into sets. Each description contains the coarsest level mesh and a number of tree sets coded with different rates. The original 3-D geometry can be reconstructed with acceptable quality from any received description. More descriptions provide better reconstruction quality. The proposed algorithm provides flexible number of descriptions and is optimized for varying packet loss rates (PLR) and channel bandwidth. Andrey Norkin, M. Oguz Bici, Gozde Bozdagi Akar, Atanas P. Gotchev, Jaakko Astola |
VCIP | 1 |
| 2006 | Two-stage multiple description image coders: Analysis and comparative study
Andrey Norkin, Atanas P. Gotchev, Karen Egiazarian, Jaakko Astola |
Signal Process. Image Commun. | 1 |