EDBT 2026 Demo / reviewers in the wild / expert
Elena Alshina
dblp:55/9463
· DBLP profile ↗
37ranked-venue papers
2as first author
19since 2021 · last 2025
0000-0001-7099-5371ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 6Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Suitability of Perceptual Quality Metrics for Learning-Based Screen Content CompressionabstractLearned image compression methods tailored for screen content demand perceptual quality metrics that are both accurate and computationally efficient. Traditional metrics such as PSNR and SSIM often fail to capture perceptual distortions specific to screen content, while many advanced learned perceptual metrics are too complex or non-differentiable for practical use with implicit neural codecs. These codecs require loss functions to be evaluated tens of thousands of times during training, imposing a strict complexity limit of a few hundred multiply-accumulate operations per pixel. In this paper, we evaluate several differentiable, low-complexity perceptual metrics on three screen content quality datasets, including a newly collected compression-focused crowdsourced dataset, PerceptualSCC. Our results identify VIF and NLPD as the best-performing metrics; however, only NLPD demonstrates stable behavior in gradient-based optimization experiments. These findings suggest NLPD as a strong candidate for guiding learned screen content compression, balancing perceptual fidelity, computational cost, and optimization stability. H. Burak Dogaroglu, Hongjie You, Atanas Boev, Elena Alshina, Eckehard G. Steinbach |
ISM | 4 |
| 2025 | A Method for Rate Point Determination for Visual Evaluation of Video Sequences
Mathias Wien, Adam Wieckowski, Elena Alshina, Edouard François, Pavel Nikitin, Kenneth Andersson |
PCS | 3 |
| 2025 | Overview of Variable Rate Coding in JPEG AIabstractEmpirical evidence has demonstrated that learning-based image compression can outperform classical compression frameworks. This has led to the ongoing standardization of learned-based image codecs, namely Joint Photographic Experts Group (JPEG) AI. The objective of JPEG AI is to enhance compression efficiency and provide a software and hardware-friendly solution. Based on our research, JPEG AI represents the first standardization that can facilitate the implementation of a learned image codec on a mobile device. This article presents an overview of the variable rate coding functionality in JPEG AI, which includes three variable rate adaptations: a three-dimensional quality map, a fast bit rate matching algorithm, and a training strategy. The variable rate adaptations offer a continuous rate function up to 2.0 bpp, exhibiting a high level of performance, a flexible bit allocation between different color components, and a region of interest function for the specified use case. The evaluation of performance encompasses both objective and subjective results. With regard to the objective bit rate matching, the main profile with low complexity yielded a 13.1% BD-rate gain over VVC intra, while the high profile with high complexity achieved a 19.2% BD-rate gain over VVC intra. The BD-rate result is calculated as the mean of the seven perceptual metrics defined in the JPEG AI common test conditions. With respect to subjective results, the example of improving the quality of the region of interest is illustrated. Panqi Jia, Fabian Brand, Dequan Yu, Alexander Karabutov, Elena Alshina, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Quantized Decoder in Learned Image Compression for Deterministic ReconstructionabstractLearned image compression has a problem of non-bit-exact reconstruction due to different calculations of floating point arithmetic on different devices. This paper shows a method to achieve a deterministic reconstructed image by quantizing only the decoder of the learned image compression model. From the implementation perspective of an image codec, it is beneficial to have the results reproducible when decoded on different devices. In this paper, we study quantization of weights and activations without overflow of accumulator in all decoder subnetworks. We show that the results are bit-exact at the output, and the resulting BD-rate loss of quantization of decoder is 0.5% in the case of 16-bit weights and 16-bit activations, and 7.9% in the case of 8-bit weights and 16-bit activations. Esin Koyuncu, Timofey Solovyev, Johannes Sauer, Elena Alshina, André Kaup |
ICASSP | 4 |
| 2024 | Adapting Learned Image Codecs To Screen Content Via Adjustable TransformationsabstractAs learned image codecs (LICs) become more prevalent, their low coding efficiency for out-of-distribution data becomes a bottleneck for some applications. To improve the performance of LICs for screen content (SC) images without breaking backwards compatibility, we propose to introduce parameterized and invertible linear transformations into the coding pipeline without changing the underlying baseline codec’s operation flow. We design two neural networks to act as prefilters and postfilters in our setup to increase the coding efficiency and help with the recovery from coding artifacts. Our end-to-end trained solution achieves up to $10 \%$ bitrate savings on SC compression compared to the baseline LICs while introducing only $1 \%$ extra parameters. H. Burak Dogaroglu, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach |
ICIP | 4 |
| 2024 | Adaptive Variance-Threshold-Based Skip Modes for Learned Video Compression Using a Motion Complexity CriterionabstractSkip modes are a powerful tool to reduce the rate in video compression. The main idea is that the residual areas where the prediction performs well are not transmitted since the prediction quality is good enough that the prediction signal itself can be used as the reconstruction signal. This is commonly used, e.g., in the compression standard VVC, where a skip flag can be transmitted for inter blocks under certain conditions. The skipped residual block is then not transmitted and the content is instead inferred to be zero at the decoder. Current learning-based methods use different kinds of skip modes. One possibility here arises from the fact that the coders estimate and transmit the variance for each transmitted symbol. It has been proposed to use this estimated variance to derive a skip mode. When the variance falls below a threshold, the symbol is not transmitted. In this paper we propose an extension to this method. By classifying each position in the latent space according to the local motion complexity, we can transmit adaptive thresholds for each class. That way, we can employ motion information to refine the granularity of the skip mode. When we implement this method in FVC, we are able to save up 2.11% rate on a GOP 20 sequence. We also discuss the behavior of increasingly adaptive skip modes in scenarios with larger GOP size, where error-propagation becomes a larger issue. Fabian Brand, Jürgen Seiler, Johannes Sauer, Elena Alshina, André Kaup |
PCS | 4 |
| 2024 | Bit Rate Matching Algorithm Optimization in JPEG-AI Verification ModelabstractThe research on neural network (NN) based image compression has shown superior performance compared to classical compression frameworks. Unlike the hand-engineered transforms in the classical frameworks, NN-based models learn the non-linear transforms providing more compact bit represen-tations, and achieve faster coding speed on parallel devices over their classical counterparts. Those properties evoked the attention of both scientific and industrial communities, resulting in the standardization activity JPEG-AI. The verification model for the standardization process of JPEG-AI is already in development and has surpassed the advanced VVC intra codec. To generate reconstructed images with the desired bits per pixel and assess the BD-rate performance of both the JPEG-AI verification model and VVC intra, bit rate matching is employed. However, the current state of the JPEG-AI verification model experiences significant slowdowns during bit rate matching, resulting in suboptimal performance due to an unsuitable model. The proposed methodology offers a gradual algorithmic optimization for matching bit rates, resulting in a fourfold acceleration and over 1% improvement in BD-rate at the base operation point. At the high operation point, the acceleration increases up to sixfold. Panqi Jia, Ahmet Burakhan Koyuncu, Jue Mao, Ze Cui, Tiansheng Guo, Timofey Solovyev, Alexander Karabutov, Yin Zhao, Jing Wang 0194, Elena Alshina, André Kaup |
PCS | 11 |
| 2024 | Bit Distribution Study and Implementation of Spatial Quality Map in the JPEG-AI StandardizationabstractCurrently, there is a high demand for neural network-based image compression codecs. These codecs employ non-linear transforms to create compact bit representations and facilitate faster coding speeds on devices compared to the handcrafted transforms used in classical frameworks. The scientific and industrial communities are highly interested in these properties, leading to the standardization effort of JPEG-AI. The JPEG-AI verification model has been released and is currently under development for standardization. Utilizing neural networks, it can outperform the classic codec VVC intra by over 10% BD-rate operating at base operation point. Researchers attribute this success to the flexible bit distribution in the spatial domain, in contrast to VVC intra’s anchor that is generated with a constant quality point. However, our study reveals that VVC intra displays a more adaptable bit distribution structure through the implementation of various block sizes. As a result of our observations, we have proposed a spatial bit allocation method to optimize the JPEG-AI verification model’s bit distribution and enhance the visual quality. Furthermore, by applying the VVC bit distribution strategy, the objective performance of JPEG-AI verification mode can be further improved, resulting in a maximum gain of 0.45 dB in PSNR-Y. Panqi Jia, Jue Mao, Esin Koyuncu, Ahmet Burakhan Koyuncu, Timofey Solovyev, Alexander Karabutov, Yin Zhao, Elena Alshina, André Kaup |
VCIP | 8 |
| 2024 | Judder Modelling Framework with Perceptual Quality Score Prediction for HDR VideosabstractJudder is a motion artifact describing a perceptual mismatch between human visual system and discrete movements on a display. Judder is correlated with video frame rates, motion speeds and brightness. It significantly degrades the perceived video quality. We present a framework for modeling judder, where the input frames are processed to generate optical flow and a sensitivity map. Concurrently, a specifically designed attention map is integrated into the process. These components collectively contribute to the creation of a feature map, which is utilized to compute judder scores. The proposed approach explicitly takes into account the video frame rate, resulting in an accurate judder score to predict the subjective mean opinion scores ultimately. Our assessment of the proposed framework on an HDR video dataset shows that judderness is highly influenced by the frame rate and, to some extent, by motion speed and brightness. The prediction of perceived quality scores shows improvement compared to the baseline framework. Hongjie You, Nicola Giuliani, Atanas Boev, Elena Alshina, Eckehard G. Steinbach |
VCIP | 5 |
| 2024 | Efficient Contextformer: Spatio-Channel Window Attention for Fast Context Modeling in Learned Image CompressionabstractEntropy estimation is essential for the performance of learned image compression. It has been demonstrated that a transformer-based entropy model is of critical importance for achieving a high compression ratio, however, at the expense of a significant computational effort. In this work, we introduce the Efficient Contextformer (eContextformer) – a computationally efficient transformer-based autoregressive context model for learned image compression. The eContextformer efficiently fuses the patch-wise, checkered, and channel-wise grouping techniques for parallel context modeling, and introduces a shifted window spatio-channel attention mechanism. We explore better training strategies and architectural designs and introduce additional complexity optimizations. During decoding, the proposed optimization techniques dynamically scale the attention span and cache the previous attention computations, drastically reducing the model and runtime complexity. Compared to the non-parallel approach, our proposal has ~145x lower model complexity and ~210x faster decoding speed, and achieves higher average bit savings on Kodak, CLIC2020, and Tecnick datasets. Additionally, the low complexity of our context model enables online rate-distortion algorithms, which further improve the compression performance. We achieve up to 17% bitrate savings over the intra coding of Versatile Video Coding (VVC) Test Model (VTM) 16.2 and surpass various learning-based compression models. Ahmet Burakhan Koyuncu, Panqi Jia, Atanas Boev, Elena Alshina, Eckehard G. Steinbach |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | CALC-VFS: Content-adaptive low-complexity Video Frame SynthesisabstractWe present a content-adaptive, low-complexity video frame synthesis algorithm. Our approach applies the dynamic convolutions content adaptation approach to the widely used frame synthesis algorithm IFRNet. By introducing dynamic convolutions into both the pyramid encoder and the coarse-to-fine decoders of IFRNet, we enforce sparsity, thereby limiting the computationally expensive operations to only the necessary pixels. Training for specific sparsity targets allows us to achieve overall less computational complexity compared to IFRNet while having similar performance. We demonstrate the performance and content adaptivity in two test scenarios and show the savings in computational budget (approximately 20-40%) compared to the baseline IFRNet. Nicola Giuliani, Hongjie You, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach |
ISM | 5 |
| 2023 | Subjective video quality assessment of immersive HDR content on head-mounted displaysabstractHigh dynamic range (HDR) videos are known to provide better visual quality on HDR TV displays. Head-mounted displays (HMDs) are an integral part of immersive visual experiences. However, typical HMDs are equipped with standard dynamic range (SDR) displays, failing to show details in bright and dark areas of HDR content. Therefore, we aim to evaluate the perceptual quality improvement, in terms of mean opinion score (MOS), when observers view HDR instead of SDR content on immersive displays. We developed a pipeline to render 2D HDR and tone-mapped SDR immersive videos for HMDs. We conducted two single-stimulus subjective evaluation experiments to evaluate (1) the perceived visual difference when comparing one HDR scene with three tone-mapped SDR versions of it, and (2) how frame rates impact the perceptual quality of HDR immersive videos. Our results in MOS show that (1) there is a significant improvement of perceptual quality in HDR compared to tone-mapped SDR content, and (2) HDR immersive videos benefit much more from higher frame rates than SDR videos. Hongjie You, Nicola Giuliani, Atanas Boev, Elena Alshina, Eckehard G. Steinbach |
VCIP | 5 |
| 2022 | Contextformer: A Transformer with Spatio-Channel Attention for Context Modeling in Learned Image Compression
Ahmet Burakhan Koyuncu, Han Gao 0001, Atanas Boev, Georgii Gaikov, Elena Alshina, Eckehard G. Steinbach |
ECCV (19) | 5 |
| 2022 | Learning-Based Conditional Image Coder Using Color SeparationabstractRecently, image compression codecs based on Neural Networks (NN) outperformed the state-of-art classic ones such as BPG, an image format based on HEVC intra. However, the typical NN codec has high complexity, and it has limited options for parallel data processing. In this work, we propose a conditional separation principle that aims to improve parallelization and lower the computational requirements of an NN codec. We present a Conditional Color Separation (CCS) codec which follows this principle. The color components of an image are split into primary and non-primary ones. The processing of each component is done separately, by jointly trained networks. Our approach allows parallel processing of each component, flexibility to select different channel numbers, and an overall complexity reduction. The CCS codec uses over 40% less memory, has 2x faster encoding and 22% faster decoding speed, with only 4% BD-rate loss in RGB PSNR compared to our baseline model over BPG. Panqi Jia, Ahmet Burakhan Koyuncu, Georgii Gaikov, Alexander Karabutov, Elena Alshina, André Kaup |
PCS | 5 |
| 2022 | Device Interoperability for Learned Image Compression with Weights and Activations QuantizationabstractLearning-based image compression has improved to a level where it can outperform traditional image codecs such as HEVC and VVC in terms of coding performance. In addition to good compression performance, device interoperability is essential for a compression codec to be deployed, i.e., encoding and decoding on different CPUs or GPUs should be error-free and with negligible performance reduction. In this paper, we present a method to solve the device interoperability problem of a state-of-the-art image compression network. We implement quantization to entropy networks which output entropy parameters. We suggest a simple method which can ensure cross-platform encoding and decoding, and can be implemented quickly with minor performance deviation, of 0.3% BD-rate, from floating point model results. Esin Koyuncu, Timofey Solovyev, Elena Alshina, André Kaup |
PCS | 3 |
| 2021 | Quality-Blind Compressed Color Image Enhancement with Convolutional Neural NetworksabstractLossy compressed images and videos suffer from visible compression artifacts, especially when the bit-rate is low. To improve the quality of the compressed image while keeping the same bit-rate, decoder-side compression artifacts reduction (CAR) becomes important. Recently, convolutional neural networks are adopted for CAR tasks and achieve the state-of-the-art performance. However, most CAR algorithms only focus on the reconstruction of the luminance channel. Also, a separate model usually needs to be trained for each quality factor (QF), which makes these approaches not practical in existing codecs. In this paper, we analyze a quality-blind training strategy and compare it with training separate models for each QF. The testing results with three representative CAR algorithms show the superiority of the quality-blind training compared to separate training. The results for pseudo and real quality-blind CAR tests further prove the generalizability of the quality-blind training for practical CAR tasks. Kai Cui 0003, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach |
ISCAS | 4 |
| 2021 | Convolutional neural network-based post-filtering for compressed YUV420 images and videoabstractImages and videos compressed with lossy compression algorithms usually suffer from visible distortions, especially when the bitrate is low. To improve the quality without spending extra bitrate, many image and video codecs have built-in filters to mitigate these artifacts. However, most of them are only applied on the luminance channel, while the chrominance channels remain unmodified. While this is partly justified by the observation that the luminance channel usually contains more details and has higher-resolution than the chrominance channels. We observe that the luminance and chrominance channels still have latent correlations. Therefore, the post-filtering of the chrominance channels is also beneficial and can be driven by the information from the luminance channel. In this paper, we propose a 3-stage YUV post-filtering network for compressed YUV420 images and video. The proposed 3-stage structure not only improves the quality of the luminance channel, but also exploits the luma-chroma correlations to improve the quality of the chrominance channels. Our experimental results show that the proposed approach achieves 3.60%/12.75%/14.93% Bj⊘ntegaard Delta bitrate improvement for the Y, U and V channels over the VVC 10.0 codec for All-Intra configuration. Kai Cui 0003, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach |
PCS | 4 |
| 2021 | Geometric Partitioning Mode in Versatile Video Coding: Algorithm Review and AnalysisabstractThis paper presents an overview of the geometric partitioning mode (GPM) algorithm that is a part of the most recent Versatile Video Coding (VVC) standard. The GPM algorithm aims to increase the partitioning precision of moving objects using non-rectangular and asymmetric rectangular partitions on top of the conventional rectangular block partitioning structure of VVC. Novel features of GPM contributing to the increase in coding efficiency and the reduction in encoder and decoder complexity are detailed and analyzed in this paper. Evaluated with VVC test model version 8.0 under the joint video experts team common test conditions, experimental results show that the presented GPM algorithm provides luma Bjøntegaard Delta rate reduction of 0.70% for random access and of 1.55% for low delay with B slices configurations, with roughly 3% to 5% additional encoding time and negligible decoder runtime change. Furthermore, as GPM provides more precise partitions for the boundaries of the moving objects, an improvement of visual quality is seen in GPM coded sequences. Han Gao 0001, Semih Esenlik, Elena Alshina, Eckehard G. Steinbach |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Subblock-Based Motion Derivation and Inter Prediction Refinement in the Versatile Video Coding StandardabstractEfficient representation and coding of fine-granular motion information is one of the key research areas for exploiting inter-frame correlation in video coding. Representative techniques towards this direction are affine motion compensation (AMC), decoder-side motion vector refinement (DMVR), and subblock-based temporal motion vector prediction (SbTMVP). Fine-granular motion information is derived at subblock level for all the three coding tools. In addition, the obtained inter prediction can be further refined by two optical flow-based coding tools, the bi-directional optical flow (BDOF) for bi-directional inter prediction and the prediction refinement with optical flow (PROF) exclusively used in combination with AMC. The aforementioned five coding tools have been extensively studied and finally adopted in the Versatile Video Coding (VVC) standard. This paper presents technical details of each tool and highlights the design elements with the consideration of typical hardware implementations. Following the common test conditions defined by Joint Video Experts Team (JVET) for the development of VVC, 5.7% bitrate reduction on average is achieved by the five tools. For test sequences characterized by large and complex motion, up to 13.4% bitrate reduction is observed. Additionally, visual quality improvement is demonstrated and analyzed. Haitao Yang 0001, Huanbang Chen, Jianle Chen, Semih Esenlik, Sriram Sethuraman, Xiaoyu Xiu, Elena Alshina, Jiancong Luo |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2020 | Intra Prediction in the Emerging VVC Video Coding StandardabstractThe focus of this work is on intra-prediction tools that distinguish Versatile Video Coding (VVC) [1] from its predecessors and do not add RD-checks on the encoder side, i.e. their coding efficiency is achieved due to their inherent properties but not by additional encoder-side complexity increase. Alexey Filippov, Vasily Rufitskiy, Jianle Chen, Elena Alshina |
DCC | 4 |
| 2020 | Advanced Geometric-Based Inter Prediction for Versatile Video CodingabstractBlock-based partitioning is one of the fundamental techniques in video coding. Geometric-based block partitioning is a well-studied method to enable better spatial adaptation to the signal properties. This paper introduces the most recent proposal of advanced geometric-based inter prediction (GIP) made to the state-of-the-art are video coding standard - Versatile Video Coding (VVC). Implemented in the latest test model VTM-6.0 to generalize the existing triangle partition mode (TPM) and evaluated with the Joint Video Experts Team (JVET) Common Test Conditions (CTC) sequences, the proposed advanced GIP scheme provides luma BD-rate reduction of 0.56% for random access (RA) and 1.37% for low-delay (LB) test cases with 2% encoder runtime increase and negligible decoder runtime increase. Furthermore, BD-rate reductions up to 2.92% and 3.49% for RA and LB test cases can be achieved in the absence of multiple related VVC inter prediction tools. Han Gao 0001, Ru-Ling Liao, Kevin Reuze, Semih Esenlik, Elena Alshina, Yan Ye 0003, Jie Chen 0006, Jiancong Luo, Chun-Chi Chen, Han Huang 0001, Wei-Jung Chien, Vadim Seregin, Marta Karczewicz |
DCC | 5 |
| 2020 | A Triangulation-Based Backward Adaptive Motion Field Subsampling SchemeabstractOptical flow procedures are used to generate dense motion fields which approximate true motion. Such fields contain a large amount of data and if we need to transmit such a field, the raw data usually exceeds the raw data of the two images it was computed from. In many scenarios, however, it is of interest to transmit a dense motion field efficiently. Most prominently this is the case in inter prediction for video coding. In this paper we propose a transmission scheme based on subsampling the motion field. Since a field which was subsampled with a regularly spaced pattern usually yields suboptimal results, we propose an adaptive subsampling algorithm that preferably samples vectors at positions where changes in motion occur. The subsampling pattern is fully reconstructable without the need for signaling of position information. We show an average gain of 2.95 dB in average end point error compared to regular subsampling. Furthermore we show that an additional prediction stage can improve the results by an additional 0.43 dB, gaining 3.38 dB in total. Fabian Brand, Jürgen Seiler, Elena Alshina, André Kaup |
MMSP | 3 |
| 2019 | Low-Complexity Geometric Inter-Prediction for Versatile Video CodingabstractNon-rectangular block partitioning is a well-known method for improved inter-picture prediction in video coding, enabling better spatial adaptation to the signal properties. This contribution presents the most recent proposal of geometric inter-prediction (GIP) made to the Versatile Video Coding (VVC) standardization activity led by the Joint Video Experts Team (JVET). Implemented in the latest test model VTM-5.0 and evaluated according to the JVET Common Test Conditions, the proposed low-complexity GIP scheme provides objective luma BD-rate reductions of 0.22 % for random access and 0.44 % for low-delay test cases at 7% encoder runtime increase and negligible decoder runtime increase. The coding gain is provided by non-triangular partitioned blocks and in the presence of multiple other VVC coding tools. Furthermore, BD-rate reductions of 2.58 % and 2.78 % can be achieved specifically for pure screen content by employing an adaptive blending filter. Max Bläser, Han Gao 0001, Semih Esenlik, Elena Alshina, Zhijie Zhao, Christian Rohlfing, Eckehard G. Steinbach |
PCS | 4 |
| 2017 | Omnidirectional Video Quality Metrics and Evaluation ProcessabstractWidespread of virtual reality technologies across entertainment formats has created a diverse infrastructure of related technologies as head-mount displays, dome screens and virtual reality multi-camera platforms. As omnidirectional content is processing pipeline is completely different form conventional planar video and involves multiple conversion steps which affect quality in a different way. As a result of our research we propose objective quality estimation methodology and a set of tools to evaluate different projection methods and coding tools for omnidirectional video content. Vladyslav Zakharchenko, Kwangpyo Choi, Elena Alshina, Jeonghoon Park |
DCC | 3 |
| 2016 | Bi-directional Pptical Flow for Future Video CodecabstractPaper presents theoretical explanation for bi-directional optical flow technique in generic case. Both non-equal distance to reference frames and two reference frames from the same side of predicted frame are allowed. Dynamic range analysis during bi-directional optical flow calculations is provided. Limits for refinement motion vector are recommended. Alexander Alshin, Elena Alshina |
DCC | 2 |
| 2016 | BIO performance complexity trade-offabstractBi-directional optical flow (so-called BIO) is part of Joint Exploration Model (JEM) which explores potential coding efficiency improvement over state-of-the-art video codec. BIO allows fine motion compensation on a sample level without additional signaling, since refinement is explicitly calculated using just texture information from both reference frames under assumption the validity of optical flow equation. BIO reduces BD-rate in average by more than 2% (up to 5% for some test video), but computational complexity is rather high. Two simplifications for BIO are studied in this paper. First is redesign chain for MC prediction and gradients calculation scheme. Simplified scheme has slightly high latency but reduces amount of multiplications in bi-predicted blocks by factor 2. Another simplification is clustering samples in order to perform motion refinement in BIO not per sample but for group of samples. This allows reduction of division operation in BIO by factor 9.7 in average (up to 256 times in largest blocks). Both modifications enabled together maintain the same performance for BIO in JEM while reduce encoding and decoding run-time significantly. Alexander Alshin, Elena Alshina |
PCS | 2 |
| 2015 | Resampling Process of the Scalable High Efficiency Video CodingabstractSHVC is the scalable extension of the latest video coding standard High Efficiency Video Coding (HEVC) and spatial resampling process is inevitable module to support spatial scalability. This paper describes in details the resampling process, including both texture and motion data resampling in SHVC, and using experimental evidence, demonstrate their benefits in terms of coding efficiency. Jianle Chen, Elena Alshina, Xiang Li 0003, Marta Karczewicz, Alexander Alshin |
DCC | 2 |
| 2014 | Region based inter-layer cross-color filtering for scalable extension of HEVCabstractInter-layer filtering is a key module of the emerging Scalable Extension of High Efficiency Video Coding Standard (SHVC). In SHVC, up-sampled based layer reconstructed pictures are used as inter-layer references to predict enhancement layer frames such that inter-layer redundancy is reduced. To improve the coding performance of inter-layer filtering, luma plane based chroma plane enhancement was proposed at picture level. However, the efficiency of the picture level adaptation is not very promising when picture resolution is high. To address this issue, region based inter-layer cross-color filtering is proposed in this paper. Simulations under the common test conditions defined by Joint Collaborative Team on Video Coding (JCT-VC) showed that significant chroma coding gain and moderate luma improvement were achieved by the proposed method. When compared to the luma plane based chroma plane enhancement method, the coding gain over SHVC reference software SHM-2.0 is about doubled while the decoding complexity is kept even lower. Moreover, the proposed method outperforms other tools studied in SHVC core experiment on inter-layer filtering. Xiang Li 0003, Jianle Chen, Marta Karczewicz, Elena Alshina, Alexander Alshin, Yongjin Cho |
ICIP | 5 |
| 2014 | Sample adaptive offset in AVS2 video standardabstractAVS2 video standard is the next-generation video coding standard under the development of Audio Video coding Standard (AVS) workgroup of China. In this paper, the design of Sample Adaptive Offset (SAO) in AVS2 is presented. Considering the implementation issues, a shifted structure in which the SAO parameter region is shifted from the Largest Coding Unit (LCU) to the upper-left is adopted to make the SAO parameter region consistent with the processing region in implementation. Moreover, the category dependent offset is introduced in the edge type based on the statistical results to improve the offset coding and non-consecutive offset bands are adopted in the band type to optimize offset bands. The test results show that SAO achieves on average 0.3% to 1.4% luma coding gain in AVS2 common test conditions. Sunil Lee, Elena Alshina, Yinji Piao |
VCIP | 3 |
| 2013 | Sample Adaptive Offset Design in HEVCabstractThis paper is devoted to Sample Adaptive Offset (SAO). This technique was recently added into High Efficiency Video Coding (HEVC) standard. The concept of SAO is to reduce sample distortion of a region by classifying the region samples into multiple categories, obtaining an offset for each category, and then adding the offset to each sample, where the classifier index and the offsets are coded in the bit stream. Alexander Alshin, Elena Alshina, Jeong-Hoon Park |
DCC | 2 |
| 2013 | Interpolation filter design in HEVC and its coding efficiency - complexity analysisabstractCoding efficiency gains in the High Efficiency Video Coding (H.265/HEVC) standard are achieved by improving many aspects of the traditional hybrid coding framework. Motion compensated prediction, and in particular the interpolation filter, is one of the areas that was improved significantly over H.264/AVC. This paper presents the details of the motion compensation interpolation filter design of the H.265/HEVC standard and its improvements over the interpolation filter design of H.264/AVC. These improvements include discrete cosine transform based filter coefficient design, utilizing longer filter taps for luma and chroma interpolation and using higher precision operations in the intermediate computations. The computational complexity of HEVC interpolation filter is also analyzed both from theoretical and practical perspectives. Experimental results show that a 4.5% average bitrate reduction for the luma component and 13.0% average bitrate reduction for the chroma components are achieved compared to interpolation filter of H.264/AVC. The coding efficiency gains are significant for some video sequences and can reach up to 21.7%. Kemal Ugur, Alexander Alshin, Elena Alshina, Frank Bossen, Woojin Han 0001, Jeong-Hoon Park, Jani Lainema |
ICASSP | 3 |
| 2013 | Inter-layer filtering for scalable extension of HEVCabstractThis paper introduces inter-layer filters for the scalable extension of High Efficiency Video Coding (SHVC) standard, which is being developed by the Joint Collaborative Team on Video Coding (JCT-VC). The major new coding tool in SHVC is inter-layer texture prediction. It provides about 18% average BD-rate reduction compared with HEVC two-layer simulcast. In the case of spatial scalability, base layer reconstructed pictures are up-sampled to the enhancement layer resolution to generate inter-layer texture prediction. A set of 2D separable 8 taps (luma) and 4 taps (chroma) DCT based interpolation filters, which follow the design principles of HEVC motion compensation interpolation filter, are used in the up-sampling process. In the case of SNR scalability, the up-sampling process is not needed since the reference layer has the same spatial resolution as the current layer but encoded with lower quality. This paper proposes a novel inter-layer filter with denoising effect for SNR scalability to improve enhancement layer coding efficiency and equalize the number of stages in inter-layer processing between SNR and spatial scalabilities. Experimental results show that the usage of inter-layer de-noising filter in SNR scalability provides up to 7.5% BD-rate reduction and has observable improvement on subjective visual quality. Elena Alshina, Alexander Alshin, Yongjin Cho, Jeong-Hoon Park, Jianle Chen, Xiang Li 0003, Vadim Seregin, Marta Karczewicz |
PCS | 1 |
| 2013 | High precision probability estimation for CABACabstractEntropy coding is the main important part of all advanced video compression schemes. Context-adaptive binary arithmetic coding (CABAC) is entropy coding used in H.264/MPEG-4 AVC and H.265/HEVC standards. Probability estimation is the key factor of CABAC performance efficiency. In this paper high accuracy probability estimation for CABAC is presented. This technique is based on multiple estimations using different models. Proposed method was efficiently realized in integer arithmetic. High precision probability estimation for CABAC provides up-to 1,4% BD-rate gain. Alexander Alshin, Elena Alshina, Jeong-Hoon Park |
VCIP | 2 |
| 2012 | Sample Adaptive Offset in the HEVC StandardabstractThis paper provides a technical overview of a newly added in-loop filtering technique, sample adaptive offset (SAO), in High Efficiency Video Coding (HEVC). The key idea of SAO is to reduce sample distortion by first classifying reconstructed samples into different categories, obtaining an offset for each category, and then adding the offset to each sample of the category. The offset of each category is properly calculated at the encoder and explicitly signaled to the decoder for reducing sample distortion effectively, while the classification of each sample is performed at both the encoder and the decoder for saving side information significantly. To achieve low latency of only one coding tree unit (CTU), a CTU-based syntax design is specified to adapt SAO parameters for each CTU. A CTU-based optimization algorithm can be used to derive SAO parameters of each CTU, and the SAO parameters of the CTU are inter leaved into the slice data. It is reported that SAO achieves on average 3.5% BD-rate reduction and up to 23.5% BD-rate reduction with less than 1% encoding time increase and about 2.5% decoding time increase under common test conditions of HEVC reference software version 8.0. Chih-Ming Fu, Elena Alshina, Alexander Alshin, Yu-Wen Huang, Ching-Yeh Chen, Chia-Yang Tsai, Shawmin Lei, Jeong-Hoon Park, Woojin Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Rotational transform for image and video compressionabstractTo improve video coding efficiency, the Rotational Transform (ROT) was proposed for adaptive switching between different transforms cores. The Karhunen Loeve Transform (KLT) is known to be optimal for given residual but requires much side information to be signaled to the decoder. The Discrete Cosine transform (DCT) is known to be close to optimal but for strongly directional components, it is sub-optimal. The main idea of ROT is that small modification of DCT coefficients can improve energy compaction. The ROT is implemented as a secondary transform applied after the primary DCT. The ROT matrix is sparse and thus enjoys relatively small computational complexity and memory usage increment. The encoder tries every rotational transform from the dictionary. Only one number, the ROT index, needs to be signaled to the decoder. Because the ROT is an orthogonal transform, encoder search is greatly simplified: distortion can be estimated in frequency domain and no inverse transformation is needed. This makes the ROT an efficient way to improve image/video compression. The ROT Coding gain for Intra slice is 2-3% in the HM 1.0 software implementation. Elena Alshina, Alexander Alshin, Felix C. A. Fernandes |
ICIP | 1 |
| 2010 | Bi-directional optical flow for improving motion compensationabstractNew method improving B-slice prediction is proposed. By combining the optical flow concept and high accuracy gradients evaluation we construct the algorithm which allows pixel-wise refinement of motion. This approach does not require any signaling for decoder. According to tests with WQVGA sequences bit-saving of 2%-6% can be achieved using this tool. Alexander Alshin, Elena Alshina, Tammy Lee |
PCS | 2 |
| 2010 | Improved Video Compression Efficiency Through Flexible Unit Representation and Corresponding Extension of Coding ToolsabstractThis paper proposes a novel video compression scheme based on a highly flexible hierarchy of unit representation which includes three block concepts: coding unit (CU), prediction unit (PU), and transform unit (TU). This separation of the block structure into three different concepts allows each to be optimized according to its role; the CU is a macroblock-like unit which supports region splitting in a manner similar to a conventional quadtree, the PU supports nonsquare motion partition shapes for motion compensation, while the TU allows the transform size to be defined independently from the PU. Several other coding tools are extended to arbitrary unit size to maintain consistency with the proposed design, e.g., transform size is extended up to 64 × 64 and intraprediction is designed to support an arbitrary number of angles for variable block sizes. Other novel techniques such as a new noncascading interpolation Alter design allowing arbitrary motion accuracy and a leaky prediction technique using both open-loop and closed-loop predictors are also introduced. The video codec described in this paper was a candidate in the competitive phase of the high-efficiency video coding (HEVC) standardization work. Compared to H.264/AVC, it demonstrated bit rate reductions of around 40% based on objective measures and around 60% based on subjective testing with 1080 p sequences. It has been partially adopted into the first standardization model of the collaborative phase of the HEVC effort. Woojin Han 0001, Junghye Min, Il-Koo Kim, Elena Alshina, Alexander Alshin, Tammy Lee, Jianle Chen, Vadim Seregin, Sunil Lee, Yoon Mi Hong, Min-Su Cheon, Nikolay Shlyakhov, Ken McCann, Thomas Davies 0002, Jeong-Hoon Park |
IEEE Trans. Circuits Syst. Video Technol. | 4 |