Atanas Boev

dblp:36/10699 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0003-0863-4000ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 On the Suitability of Perceptual Quality Metrics for Learning-Based Screen Content Compression
abstract
Learned image compression methods tailored for screen content demand perceptual quality metrics that are both accurate and computationally efficient. Traditional metrics such as PSNR and SSIM often fail to capture perceptual distortions specific to screen content, while many advanced learned perceptual metrics are too complex or non-differentiable for practical use with implicit neural codecs. These codecs require loss functions to be evaluated tens of thousands of times during training, imposing a strict complexity limit of a few hundred multiply-accumulate operations per pixel. In this paper, we evaluate several differentiable, low-complexity perceptual metrics on three screen content quality datasets, including a newly collected compression-focused crowdsourced dataset, PerceptualSCC. Our results identify VIF and NLPD as the best-performing metrics; however, only NLPD demonstrates stable behavior in gradient-based optimization experiments. These findings suggest NLPD as a strong candidate for guiding learned screen content compression, balancing perceptual fidelity, computational cost, and optimization stability.
H. Burak Dogaroglu, Hongjie You, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
ISM3
2025 Fine-Grained HDR Image Quality Assessment From Noticeably Distorted to Very High Fidelity
abstract
High dynamic range (HDR) and wide color gamut (WCG) technologies significantly improve color reproduction compared to standard dynamic range (SDR) and standard color gamuts, resulting in more accurate, richer, and more immersive images. However, HDR increases data demands, posing challenges for bandwidth efficiency and compression techniques. Advances in compression and display technologies require more precise image quality assessment, particularly in the high-fidelity range where perceptual differences are subtle. To address this gap, we introduce AIC-HDR2025, the first such HDR dataset, comprising 100 test images generated from five HDR sources, each compressed using four codecs at five compression levels. It covers the high-fidelity range, from visible distortions to compression levels below the visually lossless threshold. A subjective study was conducted using the JPEG AIC-3 test methodology, combining plain and boosted triplet comparisons. In total, 34,560 ratings were collected from 151 participants across four fully controlled labs. The results confirm that AIC-3 enables precise HDR quality estimation, with 95% confidence intervals averaging a width of 0.27 at 1 JND. In addition, several recently proposed objective metrics were evaluated based on their correlation with subjective ratings. The dataset is publicly available1.
Mohsen Jenadeleh, Jon Sneyers, Davi Lazzarotto, Shima Mohammadi, Dominik Keller, Atanas Boev, Rakesh Rao Ramachandra Rao, António M. G. Pinheiro, Thomas Richter 0005, Alexander Raake, Touradj Ebrahimi, João Ascenso, Dietmar Saupe
QoMEX6
2024 Adapting Learned Image Codecs To Screen Content Via Adjustable Transformations
abstract
As learned image codecs (LICs) become more prevalent, their low coding efficiency for out-of-distribution data becomes a bottleneck for some applications. To improve the performance of LICs for screen content (SC) images without breaking backwards compatibility, we propose to introduce parameterized and invertible linear transformations into the coding pipeline without changing the underlying baseline codec’s operation flow. We design two neural networks to act as prefilters and postfilters in our setup to increase the coding efficiency and help with the recovery from coding artifacts. Our end-to-end trained solution achieves up to $10 \%$ bitrate savings on SC compression compared to the baseline LICs while introducing only $1 \%$ extra parameters.
H. Burak Dogaroglu, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
ICIP3
2024 Judder Modelling Framework with Perceptual Quality Score Prediction for HDR Videos
abstract
Judder is a motion artifact describing a perceptual mismatch between human visual system and discrete movements on a display. Judder is correlated with video frame rates, motion speeds and brightness. It significantly degrades the perceived video quality. We present a framework for modeling judder, where the input frames are processed to generate optical flow and a sensitivity map. Concurrently, a specifically designed attention map is integrated into the process. These components collectively contribute to the creation of a feature map, which is utilized to compute judder scores. The proposed approach explicitly takes into account the video frame rate, resulting in an accurate judder score to predict the subjective mean opinion scores ultimately. Our assessment of the proposed framework on an HDR video dataset shows that judderness is highly influenced by the frame rate and, to some extent, by motion speed and brightness. The prediction of perceived quality scores shows improvement compared to the baseline framework.
Hongjie You, Nicola Giuliani, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
VCIP4
2024 Efficient Contextformer: Spatio-Channel Window Attention for Fast Context Modeling in Learned Image Compression
abstract
Entropy estimation is essential for the performance of learned image compression. It has been demonstrated that a transformer-based entropy model is of critical importance for achieving a high compression ratio, however, at the expense of a significant computational effort. In this work, we introduce the Efficient Contextformer (eContextformer) – a computationally efficient transformer-based autoregressive context model for learned image compression. The eContextformer efficiently fuses the patch-wise, checkered, and channel-wise grouping techniques for parallel context modeling, and introduces a shifted window spatio-channel attention mechanism. We explore better training strategies and architectural designs and introduce additional complexity optimizations. During decoding, the proposed optimization techniques dynamically scale the attention span and cache the previous attention computations, drastically reducing the model and runtime complexity. Compared to the non-parallel approach, our proposal has ~145x lower model complexity and ~210x faster decoding speed, and achieves higher average bit savings on Kodak, CLIC2020, and Tecnick datasets. Additionally, the low complexity of our context model enables online rate-distortion algorithms, which further improve the compression performance. We achieve up to 17% bitrate savings over the intra coding of Versatile Video Coding (VVC) Test Model (VTM) 16.2 and surpass various learning-based compression models.
Ahmet Burakhan Koyuncu, Panqi Jia, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.3
2023 CALC-VFS: Content-adaptive low-complexity Video Frame Synthesis
abstract
We present a content-adaptive, low-complexity video frame synthesis algorithm. Our approach applies the dynamic convolutions content adaptation approach to the widely used frame synthesis algorithm IFRNet. By introducing dynamic convolutions into both the pyramid encoder and the coarse-to-fine decoders of IFRNet, we enforce sparsity, thereby limiting the computationally expensive operations to only the necessary pixels. Training for specific sparsity targets allows us to achieve overall less computational complexity compared to IFRNet while having similar performance. We demonstrate the performance and content adaptivity in two test scenarios and show the savings in computational budget (approximately 20-40%) compared to the baseline IFRNet.
Nicola Giuliani, Hongjie You, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
ISM4
2023 Subjective video quality assessment of immersive HDR content on head-mounted displays
abstract
High dynamic range (HDR) videos are known to provide better visual quality on HDR TV displays. Head-mounted displays (HMDs) are an integral part of immersive visual experiences. However, typical HMDs are equipped with standard dynamic range (SDR) displays, failing to show details in bright and dark areas of HDR content. Therefore, we aim to evaluate the perceptual quality improvement, in terms of mean opinion score (MOS), when observers view HDR instead of SDR content on immersive displays. We developed a pipeline to render 2D HDR and tone-mapped SDR immersive videos for HMDs. We conducted two single-stimulus subjective evaluation experiments to evaluate (1) the perceived visual difference when comparing one HDR scene with three tone-mapped SDR versions of it, and (2) how frame rates impact the perceptual quality of HDR immersive videos. Our results in MOS show that (1) there is a significant improvement of perceptual quality in HDR compared to tone-mapped SDR content, and (2) HDR immersive videos benefit much more from higher frame rates than SDR videos.
Hongjie You, Nicola Giuliani, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
VCIP4
2022 Contextformer: A Transformer with Spatio-Channel Attention for Context Modeling in Learned Image Compression
Ahmet Burakhan Koyuncu, Han Gao 0001, Atanas Boev, Georgii Gaikov, Elena Alshina, Eckehard G. Steinbach
ECCV (19)3
2021 Quality-Blind Compressed Color Image Enhancement with Convolutional Neural Networks
abstract
Lossy compressed images and videos suffer from visible compression artifacts, especially when the bit-rate is low. To improve the quality of the compressed image while keeping the same bit-rate, decoder-side compression artifacts reduction (CAR) becomes important. Recently, convolutional neural networks are adopted for CAR tasks and achieve the state-of-the-art performance. However, most CAR algorithms only focus on the reconstruction of the luminance channel. Also, a separate model usually needs to be trained for each quality factor (QF), which makes these approaches not practical in existing codecs. In this paper, we analyze a quality-blind training strategy and compare it with training separate models for each QF. The testing results with three representative CAR algorithms show the superiority of the quality-blind training compared to separate training. The results for pseudo and real quality-blind CAR tests further prove the generalizability of the quality-blind training for practical CAR tasks.
Kai Cui 0003, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
ISCAS3
2021 Convolutional neural network-based post-filtering for compressed YUV420 images and video
abstract
Images and videos compressed with lossy compression algorithms usually suffer from visible distortions, especially when the bitrate is low. To improve the quality without spending extra bitrate, many image and video codecs have built-in filters to mitigate these artifacts. However, most of them are only applied on the luminance channel, while the chrominance channels remain unmodified. While this is partly justified by the observation that the luminance channel usually contains more details and has higher-resolution than the chrominance channels. We observe that the luminance and chrominance channels still have latent correlations. Therefore, the post-filtering of the chrominance channels is also beneficial and can be driven by the information from the luminance channel. In this paper, we propose a 3-stage YUV post-filtering network for compressed YUV420 images and video. The proposed 3-stage structure not only improves the quality of the luminance channel, but also exploits the luma-chroma correlations to improve the quality of the chrominance channels. Our experimental results show that the proposed approach achieves 3.60%/12.75%/14.93% Bj⊘ntegaard Delta bitrate improvement for the Y, U and V channels over the VVC 10.0 codec for All-Intra configuration.
Kai Cui 0003, Ahmet Burakhan Koyuncu, Atanas Boev, Elena Alshina, Eckehard G. Steinbach
PCS3
2021 Parallelized Context Modeling for Faster Image Coding
abstract
Learning-based image compression has reached the performance of classical methods such as BPG. One common approach is to use an autoencoder network to map the pixel information to a latent space and then approximate the symbol probabilities in that space with a context model. During inference, the learned context model provides symbol probabilities, which are used by the entropy encoder to obtain the bitstream. Currently, the most effective context models use autoregression, but autoregression results in a very high decoding complexity due to the serialized data processing. In this work, we propose a method to parallelize the autoregressive process used for image compression. In our experiments, we achieve a decoding speed that is over 8 times faster than the standard autoregressive context model almost without compression performance reduction.
Ahmet Burakhan Koyuncu, Kai Cui 0003, Atanas Boev, Eckehard G. Steinbach
VCIP3
2014 Measurement of perceived spatial resolution in 3D light-field displays
abstract
Effective spatial resolution of projection-based 3D light-field (LF) displays is an important quantity, which is informative about the capabilities of the display to recreate views in space and is important for content creation. We propose a subjective experiment to measure the spatial resolution of LF displays and compare it to our objective measurement technique. The subjective experiment determines the limit of visibility on the screen as perceived by viewers. The test involves subjects determining the direction of patterns that resemble tumbling E eye test charts. These results are checked against the LF display resolution determined by objective means. The objective measurement models the display as a signal-processing channel. It characterizes the display throughput in terms of passband, quantified by spatial resolution measurements in multiple directions. We also explore the effect of viewing angle and motion parallax on the spatial resolution.
Péter Tamás Kovács, Kristóf Lackner, Attila Barsi, Ákos Balázs, Atanas Boev, Robert Bregovic, Atanas P. Gotchev
ICIP5
2011 3D-DCT based perceptual quality assessment of stereo video
abstract
In this paper, we present a novel stereoscopic video quality assessment method based on 3D-DCT transform. In our approach, similar blocks from left and right views of stereoscopic video frames are found by block-matching, grouped into 3D stack and then analyzed by 3D-DCT. Comparison between reference and distorted images are made in terms of MSE calculated within the 3D-DCT domain and modified to reflect the contrast sensitive function and luminance masking. We validate our quality assessment method using test videos annotated with results from subjective tests. The results show that the proposed algorithm outperforms current popular metrics over a wide range of distortion levels.
Lina Jin, Atanas Boev, Atanas P. Gotchev, Karen Egiazarian
ICIP2
2011 Three-Dimensional Media for Mobile Devices
abstract
This paper aims at providing an overview of the core technologies enabling the delivery of 3-D Media to next-generation mobile devices. To succeed in the design of the corresponding system, a profound knowledge about the human visual system and the visual cues that form the perception of depth, combined with understanding of the user requirements for designing user experience for mobile 3-D media, are required. These aspects are addressed first and related with the critical parts of the generic system within a novel user-centered research framework. Next-generation mobile devices are characterized through their portable 3-D displays, as those are considered critical for enabling a genuine 3-D experience on mobiles. Quality of 3-D content is emphasized as the most important factor for the adoption of the new technology. Quality is characterized through the most typical, 3-D-specific visual artifacts on portable 3-D displays and through subjective tests addressing the acceptance and satisfaction of different 3-D video representation, coding, and transmission methods. An emphasis is put on 3-D video broadcast over digital video broadcasting-handheld (DVB-H) in order to illustrate the importance of the joint source-channel optimization of 3-D video for its efficient compression and robust transmission over error-prone channels. The comparative results obtained identify the best coding and transmission approaches and enlighten the interaction between video quality and depth perception along with the influence of the context of media use. Finally, the paper speculates on the role and place of 3-D multimedia mobile devices in the future internet continuum involving the users in cocreation and refining of rich 3-D media content.
Atanas P. Gotchev, Gozde Bozdagi Akar, Tolga K. Çapin, Dominik Strohmeier, Atanas Boev
Proc. IEEE5