Cosmin Stejerean

dblp:314/7052 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2026
0009-0006-3523-2479ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021
YearPublicationVenuePosition
2026 Joint Quality Assessment and Example-Guided Tone Mapping by Disentangling Picture Appearance From Content
abstract
The deep learning revolution has strongly impacted low-level image processing tasks such as style/domain transfer, enhancement/restoration, and visual quality assessments. Despite often being treated separately, the aforementioned tasks share a common theme of understanding, editing, or enhancing the appearance of input images without modifying the underlying content. We leverage this observation to develop a novel disentangled representation learning method that decomposes inputs into content and appearance features. The model is trained in a self-supervised manner and we use the learned features to develop a new quality prediction model named DisQUE. We demonstrate through extensive evaluations that DisQUE achieves state-of-the-art accuracy across quality prediction tasks and distortion types. Moreover, we demonstrate that the same features may also be used for image processing tasks such as HDR tone mapping, where the desired output characteristics may be tuned using example input-output pairs.
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Hassene Tmar, Alan C. Bovik
IEEE Trans. Image Process.2
2025 SDART: Spatial Dart AR Simulation with Hand-Tracked Input
abstract
We present a physics-driven 3D dart-throwing interaction system for Apple Vision Pro (AVP), developed using Unity 6 engine and running in augmented reality (AR) mode on the device. The system utilizes the PolySpatial and Apple's ARKit software development kits (SDKs) to ensure hand input and tracking in order to intuitively spawn, grab, and throw virtual darts similar to real darts. The application benefits from physics simulations alongside the innovative no-controller input system of AVP to manipulate objects realistically in an unbounded spatial volume. By implementing spatial distance measurement, scoring logic, and recording user performance, this project enables user studies on quality of experience in interactive experiences. To evaluate the perceived quality and realism of the interaction, we conducted a subjective study with 10 participants using a structured questionnaire. The study measured various aspects of the user experience, including visual and spatial realism, control fidelity, depth perception, immersiveness, and enjoyment. Results indicate high mean opinion scores (MOS) across key dimensions.
Milad Ghanbari, Wei Zhou 0021, Cosmin Stejerean, Christian Timmerer, Hadi Amirpour
ACM Multimedia3
2025 STACK: Spatial Tower Assembly using Controlled Kinetics
abstract
This paper presents a block stacking simulation developed for Apple Vision Pro (AVP) using Unity’s PolySpatial framework, designed to study both depth perception in spatial computing and physics comprehension of user-driven kinetic controls in augmented reality (AR). The simulation offers two interactive modes: a tower assembly mode and a removal mode. Each game session includes four stages with the virtual table positioned at various distances to observe user adaptation across varying virtual depths. User input is captured through eye tracking and hand tracking, and block behavior is handled by real-time physics simulation, which includes collision response, gravity, and mass-based interactions. The system supports two physics configurations: raw Unity physics and a modified variant with adjusted material and rigidbody parameters for improved stability and realism. It utilizes spatial computing features such as world anchoring to preserve spatial consistency and depth perception through stereoscopic rendering and dynamic shadows, so that users can better judge the spatial coordinates between virtual blocks and their physical surroundings. The simulation is intended to evaluate how 3D spatial rendering and physically realistic interactions contribute to immersion and task performance in AR environments. To assess user performance, the system records key interaction metrics to support analysis of learning progression, control accuracy, and adaptability across varying distances and physics configurations. This work contributes to the understanding of spatial and physics-based interaction design in AR and may inform future applications in education, simulation, and spatial gaming.
Milad Ghanbari, Hadi Amirpour, Christian Timmerer, Mohammad Hossein Izadimehr, Wei Zhou 0021, Cosmin Stejerean
VCIP6
2025 Cut-FUNQUE: An objective quality model for compressed tone-mapped High Dynamic Range videos
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Hassene Tmar, Alan C. Bovik
Signal Process. Image Commun.2
2024 Bitrate Ladder Construction Using Visual Information Fidelity
abstract
Recently proposed perceptually optimized per-title video encoding methods provide better BD-rate savings than fixed bitrate-ladder approaches that have been employed in the past. However, a disadvantage of per-title encoding is that it requires significant time and energy to compute bitrate ladders. Over the past few years, a variety of methods have been proposed to construct optimal bitrate ladders including using low-level features to predict cross-over bitrates, optimal resolutions for each bitrate, predicting visual quality, etc. Here, we deploy features drawn from Visual Information Fidelity (VIF) (VIF features) extracted from uncompressed videos to predict the visual quality (VMAF) of compressed videos. We present multiple VIF feature sets extracted from different scales and subbands of a video to tackle the problem of bitrate ladder construction. Comparisons are made against a fixed bitrate ladder and a bitrate ladder obtained from exhaustive encoding using Bjontegaard delta metrics.
Krishna Srikar Durbha, Hassene Tmar, Cosmin Stejerean, Ioannis Katsavounidis, Alan C. Bovik
PCS3
2024 A FUNQUE Approach to the Quality Assessment of Compressed HDR Videos
abstract
Recent years have seen steady growth in the popularity and availability of High Dynamic Range (HDR) content, particularly videos, streamed over the internet. As a result, assessing the subjective quality of HDR videos, which are generally subjected to compression, is of increasing importance. In particular, we target the task of full-reference quality assessment of compressed HDR videos. The state-of-the-art (SOTA) approach HDRMAX involves augmenting off-the-shelf video quality models, such as VMAF, with features computed on nonlinearly transformed video frames. However, HDRMAX increases the computational complexity of models like VMAF. Here, we show that an efficient class of video quality prediction models named FUNQUE+ achieves SOTA accuracy. This shows that the FUNQUE+ models are flexible alternatives to VMAF that achieve higher HDR video quality prediction accuracy at lower computational cost.
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Alan C. Bovik
PCS2
2024 Energy-Efficient Video Streaming: A Study on Bit Depth and Color Subsampling
abstract
As video dimensions – including resolution, frame rate, and bit depth – increase, a larger bitrate is required to maintain a higher Quality of Experience (QoE). While videos are often optimized for resolution and frame rate to improve compression and energy efficiency, the impact of color space is often overlooked. Larger color spaces are essential for avoiding color banding and delivering High Dynamic Range (HDR) content with richer, more accurate colors, although this comes at the cost of higher processing energy. This paper investigates the effects of bit depth and color subsampling on video compression efficiency and energy consumption. By analyzing different bit depths and subsampling schemes, we aim to determine optimized settings that balance compression efficiency with energy consumption, ultimately contributing to more sustainable and high-quality video delivery. We evaluate both encoding and decoding energy consumption and assess the quality of videos using various metrics including PSNR, VMAF, ColorVideoVDP, and CAMBI. Our findings offer valuable insights for video codec developers and content providers aiming to improve the performance and environmental footprint of their video streaming services.
Hadi Amirpour, Lingfeng Qu, Jong Hwan Ko, Cosmin Stejerean, Christian Timmerer
VCIP4
2024 One Transform to Compute Them All: Efficient Fusion-Based Full-Reference Video Quality Assessment
abstract
The Visual Multimethod Assessment Fusion (VMAF) algorithm has recently emerged as a state-of-the-art approach to video quality prediction, that now pervades the streaming and social media industry. However, since VMAF requires the evaluation of a heterogeneous set of quality models, it is computationally expensive. Given other advances in hardware-accelerated encoding, quality assessment is emerging as a significant bottleneck in video compression pipelines. Towards alleviating this burden, we propose a novel Fusion of Unified Quality Evaluators (FUNQUE) framework, by enabling computation sharing and by using a transform that is sensitive to visual perception to boost accuracy. Further, we expand the FUNQUE framework to define a collection of improved low-complexity fused-feature models that advance the state-of-the-art of video quality performance with respect to both accuracy, by 4.2% to 5.3%, and computational efficiency, by factors of 3.8 to 11 times!.
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Ioannis Katsavounidis, Alan C. Bovik
IEEE Trans. Image Process.2
2023 Estimating Uncertainty On Video Quality Metrics
abstract
Video Quality Metrics (VQM) are models used to predict the score that a user would give to the quality of a video visualization. They are widely used in video processing systems, for monitoring end-to-end quality or system troubleshooting for example. In these scenarios, the improvement is quantified based on a certain enhancement of a VQM score and trouble-detection is done based on a certain drop or threshold computed based on a VQM. Yet, whether such improvement or fault-detection is worth a significant increase in power consumption is questionable. Therefore, the goal of this work is to propose a method to predict the uncertainty of the quality metric. In this paper, we propose a framework to evaluate the confidence interval of a VQM for a given content using simple video features. We assess the performance of the framework by using the confidence intervals to predict if two videos are of similar or different quality and show that in most cases our approach performs better than just using a constant confidence interval.
Patrick Le Callet, Suiyi Ling, Haixiong Wang, Ioannis Katsavounidis, Zafar Shahid, Cosmin Stejerean
ICASSP7
2022 Funque: Fusion of Unified Quality Evaluators
abstract
Fusion-based quality assessment has emerged as a powerful method for developing high-performance quality models from quality models that individually achieve lower performances. A prominent example of such an algorithm is VMAF, which has been widely adopted as an industry standard for video quality prediction along with SSIM. In addition to advancing the state-of-the-art, it is imperative to alleviate the computational burden presented by the use of a heterogeneous set of quality models. In this paper, we unify "atom" quality models by computing them on a common transform domain that accounts for the Human Visual System, and we propose FUNQUE, a quality model that fuses unified quality evaluators. We demonstrate that in comparison to the state-of-the-art, FUNQUE offers significant improvements in both correlation against subjective scores and efficiency, due to computation sharing.
Abhinau Kumar Venkataramanan, Cosmin Stejerean, Alan C. Bovik
ICIP2
2022 Advances in Quality Assessment Of Video Streaming Systems: Algorithms, Methods, Tools
abstract
Quality assessment of video has matured significantly in the last 10 years due to a flurry of relevant developments in academia and industry, with relevant initiatives in VQEG, AOMedia, MPEG, ITU-T P.910, and other standardization and advisory bodies . Most advanced video streaming systems are now clearly moving away from good old-fashioned' PSNR and structural similarity type of assessment towards metrics that align better to mean opinion scores from viewers. Several of these algorithms, methods and tools have only been developed in the last 3-5 years and, while they are of significant interest to the research community, their advantages and limitations are not widely known in the research community. This tutorial provides this overview, but also focuses on practical aspects and how to design quality assessment tests that can scale to large datasets.
Yiannis Andreopoulos, Cosmin Stejerean
ACM Multimedia2
2022 Domain-Specific Fusion Of Objective Video Quality Metrics
abstract
Video processing algorithms like video upscaling, denoising, and compression are now increasingly optimized for perceptual quality metrics instead of signal distortion. This means that they may score well for metrics like video multi-method assessment fusion (VMAF), but this may be because of metric overfitting. This imposes the need for costly subjective quality assessments that cannot scale to large datasets and large parameter explorations. We propose a methodology that fuses multiple quality metrics based on small scale subjective testing in order to unlock their use at scale for specific application domains of interest. This is achieved by employing pseudo-random sampling of the resolution, quality range and test video content available, which is initially guided by quality metrics in order to cover the quality range useful to each application. The selected samples then undergo a subjective test, such as ITU-T P.910 absolute categorical rating, with the results of the test postprocessed and used as the means to derive the best combination of multiple objective metrics using support vector regression. We showcase the benefits of this approach in two applications: video encoding with and without perceptual preprocessing, and deep video denoising & upscaling of compressed content. For both applications, the derived fusion of metrics allows for a more robust alignment to mean opinion scores than a perceptually-uninformed combination of the original metrics themselves. The dataset and code is available at https://github.com/isize-tech/VideoQualityFusion.
Aaron Chadha, Ioannis Katsavounidis, Ayan Kumar Bhunia, Cosmin Stejerean, Muhammad Umar Karim Khan, Yiannis Andreopoulos
ACM Multimedia4