EDBT 2026 Demo / reviewers in the wild / expert
Emin Zerman
dblp:141/9922
· DBLP profile ↗
26ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0002-3210-8978ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 8 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 9 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Real-Time View Synthesis with Multiplane Image Network using Multimodal SupervisionabstractRecent advances in view synthesis from a single image have increased visual quality of the newly synthesized viewpoints significantly. However, the high computational cost of state-of-the-art methods remains a critical bottleneck, limiting their adoption in real-time applications such as immersive telepresence. To address this limitation, we present a multiplane image (MPI) network that achieves real-time view synthesis. Unlike existing approaches that often rely on a separate depth estimation network to guide the network for estimating MPI parameters, our framework directly predicts the MPI parameters from a single RGB image. To guide the network for estimating correct parameters, we introduce a training strategy that leverages joint supervision from both view synthesis and depth estimation losses to ensure visual fidelity. During inference, our method exclusively utilizes the optimized view synthesis branch, while the depth decoder is only used for training. Our end-to-end approach renders views from a single input image in real-time. Extensive experiments validate that our method provides a compelling rendering speed with visual quality in par with state-of-the-art methods, highlighting its suitability for live, interactive applications. Manu Gond, Mohammadreza Shamshirgarha, Emin Zerman, Sebastian Knorr, Mårten Sjöström |
MMSP | 3 |
| 2025 | Subjective Visual Quality Assessment of Compressed Light Field Images: Learning-based vs. Conventional MethodsabstractLight fields (LF) technology enables the capture and reproduction of a 3D scene accurately, which enhances visual experience in various applications. The sheer volume of multiplicity of the captured views creates logistical problems in both storage and data transmission, which makes LF compression crucial. Even though many LF compression techniques have been proposed and evaluated in recent years, the leading-edge learning-based approaches have not been subject to the same level of scrutiny. This paper presents a subjective quality assessment study on four different LF compression methods, including two learning-based LF compression methods, which have not been studied before from a subjective quality point of view. For this purpose, subjective opinion scores were collected from viewers in two different universities for a cross-lab study. The results indicate that the learning-based compression methods have different behavior in their rate-distortion curves, and that there is room for improvement for learning-based methods. A qualitative analysis also shows that their artifact structures are different from conventional ones. The results highlight the need for a perceptual objective quality metric that takes different types of artifacts into account. The obtained subjective quality database (MiX-LFQDB) is made public to support further research in this area: https://doi.org/10.5281/zenodo.16778670 Emin Zerman, Soheib Takhtardeshir, Anthony Trioux, Jianlong Qin, Roger Olsson, Mårten Sjöström |
MMSP | 1 |
| 2025 | A Visual Quality of Experience Toolkit for Realistic Immersive Telepresence ApplicationsabstractImmersive imaging applications gained a lot of traction in the last decade with the advances in capture, processing, compression, transmission, and display technologies. Recent works on spherical light fields and new view synthesis bring new challenges with respect to fast rendering of new views and a smooth quality of experience (QoE). The real-time rendering capabilities enable realistic immersive telepresence applications, which extend the traditional telepresence systems with improved visual realism and possible addition of depth cues. The recent efforts on the spherical light fields and new view synthesis approaches show that a combined light field and spherical data visualization will be beneficial for the scientific community. In this paper, we provide a WebGL- and WebXR-based visual QoE toolkit which can be used both on traditional displays and extended reality headsets, presenting various visual modalities. To validate the usefulness of the proposed toolkit, we conducted a pilot test on our publicly available spherical light field database. This toolkit can be used for easier and faster quality assessment and can support the scientific community for subjective visual QoE studies focusing on telecommunication, telepresence, and augmented telepresence applications. Manu Gond, Emin Zerman, Mohammadreza Shamshirgarha, Sebastian Knorr, Mårten Sjöström |
QoMEX | 2 |
| 2024 | A Toolkit to Benchmark Point Cloud Quality Metrics with Multi-Track Evaluation CriteriaabstractPoint clouds (PCs) gained popularity as a representation for 3D objects and scenes and are widely used in numerous applications in augmented and virtual reality domains. Concurrently, quality assessment of PCs became even more relevant to improve various aspects of these imaging pipelines. To stimulate further growth and interest in point cloud quality assessment (PCQA), we created a large-scale PCQA dataset (called “BASICS”) which provides the research community with a relevant and challenging dataset to develop reliable objective quality metrics, and we organized the PCVQA grand challenge at ICIP 2023. In this paper, we provide a track-based evaluation methodology for benchmarking visual quality metrics, mirroring the PCVQA grand challenge evaluation scenarios designed to mimic real-life applications. Furthermore, we provide a state-of-the-art benchmark for the point cloud quality metrics. The track-based benchmarking approach shows that there is room for improvement in certain research directions, drawing attention to open problems in the PCQA domain. Ali Ak, Emin Zerman, Maurice Quach, Aladine Chetouani, Giuseppe Valenzise, Patrick Le Callet |
ICIP | 2 |
| 2024 | A Spherical Light Field Database for Immersive Telecommunication and Telepresence ApplicationsabstractImmersive imaging technologies provide an enhanced user experience for visual applications and are getting ready for commonplace use by the industry and the general populace. In particular, light field is a promising technology that enables the capture and reproduction of real light rays from the scene, which can provide a backbone for immersive telecommunication and telepresence applications. Nevertheless, there are still many challenges in transmitting and reproducing light field data. This paper proposes a spherical light field dataset that can be used as a foundation for developing telepresence applications. The Spherical Light Field Database (SLFDB) consists of a light field of 60 views captured with an omnidirectional camera in 20 scenes. To show the usefulness of the proposed database, we provide two use cases: compression and viewpoint estimation. The initial results validate that the publicly available SLFDB will benefit the scientific community. Emin Zerman, Manu Gond, Soheib Takhtardeshir, Roger Olsson, Mårten Sjöström |
QoMEX | 1 |
| 2024 | BASICS: Broad Quality Assessment of Static Point Clouds in a Compression ScenarioabstractPoint clouds have become increasingly prevalent in representing 3D scenes within virtual environments, alongside 3D meshes. Their ease of capture has facilitated a wide array of applications on mobile devices, from smartphones to autonomous vehicles. Notably, point cloud compression has reached an advanced stage and has been standardized. However, the availability of quality assessment datasets, which are essential for developing improved objective quality metrics, remains limited. In this paper, we introduce BASICS, a large-scale quality assessment dataset tailored for static point clouds. The BASICS dataset comprises 75 unique point clouds, each compressed with four different algorithms including a learning-based method, resulting in the evaluation of nearly 1500 point clouds by 3500 unique participants. Furthermore, we conduct a comprehensive analysis of the gathered data, benchmark existing point cloud quality assessment metrics and identify their limitations. By publicly releasing the BASICS dataset, we lay the foundation for addressing these limitations and fostering the development of more precise quality metrics. Ali Ak, Emin Zerman, Maurice Quach, Aladine Chetouani, Aljoscha Smolic, Giuseppe Valenzise, Patrick Le Callet |
IEEE Trans. Multim. | 2 |
| 2021 | Deep Color Mismatch Correction In Stereoscopic 3d ImagesabstractColor mismatch in stereoscopic 3D (S3D) images can create visual discomfort and affect the performance of S3D image processing algorithms, e.g., for depth estimation. In this paper, we propose a new deep learning-based solution for the problem of color mismatch correction. The proposed solution consists of a multi-task convolutional neural network, where color correction is the primary task and correspondence estimation is the secondary task. For the training and evaluation of the proposed network, a new S3D image dataset with color mismatch was created. Based on this dataset, experiments were conducted showing the effectiveness of our solution. Simone Croci, Cagri Ozcinar, Emin Zerman, Roman Dudek, Sebastian Knorr, Aljoscha Smolic |
ICIP | 3 |
| 2021 | Light Field Visual Attention Prediction Using Fourier Disparity LayersabstractIn this paper, we present a novel layered saliency model for Light Fields (LF) called Fourier Disparity Layer Saliency Estimation (FDLSE). The layers are constructed from the existing Fourier Disparity Layer (FDL) LF representation. Our FDLSE model can be used to predict the visual attention (VA) of any LF rendering with arbitrary viewpoint, aperture and depth-of-focus, without the need to generate the rendered image itself. The proposed model surpasses previous work in the following areas. Our method requires the estimation of the saliency map of only one sub-aperture image instead of the full view array. Furthermore, this model does not require pre-estimated disparity maps, but instead relies on the FDL model whose computation fully takes advantage of GPU parallelisation. Finally, FDLSE shows visual improvements and performs quantitatively on par with our previous FGSE model when evaluated on VA prediction of refocus renderings. To our knowledge these are the only two models which can be used to predict LF VA. Ailbhe Gill, Mikael Le Pendu, Martin Alain, Emin Zerman, Aljoscha Smolic |
MMSP | 4 |
| 2021 | The Effect of Temporal Sub-sampling on the Accuracy of Volumetric Video Quality AssessmentabstractVolumetric video content has attracted increasing research interests over the last decade, as it facilitates the integration of dynamic real world content in virtual environments. Point cloud is one of the most common alternatives to represent volumetric video content. Yet, such representation requires an enormous data storage and pose significant greater pressures on compression algorithms compared to the standard 2D video. This challenge has unleashed a new wave in the development of novel point cloud compression technologies, which need to be evaluated in terms of production quality. Due to the high dimensionality of the data, evaluating the performances of relevant coding algorithms can be time consuming. This puts a barrier on optimizing coding algorithms with complex, but perceptually accurate, objective quality metrics. In this study, we thus explore the possibility of reducing temporal-dimension of the content under-evaluation, i.e., temporal sub-sampling, for objective quality evaluation without sacrificing from the correlation with the subjective opinion. In addition, we exploit different temporal pooling methods to further make the quality evaluation procedure more efficient. In total 30 different objective quality metrics were tested on the the V-SENSE volumetric video quality database. According to experimental results, there is no need to employ full frame-rate (30 fps) assessment to reach the meaningful correlation for the considered quality metrics. These observations could be referred to reduce the computation complexity regarding the evaluation and optimization of the relevant compression algorithms. Ali Ak, Emin Zerman, Suiyi Ling, Patrick Le Callet, Aljoscha Smolic |
PCS | 2 |
| 2021 | Focus Guided Light Field Saliency EstimationabstractLight field imaging enables us to capture all light rays in a visual scene. As light fields are four-dimensional, their captures come with an increased amount of information to take advantage of. This has stimulated ongoing light field specific research into virtual viewpoints and shallow depth of field rendering, commonly called refocusing. However, the computation time and memory required to perform these operations can make tasks such as real-time rendering impractical. One solution is to exploit the salient information of light fields to focus resources on regions that attract visual attention when using these algorithms. Although saliency estimation methods for light fields have been previously explored, these focus mainly on salient object segmentation with the goal of generating one saliency map per light field.Aiming to create a basis for a 4D saliency prediction model analogous to light fields, this paper proposes a saliency estimation method specific to light fields that considers the refocusing operation. The proposed method modifies an existing view rendering algorithm with focus guidance, obtained from the light field disparity. This facilitates the construction of saliency maps without the need to render the corresponding view itself, which will help to speed up processing operations that are compatible. The results show that the proposed saliency estimation approach yields very good predictions of visual attention across multiple planes of the light field. We anticipate that this approach can be extended for a range of rendering applications. Ailbhe Gill, Emin Zerman, Martin Alain, Mikael Le Pendu, Aljoscha Smolic |
QoMEX | 2 |
| 2021 | User Behaviour Analysis of Volumetric Video in Augmented RealityabstractAugmented reality (AR) is getting popular, and among other content creation techniques, volumetric video allows to bring dynamic real world content as captured by cameras into such applications. To develop efficient algorithms for compression and transmission of the volumetric media, it is important to understand how users will consume this new form of dynamic 3D content. In this paper, we analyse the user behaviour for the volumetric video consumption in AR. In particular, we study the distribution of users' viewpoints, relative locations, and average distances from the content. For this purpose, we built an Android AR application using the volumetric video and conducted a user study remotely. The results show that users spent most of their time looking at the frontal part of the volumetric video, and this indicates the importance of the face in visual attention. The collected user behaviour data are made public to support further research. Emin Zerman, Radhika Kulkarni, Aljoscha Smolic |
QoMEX | 1 |
| 2020 | Just Noticeable Quantization Levels For High Dynamic Range ImagesabstractJust noticeable quantization levels, which are conventionally used in picture coding, have been mainly developed for standard 8-bit images and low dynamic range (LDR) typical screens. The quantization levels however have not been adapted yet for high dynamic range (HDR) imaging and its accompanied HDR displays, which can reach up to a peak luminance of 4000 cd/m2. This study proposes an experimental methodology on HDR displays to determine just noticeable quantization levels for discrete cosine transform (DCT) coefficients on high luminance images. In the first stage of the proposed method, the quantization noise patterns for different DCT frequencies at different mean luminances are rendered by predicting the LED and LCD values of the two layer HDR display. Then, a two alternative forced choice based psychovisual experimental procedure using geometric search and QUEST methodology is realized by randomly presenting the rendered quantization noise at different amplitudes to the subjects in order to determine the just noticeable levels. The experiments are performed over 3 subjects for 30 different frequencies of 8×8 DCT patterns at mean luminances of 100 cd/m2and 1000 cd/m2. The results are interpreted with respect to frequency and luminance changes and from the point of utilized methodology, namely geometric search and QUEST. Sevim Begüm Sözer, Alper Koz, Ahmet Oguz Akyüz, Emin Zerman, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 4 |
| 2020 | Immersive Imaging Technologies: From Capture to DisplayabstractNew immersive imaging technologies enable creating multimedia systems that would increase the viewer presence and provide an immersive experience. This half-day tutorial aims to give an overview of these new immersive imaging systems and help the participants understand the content creation and delivery pipeline for the immersive imaging technologies. The tutorial will go over the full imaging pipeline, from camera setup for content capture, through content compression / streaming, to content display and related perceptual studies. Martin Alain, Emin Zerman, Cagri Ozcinar |
ACM Multimedia | 2 |
| 2020 | Textured Mesh vs Coloured Point Cloud: A Subjective Study for Volumetric Video CompressionabstractVolumetric video (VV) pipelines reached a high level of maturity, creating interest to use such content in interactive visualisation scenarios. VV allows real world content to be captured and represented as 3D models, which can be viewed from any chosen viewpoint and direction. Thus, VV is ideal to be used in augmented reality (AR) or virtual reality (VR) applications. Both textured polygonal meshes and point clouds are popular methods to represent VV. Even though the signal and image processing community slightly favours the point cloud due to its simpler data structure and faster acquisition, textured polygonal meshes might have other benefits such as better visual quality and easier integration with computer graphics pipelines. To better understand the difference between them, in this study, we compare these two different representation formats for a VV compression scenario utilising state-of-the-art compression techniques. For this purpose, we build a database and collect user opinion scores for subjective quality assessment of the compressed VV. The results show that meshes provide the best quality at high bitrates, while point clouds perform better for low bitrate cases. The created VV quality database will be made available online to support further scientific studies on VV quality assessment. Emin Zerman, Cagri Ozcinar, Pan Gao 0001, Aljoscha Smolic |
QoMEX | 1 |
| 2020 | High Quality Light Field Extraction and Post-Processing for Raw Plenoptic DataabstractLight field technology has reached a certain level of maturity in recent years, and its applications in both computer vision research and industry are offering new perspectives for cinematography and virtual reality. Several methods of capture exist, each with its own advantages and drawbacks. One of these methods involves the use of handheld plenoptic cameras. While these cameras offer freedom and ease of use, they also suffer from various visual artefacts and inconsistencies. We propose in this paper an advanced pipeline that enhances their output. After extracting sub-aperture images from the RAW images with our demultiplexing method, we perform three correction steps. We first remove hot pixel artefacts, then correct colour inconsistencies between views using a colour transfer method, and finally we apply a state of the art light field denoising technique to ensure a high image quality. An in-depth analysis is provided for every step of the pipeline, as well as their interaction within the system. We compare our approach to existing state of the art sub-aperture image extracting algorithms, using a number of metrics as well as a subjective experiment. Finally, we showcase the positive impact of our system on a number of relevant light field applications. Pierre Matysiak, Mairéad Grogan, Mikael Le Pendu, Martin Alain, Emin Zerman, Aljoscha Smolic |
IEEE Trans. Image Process. | 5 |
| 2020 | From Pairwise Comparisons and Rating to a Unified Quality ScaleabstractThe goal of psychometric scaling is the quantification of perceptual experiences, understanding the relationship between an external stimulus, the internal representation and the response. In this paper, we propose a probabilistic framework to fuse the outcome of different psychophysical experimental protocols, namely rating and pairwise comparisons experiments. Such a method can be used for merging existing datasets of subjective nature and for experiments in which both measurements are collected. We analyze and compare the outcomes of both types of experimental protocols in terms of time and accuracy in a set of simulations and experiments with benchmark and real-world image quality assessment datasets, showing the necessity of scaling and the advantages of each protocol and mixing. Although most of our examples focus on image quality assessment, our findings generalize to any other subjective quality-of-experience task. María Pérez-Ortiz 0001, Aliaksei Mikhailiuk, Emin Zerman, Vedad Hulusic, Giuseppe Valenzise, Rafal Mantiuk |
IEEE Trans. Image Process. | 3 |
| 2019 | Colornet - Estimating Colorfulness in Natural ImagesabstractMeasuring the colorfulness of a natural or virtual scene is critical for many applications in image processing field ranging from capturing to display. In this paper, we propose the first deep learning-based colorfulness estimation metric. For this purpose, we develop a color rating model which simultaneously learns to extracts the pertinent characteristic color features and the mapping from feature space to the ideal colorfulness scores for a variety of natural colored images. Additionally, we propose to overcome the lack of adequate annotated dataset problem by combining/aligning two publicly available colorfulness databases using the results of a new subjective test which employs a common subset of both databases. Using the obtained subjectively annotated dataset with 180 colored images, we finally demonstrate the efficacy of our proposed model over the traditional methods, both quantitatively and qualitatively. Emin Zerman, Aakanksha Rana, Aljoscha Smolic |
ICIP | 1 |
| 2019 | Voronoi-based Objective Quality Metrics for Omnidirectional VideoabstractOmnidirectional video (ODV) represents one of the latest and most promising trends in immersive media. The success of ODV depends on the ability to deliver high-quality ODV to the viewers. For this reason, new methods are needed to measure ODV quality that takes into account the interactive look around nature and the spherical representation of ODV. In this paper, we study full-reference objective quality metrics for ODV based on typical encoding distortions in adaptive streaming systems, namely, scaling and compression. The contribution of this paper is three-fold. First, we propose new objective metrics that take into account the unique aspects of ODV. The proposed metrics are based on the subdivision of a given ODV into multiple patches using the spherical Voronoi diagram. Second, we introduce a new dataset of 75 impaired ODVs with different resolutions and compression levels, together with the subjective quality scores gathered during an experiment with 21 participants. Third, we evaluate the proposed Voronoi-based objective metrics using our dataset. The evaluation of the proposed objective metrics and the comparison with existing metrics show that the proposed metrics achieve a better correlation with the subjective scores. The ODV dataset together with the subjective quality scores and the code of the proposed quality metrics are available with this paper. Simone Croci, Cagri Ozcinar, Emin Zerman, Julián Cabrera, Aljoscha Smolic |
QoMEX | 3 |
| 2019 | Analysing the Impact of Cross-Content Pairs on Pairwise Comparison ScalingabstractPairwise comparisons (PWC) methodology is one of the most commonly used methods for subjective quality assessment, especially for computer graphics and multimedia applications. Unlike rating methods, a psychometric scaling operation is required to convert PWC results to numerical subjective quality values. Due to the nature of this scaling operation, the obtained quality scores are relative to the set they are computed in. While it is customary to compare different versions of the same content, in this work we study how cross-content comparisons may benefit psychometric scaling. For this purpose, we use two different video quality databases which have both rating and PWC experiment results. The results show that despite same-content comparisons play a major role in the accuracy of psychometric scaling, the use of a small portion of cross-content comparison pairs is indeed beneficial to obtain more accurate quality estimates. Emin Zerman, Giuseppe Valenzise, Aljoscha Smolic |
QoMEX | 1 |
| 2019 | An Adaptive Quantizer for High Dynamic Range Content: Application to Video CodingabstractIn this paper, we propose an adaptive perceptual quantization method to convert the representation of high dynamic range (HDR) content from the floating point data type to integer, which is compatible with the current image/video coding and display systems. The proposed method considers the luminance distribution of the HDR content, as well as the detectable contrast threshold of the human visual system, in order to preserve more contrast information than the perceptual quantizer (PQ) in integer representation. Aiming to demonstrate the effectiveness of this quantizer for HDR video compression, we implemented it in a mapping function on the top of the HDR video coding system based on high efficiency video coding standard. Moreover, a comparison function is also introduced to decrease the additional bit-rate of side information, generated by the mapping function. Objective quality measurements and subjective tests have been conducted in order to evaluate the quality of the reconstructed HDR videos. Subjective test results have shown that the proposed method can improve, in a significant manner, the perceived quality of some reconstructed HDR videos. In the objective assessment, the proposed method achieves improvements over PQ in terms of the average bit-rate gain for metrics used in the measurement. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Giuseppe Valenzise, Emin Zerman |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2017 | An adaptive perceptual quantization method for HDR video codingabstractThis paper presents a new adaptive perceptual quantization method for the High Dynamic Range (HDR) content. This method considers the luminance distribution of the HDR image as well as the Minimum Detectable Contrast (MDC) thresholds to preserve the contrast information during quantization. Base on this method, we develop a mapping function for HDR video compression and apply it to a HEVC Main 10 Profile-based video coding chain. Our experiments show that the proposed mapping function can efficiently improve the quality of the reconstructed HDR video in both objective and subjective assessments. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Giuseppe Valenzise, Emin Zerman |
ICIP | 6 |
| 2017 | Statistical analysis and directional coding of layer-based HDR image coding residueabstractExisting methods for layer-based backward compatible high dynamic range (HDR) image and video coding mostly focus on the rate-distortion optimization of base layer while neglecting the encoding of the residue signal in the enhancement layer. Although some recent studies handle residue coding by designing function based fixed global mapping curves for 8-bit conversion and exploiting standard codecs on the resulting 8-bit images, they do not take the local characteristics of residue blocks into account. Inspired by the local anisotropic characteristics of the residue signal and directional methods for motion compensated low dynamic range (LDR) video coding, in this paper we first investigate whether HDR image coding residue exhibits also local anisotropic characteristics. Specifically, we verify directional structures in residue blocks by means of auto-covariance analysis for different bitrates, spatial activities and dynamic ranges as the main variables in HDR image coding. Then, we compare the rate distortion performances of directional coding methods with the baseline residue coding methods in the literature along with different combinations of 8-bit conversion methods. The experiments indicate that content dependent 8-bit conversions and directional coding significantly outperforms the existing function based 8-bit conversions and typical coding for residue coding. Kutan Feyiz, Fatih Kamisli, Emin Zerman, Giuseppe Valenzise, Alper Koz, Frédéric Dufaux |
MMSP | 3 |
| 2017 | Effect of color space on high dynamic range video compression performanceabstractHigh dynamic range (HDR) technology allows for capturing and delivering a greater range of luminance levels compared to traditional video using standard dynamic range (SDR). At the same time, it has brought multiple challenges in content distribution, one of them being video compression. While there has been a significant amount of work conducted on this topic, there are some aspects that could still benefit this area. One such aspect is the choice of color space used for coding. In this paper, we evaluate through a subjective study how the performance of HDR video compression is affected by three color spaces: the commonly used Y'CbCr, and the recently introduced ITP (ICtCp) and Ypu'v'. Five video sequences are compressed at four bit rates, selected in a preliminary study, and their quality is assessed using pairwise comparisons. The results of pairwise comparisons are further analyzed and scaled to obtain quality scores. We found no evidence of ITP improving compression performance over Y'CbCr. We also found that Ypu'v' results in a moderately lower performance for some sequences. Emin Zerman, Vedad Hulusic, Giuseppe Valenzise, Rafal Mantiuk, Frédéric Dufaux |
QoMEX | 1 |
| 2016 | Video content analysis method for audiovisual quality assessmentabstractIn this study a novel, spatio-temporal characteristics based video content analysis method is presented. The proposed method has been evaluated on different video quality assessment databases, which include videos with different characteristics and distortion types. Test results obtained on different databases demonstrate the robustness and accuracy of the proposed content analysis method. Moreover, this analysis method is employed in order to examine the performance improvement in audiovisual quality assessment when the video content is taken into consideration. Baris Konuk, Emin Zerman, Gokce Nur, Gozde Bozdagi Akar |
QoMEX | 2 |
| 2014 | A parametric video quality model based on source and network characteristicsabstractThe increasing demand for streaming video raises the need for flexible and easily implemented Video Quality Assessment (VQA) metrics. Although there are different VQA metrics, most of these are either Full-Reference (FR) or Reduced-Reference (RR). Both FR and RR metrics bring challenges for on-the-fly multimedia systems due to the necessity of additional network traffic for reference data. No-Reference (NR) video metrics, on the other hand, as the name suggests, are much more flexible for user-end applications. This introduces a need for robust and efficient NR VQA metrics. In this paper, an NR VQA metric considering spatiotemporal information, bit rate, and packet loss rate characteristics of a video content is proposed. The proposed metric is evaluated on EPFL-PoliMI dataset, which includes different video content characteristics. The experimental results show that the proposed metric is a robust and accurate NR VQA metric towards diverse video content characteristics. Emin Zerman, Baris Konuk, Gokce Nur, Gozde Bozdagi Akar |
ICIP | 1 |
| 2013 | A spatiotemporal no-reference video quality assessment modelabstractMany researchers have been developing objective video quality assessment methods due to increasing demand for perceived video quality measurement results by end users to speed-up advancements of multimedia services. However, most of these methods are either Full-Reference (FR) metrics, which require the original video or Reduced-Reference (RR) metrics, which need some features extracted from the original video. No-Reference (NR) metrics, on the other hand, do not require any information about the original video; hence, are much more suitable for applications like video streaming. This paper presents a novel, objective, NR video quality assessment algorithm. The proposed algorithm is based on utilization of spatial extent of video, temporal extent of video using motion vectors, bit rate, and packet loss ratio. Test results obtained using LIVE video quality database demonstrate the accuracy and robustness of the proposed metric. Baris Konuk, Emin Zerman, Gokce Nur, Gozde Bozdagi Akar |
ICIP | 2 |