Cagri Ozcinar

dblp:158/9750 · DBLP profile ↗
← Back
40ranked-venue papers
10as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 10 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorComputer networks · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360$^{\circ }$∘ Videos
abstract
Omnidirectional videos (ODVs) are redefining viewer experiences in virtual reality (VR) by offering an unprecedented full field-of-view (FOV). This study extends the domain of saliency prediction to 360$^\circ$∘ environments, addressing the complexities of spherical distortion and the integration of spatial audio. Contextually, ODVs have transformed user experience by adding a spatial audio dimension that aligns sound direction with the viewer's perspective in spherical scenes. Motivated by the lack of comprehensive datasets for 360$^\circ$∘ audio-visual saliency prediction, our study curates YT360-EyeTracking, a new dataset of 81 ODVs, each observed under varying audio-visual conditions. Our goal is to explore how to utilize audio-visual cues to effectively predict visual saliency in 360$^\circ$∘ videos. Towards this aim, we propose two novel saliency prediction models: SalViT360, a vision-transformer-based framework for ODVs equipped with spherical geometry-aware spatio-temporal attention layers, and SalViT360-AV, which further incorporates transformer adapters conditioned on audio input. Our results on a number of benchmark datasets, including our YT360-EyeTracking, demonstrate that SalViT360 and SalViT360-AV significantly outperform existing methods in predicting viewer attention in 360$^\circ$∘ scenes. Interpreting these results, we suggest that integrating spatial audio cues in the model architecture is crucial for accurate saliency prediction in omnidirectional videos.
Mert Cokelek, Halit Ozsoy, Nevrez Imamoglu, Cagri Ozcinar, Inci Ayhan, Erkut Erdem, Aykut Erdem
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Omnidirectional image quality assessment with local-global vision transformers
Nafiseh Jabbari Tofighi, Mohamed Hedi Elfkir, Nevrez Imamoglu, Cagri Ozcinar, Aykut Erdem, Erkut Erdem
Image Vis. Comput.4
2023 Spherical Vision Transformer for 360° Video Saliency Prediction
Mert Cokelek, Nevrez Imamoglu, Cagri Ozcinar, Erkut Erdem, Aykut Erdem
BMVC3
2023 ST360IQ: No-Reference Omnidirectional Image Quality Assessment With Spherical Vision Transformers
abstract
Omnidirectional images, aka 360° images, can deliver immersive and interactive visual experiences. As their popularity has increased dramatically in recent years, evaluating the quality of 360° images has become a problem of interest since it provides insights for capturing, transmitting, and consuming this new media. However, directly adapting quality assessment methods proposed for standard natural images for omnidirectional data poses certain challenges. These models need to deal with very high-resolution data and implicit distortions due to the spherical form of the images. In this study, we present a method for no-reference 360° image quality assessment. Our proposed ST360IQ model extracts tangent viewports from the salient parts of the input omnidirectional image and employs a vision-transformers based module processing saliency selective patches/tokens that estimates a quality score from each viewport. Then, it aggregates these scores to give a final quality score. Our experiments on two benchmark datasets, namely OIQA and CVIQ datasets, demonstrate that as compared to the state-of-the-art, our approach predicts the quality of an omnidirectional image correlated with the human-perceived image quality. The code has been available on https://github.com/Nafiseh-Tofighi/ST360IQ
Nafiseh Jabbari Tofighi, Mohamed Hedi Elfkir, Nevrez Imamoglu, Cagri Ozcinar, Erkut Erdem, Aykut Erdem
ICASSP4
2023 Automatic content moderation on social media
Dogus Karabulut, Cagri Ozcinar, Gholamreza Anbarjafari
Multim. Tools Appl.2
2022 Privacy-Preserving Viewport Prediction using Federated Learning for 360° Live Video Streaming
abstract
Predicting the user's viewport scanpath is an essential task for 360° viewport-based adaptive streaming. It informs the system which parts of content should be streamed with high quality for bandwidth saving over the best-effort Internet. However, in light of growing privacy concerns among consumers and increasingly strict data privacy legislation, user data collection and storage have been constrained. This paper proposes a novel privacy-preserving framework employing Federated Learning (FL) for online viewport prediction in a live 360° video streaming scenario. In this framework, the user data is only collected and processed on the client-side in the current viewing session and not shared with external parties, e.g., servers, and other clients. We evaluated the framework over a widely-used dataset and measure the computation and transmission time of the proposed streaming system. The experiments show that our framework provides high prediction accuracy and achieves real-time computation requirements of live video streaming. On privacy preservation, our results demonstrate that in a tile-based 360° video streaming system, the user identification rate can be decreased by 18.11 percentage points in 4 x 3 tiles per frame and 9.65 percentage points in 16x9 tiles per frame. The code will be publicly available to further contribute to the community.
Fang-Yi Chao, Cagri Ozcinar, Aljoscha Smolic
MMSP2
2021 Deep Color Mismatch Correction In Stereoscopic 3d Images
abstract
Color mismatch in stereoscopic 3D (S3D) images can create visual discomfort and affect the performance of S3D image processing algorithms, e.g., for depth estimation. In this paper, we propose a new deep learning-based solution for the problem of color mismatch correction. The proposed solution consists of a multi-task convolutional neural network, where color correction is the primary task and correspondence estimation is the secondary task. For the training and evaluation of the proposed network, a new S3D image dataset with color mismatch was created. Based on this dataset, experiments were conducted showing the effectiveness of our solution.
Simone Croci, Cagri Ozcinar, Emin Zerman, Roman Dudek, Sebastian Knorr, Aljoscha Smolic
ICIP2
2021 Transformer-based Long-Term Viewport Prediction in 360° Video: Scanpath is All You Need
abstract
Virtual Reality (VR) multimedia technology has dramatically advanced in recent years. Its immersive and interactive natures enable users to view any direction in 360° content freely. Users do not see the entire 360° content at a glance, but only a portion in the viewport. Viewport-based adaptive streaming, which streams only the user’s viewport of interest with high quality, has emerged as the primary technique to save bandwidth over the best-effort Internet. Thus, users’ viewport prediction in the forthcoming seconds becomes an essential task for informing the streaming decisions in the VR system. Various viewport prediction methods based on deep neural networks have been proposed. However, typically they are composed of complex Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) that require heavy computation. To achieve high prediction accuracy in limited computation time in a streaming system, we propose a new transformer-based architecture, named 360° Viewport Prediction Transformer (VPT360), that only leverages the past viewport scanpath to predict a user’s future viewport scanpath. We evaluate VPT360 over three widely-used datasets and compare the computation complexity with the state-of-the-art methods. The experiments show that our VPT360 provides the highest accuracy for short-term and long-term prediction and achieves the lowest computation complexity. The code is publicly available at https://github.com/FannyChao/VPT360 to further contribute to the community.
Fang-Yi Chao, Cagri Ozcinar, Aljoscha Smolic
MMSP2
2020 Hierarchical Fourier Disparity Layer Transmission For Light Field Streaming
abstract
In this paper, we present a novel approach to efficiently transmit light fields in the Fourier Disparity Layer (FDL) representation using a binary hierarchical scheme. The FDL model consists of a set of additive layers which can be simply shifted and summed to render a view of the light field at any angular coordinate. In order to transmit the FDL model, we propose a method for building a binary tree where the root consists of a single compound layer obtained as the sum of all the original layers. Subsequent levels of the tree are obtained by splitting a parent layer into two children layers whose sum is equal to the parent layer. Hence, the FDL model is recursively refined with additional layers at each new level, resulting in a scalable representation. An efficient scheme is proposed to encode a single image in order to split a parent layer into its two children. Thanks to this approach, the total number of images to decode for receiving the complete tree is equal to the number of layers in the original FDL model, which is typically smaller than the number of views required in the traditional light field representation.
Mikael Le Pendu, Cagri Ozcinar, Aljoscha Smolic
ICIP2
2020 Immersive Imaging Technologies: From Capture to Display
abstract
New immersive imaging technologies enable creating multimedia systems that would increase the viewer presence and provide an immersive experience. This half-day tutorial aims to give an overview of these new immersive imaging systems and help the participants understand the content creation and delivery pipeline for the immersive imaging technologies. The tutorial will go over the full imaging pipeline, from camera setup for content capture, through content compression / streaming, to content display and related perceptual studies.
Martin Alain, Emin Zerman, Cagri Ozcinar
ACM Multimedia3
2020 Textured Mesh vs Coloured Point Cloud: A Subjective Study for Volumetric Video Compression
abstract
Volumetric video (VV) pipelines reached a high level of maturity, creating interest to use such content in interactive visualisation scenarios. VV allows real world content to be captured and represented as 3D models, which can be viewed from any chosen viewpoint and direction. Thus, VV is ideal to be used in augmented reality (AR) or virtual reality (VR) applications. Both textured polygonal meshes and point clouds are popular methods to represent VV. Even though the signal and image processing community slightly favours the point cloud due to its simpler data structure and faster acquisition, textured polygonal meshes might have other benefits such as better visual quality and easier integration with computer graphics pipelines. To better understand the difference between them, in this study, we compare these two different representation formats for a VV compression scenario utilising state-of-the-art compression techniques. For this purpose, we build a database and collect user opinion scores for subjective quality assessment of the compressed VV. The results show that meshes provide the best quality at high bitrates, while point clouds perform better for low bitrate cases. The created VV quality database will be made available online to support further scientific studies on VV quality assessment.
Emin Zerman, Cagri Ozcinar, Pan Gao 0001, Aljoscha Smolic
QoMEX2
2020 Towards Audio-Visual Saliency Prediction for Omnidirectional Video with Spatial Audio
abstract
Omnidirectional videos (ODVs) with spatial audio enable viewers to perceive 360° directions of audio and visual signals during the consumption of ODVs with head-mounted displays (HMDs). By predicting salient audio-visual regions, ODV systems can be optimized to provide an immersive sensation of audio-visual stimuli with high-quality. Despite the intense recent effort for ODV saliency prediction, the current literature still does not consider the impact of auditory information in ODVs. In this work, we propose an audio-visual saliency (AVS360) model that incorporates 360° spatial-temporal visual representation and spatial auditory information in ODVs. The proposed AVS360 model is composed of two 3D residual networks (ResNets) to encode visual and audio cues. The first one is embedded with a spherical representation technique to extract 360° visual features, and the second one extracts the features of audio using the log mel-spectrogram. We emphasize sound source locations by integrating audio energy map (AEM) generated from spatial audio description (i.e., ambisonics) and equator viewing behavior with equator center bias (ECB). The audio and visual features are combined and fused with AEM and ECB via attention mechanism. Our experimental results show that the AVS360 model has significant superiority over five state-of-the-art saliency models. To the best of our knowledge, it is the first w ork that develops the audio-visual saliency model in ODVs. The code will be publicly available to foster future research on audio-visual saliency in ODVs.
Fang-Yi Chao, Cagri Ozcinar, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges, Aljoscha Smolic
VCIP2
2020 Cycle-consistent generative adversarial neural networks based low quality fingerprint enhancement
Dogus Karabulut, Pavlo Tertychnyi, Hasan Sait Arslan, Cagri Ozcinar, Kamal Nasrollahi, Joan Valls, Joan Vilaseca, Thomas B. Moeslund, Gholamreza Anbarjafari
Multim. Tools Appl.4
2020 Correction to: Optimal image compression via block-based adaptive colour reduction with minimal contour effect
Pejman Rasti, Iiris Lüsi, Anastasia Bolotnikova, Morteza Daneshmand, Cagri Ozcinar, Gholamreza Anbarjafari
Multim. Tools Appl.5
2020 Do Users Behave Similarly in VR? Investigation of the User Influence on the System Design
abstract
With the overarching goal of developing user-centric Virtual Reality (VR) systems, a new wave of studies focused on understanding how users interact in VR environments has recently emerged. Despite the intense efforts, however, current literature still does not provide the right framework to fully interpret and predict users’ trajectories while navigating in VR scenes. This work advances the state-of-the-art on both the study of users’ behaviour in VR and the user-centric system design. In more detail, we complement current datasets by presenting a publicly available dataset that provides navigation trajectories acquired for heterogeneous omnidirectional videos and different viewing platforms—namely, head-mounted display, tablet, and laptop. We then present an exhaustive analysis on the collected data to better understand navigation in VR across users, content, and, for the first time, across viewing platforms. The novelty lies in the user-affinity metric, proposed in this work to investigate users’ similarities when navigating within the content. The analysis reveals useful insights on the effect of device and content on the navigation, which could be precious considerations from the system design perspective. As a case study of the importance of studying users’ behaviour when designing VR systems, we finally propose a user-centric server optimisation. We formulate an integer linear program that seeks the best stored set of omnidirectional content that minimises encoding and storage cost while maximising the user’s experience. This is posed while taking into account network dynamics, type of video content, and also user population interactivity. Experimental results prove that our solution outperforms common company recommendations in terms of experienced quality but also in terms of encoding and storage, achieving a savings up to 70%. More importantly, we highlight a strong correlation between the storage cost and the user-affinity metric, showing the impact of the latter in the system architecture design.
Silvia Rossi 0001, Cagri Ozcinar, Aljoscha Smolic, Laura Toni
ACM Trans. Multim. Comput. Commun. Appl.2
2019 On the effect of age perception biases for real age regression
abstract
Automatic age estimation from facial images represents an important task in computer vision. This paper analyses the effect of gender, age, ethnic, makeup and expression attributes of faces as sources of bias to improve deep apparent age prediction. Following recent works where it is shown that apparent age labels benefit real age estimation, rather than direct real to real age regression, our main contribution is the integration, in an end-to-end architecture, of face attributes for apparent age prediction with an additional loss for real age regression. Experimental results on the APPA-REAL dataset indicate the proposed network successfully take advantage of the adopted attributes to improve both apparent and real age estimation. Our model outperformed a state-of-the-art architecture proposed to separately address apparent and real age regression. Finally, we present preliminary results and discussion of a proof of concept application using the proposed model to regress the apparent age of an individual based on the gender of an external observer.
Júlio C. S. Jacques Júnior, Cagri Ozcinar, Marina Marjanovic, Xavier Baró, Gholamreza Anbarjafari, Sergio Escalera
FG2
2019 Towards Generating Ambisonics Using Audio-visual Cue for Virtual Reality
abstract
Ambisonics i.e., a full-sphere surround sound, is quintessential with 360° visual content to provide a realistic virtual reality (VR) experience. While 360° visual content capture gained a tremendous boost recently, the estimation of corresponding spatial sound is still challenging due to the required sound-field microphones or information about the sound-source locations. In this paper, we introduce a novel problem of generating Ambisonics in 360° videos using the audiovisual cue. With this aim, firstly, a novel 360° audio-visual video dataset of 265 videos is introduced with annotated sound-source locations. Secondly, a pipeline is designed for an automatic Ambisonic estimation problem. Benefiting from the deep learning based audiovisual feature-embedding and prediction modules, our pipeline estimates the 3D sound-source locations and further use such locations to encode to the B-format. To benchmark our dataset and pipeline, we additionally propose evaluation criteria to investigate the performance using different 360° input representations. Our results demonstrate the efficacy of the proposed pipeline and open up a new area of research in 360° audio-visual analysis for future investigations.
Aakanksha Rana, Cagri Ozcinar, Aljoscha Smolic
ICASSP2
2019 A Study of Light Field Streaming for An Interactive Refocusing Application
abstract
Light fields are able to capture light rays from a scene arriving at different angles, which allows post-capture rendering applications such as interactive viewpoint selection or refocusing. However, this additional angular information comes at the price of a significant increase of the data volume compared to traditional 2D images. While light field compression is still an ongoing research effort, showing impressive compression gain with the latest coding standard, light fields are in practice often stored on remote servers to avoid consuming unnecessary storage of the user devices. A typical cost-effective solution for light field visualisation is then to render the requested image on the server and transmit the result to the user. Another trivial solution would be to directly send the light field to the user and perform the rendering process directly on the client side to avoid transmission delay. While the latter solution seems instinctively less optimal and is usually discarded in previous work because of an expected unacceptable startup delay, we propose a quantitative study to compare both solutions in terms of rate-distortion (RD) performance. A counterintuitive finding of this paper is that accepting a reasonable startup delay (a few seconds) can provide a significant improvement of the RD performances.
Martin Alain, Cagri Ozcinar, Aljoscha Smolic
ICIP2
2019 Super-resolution of Omnidirectional Images Using Adversarial Learning
abstract
An omnidirectional image (ODI) enables viewers to look in every direction from a fixed point through a head-mounted display providing an immersive experience compared to that of a standard image. Designing immersive virtual reality systems with ODIs is challenging as they require high resolution content. In this paper, we study super-resolution for ODIs and propose an improved generative adversarial network based model which is optimized to handle the artifacts obtained in the spherical observational space. Specifically, we propose to use a fast PatchGAN discriminator, as it needs fewer parameters and improves the super-resolution at a fine scale. We also explore the generative models with adversarial learning by introducing a spherical-content specific loss function, called 360-SS. To train and test the performance of our proposed model we prepare a dataset of 4500 ODIs. Our results demonstrate the efficacy of the proposed method and identify new challenges in ODI super-resolution for future investigations.
Cagri Ozcinar, Aakanksha Rana, Aljoscha Smolic
MMSP1
2019 Voronoi-based Objective Quality Metrics for Omnidirectional Video
abstract
Omnidirectional video (ODV) represents one of the latest and most promising trends in immersive media. The success of ODV depends on the ability to deliver high-quality ODV to the viewers. For this reason, new methods are needed to measure ODV quality that takes into account the interactive look around nature and the spherical representation of ODV. In this paper, we study full-reference objective quality metrics for ODV based on typical encoding distortions in adaptive streaming systems, namely, scaling and compression. The contribution of this paper is three-fold. First, we propose new objective metrics that take into account the unique aspects of ODV. The proposed metrics are based on the subdivision of a given ODV into multiple patches using the spherical Voronoi diagram. Second, we introduce a new dataset of 75 impaired ODVs with different resolutions and compression levels, together with the subjective quality scores gathered during an experiment with 21 participants. Third, we evaluate the proposed Voronoi-based objective metrics using our dataset. The evaluation of the proposed objective metrics and the comparison with existing metrics show that the proposed metrics achieve a better correlation with the subjective scores. The ODV dataset together with the subjective quality scores and the code of the proposed quality metrics are available with this paper.
Simone Croci, Cagri Ozcinar, Emin Zerman, Julián Cabrera, Aljoscha Smolic
QoMEX2
2019 A novel deep network architecture for reconstructing RGB facial images from thermal for face recognition
Andre Litvin, Kamal Nasrollahi, Sergio Escalera, Cagri Ozcinar, Thomas B. Moeslund, Gholamreza Anbarjafari
Multim. Tools Appl.4
2019 Adaptive multi-view video streaming using side information over peer-to-peer networks
Cagri Ozcinar, Erhan Ekmekcioglu, Gholamreza Anbarjafari, Ahmet M. Kondoz
Multim. Tools Appl.1
2018 Changes in Facial Expression as Biometric: A Database and Benchmarks of Identification
abstract
Facial dynamics can be considered as unique signatures for discrimination between people. These have started to become important topic since many devices have the possibility of unlocking using face recognition or verification. In this work, we evaluate the efficacy of the transition frames of video in emotion as compared to the peak emotion frames for identification. For experiments with transition frames we extract features from each frame of the video from a fine-tuned VGG-Face Convolutional Neural Network (CNN) and geometric features from facial landmark points. To model the temporal context of the transition frames we train a Long-Short Term Memory (LSTM) on the geometric and the CNN features. Furthermore, we employ two fusion strategies: first, an early fusion, in which the geometric and the CNN features are stacked and fed to the LSTM. Second, a late fusion, in which the prediction of the LSTMs, trained independently on the two features, are stacked and used with a Support Vector Machine (SVM). Experimental results show that the late fusion strategy gives the best results and the transition frames give better identification results as compared to the peak emotion frames.
Rain Eric Haamer, Kaustubh Kulkarni, Nasrin Imanpour, Mohammad A. Haque, Egils Avots, Michelle Breisch, Kamal Nasrollahi, Sergio Escalera, Cagri Ozcinar, Xavier Baró, Ahmad Reza Naghsh-Nilchi, Thomas B. Moeslund, Gholamreza Anbarjafari
FG9
2018 Director's Cut - Analysis of Aspects of Interactive Storytelling for VR Films
Colm O. Fearghail, Cagri Ozcinar, Sebastian Knorr, Aljoscha Smolic
ICIDS2
2018 Optimization of Occlusion-Inducing Depth Pixels in 3-D Video Coding
abstract
The optimization of occlusion-inducing depth pixels in depth map coding has received little attention in the literature, since their associated texture pixels are occluded in the synthesized view and their effect on the synthesized view is considered negligible. However, the occlusion-inducing depth pixels still need to consume the bits to be transmitted, and will induce geometry distortion that inherently exists in the synthesized view. In this paper, we propose an efficient depth map coding scheme specifically for the occlusion-inducing depth pixels by using allowable depth distortions. Firstly, we formulate a problem of minimizing the overall geometry distortion in the occlusion subject to the bit rate constraint, for which the depth distortion is properly adjusted within the set of allowable depth distortions that introduce the same disparity error as the initial depth distortion. Then, we propose a dynamic programming solution to find the optimal depth distortion vector for the occlusion. The proposed algorithm can improve the coding efficiency without alteration of the occlusion order. Simulation results confirm the performance improvement compared to other existing algorithms.
Pan Gao 0001, Cagri Ozcinar, Aljoscha Smolic
ICIP2
2018 Visual Attention in Omnidirectional Video for Virtual Reality Applications
abstract
Understanding of visual attention is crucial for omnidirectional video (ODV) viewed for instance with a head-mounted display (HMD), where only a fraction of an ODV is rendered at a time. Transmission and rendering of ODV can be optimized by understanding how viewers consume a given ODV in virtual reality (VR) applications. In order to predict video regions that might draw the attention of viewers, saliency maps can be estimated by using computational visual attention models. As no such model currently exists for ODV, but given the importance for emerging ODV applications, we create a new visual attention user dataset for ODV, investigate behavior of viewers when consuming the content, and analyze the prediction performance of state-of-the-art visual attention models. Our developed test-bed and dataset will be publicly available with this paper, to stimulate and support research on ODV.
Cagri Ozcinar, Aljoscha Smolic
QoMEX1
2018 Omnidirectional Video Streaming Using Visual Attention-Driven Dynamic Tiling for VR
abstract
This paper proposes a new adaptive omnidirectional video (ODV) streaming system that uses visual attention (VA) maps. The proposed method benefits from a novel approach to VA-based bitrate allocation algorithm and dynamic tiling, providing enhanced virtual reality (VR) video experiences. The main contribution of this paper is the use of VA maps: (i) to distribute a given bitrate budget among a set of tiles of a given ODV and, (ii) to decide an optimal tiling structure (i.e., tile scheme) per chunk. For this, a novel objective metric is proposed: the visual attention spherical weighted (VASW) PSNR. This metric operates in the spherical domain and by means of a VA probabilistic model aims at capturing the quality of the actual areas observed by the users when navigating through the ODV content. We evaluate the proposed system performance with varying bandwidth conditions and the tracked head orientations from disjoint user experiments. Results show that the proposed system significantly outperforms the existing tiled-based streaming method.
Cagri Ozcinar, Julián Cabrera, Aljoscha Smolic
VCIP1
2018 Spatio-temporal constrained tone mapping operator for HDR video compression
Cagri Ozcinar, Paul Lauga, Giuseppe Valenzise, Frédéric Dufaux
J. Vis. Commun. Image Represent.1
2018 Imperceptible non-blind watermarking and robustness against tone mapping operation attacks for high dynamic range images
Gholamreza Anbarjafari, Cagri Ozcinar
Multim. Tools Appl.2
2018 Optimal image compression via block-based adaptive colour reduction with minimal contour effect
Iiris Lüsi, Anastasia Bolotnikova, Morteza Daneshmand, Cagri Ozcinar, Gholamreza Anbarjafari
Multim. Tools Appl.4
2017 Joint Challenge on Dominant and Complementary Emotion Recognition Using Micro Emotion Features and Head-Pose Estimation: Databases
abstract
In this work two databases for the Joint Challenge on Dominant and Complementary Emotion Recognition Using Micro Emotion Features and Head-Pose Estimation1 are introduced. Head pose estimation paired with and detailed emotion recognition have become very important in relation to human-computer interaction. The 3D head pose database, SASE, is a 3D database acquired with Microsoft Kinect 2 camera, including RGB and depth information of different head poses which is composed by a total of 30000 frames with annotated markers, including 32 male and 18 female subjects. For the dominant and complementary emotion database, iCVMEFED, includes 31250 images with different emotions of 115 subjects whose gender distribution is almost uniform. For each subject there are 5 samples. The emotions are composed by 7 basic emotions plus neutral, being defined as complementary and dominant pairs. The emotion associated to the images were labeled with the support of psychologists.
Iiris Lüsi, Júlio C. S. Jacques Júnior, Jelena Gorbova, Xavier Baró, Sergio Escalera, Hasan Demirel, Juri Allik, Cagri Ozcinar, Gholamreza Anbarjafari
FG8
2017 Viewport-aware adaptive 360° video streaming using tiles for virtual reality
abstract
360° video is attracting an increasing amount of attention in the context of Virtual Reality (VR). Owing to its very high-resolution requirements, existing professional streaming services for 360° video suffer from severe drawbacks. This paper introduces a novel end-to-end streaming system from encoding to displaying, to transmit 8K resolution 360° video and to provide an enhanced VR experience using Head Mounted Displays (HMDs). The main contributions of the proposed system are about tiling, integration of the MPEG-Dynamic Adaptive Streaming over HTTP (DASH) standard, and viewport-aware bitrate level selection. Tiling and adaptive streaming enable the proposed system to deliver very high-resolution 360° video at good visual quality. Further, the proposed viewport-aware bitrate assignment selects an optimum DASH representation for each tile in a viewport-aware manner. The quality performance of the proposed system is verified in simulations with varying network bandwidth using realistic view trajectories recorded from user experiments. Our results show that the proposed streaming system compares favorably compared to existing methods in terms of PSNR and SSIM inside the viewport.
Cagri Ozcinar, Ana De Abreu, Aljoscha Smolic
ICIP1
2017 Estimation of Optimal Encoding Ladders for Tiled 360° VR Video in Adaptive Streaming Systems
abstract
Given the significant industrial growth of demand for virtual reality (VR), 360ovideo streaming is one of the most important VR applications that require cost-optimal solutions to achieve widespread proliferation of VR technology. Because of its inherent variability of data-intensive content types and its tiled-based encoding and streaming, 360ovideo requires new encoding ladders in adaptive streaming systems to achieve cost-optimal and immersive streaming experiences. In this context, this paper targets both the provider's and client's perspectives and introduces a new content-aware encoding ladder estimation method for tiled 360oVR video in adaptive streaming systems. The proposed method first categories a given 360ovideo using its features of encoding complexity and estimates the visual distortion and resource cost of each bitrate level based on the proposed distortion and resource cost models. An optimal encoding ladder is then formed using the proposed integer linear programming (ILP) algorithm by considering practical constraints. Experimental results of the proposed method are compared with the recommended encoding ladders of professional streaming service providers. Evaluations show that the proposed encoding ladders deliver better results compared to the recommended encoding ladders in terms of objective quality for 360ovideo, providing optimal encoding ladders using a set of service provider's constraint parameters.
Cagri Ozcinar, Ana De Abreu, Sebastian Knorr, Aljoscha Smolic
ISM1
2017 Look around you: Saliency maps for omnidirectional images in VR applications
abstract
Understanding visual attention has always been a topic of great interest in the graphics, image/video processing, robotics and human-computer interaction communities. By understanding salient image regions, the compression, transmission and rendering algorithms can be optimized. This is particularly important in omnidirectional images (ODIs) viewed with a head-mounted display (HMD), where only a fraction of the captured scene is displayed at a time, namely viewport. In order to predict salient image regions, saliency maps are estimated either by using an eye tracker to collect eye fixations during subjective tests or by using computational models of visual attention. However, eye tracking developments for ODIs are still in the early stages and although a large list of saliency models are available, no particular attention has been dedicated to ODIs. Therefore, in this paper, we consider the problem of estimating saliency maps for ODIs viewed with HMDs, when the use of an eye tracker device is not possible. We collected viewport center trajectories (VCTs) of 32 participants for 21 ODIs and propose a method to transform the gathered data into saliency maps. The obtained saliency maps are compared in terms of image exposition time used to display each ODI in the subjective tests. Then, motivated by the equator bias tendency in ODIs, we propose a post-processing method, namely fused saliency maps (FSM), to adapt current saliency models to ODIs requirements. We show that the use of FSM on current models improves their performance by up to 20%. The developed database and testbed are publicly available with this paper.
Ana De Abreu, Cagri Ozcinar, Aljoscha Smolic
QoMEX2
2017 A new low-complexity patch-based image super-resolution
abstract
In this study, a novel single image super‐resolution (SR) method, which uses a generated dictionary from pairs of high‐resolution (HR) images and their corresponding low‐resolution (LR) representations, is proposed. First, HR and LR dictionaries are created by dividing HR and LR images into patches Afterwards, when performing SR, the distance between every patch of the input LR image and those of available LR patches in the LR dictionary are calculated. The minimum distance between the input LR patch and those in the LR dictionary is taken, and its counterpart from the HR dictionary will be passed through an illumination enhancement process resulting in consistency of illumination between neighbour patches. This process is applied to all patches of the LR image. Finally, in order to remove the blocking effect caused by merging the patches, an average of the obtained HR image and the interpolated image is calculated. Furthermore, it is shown that the stabe of dictionaries is reducible to a great degree. The speed of the system is improved by 62.5%. The quantitative and qualitative analyses of the experimental results show the superiority of the proposed technique over the conventional and state‐of‐the‐art methods.
Pejman Rasti, Kamal Nasrollahi, Olga Orlova, Gert Tamberg, Cagri Ozcinar, Thomas B. Moeslund, Gholamreza Anbarjafari
IET Comput. Vis.5
2016 Quality-aware adaptive delivery of multi-view video
abstract
Advances in video coding and networking technologies have paved the way for the Multi-View Video (MVV) streaming. However, large amounts of data and dynamic network conditions result in frequent network congestion, which may prevent video packets from being delivered on time. As a consequence, the 3D viewing experience may be degraded significantly, unless quality-aware adaptation methods are deployed. There is no research work to discuss the MVV adaptation of decision strategy or provide a detailed analysis of a dynamic network environment. This work addresses the mentioned issues for MVV streaming over HTTP for emerging multi-view displays. In this research work, the effect of various adaptations of decision strategies are evaluated and, as a result, a new quality-aware adaptation method is designed. The proposed method is benefiting from layer based video coding in such a way that high Quality of Experience (QoE) is maintained in a cost-effective manner. The conducted experimental results on MVV streaming using the proposed strategy are showing that the perceptual 3D video quality, under adverse network conditions, is enhanced significantly as a result of the proposed quality-aware adaptation.
Cagri Ozcinar, Erhan Ekmekcioglu, Ahmet M. Kondoz
ICASSP1
2016 Adaptive delivery of immersive 3D multi-view video over the Internet
Cagri Ozcinar, Erhan Ekmekcioglu, Janko Calic, Ahmet M. Kondoz
Multim. Tools Appl.1
2014 Adaptive 3D multi-view video streaming over P2P networks
abstract
Streaming 3D multi-view video to multiple clients simultaneously remains a highly challenging problem due to the high-volume of data involved and the inherent limitations imposed by the delivery networks. Delivery of multimedia streams over Peer-to-Peer (P2P) networks has gained great interest due to its ability to maximise link utilisation, preventing the transport of multiple copies of the same packet for many users. On the other hand, the quality of experience can still be significantly degraded by dynamic variations caused by congestions, unless content-aware precautionary mechanisms and adaptation methods are deployed. In this paper, a novel, adaptive multi-view video streaming over a P2P system is introduced which addresses the next generation high resolution multi-view users' experiences with autostereoscopic displays. The solution comprises the extraction of low-overhead supplementary metadata at the media encoding server that is distributed through the network and used by clients performing network adaptation. In the proposed concept, pre-selected views are discarded at a times of network congestion and reconstructed with high quality using the metadata and the neighbouring views. The experimental results show that the robustness of P2P multi-view streaming using the proposed adaptation scheme is significantly increased under congestion.
Cagri Ozcinar, Erhan Ekmekcioglu, Ahmet M. Kondoz
ICIP1
2011 Video resolution enhancement by using complex wavelet transform
abstract
In this paper, we propose a multi-frame video resolution enhancement technique based on dual tree complex wavelet transform (DT-CWT). Here, before registration step, the respective frames have been passed through an illumination compensation procedure which is based on singular value decomposition (SVD) and discrete wavelet transform (DWT). The frame subject to the resolution enhancement is decomposed into its different frequency subbands by using DT-CWT. The high frequency subbands have been interpolated by using bicubic interpolation. Furthermore, the compensated frames are registered by using Vandewalle registration with structure adaptive normalized convolution reconstruction. Afterwards, the interpolated subbands and the output of the registration technique have been combined by using inverse DT-CWT (IDT-CWT) in order to reconstruct the super resolved frame. For Akiyo video sequence there is 5.04 dB, 4.86 dB, 5.22 dB, and 5.05 dB improvements in the average PSNR values compared to Vandewalle, Marcel, Lucchese, and Keren registration techniques, respectively.
Hasan Demirel, Gholamreza Anbarjafari, Cagri Ozcinar, Sara Izadpanahi
ICIP3
2010 Satellite Image Contrast Enhancement Using Discrete Wavelet Transform and Singular Value Decomposition
abstract
In this letter, a new satellite image contrast enhancement technique based on the discrete wavelet transform (DWT) and singular value decomposition has been proposed. The technique decomposes the input image into the four frequency subbands by using DWT and estimates the singular value matrix of the low-low subband image, and, then, it reconstructs the enhanced image by applying inverse DWT. The technique is compared with conventional image equalization techniques such as standard general histogram equalization and local histogram equalization, as well as state-of-the-art techniques such as brightness preserving dynamic histogram equalization and singular value equalization. The experimental results show the superiority of the proposed method over conventional and state-of-the-art techniques.
Hasan Demirel, Cagri Ozcinar, Gholamreza Anbarjafari
IEEE Geosci. Remote. Sens. Lett.2