VLDB 2026 Research / reviewers in the wild / expert
Aljoscha Smolic
dblp:s/AljoschaSmolic · also Aljosa Smolic
· DBLP profile ↗
149ranked-venue papers
18as first author
26since 2021 · last 2025
0000-0001-7033-3335ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 141 · 15 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 18 · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gaussian Splatting vs. Classical Photogrammetry: A Comparison for Virtual BackdropsabstractThis paper describes and evaluates three end-to-end workflows from capture to reconstruction including photogrammetry, 3D Gaussian splatting (GS), and extraction of textured meshes from splats. The explicit aim is to enable individuals unfamiliar with 3D reconstruction to create backdrops for virtual production with a central presenter by using a smartphone as capture device. We show that GS significantly outperforms the alternatives in our experiments. Philipp Haslbauer, Mike Pullen, Carolin Reichherzer, Aljoscha Smolic |
QoMEX | 4 |
| 2024 | A Volumetric Video Application to Enhance Museum ExperiencesabstractVolumetric video (VV) is an emerging 3D format that allows the integration of real people into XR (extended reality) applications. Recent cost-effective AI-based methods have enabled VV capture using single handheld cameras or mobile phones. This study addresses the quality, integration, and acceptance of AI-based VV content creation in an augmented reality (AR) application designed to enhance museum experiences. The main result reveals that, although the current VV quality is lower than professional standards, users still find significant added value and enjoy its immersive experience. Anh Nguyen 0004, Jan Lässig, Anna F. Kälin, Ariana Huwiler, Philipp Haslbauer, Gareth W. Young, Kit-Yung Lam, Aljoscha Smolic |
VRST | 8 |
| 2024 | Exploring the impact of volumetric graphics on the engagement of broadcast media professionals
Gareth W. Young, Grace Dinan, Aljoscha Smolic, Jan Ondrej, Rafael Pagés |
Multim. Syst. | 3 |
| 2024 | BASICS: Broad Quality Assessment of Static Point Clouds in a Compression ScenarioabstractPoint clouds have become increasingly prevalent in representing 3D scenes within virtual environments, alongside 3D meshes. Their ease of capture has facilitated a wide array of applications on mobile devices, from smartphones to autonomous vehicles. Notably, point cloud compression has reached an advanced stage and has been standardized. However, the availability of quality assessment datasets, which are essential for developing improved objective quality metrics, remains limited. In this paper, we introduce BASICS, a large-scale quality assessment dataset tailored for static point clouds. The BASICS dataset comprises 75 unique point clouds, each compressed with four different algorithms including a learning-based method, resulting in the evaluation of nearly 1500 point clouds by 3500 unique participants. Furthermore, we conduct a comprehensive analysis of the gathered data, benchmark existing point cloud quality assessment metrics and identify their limitations. By publicly releasing the BASICS dataset, we lay the foundation for addressing these limitations and fostering the development of more precise quality metrics. Ali Ak, Emin Zerman, Maurice Quach, Aladine Chetouani, Aljoscha Smolic, Giuseppe Valenzise, Patrick Le Callet |
IEEE Trans. Multim. | 5 |
| 2024 | Blind Image Quality Assessment via Transformer Predicted Error Map and Perceptual Quality TokenabstractImage quality assessment is a fundamental problem in the field of image processing, and due to the lack of reference images in most practical scenarios, no-reference image quality assessment (NR-IQA), has gained increasing attention recently. With the development of deep learning technology, many deep neural network-based NR-IQA methods have been developed, which try to learn the image quality based on the understanding of database information. Currently, Transformer has achieved remarkable progress in various vision tasks. Since the characteristics of the attention mechanism in Transformer fit the global perceptual impact of artifacts perceived by a human, Transformer is thus well suited for image quality assessment tasks. In this paper, we propose a Transformer based NR-IQA model using a predicted objective error map and perceptual quality token. Specifically, we firstly generate the predicted error map by pre-training one model consisting of a Transformer encoder and decoder, in which the objective difference between the distorted and the reference images is used as supervision. Then, we freeze the parameters of the pre-trained model and design another branch using the vision Transformer to extract the perceptual quality token for feature fusion with the predicted error map. Finally, the fused features are regressed to the final image quality score. Extensive experiments have shown that our proposed method outperforms the state-of-the-art methods in both authentic and synthetic image datasets. Moreover, the attentional map extracted by the perceptual quality token also does conform to the characteristics of the human visual system. Jinsong Shi 0002, Pan Gao 0001, Aljoscha Smolic |
IEEE Trans. Multim. | 3 |
| 2023 | StylePrompter: All Styles Need Is AttentionabstractGAN inversion aims at inverting given images into corresponding latent codes for Generative Adversarial Networks (GANs), especially StyleGAN where exists a disentangled latent space that allows attribute-based image manipulation. As most inversion methods build upon Convolutional Neural Networks (CNNs), we transfer a hierarchical vision Transformer backbone innovatively to predict W+ latent codes at token level. We further apply a Style-driven Multi-scale Adaptive Refinement Transformer (SMART) in ℱ space to refine the intermediate style features of the generator. By treating style features as queries to retrieve lost identity information from the encoder's feature maps, SMART can not only produce high-quality inverted images but also surprisingly adapt to editing tasks. We then prove that StylePrompter lies in a more disentangled W+ and show the controllability of SMART. Finally, quantitative and qualitative experiments demonstrate that Style Prompter can achieve desirable performance in balancing reconstruction quality and editability, and is "smart" enough to fit into most edits, outperforming other ℱ -involved inversion methods. Our code is available at: https://github.com/I2-Multimedia-Lab/StylePrompter. Chenyi Zhuang, Pan Gao 0001, Aljoscha Smolic |
ACM Multimedia | 3 |
| 2022 | Image Aesthetics Assessment Using Graph Attention NetworkabstractAspect ratio and spatial layout are two of the principal factors influencing the aesthetic value of a photograph. However, incorporating these into the traditional convolution-based frameworks for the task of image aesthetics assessment is problematic. The aspect ratio of the photographs gets distorted while they are resized/cropped to a fixed dimension to facilitate training batch sampling. On the other hand, the convolutional filters process information locally and are limited in their ability to model the global spatial layout of a photograph. In this work, we present a two-stage framework based on graph neural networks and address both these problems jointly. First, we propose a feature-graph representation in which the input image is modelled as a graph, maintaining its original aspect ratio and resolution. Second, we propose a graph neural network architecture that takes this feature-graph and captures the semantic relationship between different regions of the input image using visual attention. Our experiments show that the proposed framework advances the state-of-the-art results in aesthetic score regression on the Aesthetic Visual Analysis (AVA) benchmark. Our code is publicly available for comparisons and further explorations.1 Koustav Ghosal, Aljoscha Smolic |
ICPR | 2 |
| 2022 | Jointformer: Single-Frame Lifting Transformer with Error Prediction and Refinement for 3D Human Pose EstimationabstractMonocular 3D human pose estimation technologies have the potential to greatly increase the availability of human movement data. The best-performing models for single-image 2D-3D lifting use graph convolutional networks (GCNs) that typically require some manual input to define the relationships between different body joints. We propose a novel transformer-based approach that uses the more generalised self-attention mechanism to learn these relationships within a sequence of tokens representing joints. We find that the use of intermediate supervision, as well as residual connections between the stacked encoders benefits performance. We also suggest that using error prediction as part of a multi-task learning framework improves performance by allowing the network to compensate for its confidence level. We perform extensive ablation studies to show that each of our contributions increases performance. Furthermore, we show that our approach outperforms the recent state of the art for single-frame 3D human pose estimation by a large margin. Our code and trained models are made publicly available on Github. Sebastian Lutz, Richard Blythman, Koustav Ghosal, Matthew Moynihan, Ciaran Simms, Aljoscha Smolic |
ICPR | 6 |
| 2022 | Privacy-Preserving Viewport Prediction using Federated Learning for 360° Live Video StreamingabstractPredicting the user's viewport scanpath is an essential task for 360° viewport-based adaptive streaming. It informs the system which parts of content should be streamed with high quality for bandwidth saving over the best-effort Internet. However, in light of growing privacy concerns among consumers and increasingly strict data privacy legislation, user data collection and storage have been constrained. This paper proposes a novel privacy-preserving framework employing Federated Learning (FL) for online viewport prediction in a live 360° video streaming scenario. In this framework, the user data is only collected and processed on the client-side in the current viewing session and not shared with external parties, e.g., servers, and other clients. We evaluated the framework over a widely-used dataset and measure the computation and transmission time of the proposed streaming system. The experiments show that our framework provides high prediction accuracy and achieves real-time computation requirements of live video streaming. On privacy preservation, our results demonstrate that in a tile-based 360° video streaming system, the user identification rate can be decreased by 18.11 percentage points in 4 x 3 tiles per frame and 9.65 percentage points in 16x9 tiles per frame. The code will be publicly available to further contribute to the community. Fang-Yi Chao, Cagri Ozcinar, Aljoscha Smolic |
MMSP | 3 |
| 2022 | A virtual reality volumetric music video: featuring new pagansabstractMusic videos are short films that integrate songs and imagery produced for artistic and promotional purposes. Modern music videos apply various media capture techniques and creative postproduction technologies to provide a myriad of stimulating and artistic approaches to audience entertainment and engagement for viewing across multiple devices. Within this domain, volumetric video capture technologies (Figure 2) have become an emerging means of recording and reproducing musical performances for new audiences to access via traditional 2D screens and emergent extended reality platforms, such as augmented and virtual reality. These 3D digital reproductions of musical performances are captured live and are enhanced to deliver cutting-edge audiovisual entertainment (Figure 1). However, the precise impact of volumetric video in music video entertainment is still in a state of flux. Gareth W. Young, Néill O'Dwyer, Aljoscha Smolic |
MMSys | 3 |
| 2022 | Extended Reality Ulysses DemoabstractThis demo paper proposes to exhibit the pilot episode of XR Ulysses, a creative project investigating the possibilities for live performance using three-dimensional volumetric video (VV) techniques via virtual reality (VR) technologies. XR Ulysses is part of a series of innovative performance experiments hybridizing theatre and extended reality (XR) technologies. Conference attendees are invited to don an HMD, embody the character of Stephen Dedalus, and engage Buck Mulligan in the famous opening scene of Joyce’s book, situated on the top of the Martello Tower at Sandycove (Dublin). This scene enables individuals to experience a live-action re-enactment of James Joyce’s Ulysses in VR. Néill O'Dwyer, Gareth W. Young, Aljoscha Smolic |
IMX | 3 |
| 2022 | Audience Experiences of a Volumetric Virtual Reality Music VideoabstractMusic videos are short films that integrate songs and imagery and are produced for artistic and promotional purposes. Modern music videos apply various media capture techniques and creative postproduction technologies to provide a myriad of stimulating and artistic approaches to audience entertainment and engagement for viewing across multiple devices. Within this domain, volumetric technologies are becoming a popular means of recording and reproducing musical performances for new audiences to access via traditional 2D screens and emergent virtual reality platforms. However, the precise impact of volumetric video in virtual reality music video entertainment has yet to be fully explored from a user’s perspective. Here we show how users responded to volumetric representations of music performance in virtual reality. Our results preliminarily demonstrate how audiences are likely to respond to music videos and offer insight into how future music videos may be developed for different user types. We anticipate our essay as a formative starting point for more sophisticated, interactive music videos that can be accessed and presented via extended-reality technologies. Gareth W. Young, Néill O'Dwyer, Matthew Moynihan, Aljoscha Smolic |
VR | 4 |
| 2022 | Exploring virtual reality for quality immersive empathy building experiencesabstractVirtual reality (VR) technology presents users with virtual environments to experience various interactive, immersive, and imaginary experiences. While traditional perspective-taking exercises rely on the participant to imagine a self-other merging process to feel connected with other people (typically using second and third-person narrative perspectives), VR can allow an individual to embody an other through first-person narratives delivered via multimodal – visual, aural, haptic – technology-mediated experiences. This process enables users to perceptually and effectively portal into somebody else's body, where they can potentially see, hear, and feel from the point of view of the protagonist and control choices on their behalf in real-time. This article explores the use of VR as an ‘empathy-making machine’ by facilitating perspective-taking and allowing users to experience another person's circumstances. An experiment was performed to compare two different types of perspective-taking VR applications. Levels of empathy, oneness, and attitudes towards a protagonist or focus group within VR materials were captured. Participants then identified the elements of the VR content that contributed to a quality experience. These measures were used to discuss methodologies and techniques for creating quality empathy-building techniques. The findings of this research will be used to inform future creative technology projects presented in VR. Gareth W. Young, Néill O'Dwyer, Aljoscha Smolic |
Behav. Inf. Technol. | 3 |
| 2022 | Spectral analysis of re-parameterized light fields
Martin Alain, Aljoscha Smolic |
Signal Process. Image Commun. | 2 |
| 2022 | Quality Assessment for Omnidirectional Video: A Spatio-Temporal Distortion Modeling ApproachabstractOmnidirectional video, also known as 360-degree video, has become increasingly popular nowadays due to its ability to provide immersive and interactive visual experiences. However, the ultra high resolution and the spherical observation space brought by the large spherical viewing range make omnidirectional video distinctly different from traditional 2D video. To date, the video quality assessment (VQA) for omnidirectional video is still an open issue. The existing VQA metrics for omnidirectional video only consider the spatial characteristics of distortions, but the temporal change of spatial distortions can also considerably influence human visual perception. In this paper, we propose a spatiotemporal modeling approach to evaluate the quality of the omnidirectional video. Firstly, we construct a spatioral quality assessment unit to evaluate the average distortion in temporal dimension at the eye fixation level, based upon which the smoothed distortion value is recursively calculated and consolidated by the characteristics of temporal variations. Then, we give a detailed solution of how to to integrate the three existing spatial VQA metrics into our approach. Besides, the cross-format omnidirectional video distortion measurement is also investigated. Finally, the spatiotemporal distortion of the whole video sequence is obtained by pooling. Based on the modeling approach, a full reference objective quality assessment metric for omnidirectional video is derived, namely OV-PSNR. The experimental results show that our proposed OV-PSNR greatly improves the prediction performance of the existing VQA metrics for omnidirectional video. Pan Gao 0001, Aljoscha Smolic |
IEEE Trans. Multim. | 3 |
| 2021 | TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding
Qi She, Aljoscha Smolic |
BMVC | 3 |
| 2021 | ACTION-Net: Multipath Excitation for Action RecognitionabstractSpatial-temporal, channel-wise, and motion patterns are three complementary and crucial types of information for video action recognition. Conventional 2D CNNs are computationally cheap but cannot catch temporal relationships; 3D CNNs can achieve good performance but are computationally intensive. In this work, we tackle this dilemma by designing a generic and effective module that can be embedded into 2D CNNs. To this end, we propose a spAtio-temporal, Channel and moTion excitatION (ACTION) module consisting of three paths: Spatio-Temporal Excitation (STE) path, Channel Excitation (CE) path, and Motion Excitation (ME) path. The STE path employs one channel 3D convolution to characterize spatio-temporal representation. The CE path adaptively recalibrates channel-wise feature responses by explicitly modeling interdependencies between channels in terms of the temporal aspect. The ME path calculates feature-level temporal differences, which is then utilized to excite motion-sensitive channels. We equip 2D CNNs with the proposed ACTION module to form a simple yet effective ACTION-Net with very limited extra computational cost. ACTION-Net is demonstrated by consistently outperforming 2D CNN counterparts on three backbones (i.e., ResNet-50, MobileNet V2 and BNInception) employing three datasets (i.e., Something-Something V2, Jester, and EgoGesture). Code is provided at https://github.com/V-Sense/ACTION-Net. Qi She, Aljoscha Smolic |
CVPR | 3 |
| 2021 | Light Field Style Transfer with Local Angular ConsistencyabstractStyle transfer involves combining the style of one image with the content of another to form a new image. Unlike traditional two-dimensional images which only capture the spatial intensity of light rays, four-dimensional light fields also capture the angular direction of the light rays. Thus, applying style transfer to a light field requires to not only render convincing style transfer for each view, but also to preserve its angular structure. In this paper, we present a novel optimization-based method for light field style transfer which iteratively propagates the style from the centre view towards the outer views while enforcing local angular consistency. For this purpose, a new initialisation method and angular loss function is proposed for the optimization process. In addition, since style transfer for light field is an emerging topic, no clear evaluation procedure is available. Thus, we investigate the use of a recently proposed metric designed to evaluate light field angular consistency, as well as a proposed variant. Dónal Egan, Martin Alain, Aljoscha Smolic |
ICASSP | 3 |
| 2021 | Deep Color Mismatch Correction In Stereoscopic 3d ImagesabstractColor mismatch in stereoscopic 3D (S3D) images can create visual discomfort and affect the performance of S3D image processing algorithms, e.g., for depth estimation. In this paper, we propose a new deep learning-based solution for the problem of color mismatch correction. The proposed solution consists of a multi-task convolutional neural network, where color correction is the primary task and correspondence estimation is the secondary task. For the training and evaluation of the proposed network, a new S3D image dataset with color mismatch was created. Based on this dataset, experiments were conducted showing the effectiveness of our solution. Simone Croci, Cagri Ozcinar, Emin Zerman, Roman Dudek, Sebastian Knorr, Aljoscha Smolic |
ICIP | 6 |
| 2021 | Transformer-based Long-Term Viewport Prediction in 360° Video: Scanpath is All You NeedabstractVirtual Reality (VR) multimedia technology has dramatically advanced in recent years. Its immersive and interactive natures enable users to view any direction in 360° content freely. Users do not see the entire 360° content at a glance, but only a portion in the viewport. Viewport-based adaptive streaming, which streams only the user’s viewport of interest with high quality, has emerged as the primary technique to save bandwidth over the best-effort Internet. Thus, users’ viewport prediction in the forthcoming seconds becomes an essential task for informing the streaming decisions in the VR system. Various viewport prediction methods based on deep neural networks have been proposed. However, typically they are composed of complex Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) that require heavy computation. To achieve high prediction accuracy in limited computation time in a streaming system, we propose a new transformer-based architecture, named 360° Viewport Prediction Transformer (VPT360), that only leverages the past viewport scanpath to predict a user’s future viewport scanpath. We evaluate VPT360 over three widely-used datasets and compare the computation complexity with the state-of-the-art methods. The experiments show that our VPT360 provides the highest accuracy for short-term and long-term prediction and achieves the lowest computation complexity. The code is publicly available at https://github.com/FannyChao/VPT360 to further contribute to the community. Fang-Yi Chao, Cagri Ozcinar, Aljoscha Smolic |
MMSP | 3 |
| 2021 | Light Field Visual Attention Prediction Using Fourier Disparity LayersabstractIn this paper, we present a novel layered saliency model for Light Fields (LF) called Fourier Disparity Layer Saliency Estimation (FDLSE). The layers are constructed from the existing Fourier Disparity Layer (FDL) LF representation. Our FDLSE model can be used to predict the visual attention (VA) of any LF rendering with arbitrary viewpoint, aperture and depth-of-focus, without the need to generate the rendered image itself. The proposed model surpasses previous work in the following areas. Our method requires the estimation of the saliency map of only one sub-aperture image instead of the full view array. Furthermore, this model does not require pre-estimated disparity maps, but instead relies on the FDL model whose computation fully takes advantage of GPU parallelisation. Finally, FDLSE shows visual improvements and performs quantitatively on par with our previous FGSE model when evaluated on VA prediction of refocus renderings. To our knowledge these are the only two models which can be used to predict LF VA. Ailbhe Gill, Mikael Le Pendu, Martin Alain, Emin Zerman, Aljoscha Smolic |
MMSP | 5 |
| 2021 | The Effect of Temporal Sub-sampling on the Accuracy of Volumetric Video Quality AssessmentabstractVolumetric video content has attracted increasing research interests over the last decade, as it facilitates the integration of dynamic real world content in virtual environments. Point cloud is one of the most common alternatives to represent volumetric video content. Yet, such representation requires an enormous data storage and pose significant greater pressures on compression algorithms compared to the standard 2D video. This challenge has unleashed a new wave in the development of novel point cloud compression technologies, which need to be evaluated in terms of production quality. Due to the high dimensionality of the data, evaluating the performances of relevant coding algorithms can be time consuming. This puts a barrier on optimizing coding algorithms with complex, but perceptually accurate, objective quality metrics. In this study, we thus explore the possibility of reducing temporal-dimension of the content under-evaluation, i.e., temporal sub-sampling, for objective quality evaluation without sacrificing from the correlation with the subjective opinion. In addition, we exploit different temporal pooling methods to further make the quality evaluation procedure more efficient. In total 30 different objective quality metrics were tested on the the V-SENSE volumetric video quality database. According to experimental results, there is no need to employ full frame-rate (30 fps) assessment to reach the meaningful correlation for the considered quality metrics. These observations could be referred to reduce the computation complexity regarding the evaluation and optimization of the relevant compression algorithms. Ali Ak, Emin Zerman, Suiyi Ling, Patrick Le Callet, Aljoscha Smolic |
PCS | 5 |
| 2021 | Focus Guided Light Field Saliency EstimationabstractLight field imaging enables us to capture all light rays in a visual scene. As light fields are four-dimensional, their captures come with an increased amount of information to take advantage of. This has stimulated ongoing light field specific research into virtual viewpoints and shallow depth of field rendering, commonly called refocusing. However, the computation time and memory required to perform these operations can make tasks such as real-time rendering impractical. One solution is to exploit the salient information of light fields to focus resources on regions that attract visual attention when using these algorithms. Although saliency estimation methods for light fields have been previously explored, these focus mainly on salient object segmentation with the goal of generating one saliency map per light field.Aiming to create a basis for a 4D saliency prediction model analogous to light fields, this paper proposes a saliency estimation method specific to light fields that considers the refocusing operation. The proposed method modifies an existing view rendering algorithm with focus guidance, obtained from the light field disparity. This facilitates the construction of saliency maps without the need to render the corresponding view itself, which will help to speed up processing operations that are compatible. The results show that the proposed saliency estimation approach yields very good predictions of visual attention across multiple planes of the light field. We anticipate that this approach can be extended for a range of rendering applications. Ailbhe Gill, Emin Zerman, Martin Alain, Mikael Le Pendu, Aljoscha Smolic |
QoMEX | 5 |
| 2021 | User Behaviour Analysis of Volumetric Video in Augmented RealityabstractAugmented reality (AR) is getting popular, and among other content creation techniques, volumetric video allows to bring dynamic real world content as captured by cameras into such applications. To develop efficient algorithms for compression and transmission of the volumetric media, it is important to understand how users will consume this new form of dynamic 3D content. In this paper, we analyse the user behaviour for the volumetric video consumption in AR. In particular, we study the distribution of users' viewpoints, relative locations, and average distances from the content. For this purpose, we built an Android AR application using the volumetric video and conducted a user study remotely. The results show that users spent most of their time looking at the frontal part of the volumetric video, and this indicates the importance of the face in visual attention. The collected user behaviour data are made public to support further research. Emin Zerman, Radhika Kulkarni, Aljoscha Smolic |
QoMEX | 3 |
| 2021 | Foreground color prediction through inverse compositingabstractIn natural image matting, the goal is to estimate the opacity of the foreground object in the image. This opacity controls the way the foreground and background is blended in transparent regions. In recent years, advances in deep learning have led to many natural image matting algorithms that have achieved outstanding performance in a fully automatic manner. However, most of these algorithms only predict the alpha matte from the image, which is not sufficient to create high-quality compositions. Further, it is not possible to manually interact with these algorithms in any way except by directly changing their input or output. We propose a novel recurrent neural network that can be used as a post-processing method to recover the foreground and background colors of an image, given an initial alpha estimation. Our method outperforms the state-of-the-art in color estimation for natural image matting and show that the recurrent nature of our method allows users to easily change candidate solutions that lead to superior color estimations. Sebastian Lutz, Aljoscha Smolic |
WACV | 2 |
| 2021 | Autonomous Tracking For Volumetric Video SequencesabstractAs a rapidly growing medium, volumetric video is gaining attention beyond academia, reaching industry and creative communities alike. This brings new challenges to reduce the barrier to entry from a technical and economical point of view. We present a system for robustly and autonomously performing temporally coherent tracking for volumetric sequences, specifically targeting those from sparse setups or with noisy output. Our system will detect and recover missing pertinent geometry across highly incoherent sequences as well as provide users the option of propagating drastic topology edits. In this way, affordable multi-view setups can leverage temporal consistency to reduce processing and compression overheads while also generating more aesthetically pleasing volumetric sequences. Matthew Moynihan, Susana Ruano, Rafael Pagés, Aljoscha Smolic |
WACV | 4 |
| 2020 | High Resolution Light Field Recovery with Fourier Disparity Layer Completion, Demosaicing, and Super-ResolutionabstractIn this paper, we present a novel approach for recovering high resolution light fields from input data with many types of degradation and challenges typically found in lenslet based plenoptic cameras. Those include the low spatial resolution, but also the irregular spatio-angular sampling and color sampling, the depth-dependent blur, and even axial chromatic aberrations. Our approach, based on the recent Fourier Disparity Layer representation of the light field, allows the construction of high resolution layers directly from the low resolution input views. High resolution light field views are then simply reconstructed by shifting and summing the layers. We show that when the spatial sampling is regular, the layer construction can be decomposed into linear optimization problems formulated in the Fourier domain for small groups of frequency components. We additionally propose a new preconditioning approach ensuring spatial consistency, and a color regularization term to simultaneously perform color demosaicing. For the general case of light field completion from an irregular sampling, we define a simple iterative version of the algorithm. Both approaches are then combined for an efficient super-resolution of the irregularly sampled data of plenoptic cameras. Finally, the Fourier Disparity Layer model naturally extends to take into account a depth-dependent blur and axial chromatic aberrations without requiring an estimation of depth or disparity maps. Mikael Le Pendu, Aljoscha Smolic |
ICCP | 2 |
| 2020 | A Spatio-Angular Binary Descriptor For Fast Light Field Inter View MatchingabstractLight fields are able to capture light rays from a scene arriving at different angles, effectively creating multiple perspective views of the same scene. Thus, one of the flagship applications of light fields is to estimate the captured scene geometry, which can notably be achieved by establishing correspondences between the perspective views, usually in the form of a disparity map. Such correspondence estimation has been a long standing research topic in computer vision, with application to stereo vision or optical flow. Research in this area has shown the importance of well designed descriptors to enable fast and accurate matching. We propose in this paper a binary descriptor exploiting the light field gradient over both the spatial and the angular dimensions in order to improve inter view matching. We demonstrate in a disparity estimation application that it can achieve comparable accuracy compared to existing descriptors while being faster to compute. Martin Alain, Aljoscha Smolic |
ICIP | 2 |
| 2020 | Soft Colour Segmentation On Light FieldsabstractIn this work we explore methods for allowing advanced colour editing on light field images to be performed. This investigation is twofold. First we look at soft colour algorithms to decompose images into colour layers and the various ways it could be applied to light field data in order to ensure spatially consistent results. Then, with the purpose of colour editing in mind, we present an object-based layer separation method so that editing a layer does not wrongly affect specific objects. We further discuss the advantages and drawbacks that light field data present over regular single-view images for this purpose. Finally we present some editing results to show that our methods allow us to obtain visually appealing images that remain consistent across all light field views and minimise the colour artefacts inherent to layer decomposition methods. Pierre Matysiak, Mairéad Grogan, Weston Aenchbacher, Aljoscha Smolic |
ICIP | 4 |
| 2020 | Hierarchical Fourier Disparity Layer Transmission For Light Field StreamingabstractIn this paper, we present a novel approach to efficiently transmit light fields in the Fourier Disparity Layer (FDL) representation using a binary hierarchical scheme. The FDL model consists of a set of additive layers which can be simply shifted and summed to render a view of the light field at any angular coordinate. In order to transmit the FDL model, we propose a method for building a binary tree where the root consists of a single compound layer obtained as the sum of all the original layers. Subsequent levels of the tree are obtained by splitting a parent layer into two children layers whose sum is equal to the parent layer. Hence, the FDL model is recursively refined with additional layers at each new level, resulting in a scalable representation. An efficient scheme is proposed to encode a single image in order to split a parent layer into its two children. Thanks to this approach, the total number of images to decode for receiving the complete tree is equal to the number of layers in the original FDL model, which is typically smaller than the number of views required in the traditional light field representation. Mikael Le Pendu, Cagri Ozcinar, Aljoscha Smolic |
ICIP | 3 |
| 2020 | Developing a Model Augmented Reality CurriculumabstractThis paper outlines the objectives of the working group on developing a model Augmented Reality curriculum for higher education. We motivate the need for the model curriculum by the growing Augmented Reality industry and subsequent demand for trained professionals. While the industry is growing, the educational offers that train the required skills remain limited and fragmented. The working group will address this challenge by surveying the state of the art in Augmented Reality education are reviewing available data on industry requirements. Based on the results, the group will develop a new model Augmented Reality curriculum. The working group will also develop future work recommendations for the design of teaching materials and integration of Augmented Reality in computing curricula. Mikhail Fominykh, Fridolin Wild, Ralf Klamma, Mark Billinghurst, Lisandra S. Costiner, Andrey Karsakov, Eleni E. Mangina, Judith Molka-Danielsen, Ian Pollock, Marius Preda, Aljoscha Smolic |
ITiCSE | 11 |
| 2020 | Self-supervised Light Field View Synthesis Using Cycle ConsistencyabstractHigh angular resolution is advantageous for practical applications of light fields. In order to enhance the angular resolution of light fields, view synthesis methods can be utilized to generate dense intermediate views from sparse light field input. Most successful view synthesis methods are learning-based approaches which require a large amount of training data paired with ground truth. However, collecting such large datasets for light fields is challenging compared to natural images or videos. To tackle this problem, we propose a self-supervised light field view synthesis framework with cycle consistency. The proposed method aims to transfer prior knowledge learned from high quality natural video datasets to the light field view synthesis task, which reduces the need for labeled light field data. A cycle consistency constraint is used to build bidirectional mapping enforcing the generated views to be consistent with the input views. Derived from this key concept, two loss functions, cycle loss and reconstruction loss, are used to fine-tune the pre-trained model of a state-of-the-art video interpolation method. The proposed method is evaluated on various datasets to validate its robustness, and results show it not only achieves competitive performance compared to supervised fine-tuning, but also outperforms state-of-the-art light field view synthesis methods, especially when generating multiple intermediate views. Besides, our generic light field view synthesis framework can be adopted to any pre-trained model for advanced video interpolation. Martin Alain, Aljoscha Smolic |
MMSP | 3 |
| 2020 | Textured Mesh vs Coloured Point Cloud: A Subjective Study for Volumetric Video CompressionabstractVolumetric video (VV) pipelines reached a high level of maturity, creating interest to use such content in interactive visualisation scenarios. VV allows real world content to be captured and represented as 3D models, which can be viewed from any chosen viewpoint and direction. Thus, VV is ideal to be used in augmented reality (AR) or virtual reality (VR) applications. Both textured polygonal meshes and point clouds are popular methods to represent VV. Even though the signal and image processing community slightly favours the point cloud due to its simpler data structure and faster acquisition, textured polygonal meshes might have other benefits such as better visual quality and easier integration with computer graphics pipelines. To better understand the difference between them, in this study, we compare these two different representation formats for a VV compression scenario utilising state-of-the-art compression techniques. For this purpose, we build a database and collect user opinion scores for subjective quality assessment of the compressed VV. The results show that meshes provide the best quality at high bitrates, while point clouds perform better for low bitrate cases. The created VV quality database will be made available online to support further scientific studies on VV quality assessment. Emin Zerman, Cagri Ozcinar, Pan Gao 0001, Aljoscha Smolic |
QoMEX | 4 |
| 2020 | Towards Audio-Visual Saliency Prediction for Omnidirectional Video with Spatial AudioabstractOmnidirectional videos (ODVs) with spatial audio enable viewers to perceive 360° directions of audio and visual signals during the consumption of ODVs with head-mounted displays (HMDs). By predicting salient audio-visual regions, ODV systems can be optimized to provide an immersive sensation of audio-visual stimuli with high-quality. Despite the intense recent effort for ODV saliency prediction, the current literature still does not consider the impact of auditory information in ODVs. In this work, we propose an audio-visual saliency (AVS360) model that incorporates 360° spatial-temporal visual representation and spatial auditory information in ODVs. The proposed AVS360 model is composed of two 3D residual networks (ResNets) to encode visual and audio cues. The first one is embedded with a spherical representation technique to extract 360° visual features, and the second one extracts the features of audio using the log mel-spectrogram. We emphasize sound source locations by integrating audio energy map (AEM) generated from spatial audio description (i.e., ambisonics) and equator viewing behavior with equator center bias (ECB). The audio and visual features are combined and fused with AEM and ECB via attention mechanism. Our experimental results show that the AVS360 model has significant superiority over five state-of-the-art saliency models. To the best of our knowledge, it is the first w ork that develops the audio-visual saliency model in ODVs. The code will be publicly available to foster future research on audio-visual saliency in ODVs. Fang-Yi Chao, Cagri Ozcinar, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges, Aljoscha Smolic |
VCIP | 6 |
| 2020 | High Quality Light Field Extraction and Post-Processing for Raw Plenoptic DataabstractLight field technology has reached a certain level of maturity in recent years, and its applications in both computer vision research and industry are offering new perspectives for cinematography and virtual reality. Several methods of capture exist, each with its own advantages and drawbacks. One of these methods involves the use of handheld plenoptic cameras. While these cameras offer freedom and ease of use, they also suffer from various visual artefacts and inconsistencies. We propose in this paper an advanced pipeline that enhances their output. After extracting sub-aperture images from the RAW images with our demultiplexing method, we perform three correction steps. We first remove hot pixel artefacts, then correct colour inconsistencies between views using a colour transfer method, and finally we apply a state of the art light field denoising technique to ensure a high image quality. An in-depth analysis is provided for every step of the pipeline, as well as their interaction within the system. We compare our approach to existing state of the art sub-aperture image extracting algorithms, using a number of metrics as well as a subjective experiment. Finally, we showcase the positive impact of our system on a number of relevant light field applications. Pierre Matysiak, Mairéad Grogan, Mikael Le Pendu, Martin Alain, Emin Zerman, Aljoscha Smolic |
IEEE Trans. Image Process. | 6 |
| 2020 | Deep Tone Mapping Operator for High Dynamic Range ImagesabstractA computationally fast tone mapping operator (TMO) that can quickly adapt to a wide spectrum of high dynamic range (HDR) content is quintessential for visualization on varied low dynamic range (LDR) output devices such as movie screens or standard displays. Existing TMOs can successfully tone-map only a limited number of HDR content and require an extensive parameter tuning to yield the best subjective-quality tone-mapped output. In this paper, we address this problem by proposing a fast, parameter-free and scene-adaptable deep tone mapping operator (DeepTMO) that yields a high-resolution and high-subjective quality tone mapped output. Based on conditional generative adversarial network (cGAN), DeepTMO not only learns to adapt to vast scenic-content (e.g., outdoor, indoor, human, structures, etc.) but also tackles the HDR related scene-specific challenges such as contrast and brightness, while preserving the fine-grained details. We explore 4 possible combinations of Generator-Discriminator architectural designs to specifically address some prominent issues in HDR related deep-learning frameworks like blurring, tiling patterns and saturation artifacts. By exploring different influences of scales, loss-functions and normalization layers under a cGAN setting, we conclude with adopting a multi-scale model for our task. To further leverage on the large-scale availability of unlabeled HDR data, we train our network by generating targets using an objective HDR quality metric, namely Tone Mapping Image Quality Index (TMQI). We demonstrate results both quantitatively and qualitatively, and showcase that our DeepTMO generates high-resolution, high-quality output images over a large spectrum of real-world scenes. Finally, we evaluate the perceived quality of our results by conducting a pair-wise subjective study which confirms the versatility of our method. Aakanksha Rana, Praveer Singh, Giuseppe Valenzise, Frédéric Dufaux, Nikos Komodakis, Aljoscha Smolic |
IEEE Trans. Image Process. | 6 |
| 2020 | Do Users Behave Similarly in VR? Investigation of the User Influence on the System DesignabstractWith the overarching goal of developing user-centric Virtual Reality (VR) systems, a new wave of studies focused on understanding how users interact in VR environments has recently emerged. Despite the intense efforts, however, current literature still does not provide the right framework to fully interpret and predict users’ trajectories while navigating in VR scenes. This work advances the state-of-the-art on both the study of users’ behaviour in VR and the user-centric system design. In more detail, we complement current datasets by presenting a publicly available dataset that provides navigation trajectories acquired for heterogeneous omnidirectional videos and different viewing platforms—namely, head-mounted display, tablet, and laptop. We then present an exhaustive analysis on the collected data to better understand navigation in VR across users, content, and, for the first time, across viewing platforms. The novelty lies in the user-affinity metric, proposed in this work to investigate users’ similarities when navigating within the content. The analysis reveals useful insights on the effect of device and content on the navigation, which could be precious considerations from the system design perspective. As a case study of the importance of studying users’ behaviour when designing VR systems, we finally propose a user-centric server optimisation. We formulate an integer linear program that seeks the best stored set of omnidirectional content that minimises encoding and storage cost while maximising the user’s experience. This is posed while taking into account network dynamics, type of video content, and also user population interactivity. Experimental results prove that our solution outperforms common company recommendations in terms of experienced quality but also in terms of encoding and storage, achieving a savings up to 70%. More importantly, we highlight a strong correlation between the storage cost and the user-affinity metric, showing the impact of the latter in the system architecture design. Silvia Rossi 0001, Cagri Ozcinar, Aljoscha Smolic, Laura Toni |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | DublinCity: Annotated LiDAR Point Cloud and its Applications
S. M. Iman Zolanvari, Susana Ruano, Aakanksha Rana, Alan Cummins, Rogério E. da Silva, Morteza Rahbar, Aljoscha Smolic |
BMVC | 7 |
| 2019 | Towards Generating Ambisonics Using Audio-visual Cue for Virtual RealityabstractAmbisonics i.e., a full-sphere surround sound, is quintessential with 360° visual content to provide a realistic virtual reality (VR) experience. While 360° visual content capture gained a tremendous boost recently, the estimation of corresponding spatial sound is still challenging due to the required sound-field microphones or information about the sound-source locations. In this paper, we introduce a novel problem of generating Ambisonics in 360° videos using the audiovisual cue. With this aim, firstly, a novel 360° audio-visual video dataset of 265 videos is introduced with annotated sound-source locations. Secondly, a pipeline is designed for an automatic Ambisonic estimation problem. Benefiting from the deep learning based audiovisual feature-embedding and prediction modules, our pipeline estimates the 3D sound-source locations and further use such locations to encode to the B-format. To benchmark our dataset and pipeline, we additionally propose evaluation criteria to investigate the performance using different 360° input representations. Our results demonstrate the efficacy of the proposed pipeline and open up a new area of research in 360° audio-visual analysis for future investigations. Aakanksha Rana, Cagri Ozcinar, Aljoscha Smolic |
ICASSP | 3 |
| 2019 | A Study of Light Field Streaming for An Interactive Refocusing ApplicationabstractLight fields are able to capture light rays from a scene arriving at different angles, which allows post-capture rendering applications such as interactive viewpoint selection or refocusing. However, this additional angular information comes at the price of a significant increase of the data volume compared to traditional 2D images. While light field compression is still an ongoing research effort, showing impressive compression gain with the latest coding standard, light fields are in practice often stored on remote servers to avoid consuming unnecessary storage of the user devices. A typical cost-effective solution for light field visualisation is then to render the requested image on the server and transmit the result to the user. Another trivial solution would be to directly send the light field to the user and perform the rendering process directly on the client side to avoid transmission delay. While the latter solution seems instinctively less optimal and is usually discarded in previous work because of an expected unacceptable startup delay, we propose a quantitative study to compare both solutions in terms of rate-distortion (RD) performance. A counterintuitive finding of this paper is that accepting a reasonable startup delay (a few seconds) can provide a significant improvement of the RD performances. Martin Alain, Cagri Ozcinar, Aljoscha Smolic |
ICIP | 3 |
| 2019 | Colornet - Estimating Colorfulness in Natural ImagesabstractMeasuring the colorfulness of a natural or virtual scene is critical for many applications in image processing field ranging from capturing to display. In this paper, we propose the first deep learning-based colorfulness estimation metric. For this purpose, we develop a color rating model which simultaneously learns to extracts the pertinent characteristic color features and the mapping from feature space to the ideal colorfulness scores for a variety of natural colored images. Additionally, we propose to overcome the lack of adequate annotated dataset problem by combining/aligning two publicly available colorfulness databases using the results of a new subjective test which employs a common subset of both databases. Using the obtained subjectively annotated dataset with 180 colored images, we finally demonstrate the efficacy of our proposed model over the traditional methods, both quantitatively and qualitatively. Emin Zerman, Aakanksha Rana, Aljoscha Smolic |
ICIP | 3 |
| 2019 | Super-resolution of Omnidirectional Images Using Adversarial LearningabstractAn omnidirectional image (ODI) enables viewers to look in every direction from a fixed point through a head-mounted display providing an immersive experience compared to that of a standard image. Designing immersive virtual reality systems with ODIs is challenging as they require high resolution content. In this paper, we study super-resolution for ODIs and propose an improved generative adversarial network based model which is optimized to handle the artifacts obtained in the spherical observational space. Specifically, we propose to use a fast PatchGAN discriminator, as it needs fewer parameters and improves the super-resolution at a fine scale. We also explore the generative models with adversarial learning by introducing a spherical-content specific loss function, called 360-SS. To train and test the performance of our proposed model we prepare a dataset of 4500 ODIs. Our results demonstrate the efficacy of the proposed method and identify new challenges in ODI super-resolution for future investigations. Cagri Ozcinar, Aakanksha Rana, Aljoscha Smolic |
MMSP | 3 |
| 2019 | Study on the Perception of Sharpness Mismatch in Stereoscopic VideoabstractIn this paper, we study an artifact of stereoscopic 3D (S3D) video called sharpness mismatch (SM), that occurs when one view is more blurred than the other. SM beyond a certain level can create visual discomfort, and consequently degrade the quality of experience. Therefore, it is important to measure the just noticeable sharpness mismatch (JNSM), i.e., the minimal level of SM that is perceived by the human visual system and creates discomfort. The knowledge of the JNSM can be used in the evaluation of the quality of S3D video, and more in general when processing S3D video, like in asymmetric compression. In this paper, we focus in particular on the detection of SM. For this goal, we organized a psychophysical experiment with 23 subjects and a crosstalk-free stereoscopic display in order to gather psychophysical data necessary for the development of a SM detection method. Based on the gathered experiment data, we propose a new SM detection method. The evaluation of this method shows that its performance is close but not better than that of the state-of-the-art methods. Therefore, our goal in the near future is to improve the proposed method. Simone Croci, Sebastian Knorr, Aljoscha Smolic |
QoMEX | 3 |
| 2019 | Voronoi-based Objective Quality Metrics for Omnidirectional VideoabstractOmnidirectional video (ODV) represents one of the latest and most promising trends in immersive media. The success of ODV depends on the ability to deliver high-quality ODV to the viewers. For this reason, new methods are needed to measure ODV quality that takes into account the interactive look around nature and the spherical representation of ODV. In this paper, we study full-reference objective quality metrics for ODV based on typical encoding distortions in adaptive streaming systems, namely, scaling and compression. The contribution of this paper is three-fold. First, we propose new objective metrics that take into account the unique aspects of ODV. The proposed metrics are based on the subdivision of a given ODV into multiple patches using the spherical Voronoi diagram. Second, we introduce a new dataset of 75 impaired ODVs with different resolutions and compression levels, together with the subjective quality scores gathered during an experiment with 21 participants. Third, we evaluate the proposed Voronoi-based objective metrics using our dataset. The evaluation of the proposed objective metrics and the comparison with existing metrics show that the proposed metrics achieve a better correlation with the subjective scores. The ODV dataset together with the subjective quality scores and the code of the proposed quality metrics are available with this paper. Simone Croci, Cagri Ozcinar, Emin Zerman, Julián Cabrera, Aljoscha Smolic |
QoMEX | 5 |
| 2019 | Analysing the Impact of Cross-Content Pairs on Pairwise Comparison ScalingabstractPairwise comparisons (PWC) methodology is one of the most commonly used methods for subjective quality assessment, especially for computer graphics and multimedia applications. Unlike rating methods, a psychometric scaling operation is required to convert PWC results to numerical subjective quality values. Due to the nature of this scaling operation, the obtained quality scores are relative to the set they are computed in. While it is customary to compare different versions of the same content, in this work we study how cross-content comparisons may benefit psychometric scaling. For this purpose, we use two different video quality databases which have both rating and PWC experiment results. The results show that despite same-content comparisons play a major role in the accuracy of psychometric scaling, the use of a small portion of cross-content comparison pairs is indeed beneficial to obtain more accurate quality estimates. Emin Zerman, Giuseppe Valenzise, Aljoscha Smolic |
QoMEX | 3 |
| 2019 | Robust global and local color matching in stereoscopic omnidirectional content
Roman Dudek, Simone Croci, Aljoscha Smolic, Sebastian Knorr |
Signal Process. Image Commun. | 3 |
| 2019 | Occlusion-Aware Depth Map Coding Optimization Using Allowable Depth Map DistortionsabstractIn depth map coding, rate-distortion optimization for those pixels that will cause occlusion in view synthesis is a rather challenging task, since the synthesis distortion estimation is complicated by the warping competition and the occlusion order can be easily changed by the adopted optimization strategy. In this paper, an efficient depth map coding approach using allowable depth map distortions is proposed for occlusion-inducing pixels. First, we derive the range of allowable depth level change for both the zero disparity error case and non-zero disparity error case with theoretic and geometrical proofs. Then, we formulate the problem of optimally selecting the depth distortion within allowable depth distortion range with the objective to minimize the overall synthesis distortion involved in the occlusion. The unicity and occlusion order invariance properties of allowable depth distortion range is demonstrated. Finally, we propose a dynamic programming based algorithm to locate the optimal depth distortion for each pixel. Simulation results illustrate the performance improvement of the proposed algorithm over the other state-of-the-art depth map coding optimization schemes. Pan Gao 0001, Aljoscha Smolic |
IEEE Trans. Image Process. | 2 |
| 2019 | A Fourier Disparity Layer Representation for Light FieldsabstractIn this paper, we present a new Light Field representation for efficient Light Field processing and rendering called Fourier Disparity Layers (FDL). The proposed FDL representation samples the Light Field in the depth (or equivalently the disparity) dimension by decomposing the scene as a discrete sum of layers. The layers can be constructed from various types of Light Field inputs, including a set of sub-aperture images, a focal stack, or even a combination of both. From our derivations in the Fourier domain, the layers are simply obtained by a regularized least square regression performed independently at each spatial frequency, which is efficiently parallelized in a GPU implementation. Our model is also used to derive a gradient descent-based calibration step that estimates the input view positions and an optimal set of disparity values required for the layer construction. Once the layers are known, they can be simply shifted and filtered to produce different viewpoints of the scene while controlling the focus and simulating a camera aperture of arbitrary shape and size. Our implementation in the Fourier domain allows real-time Light Field rendering. Finally, direct applications such as view interpolation or extrapolation and denoising are presented and evaluated. Mikael Le Pendu, Christine Guillemot, Aljoscha Smolic |
IEEE Trans. Image Process. | 3 |
| 2018 | AlphaGAN: Generative adversarial networks for natural image matting
Sebastian Lutz, Konstantinos Amplianitis, Aljoscha Smolic |
BMVC | 3 |
| 2018 | Faoladh: A Case Study in Cinematic VR Storytelling and Production
Declan Dowling, Colm O. Fearghail, Aljoscha Smolic, Sebastian Knorr |
ICIDS | 3 |
| 2018 | Director's Cut - Analysis of Aspects of Interactive Storytelling for VR Films
Colm O. Fearghail, Cagri Ozcinar, Sebastian Knorr, Aljoscha Smolic |
ICIDS | 4 |
| 2018 | Jonathan Swift: Augmented Reality Application for Trinity Library's Long Room
Néill O'Dwyer, Jan Ondrej, Rafael Pagés, Konstantinos Amplianitis, Aljoscha Smolic |
ICIDS | 5 |
| 2018 | Light Field Super-Resolution via LFBM5D Sparse CodingabstractIn this paper, we propose a spatial super-resolution method for light fields, which combines the SR-BM3D single image super-resolution filter and the recently introduced LFBM5D light field denoising filter. The proposed algorithm iteratively alternates between an LFBM5D filtering step and a back-projection step. The LFBM5D filter creates disparity compensated 4D patches which are then stacked together with similar 4D patches along a 5thdimension. The 5D patches are then filtered in the 5D transform domain to enforce a sparse coding of the high-resolution light field, which is a powerful prior to solve the ill-posed super-resolution problem. The back-projection step then impose the consistency between the known low-resolution light field and the-high resolution estimate. We further improve this step by using image guided filtering to remove ringing artifacts. Results show that significant improvement can be achieved compared to state-of-the-art methods, for both light fields captured with a lenslet camera or a gantry. Martin Alain, Aljoscha Smolic |
ICIP | 2 |
| 2018 | Sharpness Mismatch Detection in Stereoscopic Content with 360-Degree CapabilityabstractThis paper presents a novel sharpness mismatch detection method for stereoscopic images based on the comparison of edge width histograms of the left and right view. The new method is evaluated on the LIVE 3D Phase II and Ningbo 3D Phase I datasets and compared with two state-of-the-art methods. Experimental results show that the new method highly correlates with user scores of subjective tests and that it outperforms the current state-of-the-art. We then extend the method to stereoscopic omnidirectional images by partitioning the images into patches using a spherical Voronoi diagram. Furthermore, we integrate visual attention data into the detection process in order to weight sharpness mismatch according to the likelihood of its appearance in the viewport of the end-user's virtual reality device. For obtaining visual attention data, we performed a subjective experiment with 17 test subjects and 96 stereoscopic omnidirectional images. The entire dataset including the viewport trajectory data and resulting visual attention maps are publicly available with this paper. Simone Croci, Sebastian Knorr, Aljoscha Smolic |
ICIP | 3 |
| 2018 | Optimization of Occlusion-Inducing Depth Pixels in 3-D Video CodingabstractThe optimization of occlusion-inducing depth pixels in depth map coding has received little attention in the literature, since their associated texture pixels are occluded in the synthesized view and their effect on the synthesized view is considered negligible. However, the occlusion-inducing depth pixels still need to consume the bits to be transmitted, and will induce geometry distortion that inherently exists in the synthesized view. In this paper, we propose an efficient depth map coding scheme specifically for the occlusion-inducing depth pixels by using allowable depth distortions. Firstly, we formulate a problem of minimizing the overall geometry distortion in the occlusion subject to the bit rate constraint, for which the depth distortion is properly adjusted within the set of allowable depth distortions that introduce the same disparity error as the initial depth distortion. Then, we propose a dynamic programming solution to find the optimal depth distortion vector for the occlusion. The proposed algorithm can improve the coding efficiency without alteration of the occlusion order. Simulation results confirm the performance improvement compared to other existing algorithms. Pan Gao 0001, Cagri Ozcinar, Aljoscha Smolic |
ICIP | 3 |
| 2018 | A Pipeline for Lenslet Light Field Quality EnhancementabstractIn recent years, light fields have become a major research topic and their applications span across the entire spectrum of classical image processing. Among the different methods used to capture a light field are the lens let cameras, such as those developed by Lytro. While these cameras give a lot of freedom to the user, they also create light field views that suffer from a number of artefacts. As a result, it is common to ignore a significant subset of these views when doing high-level light field processing. We propose a pipeline to process light field views, first with an enhanced processing of RAW images to extract sub-aperture images, then a colour correction process using a recent colour transfer algorithm, and finally a denoising process using a state of the art light field denoising approach. We show that our method improves the light field quality on many levels, by reducing ghosting artefacts and noise, as well as retrieving more accurate and homogeneous colours across the sub-aperture images. Pierre Matysiak, Mairéad Grogan, Mikael Le Pendu, Martin Alain, Aljoscha Smolic |
ICIP | 5 |
| 2018 | High Dynamic Range Light Fields via Weighted Low Rank ApproximationabstractIn this paper, we propose a method for capturing High Dynamic Range (HDR) light fields with dense viewpoint sampling. Analogously to the traditional HDR acquisition process, several light fields are captured at varying exposures with a plenoptic camera. The RAW data is de-multiplexed to retrieve all light field viewpoints for each exposure and perform a soft detection of saturated pixels. Considering a matrix which concatenates all the vectorized views, we formulate the problem of recovering saturated areas as a Weighted Low Rank Approximation (WLRA) where the weights are defined from the soft saturation detection. We show that our algorithm successfully recovers the parallax in the over-exposed areas while the Truncated Nuclear Norm (TNN) minimization, traditionally used for single view HDR imaging, does not generalize to light fields. Advantages of our weighted approach as well as the simultaneous processing of all the viewpoints are also demonstrated in our experiments. Mikael Le Pendu, Christine Guillemot, Aljoscha Smolic |
ICIP | 3 |
| 2018 | Hydra: An Accelerator for Real-Time Edge-Aware Permeability Filtering in 65nm CMOSabstractMany modern video processing pipelines rely on edge-aware (EA) filtering methods. However, recent high-quality methods are challenging to run in real-time on embedded hardware due to their computational load. To this end, we propose an area-efficient and real-time capable hardware implementation of a high quality EA method. In particular, we focus on the recently proposed permeability filter (PF) that delivers promising quality and performance in the domains of high dynamic range (HDR) tone mapping, disparity and optical flow estimation. We present an efficient hardware accelerator that implements a tiled variant of the PF with low on-chip memory requirements and a significantly reduced external memory bandwidth (6.4× w.r.t. the non-tiled PF). The design has been taped out in 65 nm CMOS technology, is able to filter 720p grayscale video at 24.8 Hz and achieves a high compute density of 6.7GFLOPS/mm2(12× higher than embedded GPUs when scaled to the same technology node). The low area and bandwidth requirements make the accelerator highly suitable for integration into systems-on-chip (SoCs) where silicon area budget is constrained and external memory is typically a heavily contended resource. Manuel Eggimann, Christelle Gloor, Florian Scheidegger, Lukas Cavigelli, Michael Schaffner, Aljoscha Smolic, Luca Benini |
ISCAS | 6 |
| 2018 | Visual Attention in Omnidirectional Video for Virtual Reality ApplicationsabstractUnderstanding of visual attention is crucial for omnidirectional video (ODV) viewed for instance with a head-mounted display (HMD), where only a fraction of an ODV is rendered at a time. Transmission and rendering of ODV can be optimized by understanding how viewers consume a given ODV in virtual reality (VR) applications. In order to predict video regions that might draw the attention of viewers, saliency maps can be estimated by using computational visual attention models. As no such model currently exists for ODV, but given the importance for emerging ODV applications, we create a new visual attention user dataset for ODV, investigate behavior of viewers when consuming the content, and analyze the prediction performance of state-of-the-art visual attention models. Our developed test-bed and dataset will be publicly available with this paper, to stimulate and support research on ODV. Cagri Ozcinar, Aljoscha Smolic |
QoMEX | 2 |
| 2018 | Omnidirectional Video Streaming Using Visual Attention-Driven Dynamic Tiling for VRabstractThis paper proposes a new adaptive omnidirectional video (ODV) streaming system that uses visual attention (VA) maps. The proposed method benefits from a novel approach to VA-based bitrate allocation algorithm and dynamic tiling, providing enhanced virtual reality (VR) video experiences. The main contribution of this paper is the use of VA maps: (i) to distribute a given bitrate budget among a set of tiles of a given ODV and, (ii) to decide an optimal tiling structure (i.e., tile scheme) per chunk. For this, a novel objective metric is proposed: the visual attention spherical weighted (VASW) PSNR. This metric operates in the spherical domain and by means of a VA probabilistic model aims at capturing the quality of the actual areas observed by the users when navigating through the ODV content. We evaluate the proposed system performance with varying bandwidth conditions and the tracked head orientations from disjoint user experiments. Results show that the proposed system significantly outperforms the existing tiled-based streaming method. Cagri Ozcinar, Julián Cabrera, Aljoscha Smolic |
VCIP | 3 |
| 2018 | Affordable content creation for free-viewpoint video and VR/AR applications
Rafael Pagés, Konstantinos Amplianitis, David S. Monaghan, Jan Ondrej, Aljoscha Smolic |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | SalNet360: Saliency maps for omni-directional images with CNN
Rafael Monroy, Sebastian Lutz, Tejo Chalasani, Aljoscha Smolic |
Signal Process. Image Commun. | 4 |
| 2018 | Pipelines for HDR Video Coding Based on Luminance Independent Chromaticity PreprocessingabstractWe consider the chromaticity in high dynamic range (HDR) video coding and show the advantages of a constant luminance color space for encoding. For this, we introduce two constant luminance HDR video coding pipelines, which convert the source video to linear Yu'v'. A content dependent scaling of the chromaticity components serves as color quality parameter. This reduces perceivable color artifacts while remaining fully compatible with core High Efficiency Video Coding or other video coding standards. One of the pipelines further combines the scaling with a dedicated chromaticity transform to optimize the representation of the chromaticity components for encoding. We validate both pipelines with subjective user studies in addition to an objective comparison to the other state-of-the-art methods. The user studies show a significant improvement in perceived color quality at medium to high compression rates without sacrificing luminance quality compared with current standard coding pipelines. The objective evaluation suggests that both pipelines perform at least comparable to the current state-of-the-art methods. Samir Mahmalat, Tunç Ozan Aydin, Aljoscha Smolic |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Towards Edge-Aware Spatio-Temporal Filtering in Real-TimeabstractSpatio-temporal edge-aware (STEA) filtering methods have recently received increased attention due to their ability to efficiently solve or approximate important image-domain problems in a temporally consistent manner - which is a crucial property for video-processing applications. However, existing STEA methods are currently unsuited for real-time, embedded stream-processing settings due to their high processing latency, large memory, and bandwidth requirements, and the need for accurate optical flow to enable filtering along motion paths. To this end, we propose an efficient STEA filtering pipeline based on the recently proposed permeability filter (PF), which offers high quality and halo reduction capabilities. Using mathematical properties of the PF, we reformulate its temporal extension as a causal, non-linear infinite impulse response filter, which can be efficiently evaluated due to its incremental nature. We bootstrap our own accurate flow using the PF and its temporal extension by interpolating a quasi-dense nearest neighbour field obtained with an improved PatchMatch algorithm, which employs binarized octal orientation maps (BOOM) descriptors to find correspondences among subsequent frames. Our method is able to create temporally consistent results for a variety of applications such as optical flow estimation, sparse data upsampling, visual saliency computation and disparity estimation. We benchmark our optical flow estimation on the MPI Sintel dataset, where we currently achieve a Pareto optimal quality-efficiency tradeoff with an average endpoint error of 7.68 at 0.59 s single-core execution time on a recent desktop machine. Michael Schaffner, Florian Scheidegger, Lukas Cavigelli, Hubert Kaeslin, Luca Benini, Aljoscha Smolic |
IEEE Trans. Image Process. | 6 |
| 2017 | Viewport-aware adaptive 360° video streaming using tiles for virtual realityabstract360° video is attracting an increasing amount of attention in the context of Virtual Reality (VR). Owing to its very high-resolution requirements, existing professional streaming services for 360° video suffer from severe drawbacks. This paper introduces a novel end-to-end streaming system from encoding to displaying, to transmit 8K resolution 360° video and to provide an enhanced VR experience using Head Mounted Displays (HMDs). The main contributions of the proposed system are about tiling, integration of the MPEG-Dynamic Adaptive Streaming over HTTP (DASH) standard, and viewport-aware bitrate level selection. Tiling and adaptive streaming enable the proposed system to deliver very high-resolution 360° video at good visual quality. Further, the proposed viewport-aware bitrate assignment selects an optimum DASH representation for each tile in a viewport-aware manner. The quality performance of the proposed system is verified in simulations with varying network bandwidth using realistic view trajectories recorded from user experiments. Our results show that the proposed streaming system compares favorably compared to existing methods in terms of PSNR and SSIM inside the viewport. Cagri Ozcinar, Ana De Abreu, Aljoscha Smolic |
ICIP | 3 |
| 2017 | Estimation of Optimal Encoding Ladders for Tiled 360° VR Video in Adaptive Streaming SystemsabstractGiven the significant industrial growth of demand for virtual reality (VR), 360ovideo streaming is one of the most important VR applications that require cost-optimal solutions to achieve widespread proliferation of VR technology. Because of its inherent variability of data-intensive content types and its tiled-based encoding and streaming, 360ovideo requires new encoding ladders in adaptive streaming systems to achieve cost-optimal and immersive streaming experiences. In this context, this paper targets both the provider's and client's perspectives and introduces a new content-aware encoding ladder estimation method for tiled 360oVR video in adaptive streaming systems. The proposed method first categories a given 360ovideo using its features of encoding complexity and estimates the visual distortion and resource cost of each bitrate level based on the proposed distortion and resource cost models. An optimal encoding ladder is then formed using the proposed integer linear programming (ILP) algorithm by considering practical constraints. Experimental results of the proposed method are compared with the recommended encoding ladders of professional streaming service providers. Evaluations show that the proposed encoding ladders deliver better results compared to the recommended encoding ladders in terms of objective quality for 360ovideo, providing optimal encoding ladders using a set of service provider's constraint parameters. Cagri Ozcinar, Ana De Abreu, Sebastian Knorr, Aljoscha Smolic |
ISM | 4 |
| 2017 | Light field denoising by sparse 5D transform domain collaborative filteringabstractIn this paper, we propose to extend the state-of-the-art BM3D image denoising filter to light fields, and we denote our method LFBM5D. We take full advantage of the 4D nature of light fields by creating disparity compensated 4D patches which are then stacked together with similar 4D patches along a 5thdimension. We then filter these 5D patches in the 5D transform domain, obtained by cascading a 2D spatial transform, a 2D angular transform, and a 1D transform applied along the similarities. Furthermore, we propose to use the shape-adaptive DCT as the 2D angular transform to be robust to occlusions. Results show a significant improvement in synthetic noise removal compared to state-of-the-art methods, for both light fields captured with a lenslet camera or a gantry. Experiments on Lytro Illum camera noise removal also demonstrate a clear improvement of the light field quality. Martin Alain, Aljoscha Smolic |
MMSP | 2 |
| 2017 | Look around you: Saliency maps for omnidirectional images in VR applicationsabstractUnderstanding visual attention has always been a topic of great interest in the graphics, image/video processing, robotics and human-computer interaction communities. By understanding salient image regions, the compression, transmission and rendering algorithms can be optimized. This is particularly important in omnidirectional images (ODIs) viewed with a head-mounted display (HMD), where only a fraction of the captured scene is displayed at a time, namely viewport. In order to predict salient image regions, saliency maps are estimated either by using an eye tracker to collect eye fixations during subjective tests or by using computational models of visual attention. However, eye tracking developments for ODIs are still in the early stages and although a large list of saliency models are available, no particular attention has been dedicated to ODIs. Therefore, in this paper, we consider the problem of estimating saliency maps for ODIs viewed with HMDs, when the use of an eye tracker device is not possible. We collected viewport center trajectories (VCTs) of 32 participants for 21 ODIs and propose a method to transform the gathered data into saliency maps. The obtained saliency maps are compared in terms of image exposition time used to display each ODI in the subjective tests. Then, motivated by the equator bias tendency in ODIs, we propose a post-processing method, namely fused saliency maps (FSM), to adapt current saliency models to ODIs requirements. We show that the use of FSM on current models improves their performance by up to 20%. The developed database and testbed are publicly available with this paper. Ana De Abreu, Cagri Ozcinar, Aljoscha Smolic |
QoMEX | 3 |
| 2017 | Interactive high-quality green-screen keying via color unmixing
Yagiz Aksoy, Tunç Ozan Aydin, Marc Pollefeys, Aljoscha Smolic |
ACM Trans. Graph. | 4 |
| 2017 | Unmixing-Based Soft Color Segmentation for Image ManipulationabstractWe present a new method for decomposing an image into a set of soft color segments that are analogous to color layers with alpha channels that have been commonly utilized in modern image manipulation software. We show that the resulting decomposition serves as an effective intermediate image representation, which can be utilized for performing various, seemingly unrelated, image manipulation tasks. We identify a set of requirements that soft color segmentation methods have to fulfill, and present an in-depth theoretical analysis of prior work. We propose an energy formulation for producing compact layers of homogeneous colors and a color refinement procedure, as well as a method for automatically estimating a statistical color model from an image. This results in a novel framework for automatic and high-quality soft color segmentation that is efficient, parallelizable, and scalable. We show that our technique is superior in quality compared to previous methods through quantitative analysis as well as visually through an extensive set of examples. We demonstrate that our soft color segments can easily be exported to familiar image manipulation software packages and used to produce compelling results for numerous image manipulation applications without forcing the user to learn new tools and workflows. Yagiz Aksoy, Tunç Ozan Aydin, Aljoscha Smolic, Marc Pollefeys |
ACM Trans. Graph. | 3 |
| 2016 | Real-time temporally coherent local HDR tone mappingabstractSubjective studies showed that most HDR video tone mapping operators either produce disturbing temporal artifacts, or are limited in their local contrast reproduction capability. Recently, both these issues have been addressed by a novel temporally coherent local HDR tone mapping method, which has been shown, both qualitatively and through a subjective study, to be advantageous compared to previous methods. However, this method's high-quality results came at the cost of a computationally expensive workflow that could only be executed offline. In this paper, we present a modified algorithm which builds upon the previous work by redesigning key components to achieve real-time performance. We accomplish this by replacing the optical flow based per-pixel temporal coherency with a tone-curve-space alternative. This way we eliminate the main computational burden of the original method with little sacrifice in visual quality. Simone Croci, Tunç Ozan Aydin, Nikolce Stefanoski, Markus Gross 0001, Aljoscha Smolic |
ICIP | 5 |
| 2016 | Robust calibration of broadcast cameras based on ellipse and line contoursabstractProfessional TV studio footage often poses specific challenges to camera calibration due to lack of features and complex camera operation. As available algorithms often fail, we propose a novel approach based on robust tracking of ellipse and line features of a predefined logo. We further devise a predictive and iterative estimation algorithm, which incorporates confidence measures and filtering. Our results validate accuracy and reliability of our approach, demonstrated with challenging professional footage. Simone Croci, Nikolce Stefanoski, Aljoscha Smolic |
ICIP | 3 |
| 2016 | Luminance independent chromaticity preprocessing for HDR video codingabstractWe introduce a constant luminance HDR video coding pipeline, which converts the source video to linear Y u'v' color space and applies a dedicated chromaticity transformation before encoding. This reduces perceivable color artifacts without modifying the core codec itself. We validate our approach by a user study that shows a significant improvement in perceived color quality at high compression rates without sacrificing luminance quality compared to current standard coding pipelines. Samir Mahmalat, Nikolce Stefanoski, Daniel Luginbuhl, Tunç Ozan Aydin, Aljoscha Smolic |
ICIP | 5 |
| 2016 | Hybrid ASIC/FPGA System for Fully Automatic Stereo-to-Multiview Conversion Using IDWabstractRecently, multiview autostereoscopic dis-plays (MADs), which enable a limited glasses-free 3D experience, have become commercially available. The main problem of MADs is that they require several (typically eight or nine) views, while most of the 3D video content is in stereoscopic 3D today. In order to bridge this gap, the research community started to devise automatic multiview synthesis (MVS) methods. These algorithms require real-time processing and should be portable to end-user devices to develop their full potential. To this end, we revisit an algorithmic solution based on image domain warping (IDW) and devise a hardware architecture of a complete synthesis pipeline, provide insights into where the computationally challenging parts are, and present implementation results of a hybrid field programmable gate array/application-specific integrated circuit prototype, which is the first hardware implementation of a complete IDW-based MVS system. Based on these results, we also estimate the complexity and energy efficiency of a fully integrated solution in 65- and 28-nm CMOS technology and show that a full-high-definition real-time solution on a single chip is within reach. The proposed architecture could be used as a coprocessor in a system-on-chip targeting 3D TV sets, thereby enabling efficient content generation with limited user interaction (e.g., depth range adjustment) in real time. Michael Schaffner, Frank K. Gürkaynak, Pierre Greisen, Hubert Kaeslin, Luca Benini, Aljoscha Smolic |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | Interactive High-Quality Green-Screen Keying via Color UnmixingabstractDue to the widespread use of compositing in contemporary feature films, green-screen keying has become an essential part of postproduction workflows. To comply with the ever-increasing quality requirements of the industry, specialized compositing artists spend countless hours using multiple commercial software tools, while eventually having to resort to manual painting because of the many shortcomings of these tools. Due to the sheer amount of manual labor involved in the process, new green-screen keying approaches that produce better keying results with less user interaction are welcome additions to the compositing artist’s arsenal. We found that—contrary to the common belief in the research community—production-quality green-screen keying is still an unresolved problem with its unique challenges. In this article, we propose a novel green-screen keying method utilizing a new energy minimization-based color unmixing algorithm. We present comprehensive comparisons with commercial software packages and relevant methods in literature, which show that the quality of our results is superior to any other currently available green-screen keying solution. It is important to note that, using the proposed method, these high-quality results can be generated using only one-tenth of the manual editing time that a professional compositing artist requires to process the same content having all previous state-of-the-art tools at one’s disposal. Yagiz Aksoy, Tunç Ozan Aydin, Marc Pollefeys, Aljoscha Smolic |
ACM Trans. Graph. | 4 |
| 2015 | DRAM or no-DRAM?: exploring linear solver architectures for image domain warping in 28 nm CMOS
Michael Schaffner, Frank K. Gürkaynak, Aljoscha Smolic, Luca Benini |
DATE | 3 |
| 2015 | Chromatic calibration of an HDR display using 3D octree forestsabstractHigh dynamic range (HDR) display prototypes have been built and used for scientific studies for nearly a decade, and they are now on the verge of entering consumer market. However, problems exist regarding the accurate color reproduction capabilities on these displays. In this paper, we first characterize the color reproduction capability of a state-of-art HDR display through a set of measurements, and present a novel calibration method that takes into account the variation of the chrominance error over HDR display's wide luminance range. Our proposed 3D octree forest data structure for representing and querying the calibration function successfully addresses the challenges in calibrating HDR displays: (i) high computational complexity due to nonlinear chromatic distortions, and (ii) huge storage space demand for a look-up table (35GB vs 100kB). We show that our method achieves high color reproduction accuracy through both objective metrics and a controlled subjective study. Jing Liu 0053, Nikolce Stefanoski, Tunç Ozan Aydin, Anselm Grundhöfer, Aljoscha Smolic |
ICIP | 5 |
| 2015 | Video content and structure description based on keyframes, clusters and storyboardsabstractIn this paper we present a novel system to extract keyframes, shot clusters and structural storyboards for video content description, which can be used for a variety of summarization, visualization, classification, indexing and retrieval applications. The system automatically selects an appealing set of keyframes and creates meaningful clusters of shots. It further identifies sections that appear recurrently, which are called anchors, and typically divide television shows into different parts. This information about anchors can then be used to browse video content in a new fashion. Finally, our system creates a new type of interactive storyboard suitable to visualize and analyze the structure of the video in a novel way. Marc Junyent, Pablo Beltrán, Miquel A. Farre, Jordi Pont-Tuset, Alexandre Chapiro, Aljoscha Smolic |
MMSP | 6 |
| 2015 | Automatic multiview synthesis - Towards a mobile system on a chipabstractOver the last couple of years, multiview autostereoscopic displays (MADs) have become commercially available which enable a limited glasses-free 3D experience. The main problem of MADs is that they require several (typically 8 or 9) views, while most of the 3D video content is in stereoscopic 3D (S3D) today. In order to bridge this gap, the research community started to devise automatic multiview synthesis (MVS) methods. These algorithms require real-time processing and should be portable to end-user devices to develop their full potential. In this paper, we present a complete hardware system for fully automatic MVS. We give an overview of the algorithmic flow and corresponding hardware architecture, and provide implementation results of a hybrid FPGA/ASIC prototype - which is the first complete hardware system implementing image-domain-warping-based MVS. The proposed hardware IP could be used as a co-processor in a system-on-chip (SoC) targeting 3D TV sets, thereby enabling efficient content generation in real-time. Michael Schaffner, Frank K. Gürkaynak, Hubert Kaeslin, Luca Benini, Aljoscha Smolic |
VCIP | 5 |
| 2015 | Automatic multiview synthesis - Prototype demoabstractOverview. Today, most commercially available 3D display systems require the viewers to wear some sort of shutter-or polarization glasses, which is often regarded as inconvenience. Ideally, a 3D display system should not require the users to wear additional gear. In fact, the optimum would be a display that replicates the original light-field of a scene. So-called multiview aütostereoscopic displays (MADs) represent a step in this direction, as they are able to project several views of a scene simultaneously, enabling a glasses-free 3D experience and a limited motion parallax effect in horizontal direction. Michael Schaffner, Frank K. Gürkaynak, Hubert Kaeslin, Luca Benini, Aljoscha Smolic |
VCIP | 5 |
| 2015 | Art-directable Continuous Dynamic Range video
Alexandre Chapiro, Tunç Ozan Aydin, Nikolce Stefanoski, Simone Croci, Aljoscha Smolic, Markus Gross 0001 |
Comput. Graph. | 5 |
| 2015 | Automated Aesthetic Analysis of Photographic ImagesabstractWe present a perceptually calibrated system for automatic aesthetic evaluation of photographic images. Our work builds upon the concepts of no-reference image quality assessment, with the main difference being our focus on rating image aesthetic attributes rather than detecting image distortions. In contrast to the recent attempts on the highly subjective aesthetic judgment problems such as binary aesthetic classification and the prediction of an image's overall aesthetics rating, our method aims on providing a reliable objective basis of comparison between aesthetic properties of different photographs. To that end our system computes perceptually calibrated ratings for a set of fundamental and meaningful aesthetic attributes, that together form an "aesthetic signature" of an image. We show that aesthetic signatures can still be used to improve upon the current state-of-the-art in automatic aesthetic judgment, but also enable interesting new photo editing applications such as automated aesthetic analysis, HDR tone mapping evaluation, and providing aesthetic feedback during multi-scale contrast manipulation. Tunç Ozan Aydin, Aljoscha Smolic, Markus Gross 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2014 | Perceptual evaluation of cardboarding in 3D content visualizationabstractA pervasive artifact that occurs when visualizing 3D content is the so-called "cardboarding" effect, where objects appear flat due to depth compression, with relatively little research conducted to perceptually quantify its effects. Our aim is to shed light on the subjective preferences and practical perceptual limits of stereo vision with respect to cardboarding. We present three experiments that explore the consequences of displaying simple scenes with reduced depths using both subjective ratings and adjustments and objective sensitivity metrics. Our results suggest that compressing depth to 80% or above is likely to be acceptable, whereas sensitivity to the cardboarding artifact below 30% is very high. These values could be used in practice as guidelines for commonplace depth mapping operations in 3D production pipelines. Alexandre Chapiro, Olga Diamanti, Steven Poulakos, Carol O'Sullivan, Aljoscha Smolic, Markus Gross 0001 |
SAP | 5 |
| 2014 | Alternating attention in continuous stereoscopic depthabstractThe decoupling of eye vergence and accommodation (V/A) has been found to negatively impact depth interpretation, visual comfort and fatigue. In this paper, we explore a hypothesis that placement of visual cues within a scene can assist a viewer in the process of maintaining the V/A decoupling. This effect is demonstrated through the use of a continuous depth plane that connects spatially distinct scene elements. Our experimental design enables us to make the following three contributions: (1) We show that a continuous depth element can improve the time it takes to transition visual attention in depth. (2) We observe that the subjective assessment of fatigue emerges before we detect a quantitative decline in performance. (3) We aim to motivate that stereoscopic 3D content creators may learn scene composition, framing and montage from visual psychophysics. Steven Poulakos, Gerhard Röthlin, Adrian Schwaninger, Aljoscha Smolic, Markus Gross 0001 |
SAP | 4 |
| 2014 | An Approximate Computing Technique for Reducing the Complexity of a Direct-Solver for Sparse Linear Systems in Real-Time Video ProcessingabstractMany video processing algorithms are formulated as least-squares problems that result in large, sparse linear systems. Solving such systems in real time is very demanding. This paper focuses on reducing the computational complexity of a direct Cholesky-decomposition-based solver. Our approximation scheme builds on the observation that, in well-conditioned problems, many elements in the decomposition nearly vanish. Such elements may be pruned from the dependency graph with mild accuracy degradation. Using an example from image-domain warping, we show that pruning reduces the amount of operations per solve by over 75%, resulting in significant savings in computing time, area or energy. Michael Schaffner, Frank K. Gürkaynak, Aljoscha Smolic, Hubert Kaeslin, Luca Benini |
DAC | 3 |
| 2014 | MasterCam FVV: Robust registration of multiview sports video to a static high-resolution master camera for free viewpoint videoabstractFree viewpoint video enables interactive viewpoint selection in real world scenes, which is attractive for many applications such as sports visualization. Multi-camera registration is one of the difficult tasks in such systems. We introduce the concept of a static high resolution master camera for improved long-term multiview alignment. All broadcast cameras are aligned to a common reference. Our approach builds on frame-to-frame alignment, extended into a recursive long-term estimation process, which is shown to be accurate, robust and stable over long sequences. Florian Angehrn, Oliver Wang, Yagiz Aksoy, Markus Gross 0001, Aljoscha Smolic |
ICIP | 5 |
| 2014 | ColorBrush: Animated diffusion for intuitive colorization simulating water paintingabstractWater painting is an art, where the result and experience strongly depend on the process of creating it. Color is injected by the artist in a controlled way and diffuses into a final state. We present a system to simulate such a process, which builds on existing colorization approaches. The geodesic distance builds the mathematical foundation for our animated diffusion. We add a time-dependent weight function to generate a diffusion-like spreading effect. Tiled approximation is used to achieve interactive rates on modern mobile devices. The result is a colorization framework simulating water painting that allows giving real-time feedback on touch events with the limited hardware resources of a tablet computer. Nicolas Marki, Oliver Wang, Markus Gross 0001, Aljoscha Smolic |
ICIP | 4 |
| 2014 | REVERIE: Natural human interaction in virtual immersive environmentsabstractREVERIE (REal and Virtual Engagement in Realistic Immersive Environments [1]) targets novel research to address the demanding challenges involved with developing state-of-the-art technologies for online human interaction. The REVERIE framework enables users to meet, socialise and share experiences online by integrating cutting-edge technologies for 3D data acquisition and processing, networking, autonomy and real-time rendering. In this paper, we describe the innovative research that is showcased through the REVERIE integrated framework through richly defined use-cases which demonstrate the validity and potential for natural interaction in a virtual immersive and safe environment. Previews of the REVERIE demo and its key research components can be viewed at www.youtube.com/user/REVERIEFP7. Julie A. Wall, Ebroul Izquierdo, Lemonia Argyriou, David S. Monaghan, Noel E. O'Connor, Steven Poulakos, Aljoscha Smolic, Rufael Mekuria |
ICIP | 7 |
| 2014 | Optimizing stereo-to-multiview conversion for autostereoscopic displaysabstractAbstract We present a novel stereo‐to‐multiview video conversion method for glasses‐free multiview displays. Different from previous stereo‐to‐multiview approaches, our mapping algorithm utilizes the limited depth range of autostereoscopic displays optimally and strives to preserve the scene's artistic composition and perceived depth even under strong depth compression. We first present an investigation of how perceived image quality relates to spatial frequency and disparity. The outcome of this study is utilized in a two‐step mapping algorithm, where we (i) compress the scene depth using a non‐linear global function to the depth range of an autostereoscopic display and (ii) enhance the depth gradients of salient objects to restore the perceived depth and salient scene structure. Finally, an adapted image domain warping algorithm is proposed to generate the multiview output, which enables overall disparity range extension. Alexandre Chapiro, Simon Heinzle, Tunç Ozan Aydin, Steven Poulakos, Matthias Zwicker, Aljoscha Smolic, Markus Gross 0001 |
Comput. Graph. Forum | 6 |
| 2014 | Temporally coherent local tone mapping of HDR videoabstractRecent subjective studies showed that current tone mapping operators either produce disturbing temporal artifacts, or are limited in their local contrast reproduction capability. We address both of these issues and present an HDR video tone mapping operator that can greatly reduce the input dynamic range, while at the same time preserving scene details without causing significant visual artifacts. To achieve this, we revisit the commonly usedspatialbase-detail layer decomposition and extend it to thetemporal domain. We achieve high quality spatiotemporal edge-aware filtering efficiently by using a mathematically justified iterative approach that approximates a global solution. Comparison with the state-of-the-art, both qualitatively, and quantitatively through a controlled subjective experiment, clearly shows our method's advantages over previous work. We present local tone mapping results on challenging high resolution scenes with complex motion and varying illumination. We also demonstrate our method's capability of preserving scene details at user adjustable scales, and its advantages for low light video sequences with significant camera noise. Tunç Ozan Aydin, Nikolce Stefanoski, Simone Croci, Markus Gross 0001, Aljoscha Smolic |
ACM Trans. Graph. | 5 |
| 2013 | Efficient image resampling for multiview displaysabstractWe present an evaluation of different resampling strategies for autostereoscopic multiview displays. In particular, we compare the computational efficiency, memory requirements, and image quality of different resampling algorithms with focus on real-time architectures. Our assessment shows large differences in computational complexity for similar image quality, and aims at providing a recommendation for selecting a suitable resampling strategy. Michael Schaffner, Pierre Greisen, Simon Heinzle, Aljoscha Smolic |
ICASSP | 4 |
| 2013 | Depth estimation and depth enhancement by diffusion of depth featuresabstractCurrent trends in video technology indicate a significant increase in spatial and temporal resolution of video data. Recently, a linear-runtime feature diffusion algorithm was presented which aims for fast and accurate processing of such high resolution data. In this paper, we introduce this algorithm from the perspective of image-based depth estimation, expanding upon the algorithm by requiring interview consistency in the depth diffusion process. We also discuss different application scenarios and provide an in-depth analysis of the method in this context. Nikolce Stefanoski, Can Bal, Manuel Lang, Oliver Wang, Aljoscha Smolic |
ICIP | 5 |
| 2013 | Computational sports broadcasting: Automated director assistance for live sportsabstractLive sports broadcast is seeing a large increase in the number of cameras used for filming. More cameras can provide better coverage of the field and a wider range of experiences for viewers. However, choosing optimal cameras for broadcast demands a high level of concentration, awareness and experience from sports broadcast directors. We present an automatic assistant to help select likely candidates from a large array of possible cameras. Sports directors can then choose the final broadcast camera from the reduced suggestion set. Our assistant uses both widely acknowledged cinematography guidelines for sports directing, as well as a data-driven approach that learns specific styles from directors. Christine Chen, Oliver Wang, Simon Heinzle, Peter Carr 0001, Aljoscha Smolic, Markus Gross 0001 |
ICME | 5 |
| 2013 | Finite Element Image WarpingabstractAbstract We introduce a single unifying framework for a wide range of content‐aware image warping tasks using a finite element method (FEM). Existing approaches commonly define error terms over vertex finite differences and can be expressed as a special case of our general FEM model. In this work, we exploit the full generality of FEMs, gaining important advantages over prior methods. These advantages include arbitrary mesh connectivity allowing for adaptive meshing and efficient large‐scale solutions, a well‐defined continuous problem formulation that enables clear analysis of existing warping error functions and allows us to propose improved ones, and higher order basis functions that allow for smoother warps with fewer degrees of freedom. To support per‐element basis functions of varying degree and complex mesh connectivity with hanging nodes, we also introduce a novel use of discontinuous Galerkin FEM. We demonstrate the utility of our method by showing examples in video retargeting and camera stabilization applications, and compare our results with previous state of the art methods. Peter Kaufmann 0001, Oliver Wang, Alexander Sorkine-Hornung, Olga Sorkine-Hornung, Aljoscha Smolic, Markus Gross 0001 |
Comput. Graph. Forum | 5 |
| 2013 | DuctTake: Spatiotemporal Video CompositingabstractAbstract DuctTake is a system designed to enable practical compositing of multiple takes of a scene into a single video. Current industry solutions are based around object segmentation, a hard problem that requires extensive manual input and cleanup, making compositing an expensive part of the film‐making process. Our method instead composites shots together by finding optimal spatiotemporal seams using motion‐compensated 3D graph cuts through the video volume. We describe in detail the required components, decisions, and new techniques that together make a usable, interactive tool for compositing HD video, paying special attention to running time and performance of each section. We validate our approach by presenting a wide variety of examples and by comparing result quality and creation time to composites made by professional artists using current state‐of‐the‐art tools. Jan Rüegg, Oliver Wang, Aljoscha Smolic, Markus Gross 0001 |
Comput. Graph. Forum | 3 |
| 2013 | Evaluation and FPGA Implementation of Sparse Linear Solvers for Video Processing ApplicationsabstractSparse linear systems are commonly used in video processing applications, such as edge-aware filtering or video retargeting. Due to the 2-D nature of images, the involved problem sizes are large and thus solving such systems is computationally challenging. In this paper, we address sparse linear solvers for real-time video applications. We investigate several solver techniques, discuss hardware trade-offs, and provide field-programmable gate array (FPGA) architectures and implementation results of a Cholesky direct solver and of an iterative BiCGSTAB solver. The FPGA implementations solve 32 k × 32 k matrices at up to 50 f/s and outperform software implementations by at least one order of magnitude. Pierre Greisen, Marian Runo, Patrice Guillet, Simon Heinzle, Aljoscha Smolic, Hubert Kaeslin, Markus Gross 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2013 | Automatic View Synthesis by Image-Domain-WarpingabstractToday, stereoscopic 3D (S3D) cinema is already mainstream, and almost all new display devices for the home support S3D content. S3D distribution infrastructure to the home is already established partly in the form of 3D Blu-ray discs, video on demand services, or television channels. The necessity to wear glasses is, however, often considered as an obstacle, which hinders broader acceptance of this technology in the home. Multiviewautostereoscopic displays enable a glasses free perception of S3D content for several observers simultaneously, and support head motion parallax in a limited range. To support multiviewautostereoscopic displays in an already established S3D distribution infrastructure, a synthesis of new views from S3D video is needed. In this paper, a view synthesis method based on image-domain-warping (IDW) is presented that automatically synthesizes new views directly from S3D video and functions completely. IDW relies on an automatic and robust estimation of sparse disparities and image saliency information, and enforces target disparities in synthesized images using an image warping framework. Two configurations of the view synthesizer in the scope of a transmission and view synthesis framework are analyzed and evaluated. A transmission and view synthesis system that uses IDW is recently submitted to MPEG's call for proposals on 3D video technology, where it is ranked among the four best performing proposals. Nikolce Stefanoski, Oliver Wang, Manuel Lang, Pierre Greisen, Simon Heinzle, Aljoscha Smolic |
IEEE Trans. Image Process. | 6 |
| 2013 | Distinguishing Texture Edges From Object Boundaries in VideoabstractOne of the most fundamental problems in image processing and computer vision is the inherent ambiguity that exists between texture edges and object boundaries in real-world images and video. Despite this ambiguity, many applications in computer vision and image processing often use image edge strength with the assumption that these edges approximate object depth boundaries. However, this assumption is often invalidated by real world data, and this discrepancy is a significant limitation in many of today's image processing methods. We address this issue by introducing a simple, low-level, and patch-consistency assumption that leverages the extra information present in video data to resolve this ambiguity. Through analyzing how well patches can be modeled by simple transformations over time, we can obtain an indication of which image edges correspond to texture edges versus object boundaries. Our approach is simple to implement and has the potential to improve a wide range of image and video-based applications by suppressing the detrimental effects of strong texture edges on regularization terms. We validate our approach by presenting results on a variety of scene types and directly incorporating our augmented edge map into existing image segmentation and optical flow applications, showing results that better correspond to object boundaries. Oliver Wang, Martina Dümcke, Aljoscha Smolic, Markus Gross 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Image quality vs rate optimized coding of warps for view synthesis in 3D video applicationsabstractIn this paper, a method for efficient warp coding is presented. Coded warps are transmitted together with stereo or multi-view video to enable additional functionalities at the receiver-side through image-domain-warping, e.g. depth adaption for stereo displays or support of multi-view autostereoscopic displays. Warp coding is performed by partitioning warps using a resolution pyramid and predictively exploiting intra and inter partition dependencies. View synthesis is employed within the coding loop to control the overall coding process, i.e. to evaluate the contribution of coded partitions to the synthesis quality. It is shown that coded warps represent a practically negligible portion of about 3.6% of the overall (video+warp) bit rate. Furthermore, it is shown that a transmission of warps leads to a reduction of synthesis time up to a factor of 8 in comparison to a fully automatic receiver-side view synthesis which uses only decoded video as input. Nikolce Stefanoski, Manuel Lang, Aljoscha Smolic |
ICIP | 3 |
| 2012 | Analysis and VLSI Implementation of EWA Rendering for Real-Time HD Video ApplicationsabstractNonlinear image warping or image resampling is a necessary step in many current and upcoming video applications, such as video retargeting, stereoscopic 3-D mapping, and multiview synthesis. The challenges for real-time resampling include not only image quality but also available energy and computational power of the employed device. In this paper, we employ an elliptical-weighted average (EWA) rendering approach to 2-D image resampling. We extend the classical EWA framework for increased visual quality and provide a very large scale integration architecture for efficient view rendering. The resulting architecture is able to render high-quality video sequences in real time targeted for low-power applications in end-user display devices. Pierre Greisen, Michael Schaffner, Simon Heinzle, Marian Runo, Aljoscha Smolic, Andreas Peter Burg, Hubert Kaeslin, Markus Gross 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2012 | Practical temporal consistency for image-based graphics applicationsabstractWe present an efficient and simple method for introducing temporal consistency to a large class of optimization driven image-based computer graphics problems. Our method extends recent work in edge-aware filtering, approximating costly global regularization with a fast iterative joint filtering operation. Using this representation, we can achieve tremendous efficiency gains both in terms of memory requirements and running time. This enables us to process entire shots at once, taking advantage of supporting information that exists across far away frames, something that is difficult with existing approaches due to the computational burden of video data. Our method is able to filter along motion paths using an iterative approach that simultaneously uses and estimates per-pixel optical flow vectors. We demonstrate its utility by creating temporally consistent results for a number of applications including optical flow, disparity estimation, colorization, scribble propagation, sparse data up-sampling, and visual saliency computation. Manuel Lang, Oliver Wang, Tunç Ozan Aydin, Aljoscha Smolic, Markus Gross 0001 |
ACM Trans. Graph. | 4 |
| 2011 | 2D to 3D conversion of sports content using panoramasabstractGiven video from a single camera, conversion to two-view stereoscopic 3D is a challenging problem. We present a system to automatically create high quality stereoscopic video from monoscopic footage of field-based sports by exploiting context-specific priors, such as the ground plane, player size and known background. Our main contribution is a novel technique that constructs per-shot panoramas to ensure temporally consistent stereoscopic depth in video reconstructions. Players are rendered as billboards at correct depths on the ground plane. Our method uses additional sports priors to disambiguate segmentation artifacts and produce synthesized 3D shots that are in most cases, indistinguishable from stereoscopic ground truth footage. Lars Schnyder, Oliver Wang, Aljoscha Smolic |
ICIP | 3 |
| 2011 | Extending SVC by Content-adaptive Spatial ScalabilityabstractThis paper provides details on a complete integration of Content- adaptive Spatial Scalability (CASS) into the scalable video coding extension of H.264/AVC (SVC). CASS enables the efficient encoding of a high-quality bit stream that contains several versions of an original image sequence. Thereby, each such image sequence has been created by content-adaptive and art directed retargeting to different display aspect-ratios and/or resolutions. Non-linear dependencies between spatial layers, which have been introduced through content-adaptive retargeting, are exploited by a generalization of the three inter-layer prediction tools of SVC, i.e. by content-adaptive inter-layer texture, motion and residual prediction. The CASS extended SVC enables the transmission of video content which has been specifically adapted in an art- directed way to multiple display configurations (e.g. to SD and HD displays with 4:3 and 16:9 aspect-ratios, respectively) using a single compressed bit stream. With our extension, video content of higher semantic quality can be transmitted in a scalable way by introducing an average overhead in bit rate of 9.3%. Yongzhe Wang, Nikolce Stefanoski, Manuel Lang, Alexander Sorkine-Hornung, Aljoscha Smolic, Markus Gross 0001 |
ICIP | 5 |
| 2011 | Evaluation of backward mapping DIBR for FVV applicationsabstractIn this paper, we explore the challenges posed by wide baseline camera configurations for depth-image-based rendering, which should provide greater freedom for choosing the virtual viewpoint in a Free Viewpoint Video context, compared with the usual camera configurations intended for use in 3DTV settings. We implement a backward mapping approach with a custom filtering scheme based on median filters. Whilst the results back our initial assumption that this camera configuration provides good mobility, we show that the usual encoding for depth information referred to a global reference system is wrong and reference systems local to each camera should be used instead. Daniel Berjón, Alexander Sorkine-Hornung, Francisco Morán, Aljoscha Smolic |
ICME | 4 |
| 2011 | Automatic content creation for multiview autostereoscopic displays using image domain warpingabstractContent creation for autostereoscopic displays is a widely unresolved task. Typical methods rely on view synthesis based on depth image based rendering. Our method applies purely image domain warping instead. Input video is analyzed and information about sparse disparity, vertical edges and saliency is extracted. A constrained energy minimization problem is formulated and efficiently solved. The resulting image warping functions are used to synthesize novel views. Our approach is fully automatic, accurate, and reliable. Disocclusions and related artifacts are avoided due to smooth, saliency-driven warping functions. Our method also works well for extrapolation of views in a limited range, thus supporting multiview creation from stereo input, which is the most relevant use case scenario. Miquel A. Farre, Oliver Wang, Manuel Lang, Nikolce Stefanoski, Alexander Sorkine-Hornung, Aljoscha Smolic |
ICME | 6 |
| 2011 | Three-Dimensional Video Postproduction and ProcessingabstractThis paper gives an overview of the state-of-the-art in 3-D video postproduction and processing as well as an outlook to remaining challenges and opportunities. First, fundamentals of stereography are outlined that set the rules for proper 3-D content creation. Manipulation of the depth composition of a given stereo pair via view synthesis is identified as the key functionality in this context. Basic algorithms are described to adapt and correct fundamental stereo properties such as geometric distortions, color alignment, and stereo geometry. Then, depth image-based rendering is explained as the widely applied solution for view synthesis in 3-D content creation today. Recent improvements of depth estimation already provide very good results. However, in most cases, still interactive workflows dominate. Warping-based methods may become an alternative for some applications in the future, which do not rely on dense and accurate depth estimation. Finally, 2-D to 3-D conversion is covered, which is an important special area for reuse of existing legacy 2-D content in 3-D. Here various advanced algorithms are combined in interactive workflows. Aljoscha Smolic, Peter Kauff, Sebastian Knorr, Alexander Sorkine-Hornung, Matthias Kunter, Marcus Müller 0001, Manuel Lang |
Proc. IEEE | 1 |
| 2011 | 3D video and free viewpoint video - From capture to display
Aljoscha Smolic |
Pattern Recognit. | 1 |
| 2011 | Computational stereo camera system with programmable control loopabstractStereoscopic 3D has gained significant importance in the entertainment industry. However, production of high quality stereoscopic content is still a challenging art that requires mastering the complex interplay of human perception, 3D display properties, and artistic intent. In this paper, we present a computational stereo camera system that closes the control loop from capture and analysis to automatic adjustment of physical parameters. Intuitive interaction metaphors are developed that replace cumbersome handling of rig parameters using a touch screen interface with 3D visualization. Our system is designed to make stereoscopic 3D production as easy, intuitive, flexible, and reliable as possible. Captured signals are processed and analyzed in real-time on a stream processor. Stereoscopy and user settings define programmable control functionalities, which are executed in real-time on a control processor. Computational power and flexibility is enabled by a dedicated software and hardware architecture. We show that even traditionally difficult shots can be easily captured using our system. Simon Heinzle, Pierre Greisen, David Gallup, Christine Chen, Daniel Saner, Aljoscha Smolic, Andreas Peter Burg, Wojciech Matusik, Markus Gross 0001 |
ACM Trans. Graph. | 6 |
| 2010 | Non-linear warping and warp coding for content-adaptive prediction in advanced video coding applicationsabstractThis paper presents a new concept for scalable video coding, which is content adaptive and art-directable. Video retargeting is applied to scale video between different resolutions and aspect ratios without introducing inacceptable distortions or cutting off content. The non-linear warping operations are integrated into a spatial scalability framework, which includes two new building blocks, i.e. non-linear warping prediction and warp coding. Efficient algorithms for both processes are presented, tested and optimized. The presented results indicate that our non-linear scaling and warp coding algorithms provide efficient performance compared to standard linear scaling methods. Further, our advanced scaling algorithms, i.e. EWA splatting in combination with backward mapping, may be very useful for linear scaling as well. Aljoscha Smolic, Yongzhe Wang, Nikolce Stefanoski, Manuel Lang, Alexander Sorkine-Hornung, Markus Gross 0001 |
ICIP | 1 |
| 2010 | Content-adaptive spatial scalability for scalable video codingabstractThis paper presents an enhancement of the SVC extension of the H.264/AVC standard by content-adaptive spatial scalability (CASS). CASS introduces a novel functionality which is important for high quality content distribution. The video streams (spatial layers), which are used as input to the encoder, are created by content-adaptive and art-directable retargeting of existing high resolution video. Video is retargeted to resolutions and aspect ratios which are mainly dictated by target display devices. Thereby no content is cut off, but visually important content is preserved at the expense of a non-linear distortion of visually unimportant areas. The non-linear dependencies between such video streams are efficiently exploited by CASS for scalable coding. This is achieved by integrating warping-based non-linear texture prediction and warp coding into the SVC framework. The results indicate high prediction accuracy of non-linear predictors and high compression efficiency with limited increase in bit rate and complexity compared to the standard SVC for the case of INTRA only coding. Yongzhe Wang, Nikolce Stefanoski, Xiangzhong Fang, Aljoscha Smolic |
PCS | 4 |
| 2010 | Nonlinear disparity mapping for stereoscopic 3DabstractThis paper addresses the problem of remapping the disparity range of stereoscopic images and video. Such operations are highly important for a variety of issues arising from the production, live broadcast, and consumption of 3D content. Our work is motivated by the observation that the displayed depth and the resulting 3D viewing experience are dictated by a complex combination of perceptual, technological, and artistic constraints. We first discuss the most important perceptual aspects of stereo vision and their implications for stereoscopic content creation. We then formalize these insights into a set of basic disparity mapping operators. These operators enable us to control and retarget the depth of a stereoscopic scene in a nonlinear and locally adaptive fashion. To implement our operators, we propose a new strategy based on stereoscopic warping of the input video streams. From a sparse set of stereo correspondences, our algorithm computes disparity and image-based saliency estimates, and uses them to compute a deformation of the input views so as to meet the target disparities. Our approach represents a practical solution for actual stereo production and display that does not require camera calibration, accurate dense depth maps, occlusion handling, or inpainting. We demonstrate the performance and versatility of our method using examples from live action post-production, 3D display size adaptation, and live broadcast. An additional user study and ground truth comparison further provide evidence for the quality and practical relevance of the presented work. Manuel Lang, Alexander Sorkine-Hornung, Oliver Wang, Steven Poulakos, Aljoscha Smolic, Markus Gross 0001 |
ACM Trans. Graph. | 5 |
| 2009 | An Overview of 3D Video and Free Viewpoint Video
Aljoscha Smolic |
CAIP | 1 |
| 2009 | Coding and intermediate view synthesis of multiview video plus depthabstractFor advanced 3D Video (3DV) applications, efficient data representations are investigated, which only transmit a subset of the views that are required for 3D visualization. From this subset, all intermediate views are synthesized from sample-dense color and depth data. In this paper, the method for reliability-based view synthesis from compressed multi-view + depth data (MVD) is investigated and corresponding results are shown. The initial problem in such 3DV systems is the interdependency between view capturing, coding and view synthesis. For evaluating each component separately, we first generate results from the coding stage only, where color and depth coding is carried out separately. In the next step, we add the view synthesis stage with reliability-based view synthesis and show, how the separate coding results influence the view synthesis quality and what type of artifacts are produced. Efficient bit rate distribution between color and depth is investigated by objective as well as subjective evaluations. Furthermore, quality characteristics across the viewing range for different bit rate distributions are analyzed. Finally, the robustness of the reliability-based view synthesis to coding artifacts is presented. Karsten Müller 0001, Aljoscha Smolic, Kristina Dix, Philipp Merkle, Thomas Wiegand 0001 |
ICIP | 2 |
| 2009 | An overview of available and emerging 3D video formats and depth enhanced stereo as efficient generic solutionabstractRecently, popularity of 3D video has been growing significantly and it may turn into a home user mass market in the near future. However, diversity of 3D video content formats is still hampering wide success. An overview of available and emerging 3D video formats and standards is given, which are mostly related to specific types of applications and 3D displays. This includes conventional stereo video, multiview video, video plus depth, multiview video plus depth and layered depth video. Features and limitations are explained. Finally, depth enhanced stereo (DES) is introduced as a flexible, generic, and efficient 3D video format that can unify all others and serve as universal 3D video format in the future. Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Peter Kauff, Thomas Wiegand 0001 |
PCS | 1 |
| 2009 | The effects of multiview depth video compression on multiview rendering
Philipp Merkle, Yannick Morvan, Aljoscha Smolic, Dirk Farin, Karsten Müller 0001, Peter H. N. de With, Thomas Wiegand 0001 |
Signal Process. Image Commun. | 3 |
| 2008 | Intermediate view interpolation based on multiview video plus depth for advanced 3D video systemsabstractA system for video on multiscopic 3D displays is considered where the data representation consists of multiview video plus scene depth. At most, 3 multiview video signals are being transmitted and used together with the depth data to generate intermediate views at the receiver. The paper presents an approach to such an intermediate view interpolation that separates unreliable image regions along depth discontinuities from reliable image regions. These image regions are processed with different algorithms and then fused to obtain the final interpolated view. In contrast to previous layered approaches, two boundary layers and one reliable layer is used. Moreover, the presented technique does not rely on 3D graphics support but uses image-based 3D warping instead. For enhanced quality intermediate view generation, hole-filling and filtering methods are described. As a result, high quality intermediate views for an existing 9-view auto-stereoscopic display are presented, which prove the suitability of the approach for advanced 3D video (3DV) systems. Aljoscha Smolic, Karsten Müller 0001, Kristina Dix, Philipp Merkle, Peter Kauff, Thomas Wiegand 0001 |
ICIP | 1 |
| 2008 | Reliability-based generation and view synthesis in layered depth videoabstractIn this paper, a system for video rendering on multiscopic 3D displays is considered where the data is represented as layered depth video (LDV). This representation consists of one full or central video with associated per-pixel depth and additional residual layers. Thus, only one full view with additional residual data needs to be transmitted. The LDV data is used at the receiver to generate all intermediate views for the display. The paper presents the LDV layer extraction as well as the view synthesis, using a scene reliability-driven approach. Here, unreliable image regions are detected and in contrast to previous approaches the residual data is enlarged to reduce artifacts in unreliable areas during rendering. To provide maximum data coverage, the residual data remains at its original positions and will not be projected towards the central view. The view synthesis process also uses this reliability analysis to provide higher quality intermediate views than previous approaches. As a final result, high quality intermediate views for an existing 9-view auto-stereoscopic display are presented, which prove the suitability of the LDV approach for advanced 3D video (3DV) systems. Karsten Müller 0001, Aljoscha Smolic, Kristina Dix, Peter Kauff, Thomas Wiegand 0001 |
MMSP | 2 |
| 2007 | Multi-View Video Plus Depth Representation and CodingabstractA study on the video plus depth representation for multi-view video sequences is presented. Such a 3D representation enables functionalities like 3D television and free viewpoint video. Compression is based on algorithms for multi-view video coding, which exploit statistical dependencies from both temporal and inter-view reference pictures for prediction of both color and depth data. Coding efficiency of prediction structures with and without inter-view reference pictures is analyzed for multi-view video plus depth data, reporting gains in luma PSNR of up to 0.5 dB for depth and 0.3 dB for color. The main benefit from using a multi-view video plus depth representation is that intermediate views can be easily rendered. Therefore the impact on image quality of rendered arbitrary intermediate views is investigated and analyzed in a second part, comparing compressed multi-view video plus depth data at different bit rates with the uncompressed original. Philipp Merkle, Aljoscha Smolic, Karsten Müller 0001, Thomas Wiegand 0001 |
ICIP (1) | 2 |
| 2007 | Special issue on three-dimensional video and television
M. Reha Civanlar, Jörn Ostermann, Haldun M. Özaktas, Aljoscha Smolic, John Watson |
Signal Process. Image Commun. | 4 |
| 2007 | Depth map creation and image-based rendering for advanced 3DTV services providing interoperability and scalability
Peter Kauff, Nicole Atzpadin, Christoph Fehn, Marcus Müller 0001, Oliver Schreer, Aljoscha Smolic, Ralf Tanger |
Signal Process. Image Commun. | 6 |
| 2007 | Scene Representation Technologies for 3DTV - A Surveyabstract3-D scene representation is utilized during scene extraction, modeling, transmission and display stages of a 3DTV framework. To this end, different representation technologies are proposed to fulfill the requirements of 3DTV paradigm. Dense point-based methods are appropriate for free-view 3DTV applications, since they can generate novel views easily. As surface representations, polygonal meshes are quite popular due to their generality and current hardware support. Unfortunately, there is no inherent smoothness in their description and the resulting renderings may contain unrealistic artifacts. NURBS surfaces have embedded smoothness and efficient tools for editing and animation, but they are more suitable for synthetic content. Smooth subdivision surfaces, which offer a good compromise between polygonal meshes and NURBS surfaces, require sophisticated geometry modeling tools and are usually difficult to obtain. One recent trend in surface representation is point-based modeling which can meet most of the requirements of 3DTV, however the relevant state-of-the-art is not yet mature enough. On the other hand, volumetric representations encapsulate neighborhood information that is useful for the reconstruction of surfaces with their parallel implementations for multiview stereo algorithms. Apart from the representation of 3-D structure by different primitives, texturing of scenes is also essential for a realistic scene rendering. Image-based rendering techniques directly render novel views of a scene from the acquired images, since they do not require any explicit geometry or texture representation. 3-D human face and body modeling facilitate the realistic animation and rendering of human figures that is quite crucial for 3DTV that might demand real-time animation of human bodies. Physically based modeling and animation techniques produce impressive results, thus have potential for use in a 3DTV framework for modeling and animating dynamic scenes. As a concluding remark, it can be argued that 3-D scene and texture representation techniques are mature enough to serve and fulfill the requirements of 3-D extraction, transmission and display sides in a 3DTV scenario. A. Aydin Alatan, Yücel Yemez, Ugur Güdükbay, Xenophon Zabulis, Karsten Müller 0001, Çigdem Eroglu Erdem, C. Weigel, Aljoscha Smolic |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2007 | Introduction to the Special Section on Multiview Video CodingabstractThe 11 papers in this special section focus on multiview video coding. Jörn Ostermann, Masayuki Tanimoto, Aljoscha Smolic |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2007 | Efficient Prediction Structures for Multiview Video CodingabstractAn experimental analysis of multiview video coding (MVC) for various temporal and inter-view prediction structures is presented. The compression method is based on the multiple reference picture technique in the H.264/AVC video coding standard. The idea is to exploit the statistical dependencies from both temporal and inter-view reference pictures for motion-compensated prediction. The effectiveness of this approach is demonstrated by an experimental analysis of temporal versus inter-view prediction in terms of the Lagrange cost function. The results show that prediction with temporal reference pictures is highly efficient, but for 20% of a picture's blocks on average prediction with reference pictures from adjacent views is more efficient. Hierarchical B pictures are used as basic structure for temporal prediction. Their advantages are combined with inter-view prediction for different temporal hierarchy levels, starting from simulcast coding with no inter-view prediction up to full level inter-view prediction. When using inter-view prediction at key picture temporal levels, average gains of 1.4-dB peak signal-to-noise ratio (PSNR) are reported, while additionally using inter-view prediction at nonkey picture temporal levels, average gains of 1.6-dB PSNR are reported. For some cases, gains of more than 3 dB, corresponding to bit-rate savings of up to 50%, are obtained. Philipp Merkle, Aljoscha Smolic, Karsten Müller 0001, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Coding Algorithms for 3DTV - A SurveyabstractResearch efforts on 3DTV technology have been strengthened worldwide recently, covering the whole media processing chain from capture to display. Different 3DTV systems rely on different 3D scene representations that integrate various types of data. Efficient coding of these data is crucial for the success of 3DTV. Compression of pixel-type data including stereo video, multiview video, and associated depth or disparity maps extends available principles of classical video coding. Powerful algorithms and open international standards for multiview video coding and coding of video plus depth data are available and under development, which will provide the basis for introduction of various 3DTV systems and services in the near future. Compression of 3D mesh models has also reached a high level of maturity. For static geometry, a variety of powerful algorithms are available to efficiently compress vertices and connectivity. Compression of dynamic 3D geometry is currently a more active field of research. Temporal prediction is an important mechanism to remove redundancy from animated 3D mesh sequences. Error resilience is important for transmission of data over error prone channels, and multiple description coding (MDC) is a suitable way to protect data. MDC of still images and 2D video has already been widely studied, whereas multiview video and 3D meshes have been addressed only recently. Intellectual property protection of 3D data by watermarking is a pioneering research area as well. The 3D watermarking methods in the literature are classified into three groups, considering the dimensions of the main components of scene representations and the resulting components after applying the algorithm. In general, 3DTV coding technology is maturating. Systems and services may enter the market in the near future. However, the research area is relatively young compared to coding of other types of media. Therefore, there is still a lot of room for improvement and new development of algorithms. Aljoscha Smolic, Karsten Müller 0001, Nikolce Stefanoski, Jörn Ostermann, Atanas P. Gotchev, Gozde Bozdagi Akar, George A. Triantafyllidis, Alper Koz |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | An Advanced 3DTV Concept Providing Interoperability and Scalability for a Wide Range of Multi-Baseline GeometriesabstractThe paper discusses an advanced approach for 3DTV services that is based on the concept of an N× video-plus-depth data representation. It particularly considers aspects of interoperability, scalability, and adaptability for the case that different multi-baseline geometries are used for multi-view capturing and 3D reproduction. In addition, it presents a method for the creation of depth maps and an algorithm for depth-image-based rendering related to the system approach. Christoph Fehn, Nicole Atzpadin, Marcus Müller 0001, Oliver Schreer, Aljoscha Smolic, Ralf Tanger, Peter Kauff |
ICIP | 5 |
| 2006 | Rate-Distortion Optimization in Dynamic Mesh CompressionabstractRecent developments in the compression of dynamic meshes or mesh sequences have shown that the statistical dependencies within a mesh sequence can be exploited well by predictive coding approaches. Coders introduced so far use experimentally determined or heuristic thresholds for tuning the algorithms. In video coding rate-distortion (RD) optimization is often used to avoid fixing of thresholds and to select a coding mode. We applied these ideas and present here an RD-optimized mesh coder. It includes different prediction modes as well as an RD cost computation that controls the mode selection across all possible spatial partitions of a mesh to find the clustering structure together with the associated prediction modes. The structure of the RD-optimized D3DMC coder is presented, followed by comparative results with mesh sequences at different resolutions. Karsten Müller 0001, Aljoscha Smolic, Matthias Kautzner, Thomas Wiegand 0001 |
ICIP | 2 |
| 2006 | Increasing the Accuracy of the Space-Sweeping Approach to Stereo Reconstruction, using Spherical Backprojection SurfacesabstractIn this paper, interest is focused on the accurate and time-efficient stereo reconstruction, for the purpose of generating 3D animated scenes from multiple synchronized videos. The plane-sweeping approach is reviewed as relevant to the goal of time-efficiency, since its execution can be optimized on a GPU. A method compatible for optimization on the GPU is proposed as a more accurate alternative to plane sweeping and to the derived visibility computation. The method is compared to plane sweeping as to its accuracy, by evaluating the backprojected 3D model against independent views and using n-fold cross validation to estimate the peak signal to noise ratio (PSNR). Finally, the method's output is casted integratable with multicamera stereo reconstruction frameworks. Xenophon Zabulis, Georgios Kordelas, Karsten Müller 0001, Aljoscha Smolic |
ICIP | 4 |
| 2006 | Efficient Compression of Multi-View Video Exploiting Inter-View Dependencies Based on H.264/MPEG4-AVCabstractEfficient multi-view coding requires coding algorithms that exploit temporal, as well as inter-view dependencies between adjacent cameras. Based on a spatiotemporal analysis on the multi-view data set, we present a coding scheme utilizing an H.264/MPEG4-AVC codec. To handle the specific requirements of multi-view datasets, namely temporal and inter-view correlation, two main features of the coder are used: hierarchical B pictures for temporal dependencies and an adapted prediction scheme to exploit inter-view dependencies. Both features are set up in the H.264/MPEG4-AVC configuration file, such that coding and decoding is purely based on standardized software. Additionally, picture reordering before coding to optimize coding efficiency and inverse reordering after decoding to obtain individual views are applied. Finally, coding results are shown for the proposed multi-view coder and compared to simulcast anchor and simulcast hierarchical B picture coding Philipp Merkle, Karsten Müller 0001, Aljoscha Smolic, Thomas Wiegand 0001 |
ICME | 3 |
| 2006 | A Flexible 3D TV System for Different Multi-Baseline GeometriesabstractInteroperability, scalability and adaptability are important features for a successful introduction of future 3D TV services. Hence, new concepts must be able to adapt the multi-view geometry of the capturing system to the geometry of the 3D reproduction systems. An approach is discussed, which considers these adaptation issues based on the concept of an Ntimesvideo-plus-depth data representation. The core algorithms for depth map creation on the analysis side and depth image based rendering on the reproduction side are presented Oliver Schreer, Christoph Fehn, Nicole Atzpadin, Marcus Müller 0001, Aljoscha Smolic, Ralf Tanger, Peter Kauff |
ICME | 5 |
| 2006 | 3D Video and Free Viewpoint Video - Technologies, Applications and MPEG StandardsabstractAn overview of 3D and free viewpoint video is given in this paper with special focus on related standardization activities in MPEG. Free viewpoint video allows the user to freely navigate within real world visual scenes, as known from virtual worlds in computer graphics. Examples are shown, highlighting standards conform realization using MPEG-4. Then the principles of 3D video are introduced providing the user with a 3D depth impression of the observed scene. Example systems are described again focusing on their realization based on MPEG-4. Finally multi-view video coding is described as a key component for 3D and free viewpoint video systems. The conclusion is that the necessary technology including standard media formats for 3D and free viewpoint is available or will be available in the near future, and that there is a clear demand from industry and user side for such applications. 3D TV at home and free viewpoint video on DVD will be available soon, and will create huge new markets Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Christoph Fehn, Peter Kauff, Peter Eisert, Thomas Wiegand 0001 |
ICME | 1 |
| 2006 | Rate-distortion-optimized predictive compression of dynamic 3D mesh sequences
Karsten Müller 0001, Aljoscha Smolic, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
Signal Process. Image Commun. | 2 |
| 2005 | Predictive compression of dynamic 3D meshesabstractAn efficient algorithm for compression of dynamic time-consistent 3D meshes is presented. Such a sequence of meshes contains a large degree of temporal statistical dependencies that can be exploited for compression using DPCM. The vertex positions are predicted at the encoder from a previously decoded mesh. The difference vectors are further clustered in an octree approach. Only a representative for a cluster of difference vectors is further processed providing a significant reduction of data rate. The representatives are scaled and quantized and finally entropy coded using CABAC, the arithmetic coding technique used in H.264/MPEG4-AVC. The mesh is then reconstructed at the encoder for prediction of the next mesh. In our experiments we compare the efficiency of the proposed algorithm in terms of bit-rate and quality compared to static mesh coding and interpolator compression indicating a significant improvement in compression efficiency. Karsten Müller 0001, Aljoscha Smolic, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
ICIP (1) | 2 |
| 2005 | A generic and automatic content-based approach for improved H.264/MPEG4-AVC video codingabstractA new content-based approach for improved H.264/MPEG4-AVC video coding is presented. The framework is generic because it is based on a closed-loop texture analysis by synthesis algorithm that can automatically identify and recover from video quality impairments through artifact detectors and appropriate countermeasures. The algorithm is flexible, for it can in principle be integrated into any standards-compliant video codec. The fundamental assumption of our approach is that many video scenes can be classified into subjectively relevant and irrelevant textures. The texture categorization is thereby done by a texture analyzer (encoder side), while the corresponding texture synthesizer performs the replacement of the subjectively irrelevant textures (decoder side), given the side information generated by the texture analyzer. When implementing the proposed approach into an H.264/MPEG4-AVC codec, bit rate savings of up to 33.3% compared to an H.264/MPEG4-AVC video codec without our approach are reported. Patrick Ndjiki-Nya, Tobias Hinz, Aljoscha Smolic, Thomas Wiegand 0001 |
ICIP (2) | 3 |
| 2005 | Interactive 3-D Video Representation and Coding TechnologiesabstractInteractivity in the sense of being able to explore and navigate audio-visual scenes by freely choosing viewpoint and viewing direction, is an important key feature of new and emerging audio-visual media. This paper gives an overview of suitable technology for such applications, with a focus on international standards, which are beneficial for consumers, service providers, and manufacturers. We first give a general classification and overview of interactive scene representation formats as commonly used in computer graphics literature. Then, we describe popular standard formats for interactive three-dimensional (3-D) scene representation and creation of virtual environments, the virtual reality modeling language (VRML), and the MPEG-4 BInary Format for Scenes (BIFS) with some examples. Recent extensions to MPEG-4 BIFS, the Animation Framework eXtension (AFX), providing advanced computer graphics tools, are explained and illustrated. New technologies mainly targeted at reconstruction, modeling, and representation of dynamic real world scenes are further studied. The user shall be able to navigate photorealistic scenes within certain restrictions, which can be roughly defined as 3-D video. Omnidirectional video is an extension of the planar two-dimensional (2-D) image plane to a spherical or cylindrical image plane. Any 2-D view in any direction can be rendered from this overall recording to give the user the impression of looking around. In interactive stereo two views, one for each eye, are synthesized to provide the user with an adequate depth cue of the observed scene. Head motion parallax viewing can be supported in a certain operating range if sufficient depth or disparity data are delivered with the video data. In free viewpoint video, a dynamic scene is captured by a number of cameras. The input data are transformed into a special data representation that enables interactive navigation through the dynamic scene environment. Aljoscha Smolic, Peter Kauff |
Proc. IEEE | 1 |
| 2005 | 3-D Reconstruction of a Dynamic Environment With a Fully Calibrated Background for Traffic ScenesabstractVision-based traffic surveillance systems are more and more employed for traffic monitoring, collection of statistical data and traffic control. We present an extension of such a system that additionally uses the captured image content for 3-D scene modeling and reconstruction. A basic goal of surveillance systems is to get a good coverage of the observed area with as few cameras as possible to keep the costs low. Therefore, the 3-D reconstruction has to be done from only a few original views with limited overlap and different lighting conditions. To cope with these specific restrictions we developed a model-based 3-D reconstruction scheme that exploits a priori knowledge about the scene. The system is fully calibrated offline by estimating camera parameters from measured 3-D-2-D correspondences. Then the scene is divided into static parts, which are modeled offline and dynamic parts, which are processed online. Therefore, we segment all views into moving objects and static background. The background is modeled as multitexture planes using the original camera textures. Moving objects are segmented and tracked in each view. All segmented views of a moving object are combined to a 3-D object, which is positioned and tracked in 3-D. Here we use predefined geometric primitives and map the original textures onto them. Finally the static and dynamic elements are combined to create the reconstructed 3-D scene, where the user can freely navigate, i.e., choose an arbitrary viewpoint and direction. Additionally, the system allows analyzing the 3-D properties of the scene and the moving objects. Karsten Müller 0001, Aljoscha Smolic, Michael Drose, Patrick Voigt, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Free viewpoint video extraction, representation, coding, and renderingabstractFree viewpoint video provides the possibility to freely navigate within dynamic real world video scenes by choosing arbitrary viewpoints and view directions. So far, related work only considered free viewpoint video extraction, representation, and rendering methods. Compression and transmission has not yet been studied in detail and combined with the other components into one complete system. In this paper, we present such a complete system for efficient free viewpoint video extraction, representation, coding, and interactive rendering. Data representation is based on 3D mesh models and view-dependent texture mapping using video textures. The geometry extraction is based on a shape-from-silhouette algorithm. The resulting voxel models are converted into 3D meshes that are coded using MPEG-4 SNHC tools. The corresponding video textures are coded using an H.264/AVC codec. Our algorithms for view-dependent texture mapping have been adopted as an extension of MPEG-4 AFX. The presented results illustrate that based on the proposed methods a complete transmission system for efficient free viewpoint video can be built. Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Tobias Rein, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
ICIP | 1 |
| 2004 | Representation, coding, and rendering of 3D video objects with MPEG-4 and H.264/AVCabstract3D video objects provide the same functionalities as virtual computer graphics objects but depict the motion and appearance of real world moving objects. They can be viewed interactively from any direction and integrated in complete 3D scenes with other virtual and real world elements. So far, related work only considered extraction, representation, and rendering methods. Compression and transmission has not yet been studied in detail and combined with the other components into one complete system. In this paper, we present such a complete system for efficient 3D video object extraction, representation, coding, and interactive rendering. Data representation is based on 3D mesh models and view-dependent texture mapping using video textures. The geometry extraction is based on a shape-from-silhouette algorithm. The resulting voxel models are converted into 3D meshes that are coded using MPEG-4 SNHC tools. The corresponding video textures are preprocessed taking the object's shape into account and coded using an H.264/AVC codec. The presented results illustrate that based on the proposed methods a complete transmission system for 3D video objects can be built. Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Tobias Rein, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
MMSP | 1 |
| 2004 | Improved video coding using long-term global motion compensationabstractWe present a new approach to video coding that applies video analysis based on global motion features. A super-resolution mosaic is build for each frame to be encoded from a number of previously transmitted frames. This super-resolution mosaic is used to detect macroblocks that are only affected by global motion. For such macroblocks no prediction error is transmitted. They are purely reconstructed by prediction from the super-resolution mosaic which results in significant bitrate savings. Our results indicate total bitrate savings of 20% and more compared to a state of the art H.264/AVC codec at the same visual quality. Aljoscha Smolic, Yuriy Vatis, Heiko Schwarz, Peter Kauff, Ulrich Gölz, Thomas Wiegand 0001 |
VCIP | 1 |
| 2004 | 3DAV exploration of video-based rendering technology in MPEGabstractNew kinds of media are emerging that extend the functionality of available technology. The growth of immersive recording technologies has led to video-based rendering systems for photographing and reproducing environments in motion. This lends itself to new forms of interactivity for the viewer, including the ability to explore a photographic scene and interact with its features. The three-dimensional (3-D) qualities of objects in the scene can be extracted by analysis techniques and displayed by the use of stereo vision. The data types and image bandwidth needed for this type of media experience may require especially efficient formats for representation, coding, and transmission. An overview is presented here of the MPEG activity exploring the need for standardization in this area to support these new applications, under the name of 3DAV (for 3-D audio-visual). As an example, a detailed solution for omnidirectional video is presented as one of the application scenarios in 3DAV. Aljoscha Smolic, David McCutchen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Improved H.264/AVC coding using texture analysis and synthesisabstractWe assume that the textures in a video scene is classified into two categories: textures with unimportant subjective details and the remainder. We utilize this assumption for improved video coding using a texture analyzer and a texture synthesizer. The texture analyzer identifies the texture regions with unimportant subjective details and generates coarse masks as well as side information for the texture synthesizer at the decoder side. The texture synthesizer replaces the identified textures by inserting synthetic textures for the identified regions. The texture analyzer is based on MPEG-7 descriptors. Our approach is integrated into an H.264/AVC codec. Bit-rate savings up to 19.4% are shown for a semi-automatic texture analyzer given similar subjective quality as the H.264/AVC codec without the presented approach. Patrick Ndjiki-Nya, Bela Makai, Gabi Blattermann, Aljoscha Smolic, Heiko Schwarz, Thomas Wiegand 0001 |
ICIP (3) | 4 |
| 2003 | Multi-texture modeling of 3D traffic scenesabstractWe present a system for 3D reconstruction of traffic scenes. Traffic surveillance is a challenging scenario for 3D reconstruction in cases, where only a small number of views is available that do not contain much overlap. We address the possibilities and restrictions for modeling such scenarios with only a few cameras and introduce a compositor that allows rendering of the semi automatically generated 3D scenes. Some of the occurring problems concern camera images, which might show a common background area, but can still differ drastically in lighting effects. For foreground objects nearly no common visual information might be available, as angles between cameras may exceed even 90/spl deg/. Karsten Müller 0001, Aljoscha Smolic, Michael Drose, Patrick Voigt, Thomas Wiegand 0001 |
ICME | 2 |
| 2002 | Efficient representation and interactive streaming of high-resolution panoramic viewsabstractA new system for interactive streaming of high-resolution 360/spl deg/ panoramic views over the Internet is presented. The scene is represented very efficiently using MPEG-4 BIFS and displayed at the client using the HHI 3-D MPEG-4 player. The user can navigate interactively through the scene. The navigation decisions are evaluated to trigger the streaming of the data needed to build the visible screen view and to ensure a fluent visualization of the scene. The system layout is designed to support other, more complex photo realistic 3-D environments later on and, thus, will enable the provision of a variety of new, interactive services over the Internet. Carsten Grünheit, Aljoscha Smolic, Thomas Wiegand 0001 |
ICIP (3) | 2 |
| 2001 | High-resolution video mosaicingabstractThe generation of high-resolution mosaics from multiple images is presented. The estimation of the high-resolution signal utilizes a motion-compensated filtering approach for video frames that are affected by spatial aliasing. For that, global motion models and the corresponding estimation algorithms are employed to provide a very accurate and continuous description of sub-pixel motion for the background of the video signal. The motion information is exploited to generate mosaics with four or sixteen times the resolution of the constructing video sequence or known mosaicing techniques. The mosaicing results indicate that the presented multi-frame methods provide superior visual quality. The approach is further extended to the generation of high-resolution video. Aljoscha Smolic, Thomas Wiegand 0001 |
ICIP (3) | 1 |
| 2000 | Low-Complexity Global Motion Estimation from P-Frame Motion Vectors for MPEG-7 ApplicationsabstractWe present an algorithm for low-complexity global motion estimation, that works with block-coded video (e.g. MPEG-2). A superimposed global motion model is fitted to decoded P-frame motion vector fields, providing real-time performance. Therefore a robust M-estimator is applied since a lot of outliers have to be expected in the motion vector fields. The algorithm is compared to a state-of-the-art global motion estimator. The main characteristics of the global motion are captured accurately in most cases, although estimation may fail. We also report the integration of the estimator into a complete MPEG-7 system for search-and-retrieval of video based on motion characteristics. Aljoscha Smolic, Michael Hoeynck, Jens-Rainer Ohm |
ICIP | 1 |
| 2000 | Robust Global Motion Estimation Using a Simplified M-Estimator ApproachabstractGlobal motion estimation is an important task in a variety of video processing applications, such as coding, segmentation, classification/indexing or mosaicing. Due to the possible presence of differently moving foreground objects and other sources of distortions, robust methods such as M-estimators have to be applied. We present a simplified implementation of a robust M-estimator for global motion estimation that does not increase the computational complexity significantly compared to a non-robust estimator, while providing excellent results in terms of estimation accuracy. Additionally, unstructured image regions are detected and rejected for the estimation. This avoids aperture problems, that can have an bad impact especially on robust estimators that rely on a certain error measure. Aljoscha Smolic, Jens-Rainer Ohm |
ICIP | 1 |
| 2000 | A set of visual feature descriptors and their combination in a low-level description scheme
Jens-Rainer Ohm, F. Bunjamin, Wolfram Liebsch, Bela Makai, Karsten Müller 0001, Aljoscha Smolic, D. Zier |
Signal Process. Image Commun. | 6 |
| 1999 | A multi-feature description scheme for image and video database retrievalabstractThis paper reports about a description scheme for visual information content, which has been developed in the context of the forthcoming MPEG-7 standard. The system supports similarity-based retrieval of visual (image and video) data along feature axes like color, texture, shape/geometry and motion. The descriptors for these features have been developed in a way such that invariance against common transformations of visual material, e.g. filtering, contrast/color manipulation, resizing etc. is achieved, and that they are fitted to human perception properties. Furthermore, descriptors have been designed that allow a fast, hierarchical search procedure. A search engine has been developed on the basis of this description scheme, which allows similarity-based retrieval from an image or video database. The results show that efficient search and retrieval in visual database systems is possible based on a normative feature description such as MPEG-7. Jens-Rainer Ohm, F. Bunjamin, Wolfram Liebsch, Bela Makai, Karsten Müller 0001, Aljoscha Smolic, D. Zier |
MMSP | 6 |
| 1999 | Real-time estimation of long-term 3-D motion parameters for SNHC face animation and model-based coding applicationsabstractWe present two recursive methods for the real-time estimation of long-term three-dimensional (3-D) motion parameters from monocular image sequences suitable for synthetic/natural hybrid coding face animation and model-based coding applications. Based on feature point extractions in energy frame, the 3-D motion parameters of a human face are estimated with a predictive approach. The first method uses a recursive linear least squares approach and the second employs a nonlinear extended Kalman filter, which does not rely on a linearized model of the face motion. Both methods perform a prediction and correction loop at every time step. Compared to other methods described in the literature, the recursive and predictive structure of the proposed estimation process solves the problem of error accumulation in long-term motion estimation. This makes the estimation stable and consistent over long periods. Experimental results are presented for synthetic data and real image sequences, which demonstrate the performance of the estimation methods and compare the two approaches. Aljoscha Smolic, Bela Makai, Thomas Sikora |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | Long-term global motion estimation and its application for sprite coding, content description, and segmentationabstractWe present a new technique for long-term global motion estimation of image objects. The estimated motion parameters describe the continuous and time-consistent motion over the whole sequence relatively to a fixed reference coordinate system. The proposed method is suitable for the estimation of affine motion parameters as well as for higher order motion models like the parabolic model-combining the advantages of feature matching and optical flow techniques. A hierarchical strategy is applied for the estimation, first translation, affine motion, and finally higher order motion parameters, which is robust and computationally efficient. A closed-loop prediction scheme is applied to avoid the problem of error accumulation in long-term motion estimation. The presented results indicate that the proposed technique is a very accurate and robust approach for long-term global motion estimation, which can be used for applications such as MPEG-4 sprite coding or MPEG-7 motion description. We also show that the efficiency of global motion estimation can be significantly increased if a higher order motion model is applied, and we present a new sprite coding scheme for on-line applications. We further demonstrate that the proposed estimator serves as a powerful tool for segmentation of video sequences. Aljoscha Smolic, Thomas Sikora, Jens-Rainer Ohm |
IEEE Trans. Circuits Syst. Video Technol. | 1 |