EDBT 2026 Demo / reviewers in the wild / expert
Karol Myszkowski
dblp:m/KMyszkowski
· DBLP profile ↗
118ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-8505-4141ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 113 · 6 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Forget Superresolution, Sample Adaptively (when Path Tracing)abstractReal-time path tracing increasingly operates under extremely low sampling budgets, often below one sample per pixel, as rendering complexity, resolution, and frame-rate requirements continue to rise. Superresolution is widely used in production because it reduces path-tracing cost by tracing rays on a coarser image grid and reconstructing missing details. This creates a uniform tradeoff between cost and spatial detail: every image region receives the same reduced ray budget, although path-tracing noise, reconstruction difficulty, and perceptual importance vary strongly across the image. Adaptive sampling offers a compelling alternative, but existing end-to-end approaches rely on approximations that break down in sparse regimes. We introduce an end-to-end adaptive sampling and denoising pipeline explicitly designed for the sub-1-spp regime. Our method uses a stochastic formulation of sample placement that enables gradient estimation despite discrete sampling decisions, allowing stable training of a neural sampler at low sampling budgets. To better align optimization with human perception, we propose a tone-mapping-aware training pipeline that integrates differentiable filmic operators and a state-of-the-art perceptual loss, preventing oversampling of regions with low visual impact. In addition, we introduce a gather-based pyramidal denoising filter and a learnable generalization of albedo demodulation tailored to sparse sampling. Our results show consistent improvements over uniform sparse sampling, with notably better reconstruction of perceptually critical details such as specular highlights and shadow boundaries, and demonstrate that adaptive sampling remains effective in the sub-1-spp regime. Martin Bálint, Corentin Salaün, Hans-Peter Seidel, Karol Myszkowski |
ACM Trans. Graph. | 4 |
| 2026 | Overdriving Visual Depth Perception via Sound Modulation in VRabstractOur ability to perceive and navigate the spatial world is a cornerstone of human experience, relying on the integration of visual and auditory cues to form a coherent sense of depth and distance. In stereoscopic 3D vision, depth perception requires fixation of both eyes on a target object, which is achieved through vergence movements, with convergence for near objects and divergence for distant ones. In contrast, auditory cues provide complementary depth information through variations in loudness, interaural differences (IAD), and the frequency spectrum. We investigate the interaction between visual and auditory cues and examine how contradictory auditory information can overdrive visual depth perception in virtual reality (VR). When a new visual target appears, we introduce a spatial discrepancy between the visual and auditory cues: the visual target is shifted closer to the previously fixated object, while the corresponding sound localization is displaced in the opposite direction. By integrating these conflicting cues through multimodal processing, the resulting percept is biased toward the intended depth location. This audiovisual fusion counteracts depth compression, thus reducing the required vergence magnitude and enabling faster gaze retargeting. Such audio-driven depth enhancement may further help mitigate the vergence-accommodation conflict (VAC) in scenarios where physical depth must be compressed. In a series of psychophysical studies, we first assess the efficiency of depth overdriving for various VR-relevant combinations of initial fixations and shifted target locations, considering different scenarios of audio displacements and their loudness and frequency parameters. Next, we quantify the resulting speedup in gaze retargeting for target shifts that can be successfully overdriven by sound manipulations. Finally, we apply our method in a naturalistic VR scenario where user interface interactions with the scene show an extended perceptual depth. Daniel Jiménez Navarro, Colin Groth, Jorge Pina, Qi Sun 0003, Praneeth Chakravarthula, Karol Myszkowski, Hans-Peter Seidel, Ana Serrano |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | LEDiff: Latent Exposure Diffusion for HDR GenerationabstractWhile consumer displays increasingly support more than 10 stops of dynamic range, most image assets — such as internet photographs and generative AI content — remain limited to 8-bit low dynamic range (LDR), constraining their utility across high dynamic range (HDR) applications. Currently, no generative model can produce high-bit, high-dynamic range content in a generalizable way. Existing LDR-to-HDR conversion methods often struggle to produce photorealistic details and physically-plausible dynamic range in the clipped areas. We introduce LEDiff, a method that enables a generative model with HDR content generation through latent space fusion inspired by image-space exposure fusion techniques. It also functions as an LDR-to-HDR converter, expanding the dynamic range of existing low-dynamic range images. Our approach uses a small HDR dataset to enable a pretrained diffusion model to recover detail and dynamic range in clipped highlights and shadows. LEDiff brings HDR capabilities to existing generative models and converts any LDR image to HDR, creating photorealistic HDR outputs for image generation, image-based lighting (HDR environment map generation), and photographic effects such as depth of field simulation, where linear HDR data is essential for realistic quality. Chao Wang 0037, Zhihao Xia, Thomas Leimkühler, Karol Myszkowski, Xuaner Cecilia Zhang |
CVPR | 4 |
| 2025 | Why Slow Feels Fast and Fast Feels Slow: Evaluating and Predicting Speed MisperceptionabstractHuman perception of speed is largely driven by visual cues. However, our subjective estimations of speed are influenced by several factors that can lead to deceptive cues and speed misperception. While some prior studies have explored individual effects on speed perception, such as contrast, spatial frequency, and temporal frequency, their combined influence remains underexamined, particularly in immersive VR environments. In this work, we systematically investigate the influence and interplay of four visual factors—contrast, spatial frequency, temporal frequency, and eccentricity—on human perception of speed. To this end, we conduct a psychophysical study measuring subjective speed judgments across controlled stimuli and reveal significant perceptual biases induced by these factors. Based on our collected data, we learn a model to predict the underestimation or overestimation of perceived speed from visual scene properties. We apply and validate our findings in three immersive environments and demonstrate their influence on common VR scenarios. Finally, we discuss how understanding the factors that shape speed perception can drive the design of perceptually aligned virtual environments, with potential future applications such as correcting speed misperception and conceivably mitigating visual-vestibular conflicts by modulating perceived speed. Colin Groth, Daniel Jiménez Navarro, Zihao Zou, Ana Serrano, Karol Myszkowski, Qi Sun 0003, Praneeth Chakravarthula |
ISMAR | 7 |
| 2025 | Bracket Diffusion: HDR Image Generation by Consistent LDR DenoisingabstractAbstract We demonstrate generating HDR images using the concerted action of multiple black‐box, pre‐trained LDR image diffusion models. Common diffusion models are not HDR as, first, there is no sufficiently large HDR image dataset available to re‐train them, and, second, even if it was, re‐training such models is impossible for most compute budgets. Instead, we seek inspiration from the HDR image capture literature that traditionally fuses sets of LDR images, called “exposure brackets”, to produce a single HDR image. We operate multiple denoising processes to generate multiple LDR brackets that together form a valid HDR result. To this end, we introduce a brackets consistency term into the diffusion process to couple the brackets such that they agree across the exposure range they share. We demonstrate HDR versions of state‐of‐the‐art unconditional and conditional as well as restoration‐type (LDR2HDR) generative modeling. Mojtaba Bemana, Thomas Leimkühler, Karol Myszkowski, Hans-Peter Seidel, Tobias Ritschel 0001 |
Comput. Graph. Forum | 3 |
| 2025 | MILO: A Lightweight Perceptual Quality Metric for Image and Latent-Space OptimizationabstractWe present MILO (Metric for Image- and Latent-space Optimization), a lightweight, multiscale, perceptual metric for full-reference image quality assessment (FR-IQA). MILO is trained using pseudo-MOS (Mean Opinion Score) supervision, in which reproducible distortions are applied to diverse images and scored via an ensemble of recent quality metrics that account for visual masking effects. This approach enables accurate learning without requiring large-scale human-labeled datasets. Despite its compact architecture, MILO outperforms existing metrics across standard FR-IQA benchmarks and offers fast inference suitable for real-time applications. Beyond quality prediction, we demonstrate the utility of MILO as a perceptual loss in both image and latent domains. In particular, we show that spatial masking modeled by MILO, when applied to latent representations from a VAE encoder within Stable Diffusion, enables efficient and perceptually aligned optimization. By combining spatial masking with a curriculum learning strategy, we first process perceptually less relevant regions before progressively shifting the optimization to more visually distorted areas. This strategy leads to significantly improved performance in tasks like denoising, super-resolution, and face restoration, while also reducing computational overhead. MILO thus functions as both a state-of-the-art image quality metric and as a practical tool for perceptual optimization in generative pipelines. Ugur Çogalan, Mojtaba Bemana, Karol Myszkowski, Hans-Peter Seidel, Colin Groth |
ACM Trans. Graph. | 3 |
| 2024 | Enhancing image quality prediction with self-supervised visual maskingabstractAbstract Full‐reference image quality metrics (FR‐IQMs) aim to measure the visual differences between a pair of reference and distorted images, with the goal of accurately predicting human judgments. However, existing FR‐IQMs, including traditional ones like PSNR and SSIM and even perceptual ones such as HDR‐VDP, LPIPS, and DISTS, still fall short in capturing the complexities and nuances of human perception. In this work, rather than devising a novel IQM model, we seek to improve upon the perceptual quality of existing FR‐IQM methods. We achieve this by considering visual masking, an important characteristic of the human visual system that changes its sensitivity to distortions as a function of local image content. Specifically, for a given FR‐IQM metric, we propose to predict a visual masking model that modulates reference and distorted images in a way that penalizes the visual errors based on their visibility. Since the ground truth visual masks are difficult to obtain, we demonstrate how they can be derived in a self‐supervised manner solely based on mean opinion scores (MOS) collected from an FR‐IQM dataset. Our approach results in enhanced FR‐IQM metrics that are more in line with human prediction both visually and quantitatively. Ugur Çogalan, Mojtaba Bemana, Hans-Peter Seidel, Karol Myszkowski |
Comput. Graph. Forum | 4 |
| 2024 | Cinematic Gaussians: Real-Time HDR Radiance Fields with Depth of FieldabstractAbstract Radiance field methods represent the state of the art in reconstructing complex scenes from multi‐view photos. However, these reconstructions often suffer from one or both of the following limitations: First, they typically represent scenes in low dynamic range (LDR), which restricts their use to evenly lit environments and hinders immersive viewing experiences. Secondly, their reliance on a pinhole camera model, assuming all scene elements are in focus in the input images, presents practical challenges and complicates refocusing during novel‐view synthesis. Addressing these limitations, we present a lightweight method based on 3D Gaussian Splatting that utilizes multi‐view LDR images of a scene with varying exposure times, apertures, and focus distances as input to reconstruct a high‐dynamic‐range (HDR) radiance field. By incorporating analytical convolutions of Gaussians based on a thin‐lens camera model as well as a tonemapping module, our reconstructions enable the rendering of HDR content with flexible refocusing capabilities. We demonstrate that our combined treatment of HDR and depth of field facilitates real‐time cinematic rendering, outperforming the state of the art. Chao Wang 0037, Krzysztof Wolski, Bernhard Kerbl, Ana Serrano, Mojtaba Bemana, Hans-Peter Seidel, Karol Myszkowski, Thomas Leimkühler |
Comput. Graph. Forum | 7 |
| 2024 | Learning Images Across Scales Using Adversarial TrainingabstractThe real world exhibits rich structure and detail across many scales of observation. It is difficult, however, to capture and represent a broad spectrum of scales using ordinary images. We devise a novel paradigm for learning a representation that captures an orders-of-magnitude variety of scales from an unstructured collection of ordinary images. We treat this collection as a distribution of scale-space slices to be learned using adversarial training, and additionally enforce coherency across slices. Our approach relies on a multiscale generator with carefully injected procedural frequency content, which allows to interactively explore the emerging continuous scale space. Training across vastly different scales poses challenges regarding stability, which we tackle using a supervision scheme that involves careful sampling of scales. We show that our generator can be used as a multiscale generative model, and for reconstructions of scale spaces from unstructured patches. Significantly outperforming the state of the art, we demonstrate zoom-in factors of up to 256x at high quality and scale consistency. Krzysztof Wolski, Adarsh Djeacoumar, Alireza Javanmardi, Hans-Peter Seidel, Christian Theobalt, Guillaume Cordonnier, Karol Myszkowski, George Drettakis, Xingang Pan, Thomas Leimkühler |
ACM Trans. Graph. | 7 |
| 2024 | Measuring and Predicting Multisensory Reaction Latency: A Probabilistic Model for Visual-Auditory IntegrationabstractVirtual/augmented reality (VR/AR) devices offer both immersive imagery and sound. With those wide-field cues, we can simultaneously acquire and process visual and auditory signals to quickly identify objects, make decisions, and take action. While vision often takes precedence in perception, our visual sensitivity degrades in the periphery. In contrast, auditory sensitivity can exhibit an opposite trend due to the elevated interaural time difference. What occurs when these senses are simultaneously integrated, as is common in VR applications such as 360° video watching and immersive gaming? We present a computational and probabilistic model to predict VR users' reaction latency to visual-auditory multisensory targets. To this aim, we first conducted a psychophysical experiment in VR to measure the reaction latency by tracking the onset of eye movements. Experiments with numerical metrics and user studies with naturalistic scenarios showcase the model's accuracy and generalizability. Lastly, we discuss the potential applications, such as measuring the sufficiency of target appearance duration in immersive video playback, and suggesting the optimal spatial layouts for AR interface design. Daniel Jiménez Navarro, Ana Serrano, Karol Myszkowski, Qi Sun 0003 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | GlowGAN: Unsupervised Learning of HDR Images from LDR Images in the WildabstractMost in-the-wild images are stored in Low Dynamic Range (LDR) form, serving as a partial observation of the High Dynamic Range (HDR) visual world. Despite limited dynamic range, these LDR images are often captured with different exposures, implicitly containing information about the underlying HDR image distribution. Inspired by this intuition, in this work we present, to the best of our knowledge, the first method for learning a generative model of HDR images from in-the-wild LDR image collections in a fully unsupervised manner. The key idea is to train a generative adversarial network (GAN) to generate HDR images which, when projected to LDR under various exposures, are indistinguishable from real LDR images. The projection from HDR to LDR is achieved via a camera model that captures the stochasticity in exposure and camera response function. Experiments show that our method GlowGAN can synthesize photorealistic HDR images in many challenging cases such as landscapes, lightning, or windows, where previous supervised generative models produce overexposed images. With the assistance of GlowGAN, we showcase the novel application of unsupervised inverse tone mapping (GlowGAN-ITM) that sets a new paradigm in this field. Unlike previous methods that gradually complete information from LDR input, GlowGAN-ITM searches the entire HDR image manifold modeled by GlowGAN for the HDR images which can be mapped back to the LDR input. GlowGAN-ITM achieves more realistic reconstruction of overexposed regions compared to state-of-the-art supervised learning models, despite not requiring HDR images or paired multi-exposure images for training. Chao Wang 0037, Ana Serrano, Xingang Pan, Bin Chen 0019, Karol Myszkowski, Hans-Peter Seidel, Christian Theobalt, Thomas Leimkühler |
ICCV | 5 |
| 2023 | Joint Sampling and Optimisation for Inverse RenderingabstractWhen dealing with difficult inverse problems such as inverse rendering, using Monte Carlo estimated gradients to optimise parameters can slow down convergence due to variance. Averaging many gradient samples in each iteration reduces this variance trivially. However, for problems that require thousands of optimisation iterations, the computational cost of this approach rises quickly. Martin Bálint, Karol Myszkowski, Hans-Peter Seidel, Gurprit Singh |
SIGGRAPH Asia | 2 |
| 2023 | The effect of display capabilities on the gloss consistency between real and virtual objectsabstractA faithful reproduction of gloss is inherently difficult because of the limited dynamic range, peak luminance, and 3D capabilities of display devices. This work investigates how the display capabilities affect gloss appearance with respect to a real-world reference object. To this end, we employ an accurate imaging pipeline to achieve a perceptual gloss match between a virtual and real object presented side-by-side on an augmented-reality high-dynamic-range (HDR) stereoscopic display, which has not been previously attained to this extent. Based on this precise gloss reproduction, we conduct a series of gloss matching experiments to study how gloss perception degrades based on individual factors: object albedo, display luminance, dynamic range, stereopsis, and tone mapping. We support the study with a detailed analysis of individual factors, followed by an in-depth discussion on the observed perceptual effects. Our experiments demonstrate that stereoscopic presentation has a limited effect on the gloss matching task on our HDR display. However, both reduced luminance and dynamic range of the display reduce the perceived gloss. This means that the visual system cannot compensate for the changes in gloss appearance across luminance (lack of gloss constancy), and the tone mapping operator should be carefully selected when reproducing gloss on a low dynamic range (LDR) display. Bin Chen 0019, Akshay Jindal, Michal Piovarci, Chao Wang 0037, Hans-Peter Seidel, Piotr Didyk, Karol Myszkowski, Ana Serrano, Rafal Mantiuk |
SIGGRAPH Asia | 7 |
| 2023 | Perceptual error optimization for Monte Carlo animation renderingabstractIndependently estimating pixel values in Monte Carlo rendering results in a perceptually sub-optimal white-noise distribution of error in image space. Recent works have shown that perceptual fidelity can be improved significantly by distributing pixel error as blue noise instead. Most such works have focused on static images, ignoring the temporal perceptual effects of animation display. We extend prior formulations to simultaneously consider the spatial and temporal domains, and perform an analysis to motivate a perceptually better spatio-temporal error distribution. We then propose a practical error optimization algorithm for spatio-temporal rendering and demonstrate its effectiveness in various configurations. Misa Korac, Corentin Salaün, Iliyan Georgiev, Pascal Grittmann, Philipp Slusallek, Karol Myszkowski, Gurprit Singh |
SIGGRAPH Asia | 6 |
| 2023 | Video frame interpolation for high dynamic range sequences captured with dual-exposure sensorsabstractAbstract Video frame interpolation (VFI) enables many important applications such as slow motion playback and frame rate conversion. However, one major challenge in using VFI is accurately handling high dynamic range (HDR) scenes with complex motion. To this end, we explore the possible advantages of dual‐exposure sensors that readily provide sharp short and blurry long exposures that are spatially registered and whose ends are temporally aligned. This way, motion blur registers temporally continuous information on the scene motion that, combined with the sharp reference, enables more precise motion sampling within a single camera shot. We demonstrate that this facilitates a more complex motion reconstruction in the VFI task, as well as HDR frame reconstruction that so far has been considered only for the originally captured frames, not in‐between interpolated frames. We design a neural network trained in these tasks that clearly outperforms existing solutions. We also propose a metric for scene motion complexity that provides important insights into the performance of VFI methods at test time. Ugur Çogalan, Mojtaba Bemana, Hans-Peter Seidel, Karol Myszkowski |
Comput. Graph. Forum | 4 |
| 2023 | Learning GAN-Based Foveated Reconstruction to Recover Perceptually Important Image FeaturesabstractA foveated image can be entirely reconstructed from a sparse set of samples distributed according to the retinal sensitivity of the human visual system, which rapidly decreases with increasing eccentricity. The use of generative adversarial networks (GANs) has recently been shown to be a promising solution for such a task, as they can successfully hallucinate missing image information. As in the case of other supervised learning approaches, the definition of the loss function and the training strategy heavily influence the quality of the output. In this work,we consider the problem of efficiently guiding the training of foveated reconstruction techniques such that they are more aware of the capabilities and limitations of the human visual system, and thus can reconstruct visually important image features. Our primary goal is to make the training procedure less sensitive to distortions that humans cannot detect and focus on penalizing perceptually important artifacts. Given the nature of GAN-based solutions, we focus on the sensitivity of human vision to hallucination in case of input samples with different densities. We propose psychophysical experiments, a dataset, and a procedure for training foveated image reconstruction. The proposed strategy renders the generator network flexible by penalizing only perceptually important deviations in the output. As a result, the method emphasized the recovery of perceptually important image features. We evaluated our strategy and compared it with alternative solutions by using a newly trained objective metric, a recent foveated video quality metric, and user experiments. Our evaluations revealed significant improvements in the perceived image reconstruction quality compared with the standard GAN-based training approach. Luca Surace, Marek Wernikowski, Cara Tursun, Karol Myszkowski, Radoslaw Mantiuk, Piotr Didyk |
ACM Trans. Appl. Percept. | 4 |
| 2023 | An Implicit Neural Representation for the Image Stack: Depth, All in Focus, and High Dynamic RangeabstractIn everyday photography, physical limitations of camera sensors and lenses frequently lead to a variety of degradations in captured images such as saturation or defocus blur. A common approach to overcome these limitations is to resort to image stack fusion, which involves capturing multiple images with different focal distances or exposures. For instance, to obtain an all-in-focus image, a set of multi-focus images is captured. Similarly, capturing multiple exposures allows for the reconstruction of high dynamic range. In this paper, we present a novel approach that combines neural fields with an expressive camera model to achieve a unified reconstruction of an all-in-focus high-dynamic-range image from an image stack. Our approach is composed of a set of specialized implicit neural representations tailored to address specific sub-problems along our pipeline: We use neural implicits to predict flow to overcome misalignments arising from lens breathing, depth, and all-in-focus images to account for depth of field, as well as tonemapping to deal with sensor responses and saturation - all trained using a physically inspired supervision structure with a differentiable thin lens model at its core. An important benefit of our approach is its ability to handle these tasks simultaneously or independently, providing flexible post-editing capabilities such as refocusing and exposure adjustment. By sampling the three primary factors in photography within our framework (focal distance, aperture, and exposure time), we conduct a thorough exploration to gain valuable insights into their significance and impact on overall reconstruction quality. Through extensive validation, we demonstrate that our method outperforms existing approaches in both depth-from-defocus and all-in-focus image reconstruction tasks. Moreover, our approach exhibits promising results in each of these three dimensions, showcasing its potential to enhance captured image quality and provide greater control in post-processing. Chao Wang 0037, Ana Serrano, Xingang Pan, Krzysztof Wolski, Bin Chen 0019, Karol Myszkowski, Hans-Peter Seidel, Christian Theobalt, Thomas Leimkühler |
ACM Trans. Graph. | 6 |
| 2022 | Gloss management for consistent reproduction of real and virtual objectsabstractA good match of material appearance between real-world objects and their digital on-screen representations is critical for many applications such as fabrication, design, and e-commerce. However, faithful appearance reproduction is challenging, especially for complex phenomena, such as gloss. In most cases, the view-dependent nature of gloss and the range of luminance values required for reproducing glossy materials exceeds the current capabilities of display devices. As a result, appearance reproduction poses significant problems even with accurately rendered images. This paper studies the gap between the gloss perceived from real-world objects and their digital counterparts. Based on our psychophysical experiments on a wide range of 3D printed samples and their corresponding photographs, we derive insights on the influence of geometry, illumination, and the display’s brightness and measure the change in gloss appearance due to the display limitations. Our evaluation experiments demonstrate that using the prediction to correct material parameters in a rendering system improves the match of gloss appearance between real objects and their visualization on a display device. Bin Chen 0019, Michal Piovarci, Chao Wang 0037, Hans-Peter Seidel, Piotr Didyk, Karol Myszkowski, Ana Serrano |
SIGGRAPH Asia | 6 |
| 2022 | Learning HDR video reconstruction for dual-exposure sensors with temporally-alternating exposures
Ugur Çogalan, Mojtaba Bemana, Karol Myszkowski, Hans-Peter Seidel, Tobias Ritschel 0001 |
Comput. Graph. | 3 |
| 2022 | Learning a self-supervised tone mapping operator via feature contrast masking lossabstractAbstract High Dynamic Range (HDR) content is becoming ubiquitous due to the rapid development of capture technologies. Nevertheless, the dynamic range of common display devices is still limited, therefore tone mapping (TM) remains a key challenge for image visualization. Recent work has demonstrated that neural networks can achieve remarkable performance in this task when compared to traditional methods, however, the quality of the results of these learning‐based methods is limited by the training data. Most existing works use as training set a curated selection of best‐performing results from existing traditional tone mapping operators (often guided by a quality metric), therefore, the quality of newly generated results is fundamentally limited by the performance of such operators. This quality might be even further limited by the pool of HDR content that is used for training. In this work we propose a learning‐based self‐supervised tone mapping operator that is trained at test time specifically for each HDR image and does not need any data labeling. The key novelty of our approach is a carefully designed loss function built upon fundamental knowledge on contrast perception that allows for directly comparing the content in the HDR and tone mapped images. We achieve this goal by reformulating classic VGG feature maps into feature contrast maps that normalize local feature differences by their average magnitude in a local neighborhood, allowing our loss to account for contrast masking effects. We perform extensive ablation studies and exploration of parameters and demonstrate that our solution outperforms existing approaches with a single set of fixed parameters, as confirmed by both objective and subjective metrics. Chao Wang 0037, Bin Chen 0019, Hans-Peter Seidel, Karol Myszkowski, Ana Serrano |
Comput. Graph. Forum | 4 |
| 2022 | Perceptual Error Optimization for Monte Carlo RenderingabstractSynthesizing realistic images involves computing high-dimensional light-transport integrals. In practice, these integrals are numerically estimated via Monte Carlo integration. The error of this estimation manifests itself as conspicuous aliasing or noise. To ameliorate such artifacts and improve image fidelity, we propose a perception-oriented framework to optimize the error of Monte Carlo rendering. We leverage models based on human perception from the halftoning literature. The result is an optimization problem whose solution distributes the error as visually pleasing blue noise in image space. To find solutions, we present a set of algorithms that provide varying trade-offs between quality and speed, showing substantial improvements over prior state of the art. We perform evaluations using quantitative and error metrics and provide extensive supplemental material to demonstrate the perceptual improvements achieved by our methods. Vassillen Chizhov, Iliyan Georgiev, Karol Myszkowski, Gurprit Singh |
ACM Trans. Graph. | 3 |
| 2022 | Dark stereo: improving depth perception under low luminanceabstractIt is often desirable or unavoidable to display Virtual Reality (VR) or stereoscopic content at low brightness. For example, a dimmer display reduces the flicker artefacts that are introduced by low-persistence VR headsets. It also saves power, prolongs battery life, and reduces the cost of a display or projection system. Additionally, stereo movies are usually displayed at relatively low luminance due to polarization filters or other optical elements necessary to separate two views. However, the binocular depth cues become less reliable at low luminance. In this paper, we propose a model of stereo constancy that predicts the precision of binocular depth cues for a given contrast and luminance. We use the model to design a novel contrast enhancement algorithm that compensates for the deteriorated depth perception to deliver good-quality stereoscopic images even for displays of very low brightness. Krzysztof Wolski, Fangcheng Zhong, Karol Myszkowski, Rafal Mantiuk |
ACM Trans. Graph. | 3 |
| 2021 | Neural Acceleration of Scattering-Aware Color 3D PrintingabstractAbstract With the wider availability of full‐color 3D printers, color‐accurate 3D‐print preparation has received increased attention. A key challenge lies in the inherent translucency of commonly used print materials that blurs out details of the color texture. Previous work tries to compensate for these scattering effects through strategic assignment of colored primary materials to printer voxels. To date, the highest‐quality approach uses iterative optimization that relies on computationally expensive Monte Carlo light transport simulation to predict the surface appearance from subsurface scattering within a given print material distribution; that optimization, however, takes in the order of days on a single machine. In our work, we dramatically speed up the process by replacing the light transport simulation with a data‐driven approach. Leveraging a deep neural network to predict the scattering within a highly heterogeneous medium, our method performs around two orders of magnitude faster than Monte Carlo rendering while yielding optimization results of similar quality level. The network is based on an established method from atmospheric cloud rendering, adapted to our domain and extended by a physically motivated weight sharing scheme that substantially reduces the network size. We analyze its performance in an end‐to‐end print preparation pipeline and compare quality and runtime to alternative approaches, and demonstrate its generalization to unseen geometry and material values. This for the first time enables full heterogenous material optimization for 3D‐print preparation within time frames in the order of the actual printing time. Tobias Rittig, Denis Sumin, Vahid Babaei, Piotr Didyk, Alexey G. Voloboy, Alexander Wilkie, Bernd Bickel, Karol Myszkowski, Tim Weyrich, Jaroslav Krivánek |
Comput. Graph. Forum | 8 |
| 2021 | Perceptual model for adaptive local shading and refresh rateabstractWhen the rendering budget is limited by power or time, it is necessary to find the combination of rendering parameters, such as resolution and refresh rate, that could deliver the best quality. Variable-rate shading (VRS), introduced in the last generations of GPUs, enables fine control of the rendering quality, in which each 16×16 image tile can be rendered with a different ratio of shader executions. We take advantage of this capability and propose a new method for adaptive control of local shading and refresh rate. The method analyzes texture content, on-screen velocities, luminance, and effective resolution and suggests the refresh rate and a VRS state map that maximizes the quality of animated content under a limited budget. The method is based on the new content-adaptive metric of judder, aliasing, and blur, which is derived from the psychophysical models of contrast sensitivity. To calibrate and validate the metric, we gather data from literature and also collect new measurements of motion quality under variable shading rates, different velocities of motion, texture content, and display capabilities, such as refresh rate, persistence, and angular resolution. The proposed metric and adaptive shading method is implemented as a game engine plugin. Our experimental validation shows a substantial increase in preference of our method over rendering with a fixed resolution and refresh rate, and an existing motion-adaptive technique. Akshay Jindal, Krzysztof Wolski, Karol Myszkowski, Rafal Mantiuk |
ACM Trans. Graph. | 3 |
| 2021 | The effect of shape and illumination on material perception: model and applicationsabstractMaterial appearance hinges on material reflectance properties but also surface geometry and illumination. The unlimited number of potential combinations between these factors makes understanding and predicting material appearance a very challenging task. In this work, we collect a large-scale dataset of perceptual ratings of appearance attributes with more than 215,680 responses for 42,120 distinct combinations of material, shape, and illumination. The goal of this dataset is twofold. First, we analyze for the first time the effects of illumination and geometry in material perception across such a large collection of varied appearances. We connect our findings to those of the literature, discussing how previous knowledge generalizes across very diverse materials, shapes, and illuminations. Second, we use the collected dataset to train a deep learning architecture for predicting perceptual attributes that correlate with human judgments. We demonstrate the consistent and robust behavior of our predictor in various challenging scenarios, which, for the first time, enables estimating perceived material attributes from general 2D images. Since our predictor relies on the final appearance in an image, it can compare appearance properties across different geometries and illumination conditions. Finally, we demonstrate several applications that use our predictor, including appearance reproduction using 3D printing, BRDF editing by integrating our predictor in a differentiable renderer, illumination design, or material recommendations for scene design. Ana Serrano, Bin Chen 0019, Chao Wang 0037, Michal Piovarci, Hans-Peter Seidel, Piotr Didyk, Karol Myszkowski |
ACM Trans. Graph. | 7 |
| 2021 | The effect of geometry and illumination on appearance perception of different material categoriesabstractAbstract The understanding of material appearance perception is a complex problem due to interactions between material reflectance, surface geometry, and illumination. Recently, Serrano et al. collected the largest dataset to date with subjective ratings of material appearance attributes, including glossiness, metallicness, sharpness and contrast of reflections. In this work, we make use of their dataset to investigate for the first time the impact of the interactions between illumination, geometry, and eight different material categories in perceived appearance attributes. After an initial analysis, we select for further analysis the four material categories that cover the largest range for all perceptual attributes: fabric, plastic, ceramic, and metal. Using a cumulative link mixed model (CLMM) for robust regression, we discover interactions between these material categories and four representative illuminations and object geometries. We believe that our findings contribute to expanding the knowledge on material appearance perception and can be useful for many applications, such as scene design, where any particular material in a given shape can be aligned with dominant classes of illumination, so that a desired strength of appearance attributes can be achieved. Bin Chen 0019, Chao Wang 0037, Michal Piovarci, Hans-Peter Seidel, Piotr Didyk, Karol Myszkowski, Ana Serrano |
Vis. Comput. | 6 |
| 2020 | Stimulating the Human Visual System Beyond Real World Performance in Future Augmented Reality DisplaysabstractNew augmented-reality near-eye displays provide capabilities for enriching real-world visual experiences with digital content. Most current research focuses on improving both hardware and software to provide digital content that seamlessly blends with the real world. This is believed to not only contribute to the visual experience but also increase human task performance. In this work, we take a step further and ask the question of whether the capabilities of current and future display designs combined with efficient perception-inspired content optimizations can be used to improve human task performance beyond the human capabilities in the natural world. Based on an in-depth analysis of previous literature, we hypothesize here that such enhancements can be achieved when the human visual system is provided with content that optimizes the oculomotor responses. To further investigate possible gains, we present a series of perceptual experiments that built upon this idea. More specifically, we focus on speeding up accommodation response, which significantly contributes to the eye-adaptation when a new stimulus is shown. Through our experiments, we demonstrate that such speedups canbe achieved, and more importantly, they can lead to significant improvements in human task performance. While not all of our results give definite answers, we believe that they reveal plentiful opportunities for further enhancing the human experience and task performance when using new augmented-reality displays. David Dunn, Okan Tarhan Tursun, Hyeonseung Yu, Piotr Didyk, Karol Myszkowski, Henry Fuchs |
ISMAR | 5 |
| 2020 | Foreword to the special section on the international conference on computer-aided design and computer graphics (CAD/Graphics) 2019
Xin Tong 0001, Karol Myszkowski, Jin Huang 0001 |
Comput. Graph. | 2 |
| 2020 | X-Fields: implicit neural view-, light- and time-image interpolationabstractWe suggest to represent an X-Field ---a set of 2D images taken across different view, time or illumination conditions, i.e., video, lightfield, reflectance fields or combinations thereof---by learning a neural network (NN) to map their view, time or light coordinates to 2D images. Executing this NN at new coordinates results in joint view, time or light interpolation. The key idea to make this workable is a NN that already knows the "basic tricks" of graphics (lighting, 3D projection, occlusion) in a hard-coded and differentiable form. The NN represents the input to that rendering as an implicit map, that for any view, time, or light coordinate and for any pixel can quantify how it will move if view, time or light coordinates change (Jacobian of pixel position with respect to view, time, illumination, etc.). Our X-Field representation is trained for one scene within minutes, leading to a compact set of trainable parameters and hence real-time navigation in view, time and illumination. Mojtaba Bemana, Karol Myszkowski, Hans-Peter Seidel, Tobias Ritschel 0001 |
ACM Trans. Graph. | 2 |
| 2020 | Imperceptible manipulation of lateral camera motion for improved virtual reality applicationsabstractVirtual Reality (VR) systems increase immersion by reproducing users' movements in the real world. However, several works have shown that this real-to-virtual mapping does not need to be precise in order to convey a realistic experience. Being able to alter this mapping has many potential applications, since achieving an accurate real-to-virtual mapping is not always possible due to limitations in the capture or display hardware, or in the physical space available. In this work, we measure detection thresholds for lateral translation gains of virtual camera motion in response to the corresponding head motion under natural viewing, and in the absence of locomotion, so that virtual camera movement can be either compressed or expanded while these manipulations remain undetected. Finally, we propose three applications for our method, addressing three key problems in VR: improving 6-DoF viewing for captured 360° footage, overcoming physical constraints, and reducing simulator sickness. We have further validated our thresholds and evaluated our applications by means of additional user studies confirming that our manipulations remain imperceptible, and showing that (i) compressing virtual camera motion reduces visible artifacts in 6-DoF, hence improving perceived quality, (ii) virtual expansion allows for completion of virtual tasks within a reduced physical space, and (iii) simulator sickness may be alleviated in simple scenarios when our compression method is applied. Ana Serrano, Diego Gutierrez, Karol Myszkowski, Belén Masiá |
ACM Trans. Graph. | 4 |
| 2019 | Learning to Predict Image-based Rendering Artifacts with Respect to a Hidden Reference ImageabstractAbstract Image metrics predict the perceived per‐pixel difference between a reference image and its degraded (e. g., re‐rendered) version. In several important applications, the reference image is not available and image metrics cannot be applied. We devise a neural network architecture and training procedure that allows predicting the MSE, SSIM or VGG16 image difference from the distorted image alone while the reference is not observed. This is enabled by two insights: The first is to inject sufficiently many un‐distorted natural image patches, which can be found in arbitrary amounts and are known to have no perceivable difference to themselves. This avoids false positives. The second is to balance the learning, where it is carefully made sure that all image errors are equally likely, avoiding false negatives. Surprisingly, we observe that the resulting no‐reference metric, subjectively, can even perform better than the reference‐based one, as it had to become robust against mis‐alignments. We evaluate the effectiveness of our approach in an image‐based rendering context, both quantitatively and qualitatively. Finally, we demonstrate two applications which reduce light field capture time and provide guidance for interactive depth adjustment. Mojtaba Bemana, Joachim Keinert, Karol Myszkowski, Michel Bätz, Matthias Ziegler 0001, Hans-Peter Seidel, Tobias Ritschel 0001 |
Comput. Graph. Forum | 3 |
| 2019 | Selecting texture resolution using a task-specific visibility metricabstractAbstract In real‐time rendering, the appearance of scenes is greatly affected by the quality and resolution of the textures used for image synthesis. At the same time, the size of textures determines the performance and the memory requirements of rendering. As a result, finding the optimal texture resolution is critical, but also a non‐trivial task since the visibility of texture imperfections depends on underlying geometry, illumination, interactions between several texture maps, and viewing positions. Ideally, we would like to automate the task with a visibility metric, which could predict the optimal texture resolution. To maximize the performance of such a metric, it should be trained on a given task. This, however, requires sufficient user data which is often difficult to obtain. To address this problem, we develop a procedure for training an image visibility metric for a specific task while reducing the effort required to collect new data. The procedure involves generating a large dataset using an existing visibility metric followed by refining that dataset with the help of an efficient perceptual experiment. Then, such a refined dataset is used to retune the metric. This way, we augment sparse perceptual data to a large number of per‐pixel annotated visibility maps which serve as the training data for application‐specific visibility metrics. While our approach is general and can be potentially applied for different image distortions, we demonstrate an application in a game‐engine where we optimize the resolution of various textures, such as albedo and normal maps. Krzysztof Wolski, Daniele Giunchi, Shinichi Kinuwaki, Piotr Didyk, Karol Myszkowski, Anthony Steed, Rafal Mantiuk |
Comput. Graph. Forum | 5 |
| 2019 | Deep point correlation designabstractDesigning point patterns with desired properties can require substantial effort, both in hand-crafting coding and mathematical derivation. Retaining these properties in multiple dimensions or for a substantial number of points can be challenging and computationally expensive. Tackling those two issues, we suggest to automatically generate scalable point patterns from design goals using deep learning. We phrase pattern generation as a deep composition of weighted distance-based unstructured filters. Deep point pattern design means to optimize over the space of all such compositions according to a user-provided point correlation loss , a small program which measures a pattern's fidelity in respect to its spatial or spectral statistics, linear or non-linear (e. g., radial) projections, or any arbitrary combination thereof. Our analysis shows that we can emulate a large set of existing patterns (blue, green, step, projective, stair, etc.-noise), generalize them to countless new combinations in a systematic way and leverage existing error estimation formulations to generate novel point patterns for a user-provided class of integrand functions. Our point patterns scale favorably to multiple dimensions and numbers of points: we demonstrate nearly 10k points in 10-D produced in one second on one GPU. All the resources (source code and the pre-trained networks) can be found at https://sampling.mpi-inf.mpg.de/deepsampling.html. Thomas Leimkühler, Gurprit Singh, Karol Myszkowski, Hans-Peter Seidel, Tobias Ritschel 0001 |
ACM Trans. Graph. | 3 |
| 2019 | Geometry-aware scattering compensation for 3D printingabstractCommercially available full-color 3D printing allows for detailed control of material deposition in a volume, but an exact reproduction of a target surface appearance is hampered by the strong subsurface scattering that causes nontrivial volumetric cross-talk at the print surface. Previous work showed how an iterative optimization scheme based on accumulating absorptive materials at the surface can be used to find a volumetric distribution of print materials that closely approximates a given target appearance. In this work, we first revisit the assumption that pushing the absorptive materials to the surface results in minimal volumetric cross-talk. We design a full-fledged optimization on a small domain for this task and confirm this previously reported heuristic. Then, we extend the above approach that is critically limited to color reproduction on planar surfaces, to arbitrary 3D shapes. Our method enables high-fidelity color texture reproduction on 3D prints by effectively compensating for internal light scattering within arbitrarily shaped objects. In addition, we propose a content-aware gamut mapping that significantly improves color reproduction for the pathological case of thin geometric features. Using a wide range of sample objects with complex textures and geometries, we demonstrate color reproduction whose fidelity is superior to state-of-the-art drivers for color 3D printers. Denis Sumin, Tobias Rittig, Vahid Babaei, Thomas Nindel, Alexander Wilkie, Piotr Didyk, Bernd Bickel, Jaroslav Krivánek, Karol Myszkowski, Tim Weyrich |
ACM Trans. Graph. | 9 |
| 2019 | Luminance-contrast-aware foveated renderingabstractCurrent rendering techniques struggle to fulfill quality and power efficiency requirements imposed by new display devices such as virtual reality headsets. A promising solution to overcome these problems is foveated rendering, which exploits gaze information to reduce rendering quality for the peripheral vision where the requirements of the human visual system are significantly lower. Most of the current solutions model the sensitivity as a function of eccentricity, neglecting the fact that it also is strongly influenced by the displayed content. In this work, we propose a new luminance-contrast-aware foveated rendering technique which demonstrates that the computational savings of foveated rendering can be significantly improved if local luminance contrast of the image is analyzed. To this end, we first study the resolution requirements at different eccentricities as a function of luminance patterns. We later use this information to derive a low-cost predictor of the foveated rendering parameters. Its main feature is the ability to predict the parameters using only a low-resolution version of the current frame, even though the prediction holds for high-resolution rendering. This property is essential for the estimation of required quality before the full-resolution image is rendered. We demonstrate that our predictor can efficiently drive the foveated rendering technique and analyze its benefits in a series of user experiments. Okan Tarhan Tursun, Elena Arabadzhiyska-Koleva, Marek Wernikowski, Radoslaw Mantiuk, Hans-Peter Seidel, Karol Myszkowski, Piotr Didyk |
ACM Trans. Graph. | 6 |
| 2019 | A Perception-driven Hybrid Decomposition for Multi-layer Accommodative DisplaysabstractMulti-focal plane and multi-layered light-field displays are promising solutions for addressing all visual cues observed in the real world. Unfortunately, these devices usually require expensive optimizations to compute a suitable decomposition of the input light field or focal stack to drive individual display layers. Although these methods provide near-correct image reconstruction, a significant computational cost prevents real-time applications. A simple alternative is a linear blending strategy which decomposes a single 2D image using depth information. This method provides real-time performance, but it generates inaccurate results at occlusion boundaries and on glossy surfaces. This paper proposes a perception-based hybrid decomposition technique which combines the advantages of the above strategies and achieves both real-time performance and high-fidelity results. The fundamental idea is to apply expensive optimizations only in regions where it is perceptually superior, e.g., depth discontinuities at the fovea, and fall back to less costly linear blending otherwise. We present a complete, perception-informed analysis and model that locally determine which of the two strategies should be applied. The prediction is later utilized by our new synthesis method which performs the image decomposition. The results are analyzed and validated in user experiments on a custom multi-plane display. Hyeonseung Yu, Mojtaba Bemana, Marek Wernikowski, Michal Chwesiuk, Okan Tarhan Tursun, Gurprit Singh, Karol Myszkowski, Radoslaw Mantiuk, Hans-Peter Seidel, Piotr Didyk |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2018 | Light-Field Intrinsic Dataset
Sumit Shekhar 0001, Shida Kunz, Matthias Ziegler 0001, Michal Chwesiuk, Dawid Palen, Karol Myszkowski, Joachim Keinert, Radoslaw Mantiuk, Piotr Didyk |
BMVC | 6 |
| 2018 | Dataset and Metrics for Predicting Local Visible DifferencesabstractA large number of imaging and computer graphics applications require localized information on the visibility of image distortions. Existing image quality metrics are not suitable for this task as they provide a single quality value per image. Existing visibility metrics produce visual difference maps, and are specifically designed for detecting just noticeable distortions but their predictions are often inaccurate. In this work, we argue that the key reason for this problem is the lack of large image collections with a good coverage of possible distortions that occur in different applications. To address the problem, we collect an extensive dataset of reference and distorted image pairs together with user markings indicating whether distortions are visible or not. We propose a statistical model that is designed for the meaningful interpretation of such data, which is affected by visual search and imprecision of manual marking. We use our dataset for training existing metrics and we demonstrate that their performance significantly improves. We show that our dataset with the proposed statistical model can be used to train a new CNN-based metric, which outperforms the existing solutions. We demonstrate the utility of such a metric in visually lossless JPEG compression, super-resolution and watermarking. Krzysztof Wolski, Daniele Giunchi, Nanyang Ye 0001, Piotr Didyk, Karol Myszkowski, Radoslaw Mantiuk, Hans-Peter Seidel, Anthony Steed, Rafal Mantiuk |
ACM Trans. Graph. | 5 |
| 2018 | Perceptual Real-Time 2D-to-3D Conversion Using Cue FusionabstractWe propose a system to infer binocular disparity from a monocular video stream in real-time. Different from classic reconstruction of physical depth in computer vision, we compute perceptually plausible disparity, that is numerically inaccurate, but results in a very similar overall depth impression with plausible overall layout, sharp edges, fine details and agreement between luminance and disparity. We use several simple monocular cues to estimate disparity maps and confidence maps of low spatial and temporal resolution in real-time. These are complemented by spatially-varying, appearance-dependent and class-specific disparity prior maps, learned from example stereo images. Scene classification selects this prior at runtime. Fusion of prior and cues is done by means of robust MAP inference on a dense spatio-temporal conditional random field with high spatial and temporal resolution. Using normal distributions allows this in constant-time, parallel per-pixel work. We compare our approach to previous 2D-to-3D conversion systems in terms of different metrics, as well as a user study and validate our notion of perceptually plausible disparity. Thomas Leimkühler, Petr Kellnhofer, Tobias Ritschel 0001, Karol Myszkowski, Hans-Peter Seidel |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Towards a Quality Metric for Dense Light FieldsabstractLight fields become a popular representation of three-dimensional scenes, and there is interest in their processing, resampling, and compression. As those operations often result in loss of quality, there is a need to quantify it. In this work, we collect a new dataset of dense reference and distorted light fields as well as the corresponding quality scores which are scaled in perceptual units. The scores were acquired in a subjective experiment using an interactive light-field viewing setup. The dataset contains typical artifacts that occur in light-field processing chain due to light-field reconstruction, multi-view compression, and limitations of automultiscopic displays. We test a number of existing objective quality metrics to determine how well they can predict the quality of light fields. We find that the existing image quality metrics provide good measures of light-field quality, but require dense reference light-fields for optimal performance. For more complex tasks of comparing two distorted light fields, their performance drops significantly, which reveals the need for new, light-field-specific metrics. Vamsi Kiran Adhikarla, Marek Vinkler, Denis Sumin, Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel, Piotr Didyk |
CVPR | 5 |
| 2017 | Perception-driven Accelerated RenderingabstractAdvances in computer graphics enable us to create digital images of astonishing complexity and realism. However, processing resources are still a limiting factor. Hence, many costly but desirable aspects of realism are often not accounted for, including global illumination, accurate depth of field and motion blur, spectral effects, etc. especially in real-time rendering. At the same time, there is a strong trend towards more pixels per display due to larger displays, higher pixel densities or larger fields of view. Further observable trends in current display technology include more bits per pixel (high dynamic range, wider color gamut/fidelity), increasing refresh rates (better motion depiction), and an increasing number of displayed views per pixel (stereo, multi-view, all the way to holographic or lightfield displays). These developments cause significant unsolved technical challenges due to aspects such as limited compute power and bandwidth. Fortunately, the human visual system has certain limitations, which mean that providing the highest possible visual quality is not always necessary. In this report, we present the key research and models that exploit the limitations of perception to tackle visual quality and workload alike. Moreover, we present the open problems and promising future research targeting the question of how we can minimize the effort to compute and display only the necessary pixels while still offering a user full visual experience. Martin Weier, Michael Stengel, Thorsten Roth, Piotr Didyk, Elmar Eisemann, Martin Eisemann, Steve Grogorick, André Hinkenjann, Ernst Kruijff, Marcus A. Magnor, Karol Myszkowski, Philipp Slusallek |
Comput. Graph. Forum | 11 |
| 2017 | Saccade landing position prediction for gaze-contingent renderingabstractGaze-contingent rendering shows promise in improving perceived quality by providing a better match between image quality and the human visual system requirements. For example, information about fixation allows rendering quality to be reduced in peripheral vision, and the additional resources can be used to improve the quality in the foveal region. Gaze-contingent rendering can also be used to compensate for certain limitations of display devices, such as reduced dynamic range or lack of accommodation cues. Despite this potential and the recent drop in the prices of eye trackers, the adoption of such solutions is hampered by system latency which leads to a mismatch between image quality and the actual gaze location. This is especially apparent during fast saccadic movements when the information about gaze location is significantly delayed, and the quality mismatch can be noticed. To address this problem, we suggest a new way of updating images in gaze-contingent rendering during saccades. Instead of rendering according to the current gaze position, our technique predicts where the saccade is likely to end and provides an image for the new fixation location as soon as the prediction is available. While the quality mismatch during the saccade remains unnoticed due to saccadic suppression, a correct image for the new fixation is provided before the fixation is established. This paper describes the derivation of a model for predicting saccade landing positions and demonstrates how it can be used in the context of gaze-contingent rendering to reduce the influence of system latency on the perceived quality. The technique is validated in a series of experiments for various combinations of display frame rate and eye-tracker sampling rate. Elena Arabadzhiyska-Koleva, Okan Tarhan Tursun, Karol Myszkowski, Hans-Peter Seidel, Piotr Didyk |
ACM Trans. Graph. | 3 |
| 2017 | Scattering-aware texture reproduction for 3D printingabstractColor texture reproduction in 3D printing commonly ignores volumetric light transport (cross-talk) between surface points on a 3D print. Such light diffusion leads to significant blur of details and color bleeding, and is particularly severe for highly translucent resin-based print materials. Given their widely varying scattering properties, this cross-talk between surface points strongly depends on the internal structure of the volume surrounding each surface point. Existing scattering-aware methods use simplified models for light difusion, and often accept the visual blur as an immutable property of the print medium. In contrast, our work counteracts heterogeneous scattering to obtain the impression of a crisp albedo texture on top of the 3D print, by optimizing for a fully volumetric material distribution that preserves the target appearance. Our method employs an efficient numerical optimizer on top of a general Monte-Carlo simulation of heterogeneous scattering, supported by a practical calibration procedure to obtain scattering parameters from a given set of printer materials. Despite the inherent translucency of the medium, we reproduce detailed surface textures on 3D prints. We evaluate our system using a commercial, five-tone 3D print process and compare against the printer's native color texturing mode, demonstrating that our method preserves high-frequency features well without having to compromise on color gamut. Oskar Elek, Denis Sumin, Ran Zhang 0007, Tim Weyrich, Karol Myszkowski, Bernd Bickel, Alexander Wilkie, Jaroslav Krivánek |
ACM Trans. Graph. | 5 |
| 2017 | Wide Field Of View Varifocal Near-Eye Display Using See-Through Deformable Membrane MirrorsabstractAccommodative depth cues, a wide field of view, and ever-higher resolutions all present major hardware design challenges for near-eye displays. Optimizing a design to overcome one of these challenges typically leads to a trade-off in the others. We tackle this problem by introducing an all-in-one solution - a new wide field of view, gaze-tracked near-eye display for augmented reality applications. The key component of our solution is the use of a single see-through, varifocal deformable membrane mirror for each eye reflecting a display. They are controlled by airtight cavities and change the effective focal power to present a virtual image at a target depth plane which is determined by the gaze tracker. The benefits of using the membranes include wide field of view (100° diagonal) and fast depth switching (from 20 cm to infinity within 300 ms). Our subjective experiment verifies the prototype and demonstrates its potential benefits for near-eye see-through displays. David Dunn, Cary Tippets, Kent Torell, Petr Kellnhofer, Kaan Aksit, Piotr Didyk, Karol Myszkowski, David P. Luebke, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2016 | Perceptual Real-time 2D-to-3D Conversion Using Cue Fusion
Thomas Leimkühler, Petr Kellnhofer, Tobias Ritschel 0001, Karol Myszkowski, Hans-Peter Seidel |
Graphics Interface | 4 |
| 2016 | Efficient Multi-image Correspondences for On-line Light Field Video ProcessingabstractAbstract Light field videos express the entire visual information of an animated scene, but their shear size typically makes capture, processing and display anoff‐lineprocess, i. e., time between initial capture and final display is far from real‐time. In this paper we propose a solution for one of the key bottlenecks in such a processing pipeline, which is a reliable depth reconstruction possibly for many views. This is enabled by a novel correspondence algorithm converting the video streams from a sparse array of off‐the‐shelf cameras into an array of animated depth maps. The algorithm is based on a generalization of the classic multi‐resolution Lucas‐Kanade correspondence algorithm from a pair of images to an entire array. Special inter‐image confidence consolidation allows recovery from unreliable matching in some locations and some views. It can be implemented efficiently in massively parallel hardware, allowing for interactive computations. The resulting depth quality as well as the computation performance compares favorably to other state‐of‐the art light field‐to‐depth approaches, as well as stereo matching techniques. Another outcome of this work is a data set of light field videos that are captured with multiple variants of sparse camera arrays. Lukasz Dabala, Matthias Ziegler 0001, Piotr Didyk, Frederik Zilly, Joachim Keinert, Karol Myszkowski, Hans-Peter Seidel, Przemyslaw Rokita, Tobias Ritschel 0001 |
Comput. Graph. Forum | 6 |
| 2016 | Perceptually Motivated BRDF Comparison using Single ImageabstractSurface reflectance of real-world materials is now widely represented by the bidirectional reflectance distribution function (BRDF) and also by spatially varying representations such as SVBRDF and the bidirectional texture function (BTF). The raw surface reflectance measurements are typically compressed or fitted by analytical models, that always introduce a certain loss of accuracy. For its evaluation we need a distance function between a reference surface reflectance and its approximate version. Although some of the past techniques tried to reflect the perceptual sensitivity of human vision, they have neither optimized illumination and viewing conditions nor surface shape. In this paper, we suggest a new image-based methodology for comparing different anisotropic BRDFs. We use optimization techniques to generate a novel surface which has extensive coverage of incoming and outgoing light directions, while preserving its features and frequencies that are important for material appearance judgments. A single rendered image of such a surface along with simultaneously optimized lighting and viewing directions leads to the computation of a meaningful BRDF difference, by means of standard image difference predictors. A psychophysical experiments revealed that our surface provides richer information on material properties than the standard surfaces often used in computer graphics, e.g., sphere or blob. Vlastimil Havran, Jirí Filip, Karol Myszkowski |
Comput. Graph. Forum | 3 |
| 2016 | An intuitive control space for material appearanceabstractMany different techniques for measuring material appearance have been proposed in the last few years. These have produced large public datasets, which have been used for accurate, data-driven appearance modeling. However, although these datasets have allowed us to reach an unprecedented level of realism in visual appearance, editing the captured data remains a challenge. In this paper, we present an intuitive control space for predictable editing of captured BRDF data, which allows for artistic creation of plausible novel material appearances, bypassing the difficulty of acquiring novel samples. We first synthesize novel materials, extending the existing MERL dataset up to 400 mathematically valid BRDFs. We then design a large-scale experiment, gathering 56,000 subjective ratings on the high-level perceptual attributes that best describe our extended dataset of materials. Using these ratings, we build and train networks of radial basis functions to act as functionals mapping the perceptual attributes to an underlying PCA-based representation of BRDFs. We show that our functionals are excellent predictors of the perceived attributes of appearance. Our control space enables many applications, including intuitive material editing of a wide range of visual properties, guidance for gamut mapping, analysis of the correlation between perceptual attributes, or novel appearance similarity metrics. Moreover, our methodology can be used to derive functionals applicable to classic analytic BRDF representations. We release our code and dataset publicly, in order to support and encourage further research in this direction. Ana Serrano, Diego Gutierrez, Karol Myszkowski, Hans-Peter Seidel, Belén Masiá |
ACM Trans. Graph. | 3 |
| 2016 | Motion parallax in stereo 3D: model and applicationsabstractBinocular disparity is the main depth cue that makes stereoscopic images appear 3D. However, in many scenarios, the range of depth that can be reproduced by this cue is greatly limited and typically fixed due to constraints imposed by displays. For example, due to the low angular resolution of current automultiscopic screens, they can only reproduce a shallow depth range. In this work, we study the motion parallax cue, which is a relatively strong depth cue, and can be freely reproduced even on a 2D screen without any limits. We exploit the fact that in many practical scenarios, motion parallax provides sufficiently strong depth information that the presence of binocular depth cues can be reduced through aggressive disparity compression. To assess the strength of the effect we conduct psycho-visual experiments that measure the influence of motion parallax on depth perception and relate it to the depth resulting from binocular disparity. Based on the measurements, we propose a joint disparity-parallax computational model that predicts apparent depth resulting from both cues. We demonstrate how this model can be applied in the context of stereo and multiscopic image processing, and propose new disparity manipulation techniques, which first quantify depth obtained from motion parallax, and then adjust binocular disparity information accordingly. This allows us to manipulate the disparity signal according to the strength of motion parallax to improve the overall depth reproduction. This technique is validated in additional experiments. Petr Kellnhofer, Piotr Didyk, Tobias Ritschel 0001, Belén Masiá, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 5 |
| 2016 | GazeStereo3D: seamless disparity manipulationsabstractProducing a high quality stereoscopic impression on current displays is a challenging task. The content has to be carefully prepared in order to maintain visual comfort, which typically affects the quality of depth reproduction. In this work, we show that this problem can be significantly alleviated when the eye fixation regions can be roughly estimated. We propose a new method for stereoscopic depth adjustment that utilizes eye tracking or other gaze prediction information. The key idea that distinguishes our approach from the previous work is to apply gradual depth adjustments at the eye fixation stage, so that they remain unnoticeable. To this end, we measure the limits imposed on the speed of disparity changes in various depth adjustment scenarios, and formulate a new model that can guide such seamless stereoscopic content processing. Based on this model, we propose a real-time controller that applies local manipulations to stereoscopic content to find the optimum between depth reproduction and visual comfort. We show that the controller is mostly immune to the limitations of low-cost eye tracking solutions. We also demonstrate benefits of our model in off-line applications, such as stereoscopic movie production, where skillful directors can reliably guide and predict viewers' attention or where attended image regions are identified during eye tracking sessions. We validate both our model and the controller in a series of user experiments. They show significant improvements in depth perception without sacrificing the visual quality when our techniques are applied. Petr Kellnhofer, Piotr Didyk, Karol Myszkowski, Mohamed Hefeeda, Hans-Peter Seidel, Wojciech Matusik |
ACM Trans. Graph. | 3 |
| 2016 | Emulating displays with continuously varying frame ratesabstractThe visual quality of a motion picture is significantly influenced by the choice of the presentation frame rate. Increasing the frame rate improves the clarity of the image and helps to alleviate many artifacts, such as blur, strobing, flicker, or judder. These benefits, however, come at the price of losing well-established film aesthetics, often referred to as the "cinematic look". Current technology leaves artists with a sparse set of choices, e.g., 24 Hz or 48 Hz, limiting the freedom in adjusting the frame rate to artistic needs, content, and display technology. In this paper, we solve this problem by proposing a novel filtering technique which enables emulating the whole spectrum of presentation frame rates on a single-frame-rate display. The key component of our technique is a set of simple yet powerful filters calibrated and evaluated in psychophysical experiments. By varying their parameters we can achieve an impression of continuously varying presentation frame rate in both the spatial and temporal dimensions. This allows artists to achieve the best balance between the aesthetics and the objective quality of the motion picture. Furthermore, we show how our technique, informed by cinematic guidelines, can adapt to the content and achieve this balance automatically. Krzysztof Templin, Piotr Didyk, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 3 |
| 2015 | What makes 2D-to-3D stereo conversion perceptually plausible?abstractDifferent from classic reconstruction of physical depth in computer vision, depth for 2D-to-3D stereo conversion is assigned by humans using semi-automatic painting interfaces and, consequently, is often dramatically wrong. Here we seek to better understand why it still does not fail to convey a sensation of depth. To this end, four typical disparity distortions resulting from manual 2D-to-3D stereo conversion are analyzed: i) smooth remapping, ii) spatial smoothness, iii) motion-compensated, temporal smoothness, and iv) completeness. A perceptual experiment is conducted to quantify the impact of each distortion on the plausibility of the 3D impression relative to a reference without distortion. Close-to-natural videos with known depth were distorted in one of the four above-mentioned aspects and subjects had to indicate if the distortion still allows for a plausible 3D effect. The smallest amounts of distortion that result in a significant rejection suggests a conservative upper bound on the quality requirement of 2D-to-3D conversion. Petr Kellnhofer, Thomas Leimkühler, Tobias Ritschel 0001, Karol Myszkowski, Hans-Peter Seidel |
SAP | 4 |
| 2015 | Purkinje Images: Conveying Different Content for Different Luminance Adaptations in a Single ImageabstractAbstract Providing multiple meanings in a single piece of art has always been intriguing to both artists and observers. We present Purkinje images, which have different interpretations depending on the luminance adaptation of the observer. Finding such images is an optimization that minimizes the sum of the distance to one reference image in photopic conditions and the distance to another reference image in scotopic conditions. To model the shift of image perception between day and night vision, we decompose the input images into a Laplacian pyramid. Distances under different observation conditions in this representation are independent between pyramid levels and pixel positions and become matrix multiplications. The optimal pixel colour can be found by inverting a small, per‐pixel linear system in real time on a GPU. Finally, two user studies analyze our results in terms of the recognition performance and fidelity with respect to the reference images. Sami Arpa, Tobias Ritschel 0001, Karol Myszkowski, Tolga K. Çapin, Hans-Peter Seidel |
Comput. Graph. Forum | 3 |
| 2015 | Motion Aware Exposure Bracketing for HDR VideoabstractAbstract Mobile phones and tablets are rapidly gaining significance as omnipresent image and video capture devices. In this context we present an algorithm that allows such devices to capture high dynamic range (HDR) video. The design of the algorithm was informed by a perceptual study that assesses the relative importance of motion and dynamic range. We found that ghosting artefacts are more visually disturbing than a reduction in dynamic range, even if a comparable number of pixels is affected by each. We incorporated these findings into a real‐time, adaptive metering algorithm that seamlessly adjusts its settings to take exposures that will lead to minimal visual artefacts after recombination into an HDR sequence. It is uniquely suitable for real‐time selection of exposure settings. Finally, we present an off‐line HDR reconstruction algorithm that is matched to the adaptive nature of our real‐time metering approach. Yulia Gryaditskaya, Tania Pouli, Erik Reinhard, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 4 |
| 2015 | Modeling Luminance Perception at Absolute ThresholdabstractAbstract When human luminance perception operates close to its absolute threshold, i. e., the lowest perceivable absolute values, appearance changes substantially compared to common photopic or scotopic vision. In particular, most observers report perceiving temporally‐varying noise. Two reasons are physiologically plausible; quantum noise (due to the low absolute number of photons) and spontaneous photochemical reactions. Previously, static noise with a normal distribution and no account for absolute values was combined with blue hue shift and blur to simulate scotopic appearance on a photopic display for movies and interactive applications (e.g., games). We present a computational model to reproduce the specific distribution and dynamics of “scotopic noise” for specific absolute values. It automatically introduces a perceptually‐calibrated amount of noise for a specific luminance level and supports animated imagery. Our simulation runs in milliseconds at HD resolution using graphics hardware and favorably compares to simpler alternatives in a perceptual experiment. Petr Kellnhofer, Tobias Ritschel 0001, Karol Myszkowski, Elmar Eisemann, Hans-Peter Seidel |
Comput. Graph. Forum | 3 |
| 2015 | A model of local adaptationabstractThe visual system constantly adapts to different luminance levels when viewing natural scenes. The state of visual adaptation is the key parameter in many visual models. While the time-course of such adaptation is well understood, there is little known about the spatial pooling that drives the adaptation signal. In this work we propose a new empirical model of local adaptation, that predicts how the adaptation signal is integrated in the retina. The model is based on psychophysical measurements on a high dynamic range (HDR) display. We employ a novel approach to model discovery, in which the experimental stimuli are optimized to find the most predictive model. The model can be used to predict the steady state of adaptation, but also conservative estimates of the visibility (detection) thresholds in complex images. We demonstrate the utility of the model in several applications, such as perceptual error bounds for physically based rendering, determining the backlight resolution for HDR displays, measuring the maximum visible dynamic range in natural scenes, simulation of afterimages, and gaze-dependent tone mapping. Peter Vangorp, Karol Myszkowski, Erich W. Graf, Rafal Mantiuk |
ACM Trans. Graph. | 2 |
| 2014 | Depth from HDR: depth induction or increased realism?abstractMany people who first see a high dynamic range (HDR) display get the impression that it is a 3D display, even though it does not produce any binocular depth cues. Possible explanations of this effect include contrast-based depth induction and the increased realism due to the high brightness and contrast that makes an HDR display "like looking through a window". In this paper we test both of these hypotheses by comparing the HDR depth illusion to real binocular depth cues using a carefully calibrated HDR stereoscope. We confirm that contrast-based depth induction exists, but it is a vanishingly weak depth cue compared to binocular depth cues. We also demonstrate that for some observers, the increased contrast of HDR displays indeed increases the realism. However, it is highly observer-dependent whether reduced, physically correct, or exaggerated contrast is perceived as most realistic, even in the presence of the real-world reference scene. Similarly, observers differ in whether reduced, physically correct, or exaggerated stereo 3D is perceived as more realistic. To accommodate the binocular depth perception and realism concept of most observers, display technologies must offer both HDR contrast and stereo personalization. Peter Vangorp, Rafal Mantiuk, Bartosz Bazyluk, Karol Myszkowski, Radoslaw Mantiuk, Simon J. Watt, Hans-Peter Seidel |
SAP | 4 |
| 2014 | Manipulating refractive and reflective binocular disparityabstractAbstract Presenting stereoscopic content on 3D displays is a challenging task, usually requiring manual adjustments. A number of techniques have been developed to aid this process, but they account for binocular disparity of surfaces that are diffuse and opaque only. However, combinations of transparent as well as specular materials are common in the real and virtual worlds, and pose a significant problem. For example, excessive disparities can be created which cannot be fused by the observer. Also, multiple stereo interpretations become possible, e. g., for glass, that both reflects and refracts, which may confuse the observer and result in poor 3D experience. In this work, we propose an efficient method for analyzing and controlling disparities in computer‐generated images of such scenes where surface positions and a layer decomposition are available. Instead of assuming a single per‐pixel disparity value, we estimate all possibly perceived disparities at each image location. Based on this representation, we define an optimization to find the best per‐pixel camera parameters, assuring that all disparities can be easily fused by a human. A preliminary perceptual study indicates, that our approach combines comfortable viewing with realistic depiction of typical specular scenes. Lukasz Dabala, Petr Kellnhofer, Tobias Ritschel 0001, Piotr Didyk, Krzysztof Templin, Karol Myszkowski, Przemyslaw Rokita, Hans-Peter Seidel |
Comput. Graph. Forum | 6 |
| 2014 | Perceptual depth compression for stereo applicationsabstractAbstract Conventional depth video compression uses video codecs designed for color images. Given the performance of current encoding standards, this solution seems efficient. However, such an approach suffers from many issues stemming from discrepancies between depth and light perception. To exploit the inherent limitations of human depth perception, we propose a novel depth compression method that employs a disparity perception model. In contrast to previous methods, we account for disparity masking, and model a distinct relation between depth perception and contrast in luminance. Our solution is a natural extension to the H.264 codec and can easily be integrated into existing decoders. It significantly improves both the compression efficiency without sacrificing visual quality of depth of rendered content, and the output of depth‐reconstruction algorithms or depth cameras. Dawid Pajak, Robert Herzog, Radoslaw Mantiuk, Piotr Didyk, Elmar Eisemann, Karol Myszkowski, Kari Pulli |
Comput. Graph. Forum | 6 |
| 2014 | Perceptually-motivated Stereoscopic Film GrainabstractAbstract Independent management of film grain in each view of a stereoscopic video can lead to visual discomfort. The existing alternative is to project the grain onto the scene geometry. Such grain, however, looks unnatural, changes object perception, and emphasizes inaccuracies in depth arising during 2D‐to‐3D conversion. We propose an advanced method of grain positioning that scatters the grain in the scene space. In a series of perceptual experiments, we estimate the optimal parameter values for the proposed method, analyze the user preference distribution among the proposed and the two existing methods, and show influence of the method on the object perception. Krzysztof Templin, Piotr Didyk, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 3 |
| 2014 | Stereo Day-for-Night: Retargeting Disparity for Scotopic VisionabstractSeveral approaches attempt to reproduce the appearance of a scotopic low-light night scene on a photopic display (“day-for-night”) by introducing color desaturation, loss of acuity, and the Purkinje shift toward blue colors. We argue that faithful stereo reproduction of night scenes on photopic stereo displays requires manipulation of not only color but also binocular disparity. To this end, we performed a psychophysics experiment to devise a model of disparity at scotopic luminance levels. Using this model, we can match binocular disparity of a scotopic stereo content displayed on a photopic monitor to the disparity that would be perceived if the scene was actually scotopic. The model allows for real-time computation of common stereo content as found in interactive applications such as simulators or computer games. Petr Kellnhofer, Tobias Ritschel 0001, Peter Vangorp, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Appl. Percept. | 4 |
| 2014 | Modeling and optimizing eye vergence response to stereoscopic cutsabstractSudden temporal depth changes, such as cuts that are introduced by video edits, can significantly degrade the quality of stereoscopic content. Since usually not encountered in the real world, they are very challenging for the audience. This is because the eye vergence has to constantly adapt to new disparities in spite of conflicting accommodation requirements. Such rapid disparity changes may lead to confusion, reduced understanding of the scene, and overall attractiveness of the content. In most cases the problem cannot be solved by simply matching the depth around the transition, as this would require flattening the scene completely. To better understand this limitation of the human visual system, we conducted a series of eye-tracking experiments. The data obtained allowed us to derive and evaluate a model describing adaptation of vergence to disparity changes on a stereoscopic display. Besides computing user-specific models, we also estimated parameters of an average observer model. This enables a range of strategies for minimizing the adaptation time in the audience. Krzysztof Templin, Piotr Didyk, Karol Myszkowski, Mohamed Hefeeda, Hans-Peter Seidel, Wojciech Matusik |
ACM Trans. Graph. | 3 |
| 2013 | Learning to Predict Localized Distortions in Rendered ImagesabstractAbstract In this work, we present an analysis of feature descriptors for objective image quality assessment. We explore a large space of possible features including components of existing image quality metrics as well as many traditional computer vision and statistical features. Additionally, we propose new features motivated by human perception and we analyze visual saliency maps acquired using an eye tracker in our user experiments. The discriminative power of the features is assessed by means of a machine learning framework revealing the importance of each feature for image quality assessment task. Furthermore, we propose a new data‐driven full‐reference image quality metric which outperforms current state‐of‐theart metrics. The metric was trained on subjective ground truth data combining two publicly available datasets. For the sake of completeness we create a new testing synthetic dataset including experimentally measured subjective distortion maps. Finally, using the same machine‐learning framework we optimize the parameters of popular existing metrics. Martin Cadík, Robert Herzog, Rafal Mantiuk, Radoslaw Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 5 |
| 2013 | Optimizing Disparity for Motion in DepthabstractAbstract Beyond the careful design of stereo acquisition equipment and rendering algorithms, disparity post‐processing has recently received much attention, where one of the key tasks is to compress the originally large disparity range to avoid viewing discomfort. The perception of dynamic stereo content however, relies on reproducing the full disparity‐time volume that a scene point undergoes in motion. This volume can be strongly distorted in manipulation, which is only concerned with changing disparity at one instant in time, even if the temporal coherence of that change is maintained. We propose an optimization to preserve stereo motion of content that was subject to an arbitrary disparity manipulation, based on a perceptual model of temporal disparity changes. Furthermore, we introduce a novel 3D warping technique to create stereo image pairs that conform to this optimized disparity map. The paper concludes with perceptual studies of motion reproduction quality and task performance in a simple game, showing how our optimization can achieve both viewing comfort and faithful stereo motion. Petr Kellnhofer, Tobias Ritschel 0001, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 3 |
| 2012 | NoRM: No-Reference Image Quality Metric for Realistic Image SynthesisabstractAbstract Synthetically generating images and video frames of complex 3D scenes using some photo‐realistic rendering software is often prone to artifacts and requires expert knowledge to tune the parameters. The manual work required for detecting and preventing artifacts can be automated through objective quality evaluation of synthetic images. Most practical objective quality assessment methods of natural images rely on a ground‐truth reference, which is often not available in rendering applications. While general purpose no‐reference image quality assessment is a difficult problem, we show in a subjective study that the performance of a dedicated no‐reference metric as presented in this paper can match the state‐of‐the‐art metrics that do require a reference. This level of predictive power is achieved exploiting information about the underlying synthetic scene (e.g., 3D surfaces, textures) instead of merely considering color, and training our learning framework with typical rendering artifacts. We show that our method successfully detects various non‐trivial types of artifacts such as noise and clamping bias due to insufficient virtual point light sources, and shadow map discretization artifacts. We also briefly discuss an inpainting method for automatic correction of detected artifacts. Robert Herzog, Martin Cadík, Tunç Ozan Aydin, Kwang In Kim, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 5 |
| 2012 | 3D Material Style TransferabstractAbstract This work proposes a technique to transfer the material style or mood from a guide source such as an image or video onto a target 3D scene. It formulates the problem as a combinatorial optimization of assigning discrete materials extracted from the guide source to discrete objects in the target 3D scene. The assignment is optimized to fulfill multiple goals: overall image mood based on several image statistics; spatial material organization and grouping as well as geometric similarity between objects that were assigned to similar materials. To be able to use common uncalibrated images and videos with unknown geometry and lighting as guides, a material estimation derives perceptually plausible reflectance, specularity, glossiness, and texture. Finally, results produced by our method are compared to manual material assignments in a perceptual study. Chuong H. Nguyen, Tobias Ritschel 0001, Karol Myszkowski, Elmar Eisemann, Hans-Peter Seidel |
Comput. Graph. Forum | 3 |
| 2012 | New measurements reveal weaknesses of image quality metrics in evaluating graphics artifactsabstractReliable detection of global illumination and rendering artifacts in the form of localized distortion maps is important for many graphics applications. Although many quality metrics have been developed for this task, they are often tuned for compression/transmission artifacts and have not been evaluated in the context of synthetic CG-images. In this work, we run two experiments where observers use a brush-painting interface to directly mark image regions with noticeable/objectionable distortions in the presence/absence of a high-quality reference image, respectively. The collected data shows a relatively high correlation between the with-reference and no-reference observer markings. Also, our demanding per-pixel image-quality datasets reveal weaknesses of both simple (PSNR, MSE, sCIE-Lab) and advanced (SSIM, MS-SSIM, HDR-VDP-2) quality metrics. The most problematic are excessive sensitivity to brightness and contrast changes, the calibration for near visibility-threshold distortions, lack of discrimination between plausible/implausible illumination, and poor spatial localization of distortions for multi-scale metrics. We believe that our datasets have further potential in improving existing quality metrics, but also in analyzing the saliency of rendering distortions, and investigating visual equivalence given our with- and no-reference data. Martin Cadík, Robert Herzog, Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 4 |
| 2012 | A luminance-contrast-aware disparity model and applicationsabstractBinocular disparity is one of the most important depth cues used by the human visual system. Recently developed stereo-perception models allow us to successfully manipulate disparity in order to improve viewing comfort, depth discrimination as well as stereo content compression and display. Nonetheless, all existing models neglect the substantial influence of luminance on stereo perception. Our work is the first to account for the interplay of luminance contrast (magnitude/frequency) and disparity and our model predicts the human response to complex stereo-luminance images. Besides improving existing disparity-model applications (e.g., difference metrics or compression), our approach offers new possibilities, such as joint luminance contrast and disparity manipulation or the optimization of auto-stereoscopic content. We validate our results in a user study, which also reveals the advantage of considering luminance contrast and its significant impact on disparity manipulation techniques. Piotr Didyk, Tobias Ritschel 0001, Elmar Eisemann, Karol Myszkowski, Hans-Peter Seidel, Wojciech Matusik |
ACM Trans. Graph. | 4 |
| 2012 | Highlight microdisparity for improved gloss depictionabstractHuman stereo perception of glossy materials is substantially different from the perception of diffuse surfaces: A single point on a diffuse object appears the same for both eyes, whereas it appears different to both eyes on a specular object. As highlights are blurry reflections of light sources they have depth themselves, which is different from the depth of the reflecting surface. We call this difference in depth impression the "highlight disparity". Due to artistic motivation, for technical reasons, or because of incomplete data, highlights often have to be depicted on-surface, without any disparity. However, it has been shown that a lack of disparity decreases the perceived glossiness and authenticity of a material. To remedy this contradiction, our work introduces a technique for depiction of glossy materials, which improves over simple on-surface highlights, and avoids the problems of physical highlights. Our technique is computationally simple, can be easily integrated in an existing (GPU) shading system, and allows for local and interactive artistic control. Krzysztof Templin, Piotr Didyk, Tobias Ritschel 0001, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 4 |
| 2011 | Multidimensional image retargetingabstractRetargeting refers to the process by which an image or video is adapted from the display device for which it was meant (target display) to another one (retarget display). The retarget display has different features from the target one such as dynamic range, discretization levels, color gamut, multi-view, and refresh rate spatial resolution. This is a very relevant topic in graphics, given the increasing number of display devices from large, high-contrast screens to small cell phones with limited dynamic range; a lot of techniques are being published in different venues, and it's hard to keep up. For most cases retargeting can be an ill-posed problem, for example in the process of displaying Low Dynamic Range (LDR) or 8-bit content on High Dynamic Range (HDR) displays. Such a problem requires the retargeting algorithm to generate new content which is missing in the input image/frame. In this course, we will present the latest solutions and techniques for retargeting images along various dimensions such as dynamic range, colors, temporal and spatial resolutions, and for the first time offer a much-needed holistic view of the field. Moreover, we are going to show how to measure and analyze the changes applied to an image or video in terms of quality using both psychophysical experiments (subjective) and computational metrics (objective). The course should be of interest to anyone involved in graphics in a broader sense, given the almost unavoidable need to retarget results to different devices -from developers interested in implementing retargeting techniques, to users that just need an overall perspective. For researchers fully engaged in developing multi-dimensional retargeting techniques, this course will serve as a solid background for future algorithms. Francesco Banterle, Alessandro Artusi, Tunç Ozan Aydin, Piotr Didyk, Elmar Eisemann, Diego Gutierrez, Rafal Mantiuk, Karol Myszkowski |
SIGGRAPH Asia Courses | 8 |
| 2011 | Scalable Remote Rendering with Depth and Motion-flow Augmented StreamingabstractAbstract In this paper, we focus on efficient compression and streaming of frames rendered from a dynamic 3D model. Remote rendering and on‐the‐fly streaming become increasingly attractive for interactive applications. Data is kept confidential and only images are sent to the client. Even if the client's hardware resources are modest, the user can interact with state‐of‐the‐art rendering applications executed on the server. Our solution focuses on augmented video information, e.g., by depth, which is key to increase robustness with respect to data loss, image reconstruction, and is an important feature for stereo vision and other client‐side applications. Two major challenges arise in such a setup. First, the server workload has to be controlled to support many clients, second the data transfer needs to be efficient. Consequently, our contributions are twofold. First, we reduce the server‐based computations by making use of sparse sampling and temporal consistency to avoid expensive pixel evaluations. Second, our data‐transfer solution takes limited bandwidths into account, is robust to information loss, and compression and decompression are efficient enough to support real‐time interaction. Our key insight is to tailor our method explicitly for rendered 3D content and shift some computations on client GPUs, to better balance the server/client workload. Our framework is progressive, scalable, and allows us to stream augmented high‐resolution (e.g., HD‐ready) frames with small bandwidth on standard hardware. Dawid Pajak, Robert Herzog, Elmar Eisemann, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 4 |
| 2011 | A perceptual model for disparityabstractBinocular disparity is an important cue for the human visual system to recognize spatial layout, both in reality and simulated virtual worlds. This paper introduces a perceptual model of disparity for computer graphics that is used to define a metric to compare a stereo image to an alternative stereo image and to estimate the magnitude of the perceived disparity change. Our model can be used to assess the effect of disparity to control the level of undesirable distortions or enhancements (introduced on purpose). A number of psycho-visual experiments are conducted to quantify the mutual effect of disparity magnitude and frequency to derive the model. Besides difference prediction, other applications include compression, and re-targeting. We also present novel applications in form of hybrid stereo images and backward-compatible stereo. The latter minimizes disparity in order to convey a stereo impression if special equipment is used but produces images that appear almost ordinary to the naked eye. The validity of our model and difference metric is again confirmed in a study. Piotr Didyk, Tobias Ritschel 0001, Elmar Eisemann, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 4 |
| 2010 | Spatio-temporal upsampling on the GPUabstractPixel processing is becoming increasingly expensive for real-time applications due to the complexity of today's shaders and high-resolution framebuffers. However, most shading results are spatially or temporally coherent, which allows for sparse sampling and reuse of neighboring pixel values. This paper proposes a simple framework for spatio-temporal upsampling on modern GPUs. In contrast to previous work, which focuses either on temporal or spatial processing on the GPU, we exploit coherence in both. Our algorithm combines adaptive motion-compensated filtering over time and geometry-aware upsampling in image space. It is robust with respect to high-frequency temporal changes, and achieves substantial performance improvements by limiting the number of recomputed samples per frame. At the same time, we increase the quality of spatial upsampling by recovering missing information from previous frames. This temporal strategy also allows us to ensure that the image converges to a higher quality result. Robert Herzog, Elmar Eisemann, Karol Myszkowski, Hans-Peter Seidel |
SI3D | 3 |
| 2010 | Perceptually-motivated Real-time Temporal Upsampling of 3D Content for High-refresh-rate DisplaysabstractAbstract High‐refresh‐rate displays (e. g., 120 Hz) have recently become available on the consumer market and quickly gain on popularity. One of their aims is to reduce the perceived blur created by moving objects that are tracked by the human eye. However, an improvement is only achieved if the video stream is produced at the same high refresh rate (i. e. 120 Hz). Some devices, such as LCD TVs, solve this problem by converting low‐refresh‐rate content (i. e. 50 Hz PAL) into a higher temporal resolution (i. e. 200 Hz) based on two‐dimensional optical flow. In our approach, we will show how rendered three‐dimensional images produced by recent graphics hardware can be up‐sampled more efficiently resulting in higher quality at the same time. Our algorithm relies on several perceptual findings and preserves the naturalness of the original sequence. A psychophysical study validates our approach and illustrates that temporally up‐sampled video streams are preferred over the standard low‐rate input by the majority of users. We show that our solution improves task performance on high‐refresh‐rate displays. Piotr Didyk, Elmar Eisemann, Tobias Ritschel 0001, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 4 |
| 2010 | Bidirectional Texture Function Compression Based on Multi-Level Vector QuantizationabstractAbstract The Bidirectional Texture Function (BTF) is becoming widely used for accurate representation of real‐world material appearance. In this paper a novel BTF compression model is proposed. The model resamples input BTF data into a parametrization, allowing decomposition of individual view and illumination dependent texels into a set of multi‐dimensional conditional probability density functions. These functions are compressed in turn using a novel multi‐level vector quantization algorithm. The result of this algorithm is a set of index and scale code‐books for individual dimensions. BTF reconstruction from the model is then based on fast chained indexing into the nested stored code‐books. In the proposed model, luminance and chromaticity are treated separately to achieve further compression. The proposed model achieves low distortion and compression ratios 1:233–1:2040, depending on BTF sample variability. These results compare well with several other BTF compression methods with predefined compression ratios, usually smaller than 1:200. We carried out a psychophysical experiment comparing our method with LPCA method. BTF synthesis from the model was implemented on a standard GPU, yielded interactive framerates. The proposed method allows the fast importance sampling required by eye‐path tracing algorithms in image synthesis. Vlastimil Havran, Jirí Filip, Karol Myszkowski |
Comput. Graph. Forum | 3 |
| 2010 | Visually significant edgesabstractNumerous image processing and computer graphics methods make use of either explicitly computed strength of image edges, or an implicit edge strength definition that is integrated into their algorithms. In both cases, the end result is highly affected by the computation of edge strength. We address several shortcomings of the widely used gradient magnitude-based edge strength model through the computation of a hypothetical Human Visual System (HVS) response at edge locations. Contrary to gradient magnitude, the resulting “visual significance” values account for various HVS mechanisms such as luminance adaptation and visual masking, and are scaled in perceptually linear units that are uniform across images. The visual significance computation is implemented in a fast multiscale second-generation wavelet framework which we use to demonstrate the differences in image retargeting, HDR image stitching, and tone mapping applications with respect to the gradient magnitude model. Our results suggest that simple perceptual models provide qualitative improvements on applications utilizing edge strength at the cost of a modest computational burden. Tunç Ozan Aydin, Martin Cadík, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Appl. Percept. | 3 |
| 2010 | Video quality assessment for computer graphics applicationsabstractNumerous current Computer Graphics methods produce video sequences as their outcome. The merit of these methods is often judged by assessing the quality of a set of results through lengthy user studies. We present a full-reference video quality metric geared specifically towards the requirements of Computer Graphics applications as a faster computational alternative to subjective evaluation. Our metric can compare a video pair with arbitrary dynamic ranges, and comprises a human visual system model for a wide range of luminance levels, that predicts distortion visibility through models of luminance adaptation, spatiotemporal contrast sensitivity and visual masking. We present applications of the proposed metric to quality prediction of HDR video compression and temporal tone mapping, comparison of different rendering approaches and qualities, and assessing the impact of variable frame rate to perceived quality. Tunç Ozan Aydin, Martin Cadík, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 3 |
| 2010 | Apparent display resolution enhancement for moving imagesabstractLimited spatial resolution of current displays makes the depiction of very fine spatial details difficult. This work proposes a novel method applied to moving images that takes into account the human visual system and leads to an improved perception of such details. To this end, we display images rapidly varying over time along a given trajectory on a high refresh rate display. Due to the retinal integration time the information is fused and yields apparent super-resolution pixels on a conventional-resolution display. We discuss how to find optimal temporal pixel variations based on linear eye-movement and image content and extend our solution to arbitrary trajectories. This step involves an efficient method to predict and successfully treat potentially visible flickering. Finally, we evaluate the resolution enhancement in a perceptual study that shows that significant improvements can be achieved both for computer generated images and photographs. Piotr Didyk, Elmar Eisemann, Tobias Ritschel 0001, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 4 |
| 2010 | Contrast prescription for multiscale image editing
Dawid Pajak, Martin Cadík, Tunç Ozan Aydin, Makoto Okabe, Karol Myszkowski, Hans-Peter Seidel |
Vis. Comput. | 5 |
| 2009 | Predicting Display Visibility Under Dynamically Changing Lighting ConditionsabstractAbstract Display devices, more than ever, are finding their ways into electronic consumer goods as a result of recent trends in providing more functionality and user interaction. Combined with the new developments in display technology towards higher reproducible luminance range, the mobility and variation in capability of display devices are constantly increasing. Consequently, in real life usage it is now very likely that the display emission to be distorted by spatially and temporally varying reflections, and the observer's visual system to be not adapted to the particular display that she is viewing at that moment. The actual perception of the display content cannot be fully understood by only considering steady‐state illumination and adaptation conditions. We propose an objective method for display visibility analysis formulating the problem as a full‐reference image quality assessment problem, where the display emission under “ideal” conditions is used as the reference for real‐life conditions. Our work includes a human visual system model that accounts for maladaptation and temporal recovery of sensitivity. As an example application we integrate our method to a global illumination simulator and analyze the visibility of a car interior display under realistic lighting conditions. Tunç Ozan Aydin, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 2 |
| 2009 | High Dynamic Range Imaging and Low Dynamic Range Expansion for Generating HDR ContentabstractAbstract In the last few years, researchers in the field of High Dynamic Range (HDR) Imaging have focused on providing tools for expanding Low Dynamic Range (LDR) content for the generation of HDR images due to the growing popularity of HDR in applications, such as photography and rendering via Image‐Based Lighting, and the imminent arrival of HDR displays to the consumer market. LDR content expansion is required due to the lack of fast and reliable consumer level HDR capture for still images and videos. Furthermore, LDR content expansion, will allow the re‐use of legacy LDR stills, videos and LDR applications created, over the last century and more, to be widely available. The use of certain LDR expansion methods, those that are based on the inversion of Tone Mapping Operators (TMOs), has made it possible to create novel compression algorithms that tackle the problem of the size of HDR content storage, which remains one of the major obstacles to be overcome for the adoption of HDR. These methods are used in conjunction with traditional LDR compression methods and can evolve accordingly. The goal of this report is to provide a comprehensive overview on HDR Imaging, and an in depth review on these emerging topics. Francesco Banterle, Kurt Debattista, Alessandro Artusi, Sumanta N. Pattanaik, Karol Myszkowski, Patrick Ledda, Alan Chalmers |
Comput. Graph. Forum | 5 |
| 2009 | Anisotropic Radiance-Cache Splatting for Efficiently Computing High-Quality Global Illumination with LightcutsabstractAbstract Computing global illumination in complex scenes is even with todays computational power a demanding task. In this work we propose a novel irradiance caching scheme that combines the advantages of two state‐of‐the‐art algorithms for high‐quality global illumination rendering:lightcuts, an adaptive and hierarchical instant‐radiosity based algorithm and the widely used (ir)radiance caching algorithm for sparse sampling and interpolation of (ir)radiance in object space. Our adaptive radiance caching algorithm is based on anisotropic cache splatting, which adapts the cache footprints not only to the magnitude of the illumination gradient computed with light‐cuts but also to its orientation allowing larger interpolation errors along the direction of coherent illumination while reducing the error along the illumination gradient. Since lightcuts computes the direct and indirect lighting seamlessly, we use a two‐layer radiance cache, to store and control the interpolation of direct and indirect lighting individually with different error criteria. In multiple iterations our method detects cache interpolation errors above the visibility threshold of a pixel and reduces the anisotropic cache footprints accordingly. We achieve significantly better image quality while also speeding up the computation costs by one to two orders of magnitude with respect to the well‐known photon mapping with (ir)radiance caching procedure. Robert Herzog, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 2 |
| 2009 | Temporal Glare: Real-Time Dynamic Simulation of the Scattering in the Human EyeabstractAbstract Glare is a consequence of light scattered within the human eye when looking at bright light sources. This effect can be exploited for tone mapping since adding glare to the depiction of high‐dynamic range (HDR) imagery on a low‐dynamic range (LDR) medium can dramatically increase perceived contrast. Even though most, if not all, subjects report perceiving glare as a bright pattern that fluctuates in time, up to now it has only been modeled as a static phenomenon. We argue that the temporal properties of glare are a strong means to increase perceived brightness and to produce realistic and attractive renderings of bright light sources. Based on the anatomy of the human eye, we propose a model that enables real‐time simulation of dynamic glare on a GPU. This allows an improved depiction of HDR images on LDR media for interactive applications like games, feature films, or even by adding movement to initially static HDR images. By conducting psychophysical studies, we validate that our method improves perceived brightness and that dynamic glare‐renderings are often perceived as more attractive depending on the chosen scene. Tobias Ritschel 0001, Matthias Mittner, Jeppe Revall Frisvad, Joris Coppens, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 5 |
| 2009 | Guest editorialabstractNo abstract available. Sarah H. Creem-Regehr, Karol Myszkowski |
ACM Trans. Appl. Percept. | 2 |
| 2008 | Render2MPEG: A Perception-based Framework Towards Integrating Rendering and Video CompressionabstractAbstract Currently 3D animation rendering and video compression are completely independent processes even if rendered frames are streamed on‐the‐fly within a client‐server platform. In such scenario, which may involve time‐varying transmission bandwidths and different display characteristics at the client side, dynamic adjustment of the rendering quality to such requirements can lead to a better use of server resources. In this work, we present a framework where the renderer and MPEG codec are coupled through a straightforward interface that provides precise motion vectors from the rendering side to the codec and perceptual error thresholds for each pixel in the opposite direction. The perceptual error thresholds take into account bandwidth‐dependent quantization errors resulting from the lossy com‐pression as well as image content‐dependent luminance and spatial contrast masking. The availability of the discrete cosine transform (DCT) coefficients at the codec side enables to use advanced models of the human visual system (HVS) in the perceptual error threshold derivation without incurring any significant cost. Those error thresholds are then used to control the rendering quality and make it well aligned with the compressed stream quality. In our prototype system we use the lightcuts technique developed by Walter et al., which we enhance to handle dynamic image sequences, and an MPEG‐2 implementation. Our results clearly demonstrate many advantages of coupling the rendering with video compression in terms of faster rendering. Furthermore, temporally coherent rendering leads to a reduction of temporal artifacts. Robert Herzog, Shinichi Kinuwaki, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 3 |
| 2008 | Apparent Greyscale: A Simple and Fast Conversion to Perceptually Accurate Images and VideoabstractAbstract This paper presents a quick and simple method for converting complex images and video to perceptually accurate greyscale versions. We use a two‐step approach first to globally assign grey values and determine colour ordering, then second, to locally enhance the greyscale to reproduce the original contrast. Our global mapping is image independent and incorporates the Helmholtz‐Kohlrausch colour appearance effect for predicting differences between isoluminant colours. Our multiscale local contrast enhancement reintroduces lost discontinuities only in regions that insufficiently represent original chromatic contrast. All operations are restricted so that they preserve the overall image appearance, lightness range and differences, colour ordering, and spatial details, resulting in perceptually accurate achromatic reproductions of the colour original. Kaleigh Smith, Pierre-Edouard Landes, Joëlle Thollot, Karol Myszkowski |
Comput. Graph. Forum | 4 |
| 2008 | Dynamic range independent image quality assessmentabstractThe diversity of display technologies and introduction of high dynamic range imagery introduces the necessity of comparing images of radically different dynamic ranges. Current quality assessment metrics are not suitable for this task, as they assume that both reference and test images have the same dynamic range. Image fidelity measures employed by a majority of current metrics, based on the difference of pixel intensity or contrast values between test and reference images, result in meaningless predictions if this assumption does not hold. We present a novel image quality metric capable of operating on an image pair where both images have arbitrary dynamic ranges. Our metric utilizes a model of the human visual system, and its central idea is a new definition of visible distortion based on the detection and classification of visible changes in the image structure. Our metric is carefully calibrated and its performance is validated through perceptual experiments. We demonstrate possible applications of our metric to the evaluation of direct and inverse tone mapping operators as well as the analysis of the image appearance on displays with various characteristics. Tunç Ozan Aydin, Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 3 |
| 2008 | 3D unsharp masking for scene coherent enhancementabstractWe present a new approach for enhancing local scene contrast by unsharp masking over arbitrary surfaces under any form of illumination. Our adaptation of a well-known 2D technique to 3D interactive scenarios is designed to aid viewers in tasks like understanding complex or detailed geometric models, medical visualization and navigation in virtual environments. Our holistic approach enhances the depiction of various visual cues, including gradients from surface shading, surface reflectance, shadows, and highlights, to ease estimation of viewpoint, lighting conditions, shapes of objects and their world-space organization. Motivated by recent perceptual findings on 3D aspects of the Cornsweet illusion, we create scene coherent enhancements by treating cues in terms of their 3D context; doing so has a stronger effect than approaches that operate in a 2D image context and also achieves temporal coherence. We validate our unsharp masking in 3D with psychophysical experiments showing that the enhanced images are perceived to have better contrast and are preferred over unenhanced originals. Our operator runs at real-time rates on a GPU and the effect is easily controlled interactively within the rendering pipeline. Tobias Ritschel 0001, Kaleigh Smith, Matthias Mittner, Thorsten Grosch, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 5 |
| 2007 | High Dynamic Range Image and Video Compression - Fidelity Matching Human Visual PerformanceabstractVast majority of digital images and video material stored today can capture only a fraction of visual information visible to the human eye and does not offer sufficient quality to fully exploit capabilities of new display devices. High dynamic range (HDR) image and video formats encode the full visible range of luminance and color gamut, thus offering ultimate fidelity, limited only by the capabilities of the human eye and not by any existing technology. In this paper we demonstrate how existing image and video compression standards can be extended to encode HDR content efficiently. This is achieved by a custom color space for encoding HDR pixel values that is derived from the visual performance data. We also demonstrate how HDR image and video compression can be designed so that it is backward compatible with existing formats. Rafal Mantiuk, Grzegorz Krawczyk, Karol Myszkowski, Hans-Peter Seidel |
ICIP (1) | 3 |
| 2007 | Global Illumination using Photon Ray SplattingabstractAbstract We present a novel framework for efficiently computing the indirect illumination in diffuse and moderately glossy scenes using density estimation techniques. Many existing global illumination approaches either quickly compute an overly approximate solution or perform an orders of magnitude slower computation to obtain high‐quality results for the indirect illumination. The proposed method improves photon density estimation and leads to significantly better visual quality in particular for complex geometry, while only slightly increasing the computation time. We perform direct splatting of photon rays, which allows us to use simpler search data structures. Since our density estimation is carried out in ray space rather than on surfaces, as in the commonly used photon mapping algorithm, the results are more robust against geometrically incurred sources of bias. This holds also in combination with final gathering where photon mapping often overestimates the illumination near concave geometric features. In addition, we show that our photon splatting technique can be extended to handle moderately glossy surfaces and can be combined with traditional irradiance caching for sparse sampling and filtering in image space. Robert Herzog, Vlastimil Havran, Shinichi Kinuwaki, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 4 |
| 2007 | Contrast Restoration by Adaptive CountershadingabstractAbstract The ABSTRACT is to be in fully‐justified italicized text, between two horizontal lines, in one‐column format, below the author and affiliation information. Use the word “Abstract” as the title, in 9‐point Times, boldface type, left‐aligned to the text, initially capitalized. The abstract is to be in 9‐point, single‐spaced type. The abstract may be up to 3 inches (7.62 cm) long. Leave one blank line after the abstract, then add the subject categories according to the ACM Classification Index (see http://www.acm.org/class/1998/ ). Grzegorz Krawczyk, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 2 |
| 2006 | Beyond Tone Mapping: Enhanced Depiction of Tone Mapped HDR ImagesabstractAbstract High Dynamic Range (HDR) images capture the full range of luminance present in real world scenes, and unlike Low Dynamic Range (LDR) images, can simultaneously contain detailed information in the deepest of shadows and the brightest of light sources. For display or aesthetic purposes, it is often necessary to perform tone mapping, which creates LDR depictions of HDR images at the cost of contrast information loss. The purpose of this work is two‐fold: to analyze a displayed LDR image against its original HDR counterpart in terms of perceived contrast distortion, and to enhance the LDR depiction with perceptually driven colour adjustments to restore the original HDR contrast information. For analysis, we present a novel algorithm for the characterization of tone mapping distortion in terms of observed loss of global contrast, and loss of contour and texture details. We classify existing tone mapping operators accordingly. We measure both distortions with perceptual metrics that enable the automatic and meaningful enhancement of LDR depictions. For image enhancement, we identify artistic and photographic colour techniques from which we derive adjustments that create contrast with colour. The enhanced LDR image is an improved depiction of the original HDR image with restored contrast information. Categories and Subject Descriptors (according to ACM CCS): I.3.3 [Computer Graphics]: Picture/Image Generation I.4.0 [Image Processing and Computer Vision]: GeneralImage processing software Kaleigh Smith, Grzegorz Krawczyk, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 3 |
| 2006 | Analysis of Reproducing Real-World Appearance on Displays of Varying Dynamic RangeabstractAbstract We conduct a series of experiments to investigate the desired properties of a tone mapping operator (TMO) and to design such an operator based on subjective data. We propose a novel approach to the tone mapping problem, in which the tone mapping parameters are determined based on the data from subjective experiments, rather than an image processing algorithm or a visual model. To collect this data, a series of experiments are conducted in which the subjects adjust three generic TMO parameters: brightness, contrast and color saturation. In two experiments, the subjects are to find a) the most preferred image without a reference image (preference task) and b) the closest image to the real‐world scene which the subjects are confronted with (fidelity task). We analyze subjects’ choice of parameters to provide more intuitive control over the parameters of a tone mapping operator. Unlike most of the researched TMOs that focus on rendering for standard low dynamic range monitors, we consider a broad range of potential displays, each offering different dynamic range and brightness. We simulate capabilities of such displays on a high dynamic range (HDR) display. This allows us to address the question of how tone mapping needs to be adjusted to accommodate displays with drastically different dynamic ranges. Categories and Subject Descriptors (according to ACM CCS): I.3.8 [Computer Graphics]: High dynamic range images, Visual perception, Tone mapping Akiko Yoshida, Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 3 |
| 2006 | A perceptual framework for contrast processing of high dynamic range images
Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Appl. Percept. | 2 |
| 2006 | Backward compatible high dynamic range MPEG video compressionabstractTo embrace the imminent transition from traditional low-contrast video (LDR) content to superior high dynamic range (HDR) content, we propose a novel backward compatible HDR video compression (HDR MPEG) method. We introduce a compact reconstruction function that is used to decompose an HDR video stream into a residual stream and a standard LDR stream, which can be played on existing MPEG decoders, such as DVD players. The reconstruction function is finely tuned to the content of each HDR frame to achieve strong decorrelation between the LDR and residual streams, which minimizes the amount of redundant information. The size of the residual stream is further reduced by removing invisible details prior to compression using our HDR-enabled filter, which models luminance adaptation, contrast sensitivity, and visual masking based on the HDR content. Designed especially for DVD movie distribution, our HDR MPEG compression method features low storage requirements for HDR content resulting in a 30% size increase to an LDR video sequence. The proposed compression method does not impose restrictions or modify the appearance of the LDR or HDR video. This is important for backward compatibility of the LDR stream with current DVD appearance, and also enables independent fine tuning, tone mapping, and color grading of both streams. Rafal Mantiuk, Alexander Efremov, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 3 |
| 2005 | Interactive System for Dynamic Scene Lighting using Captured Video Environment MapsabstractWe present an interactive system for fully dynamic scene lighting using captured high dynamic range (HDR) video environment maps. The key component of our system is an algorithm for efficient decomposition of HDR video environment map captured over hemisphere into a set of representative directional light sources, which can be used for the direct lighting computation with shadows using graphics hardware. The resulting lights exhibit good temporal coherence and their number can be adaptively changed to keep a constant framerate while good spatial distribution (stratification) properties are maintained. We can handle a large number of light sources with shadows using a novel technique which reduces the cost of BRDF-based shading and visibility computations. We demonstrate the use of our system in a mixed reality application in which real and synthetic objects are illuminated by consistent lighting at interactive framerates. Vlastimil Havran, Miloslaw Smyk, Grzegorz Krawczyk, Karol Myszkowski, Hans-Peter Seidel |
Rendering Techniques | 4 |
| 2005 | Lightness Perception in Tone Reproduction for High Dynamic Range ImagesabstractAn anchoring theory of lightness perception comprehensively explains many characteristics of human visual system such as lightness constancy and its spectacular failures which are important in the perception of images. We present a novel approach to tone mapping of high dynamic range (HDR) images which is inspired by the anchoring theory. The key concept of this method is the decomposition of an HDR image into areas (frameworks) of consistent luminance and the local calculation of the lightness values. The net lightness of an image is calculated via the merging of the frameworks proportionally to their strength. We stress out the importance of relating the luminance to a known brightness value (anchoring) and investigate the advantages of anchoring to the luminance value perceived as white. We validate the accuracy of the lightness reproduction in the presented algorithm by simulating a well known perception experiment. Our approach does not affect the local contrast and preserves the natural colors of an HDR image due to the linear handling of luminance. Grzegorz Krawczyk, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 2 |
| 2005 | Temporally Coherent Irradiance Caching for High Quality Animation RenderingabstractIn rendering of high quality animations that include global illumination, the final gathering and irradiance caching are commonly used.However, the computational cost they incur is high enough to discourage their wide use in production rendering.We introduce a data structure called anchor, which lets us permanently link cache locations to points intersected by their final gathering rays.Consequently, we can cheaply probe and transfer the (ir)radiance by exploiting the temporal coherence of successive animation frames, resulting in half an order of magnitude acceleration and reduced temporal artifacts.Additionally, our anchor structure lets us render moderately glossy surfaces at the cost much lower than the traditional importance sampling techniques.We also describe an efficient, perceptually motivated and independent scheme for limiting the growth in the number of irradiance caches.Finally, an implementation in a practical rendering system is demonstrated. Miloslaw Smyk, Shinichi Kinuwaki, Roman Durikovic, Karol Myszkowski |
Comput. Graph. Forum | 4 |
| 2004 | Exploiting Temporal Coherence in Final Gathering for Dynamic ScenesabstractEfficient global illumination computation in dynamically changing environments is an important practical problem. In high-quality animation rendering costly "final gathering" technique is commonly used. We extend this technique into temporal domain by exploiting coherence between the subsequent frames. For this purpose we store previously computed incoming radiance samples and refresh them evenly in space and time using some aging criteria. The approach is based upon a two-pass photon mapping algorithm with irradiance cache, but it can be applied also in other gathering methods. The algorithm significantly reduces the cost of expensive indirect lighting computation and suppresses temporal aliasing with respect to the state of the art frame-by-frame rendering techniques Takehiro Tawara, Karol Myszkowski, Hans-Peter Seidel |
Computer Graphics International | 2 |
| 2004 | Spatio-Temporal Photon Density Estimation Using Bilateral FilteringabstractPhoton tracing and density estimation are well established techniques in global illumination computation and rendering of high-quality animation sequences. Using traditional density estimation techniques it is difficult to remove stochastic noise inherent for photon-based methods while avoiding overblurring lighting details. In this paper we investigate the use of bilateral filtering for lighting reconstruction based on the local density of photon hit points. Bilateral filtering is applied in spatio-temporal domain and provides control over the level-of-details in reconstructed lighting. All changes of lighting below this level are treated as stochastic noise and are suppressed. Bilateral filtering proves to be efficient in preserving sharp features in lighting which is in particular important for high-quality caustic reconstruction. Also, flickering between subsequent animation frames is substantially reduced due to extending bilateral filtering into temporal domain. Marco Milch, Karol Myszkowski, Kirill Dmitriev, Przemyslaw Rokita, Hans-Peter Seidel |
Computer Graphics International | 3 |
| 2004 | A CAVE system for interactive modeling of global illumination in car interiorabstractGlobal illumination dramatically improves realistic appearance\nof rendered scenes, but usually it is neglected in VR systems\ndue to its high costs. In this work we present an efficient \nglobal illumination solution specifically tailored for those CAVE\napplications, which require an immediate response for dynamic light\nchanges and allow for free motion of the observer, but involve scenes\nwith static geometry. As an application example we choose\nthe car interior modeling under free driving conditions.\nWe illuminate the car using dynamically changing High Dynamic Range \n(HDR) environment maps and use the Precomputed Radiance Transfer (PRT) \nmethod for the global illumination computation. We\nleverage the PRT method to handle scenes with non-trivial topology\nrepresented by complex meshes. Also, we propose a hybrid of\nPRT and final gathering approach for high-quality rendering\nof objects with complex Bi-directional Reflectance Distribution\nFunction (BRDF). We use\nthis method for predictive rendering of the navigation LCD panel\nbased on its measured BRDF. Since the global illumination \ncomputation leads to HDR images we propose a tone mapping\nalgorithm tailored specifically for the CAVE. We employ \nhead tracking to identify the observed screen region\nand derive for it proper luminance adaptation conditions, \nwhich are then used for tone mapping on all walls in the CAVE.\nWe distribute our global illumination and tone mapping computation\non all CPUs and GPUs available in the\nCAVE, which enables us to achieve interactive performance\neven for the costly final gathering approach. Kirill Dmitriev, Thomas Annen, Grzegorz Krawczyk, Karol Myszkowski, Hans-Peter Seidel |
VRST | 4 |
| 2004 | Perception-motivated high dynamic range video encodingabstractDue to rapid technological progress in high dynamic range (HDR) video capture and display, the efficient storage and transmission of such data is crucial for the completeness of any HDR imaging pipeline. We propose a new approach for inter-frame encoding of HDR video, which is embedded in the well-established MPEG-4 video compression standard. The key component of our technique is luminance quantization that is optimized for the contrast threshold perception in the human visual system. The quantization scheme requires only 10--11 bits to encode 12 orders of magnitude of visible luminance range and does not lead to perceivable contouring artifacts. Besides video encoding, the proposed quantization provides perceptually-optimized luminance sampling for fast implementation of any global tone mapping operator using a lookup table. To improve the quality of synthetic video sequences, we introduce a coding scheme for discrete cosine transform (DCT) blocks with high contrast. We demonstrate the capabilities of HDR video in a player, which enables decoding, tone mapping, and applying post-processing effects in real-time. The tone mapping algorithm as well as its parameters can be changed interactively while the video is playing. We can simulate post-processing effects such as glare, night vision, and motion blur, which appear very realistic due to the usage of HDR data. Rafal Mantiuk, Grzegorz Krawczyk, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 3 |
| 2004 | Reverse engineering approach to appearance-based design of metallic and pearlescent paints
Sergey V. Ershov, Roman Durikovic, Konstantin Kolchin, Karol Myszkowski |
Vis. Comput. | 4 |
| 2003 | Perceptual evaluation of tone mapping operatorsabstractNo abstract available. Frédéric Drago, William L. Martens, Karol Myszkowski, Hans-Peter Seidel |
SIGGRAPH | 3 |
| 2003 | An efficient spatio-temporal architecture for animation renderingabstractNo abstract available. Vlastimil Havran, Cyrille Damez, Karol Myszkowski, Hans-Peter Seidel |
SIGGRAPH | 3 |
| 2003 | State of the Art in Global Illumination for Interactive Applications and High-quality AnimationsabstractAbstract Global illumination algorithms are regarded as computationally intensive. This cost is a practical problem when producing animations or when interactions with complex models are required. Several algorithms have been proposed to address this issue. Roughly, two families of methods can be distinguished. The first one aims at providing interactive feedback for lighting design applications. The second one gives higher priority to the quality of results, and therefore relies on offline computations. Recently, impressive advances have been made in both categories. In this report, we present a survey and classification of the most up‐to‐date of these methods. ACM CSS: I.3.7 Computer Graphics—Three‐Dimensional Graphics and Realism Cyrille Damez, Kirill Dmitriev, Karol Myszkowski |
Comput. Graph. Forum | 3 |
| 2003 | Adaptive Logarithmic Mapping For Displaying High Contrast ScenesabstractAbstract We propose a fast, high quality tone mapping technique to display high contrast images on devices with limited dynamicrange of luminance values. The method is based on logarithmic compression of luminance values, imitatingthe human response to light. A bias power function is introduced to adaptively vary logarithmic bases, resultingin good preservation of details and contrast. To improve contrast in dark areas, changes to the gamma correctionprocedure are proposed. Our adaptive logarithmic mapping technique is capable of producing perceptually tunedimages with high dynamic content and works at interactive speed. We demonstrate a successful application of ourtone mapping technique with a high dynamic range video player enabling to adjust optimal viewing conditions forany kind of display while taking into account user preference concerning brightness, contrast compression, anddetail reproduction. Categories and Subject Descriptors (according to ACM CCS): I.3.3 [Image Processing and Computer Vision]: Image Representation Frédéric Drago, Karol Myszkowski, Thomas Annen, Norishige Chiba |
Comput. Graph. Forum | 2 |
| 2001 | Perception-guided global illumination solution for animation renderingabstractWe present a method for efficient global illumination computation in dynamic environments by taking advantage of temporal coherence of lighting distribution. The method is embedded in the framework of stochastic photon tracing and density estimation techniques. A locally operating energy-based error metric is used to prevent photon processing in the temporal domain for the scene regions in which lighting distribution changes rapidly. A perception-based error metric suitable for animation is used to keep noise inherent in stochastic methods below the sensitivity level of the human observer. As a result a perceptually-consistent quality across all animation frames is obtained. Furthermore, the computation cost is reduced compared to the traditional approaches operating solely in the spatial domain. Karol Myszkowski, Takehiro Tawara, Hiroyuki Akamine, Hans-Peter Seidel |
SIGGRAPH | 1 |
| 2001 | Validation proposal for global illumination and rendering techniques
Frédéric Drago, Karol Myszkowski |
Comput. Graph. | 2 |
| 2001 | Rendering Pearlescent Appearance Based On Paint-Composition ModellingabstractWe describe a new approach to modelling pearlescent paints based on decomposing paint layers into stacks of imaginary thin sublayers. The sublayers are chosen so thin that multiple scattering can be considered across different sublayers, while it can be neglected within each of the sublayers. Based on this assumption, an efficient recursive procedure of assembling the layers is developed, which enables to compute the paint BRDF at interactive speeds. Since the proposed paint model connects fundamental optical properties of multi-layer pearlescent and metallic paints with their microscopic structure, interactive prediction of the paint appearance based on its composition becomes possible. Sergey V. Ershov, Konstantin Kolchin, Karol Myszkowski |
Comput. Graph. Forum | 3 |
| 2001 | Perceptually Guided Corrective SplattingabstractOne of the basic difficulties with interactive walkthroughs is the high quality rendering of object surfaces with non-diffuse light scattering characteristics. Since full ray tracing at interactive rates is usually impossible, we render a precomputed global illumination solution using graphics hardware and use remaining computational power to correct the appearance of non-diffuse objects on-the-fly. The question arises, how to obtain the best image quality as perceived by a human observer within a limited amount of time for each frame. We address this problem by enforcing corrective computation for those non-diffuse objects that are selected using a computational model of visual attention. We consider both the saliency- and task-driven selection of those objects and benefit from the fact that shading artifacts of “unattended” objects are likely to remain unnoticed. We use a hierarchical image-space sampling scheme to control ray tracing and splat the generated point samples. The resulting image converges progressively to a ray traced solution if the viewing parameters remain unchanged. Moreover, we use a sample cache to enhance visual appearance if the time budget for correction has been too low for some frame. We check the validity of the cached samples using a novel criterion suited for non-diffuse surfaces and reproject valid samples into the current view. Jörg Haber, Karol Myszkowski, Hitoshi Yamauchi, Hans-Peter Seidel |
Comput. Graph. Forum | 2 |
| 2000 | Using the visual differences predictor to improve performance of progressive global illumination computationabstractA novel view-independent technique for progressive global illumination computing that uses prediction of visible differences to improve both efficiency and effectiveness of physically-sound lighting solutions has been developed. The technique is a mixture of stochastic (density estimation) and deterministic (adaptive mesh refinement) algorithms used in a sequence and optimized to reduce the differences between the intermediate and final images as perceived by the human observer in the course of lighting computation. The quantitive measurements of visibility were obtained using the model of human vision captured in the visible differences predictor (VDP) developed by Daly [1993]. The VDP responses were used to support the selection of the best component algorithms from a pool of global illumination solutions, and to enhance the selected algorithms for even better progressive refinement of image quality. The VDP was also used to determine the optimal sequential order of component-algorithm execution, and to choose the points at which switchover between algorithms should take place. As the VDP is computationally expensive, it was applied exclusively at the design and tuning stage of the composite technique, and so perceptual considerations are embedded into the resulting solution, though no VDP calculations were performed during lighting simulation. The proposed illumination technique is also novel, providing intermediate image solutions of high quality at unprecedented speeds, even for complex scenes. One advantage of the technique is that local estimates of global illumination are readily available at the early stages of computing, making possible the development of a more robust adaptive mesh subdivision, which is guided by local contrast information. Efficient object space filtering, also based on stochastically-derived estimates of the local illumination error, is applied to substantially reduce the visible noise inherent in stochastic solutions. Vladimir Volevich, Karol Myszkowski, Andrei Khodulev, Edward A. Kopylov |
ACM Trans. Graph. | 2 |
| 2000 | Perception-Based Fast Rendering and Antialiasing of Walkthrough SequencesabstractWe consider accelerated rendering of high quality walkthrough animation sequences along predefined paths. To improve rendering performance, we use a combination of a hybrid ray tracing and image-based rendering (IBR) technique and a novel perception-based antialiasing technique. In our rendering solution, we derive as many pixels as possible using inexpensive IBR techniques without affecting the animation quality. A perception-based spatiotemporal animation quality metric (AQM) is used to automatically guide such a hybrid rendering. The image flow (IF) obtained as a byproduct of the IBR computation is an integral part of the AQM. The final animation quality is enhanced by an efficient spatiotemporal antialiasing which utilizes the IF to perform a motion-compensated filtering. The filter parameters have been tuned using the AQM predictions of animation quality as perceived by the human observer. These parameters adapt locally to the visual pattern velocity. Karol Myszkowski, Przemyslaw Rokita, Takehiro Tawara |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2000 | A case study towards validation of global illumination algorithms: progressive hierarchical radiosity with clustering
Karol Myszkowski, Tosiyasu L. Kunii |
Vis. Comput. | 1 |
| 1997 | Modelling of Human Jaw Motion in Sliding ContactabstractDental CAD/CAM requires appropriate modelling of human jaw motion in contact with the opposite jaw. Such modelling is necessary for planning orthodontic treatment and robust design of dental restorations, especially their occlusal surfaces which must fit the existing articulation patterns. We propose a fast and purely geometrical approach to model sliding of the lower jaw over the surface of the upper, fixed jaw. For every discrete step of the motion, a new position and orientation of the sliding object is found by maximizing its displacement toward the fixed object, to obtain stable contact without interpenetrations. We impose constraints on the range of rotation angles to justify simplifications of our linear model handling the distance from the fixed object to chosen points on the surface of the sliding object as a function of its configuration. In such formulation, we reduce the complexity of the problem of object placement in contact to that of a linear optimization task. Expensive multi-point collision detection and distance computation are handled by rasterizing graphics hardware, which supports generation of valid configurations for complex objects at stable and interactive speeds. Our model of human jaw sliding compares favourably with experimental motion data. © 1997 John Wiley & Sons, Ltd. Karol Myszkowski, Oleg G. Okunev, Tosiyasu L. Kunii, Masumi Ibusuki |
Comput. Animat. Virtual Worlds | 1 |
| 1996 | Computer Modeling For The Occlusal Surface Of TeethabstractModeling of the occlusal surface of teeth is an important problem in computer-aided design of dental restorations. The designed shape must fit the existing jaw articulation. Also, the design process must be fast to be practical in clinical applications. In this paper we present techniques for automatic adjustment of the occlusal surface of restorations based on the results of articulation simulation. The shape of restorations is changed to avoid interpenetrations with the opponent teeth during functional jaw movements. The 3-D space mapping method is used which guarantees the surface smoothness and preservation of the main topological features of the occlusal surface. To speedup calculations we use rasterizing graphics hardware for computationally involved collision detection between complex surfaces of teeth. Karol Myszkowski, Vladimir V. Savchenko, Tosiyasu L. Kunii |
Computer Graphics International | 1 |
| 1995 | Evaluation of human jaw articulation [computer animation]abstractComputer-aided diagnosis of occlusal disorders and design of dental restorations requires an automated evaluation of jaw occlusion and chewing ability. This requires simulation of the motion of the jaws and characterization of contacts between the surfaces of teeth. We propose approaches to evaluation of the load on teeth and of the grinding process. These characteristics are derived in interactive time, and are based on distance maps and topological structure of the contact zones. The proposed approaches are general and usable in applications where modeling of contact between objects with complex geometry is required.> Tosiyasu L. Kunii, Karol Myszkowski, Oleg G. Okunev, Hirobumi Nishida, Yoshihisa Shinagawa, Masumi Ibusuki |
CA | 2 |
| 1995 | Fast collision detection between complex solids using rasterizing graphics hardware
Karol Myszkowski, Oleg G. Okunev, Tosiyasu L. Kunii |
Vis. Comput. | 1 |