Mojtaba Bemana

dblp:232/2391 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-0202-6266ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Bracket Diffusion: HDR Image Generation by Consistent LDR Denoising
abstract
Abstract We demonstrate generating HDR images using the concerted action of multiple black‐box, pre‐trained LDR image diffusion models. Common diffusion models are not HDR as, first, there is no sufficiently large HDR image dataset available to re‐train them, and, second, even if it was, re‐training such models is impossible for most compute budgets. Instead, we seek inspiration from the HDR image capture literature that traditionally fuses sets of LDR images, called “exposure brackets”, to produce a single HDR image. We operate multiple denoising processes to generate multiple LDR brackets that together form a valid HDR result. To this end, we introduce a brackets consistency term into the diffusion process to couple the brackets such that they agree across the exposure range they share. We demonstrate HDR versions of state‐of‐the‐art unconditional and conditional as well as restoration‐type (LDR2HDR) generative modeling.
Mojtaba Bemana, Thomas Leimkühler, Karol Myszkowski, Hans-Peter Seidel, Tobias Ritschel 0001
Comput. Graph. Forum1
2025 MILO: A Lightweight Perceptual Quality Metric for Image and Latent-Space Optimization
abstract
We present MILO (Metric for Image- and Latent-space Optimization), a lightweight, multiscale, perceptual metric for full-reference image quality assessment (FR-IQA). MILO is trained using pseudo-MOS (Mean Opinion Score) supervision, in which reproducible distortions are applied to diverse images and scored via an ensemble of recent quality metrics that account for visual masking effects. This approach enables accurate learning without requiring large-scale human-labeled datasets. Despite its compact architecture, MILO outperforms existing metrics across standard FR-IQA benchmarks and offers fast inference suitable for real-time applications. Beyond quality prediction, we demonstrate the utility of MILO as a perceptual loss in both image and latent domains. In particular, we show that spatial masking modeled by MILO, when applied to latent representations from a VAE encoder within Stable Diffusion, enables efficient and perceptually aligned optimization. By combining spatial masking with a curriculum learning strategy, we first process perceptually less relevant regions before progressively shifting the optimization to more visually distorted areas. This strategy leads to significantly improved performance in tasks like denoising, super-resolution, and face restoration, while also reducing computational overhead. MILO thus functions as both a state-of-the-art image quality metric and as a practical tool for perceptual optimization in generative pipelines.
Ugur Çogalan, Mojtaba Bemana, Karol Myszkowski, Hans-Peter Seidel, Colin Groth
ACM Trans. Graph.2
2024 Enhancing image quality prediction with self-supervised visual masking
abstract
Abstract Full‐reference image quality metrics (FR‐IQMs) aim to measure the visual differences between a pair of reference and distorted images, with the goal of accurately predicting human judgments. However, existing FR‐IQMs, including traditional ones like PSNR and SSIM and even perceptual ones such as HDR‐VDP, LPIPS, and DISTS, still fall short in capturing the complexities and nuances of human perception. In this work, rather than devising a novel IQM model, we seek to improve upon the perceptual quality of existing FR‐IQM methods. We achieve this by considering visual masking, an important characteristic of the human visual system that changes its sensitivity to distortions as a function of local image content. Specifically, for a given FR‐IQM metric, we propose to predict a visual masking model that modulates reference and distorted images in a way that penalizes the visual errors based on their visibility. Since the ground truth visual masks are difficult to obtain, we demonstrate how they can be derived in a self‐supervised manner solely based on mean opinion scores (MOS) collected from an FR‐IQM dataset. Our approach results in enhanced FR‐IQM metrics that are more in line with human prediction both visually and quantitatively.
Ugur Çogalan, Mojtaba Bemana, Hans-Peter Seidel, Karol Myszkowski
Comput. Graph. Forum2
2024 Cinematic Gaussians: Real-Time HDR Radiance Fields with Depth of Field
abstract
Abstract Radiance field methods represent the state of the art in reconstructing complex scenes from multi‐view photos. However, these reconstructions often suffer from one or both of the following limitations: First, they typically represent scenes in low dynamic range (LDR), which restricts their use to evenly lit environments and hinders immersive viewing experiences. Secondly, their reliance on a pinhole camera model, assuming all scene elements are in focus in the input images, presents practical challenges and complicates refocusing during novel‐view synthesis. Addressing these limitations, we present a lightweight method based on 3D Gaussian Splatting that utilizes multi‐view LDR images of a scene with varying exposure times, apertures, and focus distances as input to reconstruct a high‐dynamic‐range (HDR) radiance field. By incorporating analytical convolutions of Gaussians based on a thin‐lens camera model as well as a tonemapping module, our reconstructions enable the rendering of HDR content with flexible refocusing capabilities. We demonstrate that our combined treatment of HDR and depth of field facilitates real‐time cinematic rendering, outperforming the state of the art.
Chao Wang 0037, Krzysztof Wolski, Bernhard Kerbl, Ana Serrano, Mojtaba Bemana, Hans-Peter Seidel, Karol Myszkowski, Thomas Leimkühler
Comput. Graph. Forum5
2023 Video frame interpolation for high dynamic range sequences captured with dual-exposure sensors
abstract
Abstract Video frame interpolation (VFI) enables many important applications such as slow motion playback and frame rate conversion. However, one major challenge in using VFI is accurately handling high dynamic range (HDR) scenes with complex motion. To this end, we explore the possible advantages of dual‐exposure sensors that readily provide sharp short and blurry long exposures that are spatially registered and whose ends are temporally aligned. This way, motion blur registers temporally continuous information on the scene motion that, combined with the sharp reference, enables more precise motion sampling within a single camera shot. We demonstrate that this facilitates a more complex motion reconstruction in the VFI task, as well as HDR frame reconstruction that so far has been considered only for the originally captured frames, not in‐between interpolated frames. We design a neural network trained in these tasks that clearly outperforms existing solutions. We also propose a metric for scene motion complexity that provides important insights into the performance of VFI methods at test time.
Ugur Çogalan, Mojtaba Bemana, Hans-Peter Seidel, Karol Myszkowski
Comput. Graph. Forum2
2022 Learning HDR video reconstruction for dual-exposure sensors with temporally-alternating exposures
Ugur Çogalan, Mojtaba Bemana, Karol Myszkowski, Hans-Peter Seidel, Tobias Ritschel 0001
Comput. Graph.2
2020 X-Fields: implicit neural view-, light- and time-image interpolation
abstract
We suggest to represent an X-Field ---a set of 2D images taken across different view, time or illumination conditions, i.e., video, lightfield, reflectance fields or combinations thereof---by learning a neural network (NN) to map their view, time or light coordinates to 2D images. Executing this NN at new coordinates results in joint view, time or light interpolation. The key idea to make this workable is a NN that already knows the "basic tricks" of graphics (lighting, 3D projection, occlusion) in a hard-coded and differentiable form. The NN represents the input to that rendering as an implicit map, that for any view, time, or light coordinate and for any pixel can quantify how it will move if view, time or light coordinates change (Jacobian of pixel position with respect to view, time, illumination, etc.). Our X-Field representation is trained for one scene within minutes, leading to a compact set of trainable parameters and hence real-time navigation in view, time and illumination.
Mojtaba Bemana, Karol Myszkowski, Hans-Peter Seidel, Tobias Ritschel 0001
ACM Trans. Graph.1
2019 Learning to Predict Image-based Rendering Artifacts with Respect to a Hidden Reference Image
abstract
Abstract Image metrics predict the perceived per‐pixel difference between a reference image and its degraded (e. g., re‐rendered) version. In several important applications, the reference image is not available and image metrics cannot be applied. We devise a neural network architecture and training procedure that allows predicting the MSE, SSIM or VGG16 image difference from the distorted image alone while the reference is not observed. This is enabled by two insights: The first is to inject sufficiently many un‐distorted natural image patches, which can be found in arbitrary amounts and are known to have no perceivable difference to themselves. This avoids false positives. The second is to balance the learning, where it is carefully made sure that all image errors are equally likely, avoiding false negatives. Surprisingly, we observe that the resulting no‐reference metric, subjectively, can even perform better than the reference‐based one, as it had to become robust against mis‐alignments. We evaluate the effectiveness of our approach in an image‐based rendering context, both quantitatively and qualitatively. Finally, we demonstrate two applications which reduce light field capture time and provide guidance for interactive depth adjustment.
Mojtaba Bemana, Joachim Keinert, Karol Myszkowski, Michel Bätz, Matthias Ziegler 0001, Hans-Peter Seidel, Tobias Ritschel 0001
Comput. Graph. Forum1
2019 A Perception-driven Hybrid Decomposition for Multi-layer Accommodative Displays
abstract
Multi-focal plane and multi-layered light-field displays are promising solutions for addressing all visual cues observed in the real world. Unfortunately, these devices usually require expensive optimizations to compute a suitable decomposition of the input light field or focal stack to drive individual display layers. Although these methods provide near-correct image reconstruction, a significant computational cost prevents real-time applications. A simple alternative is a linear blending strategy which decomposes a single 2D image using depth information. This method provides real-time performance, but it generates inaccurate results at occlusion boundaries and on glossy surfaces. This paper proposes a perception-based hybrid decomposition technique which combines the advantages of the above strategies and achieves both real-time performance and high-fidelity results. The fundamental idea is to apply expensive optimizations only in regions where it is perceptually superior, e.g., depth discontinuities at the fovea, and fall back to less costly linear blending otherwise. We present a complete, perception-informed analysis and model that locally determine which of the two strategies should be applied. The prediction is later utilized by our new synthesis method which performs the image decomposition. The results are analyzed and validated in user experiments on a custom multi-plane display.
Hyeonseung Yu, Mojtaba Bemana, Marek Wernikowski, Michal Chwesiuk, Okan Tarhan Tursun, Gurprit Singh, Karol Myszkowski, Radoslaw Mantiuk, Hans-Peter Seidel, Piotr Didyk
IEEE Trans. Vis. Comput. Graph.2