EDBT 2026 Demo / reviewers in the wild / expert
Joerg H. Mueller
dblp:195/1439 · also Jörg H. Müller
· DBLP profile ↗
14ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-6368-6340ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 8 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sparse Cache Updates for Scalable Distributed Effect-Based RenderingabstractCloud computing has seen rapid growth in recent years, accompanied by the increasing popularity of game streaming services that allow users to play high-end games on low-end devices, across platforms, and from virtually anywhere. The rise of multiplayer games, shared immersive experiences, and metaverse-style applications—such as exhibitions or social virtual spaces—presents unique opportunities for improving rendering efficiency. In particular, the presence of multiple viewers within the same virtual environment opens the door for computation reuse across rendering instances. We propose a scalable, multi-GPU cloud rendering system tailored for multi-viewer scenarios. Built on top of on-surface caches (OSC), our system extends the core idea of decoupling shading from viewpoints to enable efficient reuse of shading information across multiple users. Our system is designed to scale with an increasing number of viewers by dynamically distributing rendering workloads across multiple GPUs. We further enhance scalability and significantly reduce inter-GPU bandwidth requirements from 6 × up to 65 × —through a novel sparse cache update strategy. Instead of copying full frames between GPUs, our method selectively propagates only relevant cache updates, enabling efficient data sharing while minimizing redundant transfers. Wolfgang Tatzgern, Pascal Stadlbauer, Joerg H. Mueller, Martin Sattlecker, Markus Steinberger |
SIGGRAPH Asia | 3 |
| 2025 | Adaptive Multi-view Radiance Caching for Heterogeneous Participating MediaabstractAbstract Achieving lifelike atmospheric effects, such as fog, is essential in creating immersive environments and poses a formidable challenge in real‐time rendering. Highly realistic rendering of complex lighting interacting with dynamic fog can be very resource‐intensive, due to light bouncing through a complex participating media multiple times. We propose an approach that uses a multi‐layered spherical harmonics probe grid to share computations temporarily. In addition, this world‐space storage enables the sharing of radiance data between multiple viewers. In the context of cloud rendering this means faster rendering and a significant enhancement in overall rendering quality with efficient resource utilization. Pascal Stadlbauer, Wolfgang Tatzgern, Joerg H. Mueller, Robert Stojanovic, Alexander Weinrauch, Markus Steinberger |
Comput. Graph. Forum | 3 |
| 2024 | Real-time Neural Rendering of Dynamic Light FieldsabstractAbstract Synthesising high‐quality views of dynamic scenes via path tracing is prohibitively expensive. Although caching offline‐quality global illumination in neural networks alleviates this issue, existing neural view synthesis methods are limited to mainly static scenes, have low inference performance or do not integrate well with existing rendering paradigms. We propose a novel neural method that is able to capture a dynamic light field, renders at real‐time frame rates at 1920×1080 resolution and integrates seamlessly with Monte Carlo ray tracing frameworks. We demonstrate how a combination of spatial, temporal and a novel surface‐space encoding are each effective at capturing different kinds of spatio‐temporal signals. Together with a compact fully‐fused neural network and architectural improvements, we achieve a twenty‐fold increase in network inference speed compared to related methods at equal or better quality. Our approach is suitable for providing offline‐quality real‐time rendering in a variety of scenarios, such as free‐viewpoint video, interactive multi‐view rendering, or streaming rendering. Finally, our work can be integrated into other rendering paradigms, e.g., providing a dynamic background for interactive scenarios where the foreground is rendered with traditional methods. Arno Coomans, Edoardo A. Dominici, Christian Döring, Joerg H. Mueller, Jozef Hladky, Markus Steinberger |
Comput. Graph. Forum | 4 |
| 2023 | Trim Regions for Online Computation of From-Region Potentially Visible SetsabstractVisibility computation is a key element in computer graphics applications. More specifically, a from-region potentially visible set (PVS) is an established tool in rendering acceleration, but its high computational cost means a from-region PVS is almost always precomputed. Precomputation restricts the use of PVS to static scenes and leads to high storage cost, in particular, if we need fine-grained regions. For dynamic applications, such as streaming content over a variable-bandwidth network, online PVS computation with configurable region size is required. We address this need with trim regions, a new method for generating from-region PVS for arbitrary scenes in real time. Trim regions perform controlled erosion of object silhouettes in image space, implicitly applying the shrinking theorem known from previous work. Our algorithm is the first that applies automatic shrinking to unconstrained 3D scenes, including non-manifold meshes, and does so in real time using an efficient GPU execution model. We demonstrate that our algorithm generates a tight PVS for complex scenes and outperforms previous online methods for from-viewpoint and from-region PVS. It runs at 60 Hz for realistic game scenes consisting of millions of triangles and computes PVS with a tightness matching or surpassing existing approaches. Philip Voglreiter, Bernhard Kerbl, Alexander Weinrauch, Joerg H. Mueller, Thomas Neff, Markus Steinberger, Dieter Schmalstieg |
ACM Trans. Graph. | 4 |
| 2023 | Effect-based Multi-viewer Caching for Cloud-native RenderingabstractWith cloud computing becoming ubiquitous, it appears as virtually everything can be offered as-a-service. However, real-time rendering in the cloud forms a notable exception, where the cloud adoption stops at running individual game instances in compute centers. In this paper, we explore whether a cloud-native rendering architecture is viable and scales to multi-client rendering scenarios. To this end, we propose world-space and on-surface caches to share rendering computations among viewers placed in the same virtual world. We discuss how caches can be utilized on an effect-basis and demonstrate that a large amount of computations can be saved as the number of viewers in a scene increases. Caches can easily be set up for various effects, including ambient occlusion, direct illumination, and diffuse global illumination. Our results underline that the image quality using cached rendering is on par with screen-space rendering and due to its simplicity and inherent coherence, cached rendering may even have advantages in single viewer setups. Analyzing the runtime and communication costs, we show that cached rendering is already viable in multi-GPU systems. Building on top of our research, cloud-native rendering may be just around the corner. Alexander Weinrauch, Wolfgang Tatzgern, Pascal Stadlbauer, Alexis Crickx, Jozef Hladky, Arno Coomans, Joerg H. Mueller, Markus Steinberger |
ACM Trans. Graph. | 8 |
| 2022 | Meshlets and How to Shade Them: A Study on Texture-Space ShadingabstractAbstract Commonly used image‐space layouts of shading points, such as used in deferred shading, are strictly view‐dependent, which restricts efficient caching and temporal amortization. In contrast, texture‐space layouts can represent shading on all surface points and can be tailored to the needs of a particular application. However, the best grouping of shading points—which we call a shading unit—in texture space remains unclear. Choices of shading unit granularity (how many primitives or pixels per unit) and in shading unit parametrization (how to assign texture coordinates to shading points) lead to different outcomes in terms of final image quality, overshading cost, and memory consumption. Among the possible choices, shading units consisting of larger groups of scene primitives, so‐called meshlets, remain unexplored as of yet. In this paper, we introduce a taxonomy for analyzing existing texture‐space shading methods based on the group size and parametrization of shading units. Furthermore, we introduce a novel texture‐space layout strategy that operates on large shading units: the meshlet shading atlas. We experimentally demonstrate that the meshlet shading atlas outperforms previous approaches in terms of image quality, run‐time performance and temporal upsampling for a given number of fragment shader invocations. The meshlet shading atlas lends itself to work together with popular cluster‐based rendering of meshes with high geometric detail. Thomas Neff, Joerg H. Mueller, Markus Steinberger, Dieter Schmalstieg |
Comput. Graph. Forum | 2 |
| 2021 | DONeRF: Towards Real-Time Rendering of Compact Neural Radiance Fields using Depth Oracle NetworksabstractAbstract The recent research explosion around implicit neural representations, such as NeRF, shows that there is immense potential for implicitly storing high‐quality scene and lighting information in compact neural networks. However, one major limitation preventing the use of NeRF in real‐time rendering applications is the prohibitive computational cost of excessive network evaluations along each view ray, requiring dozens of petaFLOPS. In this work, we bring compact neural representations closer to practical rendering of synthetic content in real‐time applications, such as games and virtual reality. We show that the number of samples required for each view ray can be significantly reduced when samples are placed around surfaces in the scene without compromising image quality. To this end, we propose a depth oracle network that predicts ray sample locations for each view ray with a single network evaluation. We show that using a classification network around logarithmically discretized and spherically warped depth values is essential to encode surface locations rather than directly estimating depth. The combination of these techniques leads to DONeRF, our compact dual network design with a depth oracle network as its first step and a locally sampled shading network for ray accumulation. With DONeRF, we reduce the inference costs by up to 48× compared to NeRF when conditioning on available ground truth depth information. Compared to concurrent acceleration methods for raymarching‐based neural representations, DONeRF does not require additional memory for explicit caching or acceleration structures, and can render interactively (20 frames per second) on a single GPU. Thomas Neff, Pascal Stadlbauer, Mathias Parger, Andreas Kurz, Joerg H. Mueller, Chakravarty R. Alla Chaitanya, Anton Kaplanyan, Markus Steinberger |
Comput. Graph. Forum | 5 |
| 2021 | Temporally Adaptive Shading Reuse for Real-Time Rendering and Virtual RealityabstractTemporal coherence has the potential to enable a huge reduction of shading costs in rendering. Existing techniques focus either only on spatial shading reuse or cannot adaptively choose temporal shading frequencies. We find that temporal shading reuse is possible for extended periods of time for a majority of samples, and we show under which circumstances users perceive temporal artifacts. Our analysis implies that we can approximate shading gradients to efficiently determine when and how long shading can be reused. Whereas visibility usually stays temporally coherent from frame to frame for more than 90%, we find that even in heavily animated game scenes with advanced shading, typically more than 50% of shading is also temporally coherent. To exploit this potential, we introduce a temporally adaptive shading framework and apply it to two real-time methods. Its application saves more than 57% of the shader invocations, reducing overall rendering times up to in virtual reality applications without a noticeable loss in visual quality. Overall, our work shows that there is significantly more potential for shading reuse than currently exploited. Joerg H. Mueller, Thomas Neff, Philip Voglreiter, Markus Steinberger, Dieter Schmalstieg |
ACM Trans. Graph. | 1 |
| 2018 | The Broker Queue: A Fast, Linearizable FIFO Queue for Fine-Granular Work Distribution on the GPUabstractHarnessing the power of massively parallel devices like the graphics processing unit (GPU) is difficult for algorithms that show dynamic or inhomogeneous workloads. To achieve high performance, such advanced algorithms require scalable, concurrent queues to collect and distribute work. We show that previous queuing approaches are unfit for this task, as they either (1) do not work well in a massively parallel environment, or (2) obstruct the use of individual threads on top of single-instruction-multiple-data (SIMD) cores, or (3) block during access, thus prohibiting multi-queue setups. With these issues in mind, we present the Broker Queue, a highly efficient, fully linearizable FIFO queue for fine-granular parallel work distribution on the GPU. We evaluate its performance and usability on modern GPU models against a wide range of existing algorithms. The Broker Queue is up to three orders of magnitude faster than nonblocking queues and can even outperform significantly simpler techniques that lack desired properties for fine-granular work distribution. Bernhard Kerbl, Michael Kenzel, Joerg H. Mueller, Dieter Schmalstieg, Markus Steinberger |
ICS | 3 |
| 2018 | A scalable queue for work distribution on GPUsabstractHarnessing the power of massively parallel devices like the graphics processing unit (GPU) is difficult for algorithms that show dynamic or inhomogeneous workloads. To achieve high performance, such advanced algorithms require scalable, concurrent queues to collect and distribute work. We present a new concurrent work queue, the Broker Queue, a highly efficient, linearizable queue for fine-granular work distribution on the GPU. We evaluate its usability and benefits in contrast to existing queuing algorithms. Our queue is up to one order of magnitude faster than non-blocking queues, and outperforms simpler queue designs that are unfit for fine-granular work distribution. Bernhard Kerbl, Joerg H. Mueller, Michael Kenzel, Dieter Schmalstieg, Markus Steinberger |
PPoPP | 2 |
| 2018 | Human upper-body inverse kinematics for increased embodiment in consumer-grade virtual realityabstractHaving a virtual body can increase embodiment in virtual reality (VR) applications. However, comsumer-grade VR falls short of delivering sufficient sensory information for full-body motion capture. Consequently, most current VR applications do not even show arms, although they are often in the field of view. We address this shortcoming with a novel human upper-body inverse kinematics algorithm specifically targeted at tracking from head and hand sensors only. We present heuristics for elbow positioning depending on the shoulder-to-hand distance and for avoiding reaching unnatural joint limits. Our results show that our method increases the accuracy compared to general inverse kinematics applied to human arms with the same tracking input. In a user study, participants preferred our method over displaying disembodied hands without arms, but also over a more expensive motion capture system. In particular, our study shows that virtual arms animated with our inverse kinematics system can be used for applications involving heavy arm movement. We demonstrate that our method can not only be used to increase embodiment, but can also support interaction involving arms or shoulders, such as holding up a shield. Mathias Parger, Joerg H. Mueller, Dieter Schmalstieg, Markus Steinberger |
VRST | 2 |
| 2018 | Shading atlas streamingabstractStreaming high quality rendering for virtual reality applications requires minimizing perceived latency. We introduce Shading Atlas Streaming (SAS), a novel object-space rendering framework suitable for streaming virtual reality content. SAS decouples server-side shading from client-side rendering, allowing the client to perform framerate upsampling and latency compensation autonomously for short periods of time. The shading information created by the server in object space is temporally coherent and can be efficiently compressed using standard MPEG encoding. Our results show that SAS compares favorably to previous methods for remote image-based rendering in terms of image quality and network bandwidth efficiency. SAS allows highly efficient parallel allocation in a virtualized-texture-like memory hierarchy, solving a common efficiency problem of object-space shading. With SAS, untethered virtual reality headsets can benefit from high quality rendering without paying in increased latency. Joerg H. Mueller, Philip Voglreiter, Mark Dokter, Thomas Neff, Mina Makar, Markus Steinberger, Dieter Schmalstieg |
ACM Trans. Graph. | 1 |
| 2016 | PanoVC: Pervasive telepresence using mobile phonesabstractWe are presenting PanoVC - a mobile telepresence system based on continuously updated panoramic images. We are showing that the experience of telepresence, i.e. the sense of "being there together" at a distant location can be achieved with standard state-of-the-art mobile phones. Because mobile phones are always on hand users can share their environments with others in a pervasive way. Our approach is opening up the pathway for applications in a variety of domains such as the exploration of remote environments or novel forms of videoconferencing. We present implementation details, technical evaluation results, and the findings of a user study of an indoor-outdoor environments sharing task as proof of concept. Joerg H. Mueller, Tobias Langlotz, Holger Regenbrecht |
PerCom | 1 |
| 2014 | Parallel generation of architecture on the GPUabstractAbstract In this paper, we present a novel approach for the parallel evaluation of procedural shape grammars on the graphics processing unit (GPU). Unlike previous approaches that are either limited in the kind of shapes they allow, the amount of parallelism they can take advantage of, or both, our method supports state of the art procedural modeling including stochasticity and context‐sensitivity. To increase parallelism, we explicitly express independence in the grammar, reduce inter‐rule dependencies required for context‐sensitive evaluation, and introduce intra‐rule parallelism. Our rule scheduling scheme avoids unnecessary back and forth between CPU and GPU and reduces round trips to slow global memory by dynamically grouping rules in on‐chip shared memory. Our GPU shape grammar implementation is multiple orders of magnitude faster than the standard in CPU‐based rule evaluation, while offering equal expressive power. In comparison to the state of the art in GPU shape grammar derivation, our approach is nearly 50 times faster, while adding support for geometric context‐sensitivity. Markus Steinberger, Michael Kenzel, Bernhard Kainz, Joerg H. Mueller, Peter Wonka, Dieter Schmalstieg |
Comput. Graph. Forum | 4 |