Sebastian Knorr

dblp:98/397 · DBLP profile ↗
← Back
25ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0001-9745-8605ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Real-Time View Synthesis with Multiplane Image Network using Multimodal Supervision
abstract
Recent advances in view synthesis from a single image have increased visual quality of the newly synthesized viewpoints significantly. However, the high computational cost of state-of-the-art methods remains a critical bottleneck, limiting their adoption in real-time applications such as immersive telepresence. To address this limitation, we present a multiplane image (MPI) network that achieves real-time view synthesis. Unlike existing approaches that often rely on a separate depth estimation network to guide the network for estimating MPI parameters, our framework directly predicts the MPI parameters from a single RGB image. To guide the network for estimating correct parameters, we introduce a training strategy that leverages joint supervision from both view synthesis and depth estimation losses to ensure visual fidelity. During inference, our method exclusively utilizes the optimized view synthesis branch, while the depth decoder is only used for training. Our end-to-end approach renders views from a single input image in real-time. Extensive experiments validate that our method provides a compelling rendering speed with visual quality in par with state-of-the-art methods, highlighting its suitability for live, interactive applications.
Manu Gond, Mohammadreza Shamshirgarha, Emin Zerman, Sebastian Knorr, Mårten Sjöström
MMSP4
2025 A Visual Quality of Experience Toolkit for Realistic Immersive Telepresence Applications
abstract
Immersive imaging applications gained a lot of traction in the last decade with the advances in capture, processing, compression, transmission, and display technologies. Recent works on spherical light fields and new view synthesis bring new challenges with respect to fast rendering of new views and a smooth quality of experience (QoE). The real-time rendering capabilities enable realistic immersive telepresence applications, which extend the traditional telepresence systems with improved visual realism and possible addition of depth cues. The recent efforts on the spherical light fields and new view synthesis approaches show that a combined light field and spherical data visualization will be beneficial for the scientific community. In this paper, we provide a WebGL- and WebXR-based visual QoE toolkit which can be used both on traditional displays and extended reality headsets, presenting various visual modalities. To validate the usefulness of the proposed toolkit, we conducted a pilot test on our publicly available spherical light field database. This toolkit can be used for easier and faster quality assessment and can support the scientific community for subjective visual QoE studies focusing on telecommunication, telepresence, and augmented telepresence applications.
Manu Gond, Emin Zerman, Mohammadreza Shamshirgarha, Sebastian Knorr, Mårten Sjöström
QoMEX4
2025 3D SMoE Splatting for Edge-aware Realtime Radiance Field Rendering
abstract
Steered Mixtures-of-Experts (SMoE) is an existing regression framework that has previously been applied for modeling and compression of 2D images and higher-dimensional imagery, including compression of light fields and light-field video. SMoE models are sparse, edge-aware representations that allow rendering of imagery with few Gaussians with excellent quality. In this paper a novel, edge-aware "3D SMoE Splatting" (3DSMoES) framework for 3D rendering is introduced, adopted to fit into the existing "3D Gaussian Splatting" (3DGS) CUDA optimization pipeline. Here, SMoE regression serves as a "plug-and-play" solution that replaces the established 3DGS regression as a novel workhorse. 3DSMoES achieves significant visual quality gains with drastically fewer Gaussian kernels compared to 3DGS. We observe up to approximately 4dB improvement in PSNR on individual scenes with kernel reductions between 20 to 50 percent. The sparse models are significantly faster to train and allow up to 30-50 percent improved rendering speeds.
Yi-Hsin Li, Thomas Sikora, Sebastian Knorr, Mårten Sjöström
SIGGRAPH Asia3
2025 Adaptive Segmentation-Based Initialization for Steered Mixture of Experts Image Regression
abstract
Kernel image regression methods have demonstrated excellent efficiency in various image processing tasks, including image and light-field compression, Gaussian Splatting, denoising and super-resolution. The estimation of parameters for these methods commonly employs gradient descent iterative optimization, which poses a significant computational burden for many applications. In this paper, we introduce a novel adaptive segmentation-based initialization method targeted for optimizing Steered-Mixture-of Experts (SMoE) gating networks and RadialBasis-Function (RBF) networks with steering kernels. The novel initialization method allocates kernels into pre-calculated image segments. The optimal number of kernels, kernel positions, and steering parameters are derived per segment in an iterative optimization and kernel sparsification procedure. The kernel information from local segments is then transferred into a global initialization, ready for use in iterative optimization of SMoE, RBF, and related kernel image regression methods. Results demonstrate significant improvements in both objective and subjective quality compared to regular grid, K-Means, deeplearning-based, and previous segmentation-based initialization methods. The proposed initialization method reduces kernel usage by 70% compared to other initialization methods while maintaining the same reconstruction quality. Furthermore, by generating initial parameters closer to optimized results, convergence time is reduced, achieving overall runtime savings of up to 50% compared to prior methods. Additionally, the method supports parallel computation, with initialization time halved when using four GPUs compared to one
Yi-Hsin Li, Sebastian Knorr, Mårten Sjöström, Thomas Sikora
IEEE Trans. Multim.2
2024 Adaptive Model Predictive Control-Driven Approach for Visual Detection of Micro- UAVs
abstract
Unmanned Aerial Vehicles (UAVs) provide a good base platform for dynamic vision-based target detection. The detection process can be enhanced by utilizing multiple UAVs. In such scenarios, precise trajectory planning is essential to achieve objectives while avoiding collisions with obstacles. Furthermore, in most cases, Air-to-Air (A2A) detection of micro-UAVs in large-scale environments is challenging due to their dynamic movements and other complex parameters, such as poor light, motion blur, and occlusion. To solve these challenges, developing an algorithm that can adapt the control strategy based on changing system and environmental parameters is essential. This paper proposes an Adaptive Model Predictive Control (AMPC)-driven hybrid GANYOLOX framework to mitigate these navigation and detection challenges. The proposed AMPC allows model updates during multi-UAV operations, which enables estimation techniques to predict changes in the system model and helps to compute the target UAV position. Alternatively, the framework constructs a hybrid GAN-YOLOX model to overcome various A2A detection challenges. Extensive comparative analysis of the hybrid GANYOLOX framework achieved 94% detection accuracy and out-performed existing DL frameworks by 8%.
Gunasekaran Raja, Selvam Essaky, Deepak Suresh Rajendran, Sai Ganesh Senthivel, Sebastian Knorr, Kapal Dev
ICC5
2023 Vanishing Point Aided Hash-Frequency Encoding for Neural Radiance Fields (NeRF) from Sparse 360°Input
abstract
Neural Radiance Fields (NeRF) enable novel view synthesis of 3D scenes when trained with a set of 2D images. One of the key components of NeRF is the input encoding, i.e. mapping the coordinates to higher dimensions to learn high-frequency details, which has been proven to increase the quality. Among various input mappings, hash encoding is gaining increasing attention for its efficiency. However, its performance on sparse inputs is limited. To address this limitation, we propose a new input encoding scheme that improves hash-based NeRF for sparse inputs, i.e. few and distant cameras, specifically for 360° view synthesis. In this paper, we combine frequency encoding and hash encoding and show that this combination can increase dramatically the quality of hash-based NeRF for sparse inputs. Additionally, we explore scene geometry by estimating vanishing points in omnidirectional images (ODI) of indoor and city scenes in order to align frequency encoding with scene structures. We demonstrate that our vanishing point-aided scene alignment further improves deterministic and non-deterministic encodings on image regression and NeRF tasks where sharper textures and more accurate geometry of scene structures can be reconstructed.
Thomas Maugey, Sebastian Knorr, Christine Guillemot
ISMAR3
2023 Segmentation-based Initialization for Steered Mixture of Experts
abstract
The Steered-Mixture-of-Experts (SMoE) model is an edge-aware kernel representation that has successfully been explored for the compression of images, video, and higher-dimensional data such as light fields. The present work aims to leverage the potential for enhanced compression gains through efficient kernel reduction. We propose a fast segmentation-based strategy to identify a sufficient number of kernels for representing an image and giving initial kernel parametrization. The strategy implies both reduced memory footprint and reduced computational complexity for the subsequent parameter optimization, resulting in an overall faster processing time. Fewer kernels, when combined with the inherent sparsity of the SMoEs, further enhance the overall compression performance. Empirical evaluations demonstrate a gain of 0.3-1.0 dB in PSNR for a constant number of kernels, and the use of 23 % less kernels and 25 % less time for constant PSNR. The results highlight the feasibility and practicality of the approach, positioning it as a valuable solution for various image-related applications, including image compression.
Yi-Hsin Li, Mårten Sjöström, Sebastian Knorr, Thomas Sikora
VCIP3
2023 HEADSET: Human Emotion Awareness under Partial Occlusions Multimodal DataSET
abstract
The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended Reality (XR) applications, this volumetric data has proven to be an essential technology for future XR elaboration. In this work, we present a new multimodal database to help advance the development of immersive technologies. Our proposed database provides ethically compliant and diverse volumetric data, in particular 27 participants displaying posed facial expressions and subtle body movements while speaking, plus 11 participants wearing head-mounted displays (HMDs). The recording system consists of a volumetric capture (VoCap) studio, including 31 synchronized modules with 62 RGB cameras and 31 depth cameras. In addition to textured meshes, point clouds, and multi-view RGB-D data, we use one Lytro Illum camera for providing light field (LF) data simultaneously. Finally, we also provide an evaluation of our dataset employment with regard to the tasks of facial expression classification, HMDs removal, and point cloud reconstruction. The dataset can be helpful in the evaluation and performance testing of various XR algorithms, including but not limited to facial expression recognition and reconstruction, facial reenactment, and volumetric video. HEADSET and its all associated raw data and license agreement will be publicly available for research purposes.
Fatemeh Ghorbani Lohesara, Davi Rabbouni Freitas, Christine Guillemot, Karen Egiazarian, Sebastian Knorr
IEEE Trans. Vis. Comput. Graph.5
2022 Omni-NeRF: Neural Radiance Field from 360° Image Captures
abstract
This paper tackles the problem of novel view synthesis (NVS) from 360° images with imperfect camera poses or intrinsic parameters. We propose a novel end-to-end framework for training Neural Radiance Field (NeRF) models given only 360° RGB images and their rough poses, which we refer to as Omni-NeRF. We extend the pinhole camera model of NeRF to a more general camera model that better fits omni-directional fish-eye lenses. The approach jointly learns the scene geometry and optimizes the camera parameters without knowing the fisheye projection.
Thomas Maugey, Sebastian Knorr, Christine Guillemot
ICME3
2021 Deep Color Mismatch Correction In Stereoscopic 3d Images
abstract
Color mismatch in stereoscopic 3D (S3D) images can create visual discomfort and affect the performance of S3D image processing algorithms, e.g., for depth estimation. In this paper, we propose a new deep learning-based solution for the problem of color mismatch correction. The proposed solution consists of a multi-task convolutional neural network, where color correction is the primary task and correspondence estimation is the secondary task. For the training and evaluation of the proposed network, a new S3D image dataset with color mismatch was created. Based on this dataset, experiments were conducted showing the effectiveness of our solution.
Simone Croci, Cagri Ozcinar, Emin Zerman, Roman Dudek, Sebastian Knorr, Aljoscha Smolic
ICIP5
2021 Collision-free Path Planning for UAVs using Efficient Artificial Potential Field Algorithm
abstract
Unmanned Aerial Vehicles (UAVs), a new emerging form of Internet of Things (IoT), is a promising technology to be widely used in both civil and military applications. On the fly, the UAVs need to find an efficient and safe path by avoiding both static and dynamic obstacles to carry out any mission successfully. The Artificial Potential Field (APF) algorithm is one of the conventional catalysts in UAV path planning. However, APF-aided UAVs can be easily trapped into a local minimum solution before reaching the destination. Therefore, this paper proposes an efficient APF algorithm for Collision-free Path Planning (eAPF-CPP) in UAVs. In eAPF-CPP, the attractive and repulsive potentials evaluate the quadratic distance to the destination and the obstacle respectively. The evaluation aids the UAV to select the optimal path in navigation. The eAPF-CPP mechanism is simulated in the Software-In-The-Loop (SITL) setup, and the experimental results show that the eAPF-CPP mechanism utilizes an average of 24.4 seconds to track a safe path and has a lower collision rate of 8.56% compared with Artifical Potential Field Approach (APFA).
Praveen Kumar Selvam, Gunasekaran Raja, Vasantharaj Rajagopal, Kapal Dev, Sebastian Knorr
VTC Spring5
2019 Study on the Perception of Sharpness Mismatch in Stereoscopic Video
abstract
In this paper, we study an artifact of stereoscopic 3D (S3D) video called sharpness mismatch (SM), that occurs when one view is more blurred than the other. SM beyond a certain level can create visual discomfort, and consequently degrade the quality of experience. Therefore, it is important to measure the just noticeable sharpness mismatch (JNSM), i.e., the minimal level of SM that is perceived by the human visual system and creates discomfort. The knowledge of the JNSM can be used in the evaluation of the quality of S3D video, and more in general when processing S3D video, like in asymmetric compression. In this paper, we focus in particular on the detection of SM. For this goal, we organized a psychophysical experiment with 23 subjects and a crosstalk-free stereoscopic display in order to gather psychophysical data necessary for the development of a SM detection method. Based on the gathered experiment data, we propose a new SM detection method. The evaluation of this method shows that its performance is close but not better than that of the state-of-the-art methods. Therefore, our goal in the near future is to improve the proposed method.
Simone Croci, Sebastian Knorr, Aljoscha Smolic
QoMEX2
2019 Robust global and local color matching in stereoscopic omnidirectional content
Roman Dudek, Simone Croci, Aljoscha Smolic, Sebastian Knorr
Signal Process. Image Commun.4
2018 Faoladh: A Case Study in Cinematic VR Storytelling and Production
Declan Dowling, Colm O. Fearghail, Aljoscha Smolic, Sebastian Knorr
ICIDS4
2018 Director's Cut - Analysis of Aspects of Interactive Storytelling for VR Films
Colm O. Fearghail, Cagri Ozcinar, Sebastian Knorr, Aljoscha Smolic
ICIDS3
2018 Sharpness Mismatch Detection in Stereoscopic Content with 360-Degree Capability
abstract
This paper presents a novel sharpness mismatch detection method for stereoscopic images based on the comparison of edge width histograms of the left and right view. The new method is evaluated on the LIVE 3D Phase II and Ningbo 3D Phase I datasets and compared with two state-of-the-art methods. Experimental results show that the new method highly correlates with user scores of subjective tests and that it outperforms the current state-of-the-art. We then extend the method to stereoscopic omnidirectional images by partitioning the images into patches using a spherical Voronoi diagram. Furthermore, we integrate visual attention data into the detection process in order to weight sharpness mismatch according to the likelihood of its appearance in the viewport of the end-user's virtual reality device. For obtaining visual attention data, we performed a subjective experiment with 17 test subjects and 96 stereoscopic omnidirectional images. The entire dataset including the viewport trajectory data and resulting visual attention maps are publicly available with this paper.
Simone Croci, Sebastian Knorr, Aljoscha Smolic
ICIP2
2017 Estimation of Optimal Encoding Ladders for Tiled 360° VR Video in Adaptive Streaming Systems
abstract
Given the significant industrial growth of demand for virtual reality (VR), 360ovideo streaming is one of the most important VR applications that require cost-optimal solutions to achieve widespread proliferation of VR technology. Because of its inherent variability of data-intensive content types and its tiled-based encoding and streaming, 360ovideo requires new encoding ladders in adaptive streaming systems to achieve cost-optimal and immersive streaming experiences. In this context, this paper targets both the provider's and client's perspectives and introduces a new content-aware encoding ladder estimation method for tiled 360oVR video in adaptive streaming systems. The proposed method first categories a given 360ovideo using its features of encoding complexity and estimates the visual distortion and resource cost of each bitrate level based on the proposed distortion and resource cost models. An optimal encoding ladder is then formed using the proposed integer linear programming (ILP) algorithm by considering practical constraints. Experimental results of the proposed method are compared with the recommended encoding ladders of professional streaming service providers. Evaluations show that the proposed encoding ladders deliver better results compared to the recommended encoding ladders in terms of objective quality for 360ovideo, providing optimal encoding ladders using a set of service provider's constraint parameters.
Cagri Ozcinar, Ana De Abreu, Sebastian Knorr, Aljoscha Smolic
ISM3
2013 Adaptive Image Warping for Hole Prevention in 3D View Synthesis
abstract
Increasing popularity of 3D videos calls for new methods to ease the conversion process of existing monocular video to stereoscopic or multi-view video. A popular way to convert video is given by depth image-based rendering methods, in which a depth map that is associated with an image frame is used to generate a virtual view. Because of the lack of knowledge about the 3D structure of a scene and its corresponding texture, the conversion of 2D video, inevitably, however, leads to holes in the resulting 3D image as a result of newly-exposed areas. The conversion process can be altered such that no holes become visible in the resulting 3D view by superimposing a regular grid over the depth map and deforming it. In this paper, an adaptive image warping approach as an improvement to the regular approach is proposed. The new algorithm exploits the smoothness of a typical depth map to reduce the complexity of the underlying optimization problem that is necessary to find the deformation, which is required to prevent holes. This is achieved by splitting a depth map into blocks of homogeneous depth using quadtrees and running the optimization on the resulting adaptive grid. The results show that this approach leads to a considerable reduction of the computational complexity while maintaining the visual quality of the synthesized views.
Nils Plath, Sebastian Knorr, Lutz Goldmann, Thomas Sikora
IEEE Trans. Image Process.2
2011 Three-Dimensional Video Postproduction and Processing
abstract
This paper gives an overview of the state-of-the-art in 3-D video postproduction and processing as well as an outlook to remaining challenges and opportunities. First, fundamentals of stereography are outlined that set the rules for proper 3-D content creation. Manipulation of the depth composition of a given stereo pair via view synthesis is identified as the key functionality in this context. Basic algorithms are described to adapt and correct fundamental stereo properties such as geometric distortions, color alignment, and stereo geometry. Then, depth image-based rendering is explained as the widely applied solution for view synthesis in 3-D content creation today. Recent improvements of depth estimation already provide very good results. However, in most cases, still interactive workflows dominate. Warping-based methods may become an alternative for some applications in the future, which do not rely on dense and accurate depth estimation. Finally, 2-D to 3-D conversion is covered, which is an important special area for reuse of existing legacy 2-D content in 3-D. Here various advanced algorithms are combined in interactive workflows.
Aljoscha Smolic, Peter Kauff, Sebastian Knorr, Alexander Sorkine-Hornung, Matthias Kunter, Marcus Müller 0001, Manuel Lang
Proc. IEEE3
2008 Camera motion-constraint video codec selection
abstract
In recent years advanced video codecs have been developed, such as standardized in MPEG-4. The latest video codec H.264/AVC provides compression performance superior to previous standards, but is based on the same basic motion-compensated-DCT architecture. However, for certain types of video, it has been shown that it is possible to outperform the H.264/AVC using an object-based video codec. Towards a general-purpose object-based video coding system we present an automated approach to separate a video sequences into sub-sequences regarding its camera motion type. Then, the sub-sequences are coded either with an object-based codec or the common H.264/AVC. Applying different video codecs for different kinds of camera motion, we achieve a higher overall coding gain for the video sequence. In first experimental evaluations, we demonstrate the excellence performance of this approach on two test sequences.
Andreas Krutz, Sebastian Knorr, Matthias Kunter, Thomas Sikora
MMSP2
2008 Stereoscopic 3D from 2D video with super-resolution capability
Sebastian Knorr, Matthias Kunter, Thomas Sikora
Signal Process. Image Commun.1
2007 An Image-Based Rendering (IBR) Approach for Realistic Stereo View Synthesis of TV Broadcast Based on Structure from Motion
abstract
In the past years, the 3D display technology has become a booming branch of research with fast technical progress. Hence, the 3D conversion of already existing 2D video material increases more and more in popularity. In this paper, a new approach for realistic stereo view synthesis (RSVS) of existing 2D video material is presented. The intention of our work is not a real-time conversion of existing video material with a deduction in stereo perception, but rather a more realistic off-line conversion with high accuracy. Our approach is based on structure from motion techniques and uses image-based rendering to reconstruct the desired stereo views for each video frame. The algorithm is tested on several TV broadcast videos, as well as on sequences captured with a single handheld camera. Finally, some simulation results will show the remarkable performance of this approach.
Sebastian Knorr, Thomas Sikora
ICIP (6)1
2007 Towards 3-D scene reconstruction from broadcast video
Evren Imre, Sebastian Knorr, Burak Özkalayci, Ugur Topay, A. Aydin Alatan, Thomas Sikora
Signal Process. Image Commun.2
2006 Prioritized Sequential 3D Reconstruction in Video Sequences with Multiple Motions
abstract
In this study, an algorithm is proposed to solve the multi-frame structure from motion (MFSfM) problem for monocular video sequences in dynamic scenes. The algorithm uses the epipolar criterion to segment the features belonging to independently moving objects. Once the features are segmented, corresponding objects are reconstructed individually by using a sequential algorithm, which is also capable of prioritizing the frame pairs with respect to their reliability and information content, thus achieving a fast and accurate reconstruction through efficient processing of the available data. A tracker is utilized to increase the baseline distance between views and to improve the F-matrix estimation, which is beneficial to both the segmentation and the 3D structure estimation processes. The experimental results demonstrate that our approach has the potential to effectively deal with the multi-body MFSfM problem in a generic video sequence.
Evren Imre, Sebastian Knorr, A. Aydin Alatan, Thomas Sikora
ICIP2
2004 A gradient based approach for stereoscopic error concealment
abstract
Error concealment is an important field of research in image processing. Many methods have been applied to conceal block losses in monocular images. We present a concealment strategy for block loss in stereoscopic image pairs. Unlike the error concealment techniques used for monocular images, the information of the associated image is utilized, i.e., by means of a projective transformation model, pixel values from the associated stereo image are warped to their corresponding positions in the lost block. The stereoscopic depth perception is much less affected in our approach than using monoscopic error concealment techniques.
Matthias Kunter, Sebastian Knorr, Carsten Clemens, Thomas Sikora
ICIP2