Sebastian Bosse

dblp:35/9463 · DBLP profile ↗
← Back
45ranked-venue papers
13as first author
16since 2021 · last 2026
0000-0002-8104-3904ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 12 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A-PLR: An Activity-Aware Assistant for Accurate and Usable Passive Leg Raise Tests in the Smart ICU
abstract
The Passive Leg Raise (PLR) test serves as an essential diagnostic instrument for evaluating fluid responsiveness within intensive care units (ICUs). Nevertheless, its execution and interpretation are susceptible to errors resulting from procedural complexity and clinician workload. We introduce A-PLR, an activity-aware assistant integrated into a smart ICU setting to mitigate these challenges. The system employs computer vision and spatial activity recognition to autonomously identify PLR phases, ensure procedural compliance, and provide real-time decision support via an intuitive bedside interface. This paper reports a mixed-method evaluation of A-PLR, comprising both quantitative performance metrics and qualitative insights from clinical professionals (N = 30). Within the controlled experimental setting, the results indicate improved usability and procedural support, as well as indications of enhanced diagnostic assistance and reduced cognitive load when clinicians were supported by A-PLR. In addition, user feedback highlights the potential of activity-aware, AI-driven systems to support decision-making and streamline workflow integration in (simulated) critical care environments.
Paul Chojecki, David Przewozny, Felix Tirschmann, Falko Schmid, Utz Jerzembeck, Detlef Runde, Kirsten Brukamp, Hug Aubin, Sebastian Bosse
IUI9
2025 Batch-Aware Active Learning for Object Detection
abstract
We propose a Batch-Aware Active Learning (BAAL) framework to optimize the training of object detection models, reducing annotation costs while maintaining strong model performance. The framework adapts different uncertainty sampling strategies to the specific challenges of object detection, including multi-class labelling and spatial localization. By combining uncertainty with diversity, leveraging feature representations and clustering, our method ensures diverse and informative batch selection. The non-invasive, plug-and-play design supports seamless integration with any object detection model without architectural modifications. Evaluations on COCO and Pascal VOC datasets with SSD, Faster R-CNN, YOLOv8, and RetinaNet demonstrate that our approach is not only efficient and robust but also comparable to, and in some cases exceeds, current state-of-the-art solutions.
Mykyta Kovalenko, Peter Eisert, Anna Hilsmann, Sebastian Bosse
ICIP4
2025 Fine-Tuning the Prithvi Foundation Model for Crop-Type and Health Mapping in Smallholder Farms
abstract
Artificial intelligence is rapidly reshaping how we monitor and manage agriculture from space. In this study, we fine-tune and evaluate a state-of-the-art vision foundation model, Prithvi-EO-2.0, for crop type classification and crop health monitoring across smallholder farms in Telangana, India. Leveraging temporally stacked satellite imagery, we assess the effectiveness of different fine-tuning strategies and address challenges related to fragmented landholdings and limited ground truth data. Our results demonstrate that fully fine-tuning the model significantly improves crop classification accuracy but requires careful handling of class imbalance to prevent overfitting. We discuss limitations arising from spatial resolution constraints and propose future directions using higher-resolution alternatives, emphasizing the potential of adapting large-scale pretrained models to deliver scalable and practical precision agriculture solutions in data-limited, heterogeneous smallholder systems.
Karam Tomotaki-Dawoud, Raghu Chaliganti, Shreya Champakbhai Chauhan, Pia Wolffram, Ozan Türkes, Sebastian Bosse
KES6
2025 AI-based Denoising and Interpolation of Magnetic UXO Data
abstract
Magnetic surveys are a key tool in detecting buried objects such as unexploded ordnance (UXO), where dense magnetic maps must be reconstructed from sparsely sampled gradiometer data. We present a deep learning-based approach that outperforms classical interpolation methods in both accuracy and speed. Trained on synthetic magnetic fields simulating realistic UXO signatures and measurement noise, our modified U-Net with ResNet-34 encoding reconstructs high-resolution magnetic maps from sparse inputs. Compared to state-of-the-art gridding methods, our model achieves 3–5% higher reconstruction accuracy on average while operating up to 80× faster than SOTA algorithms, enabling more efficient and interpretable UXO detection in real-world survey conditions.
Mykyta Kovalenko, David Przewozny, Paul Chojecki, Anna Hilsmann, Peter Eisert, Sebastian Bosse
SMC6
2025 Too Close for Comfort? Investigating Virtual Professor Distance and Student Learning in VR
abstract
This study investigates how virtual professor proximity influences student comprehension and attention in an immersive learning environment. 27 participants experienced three lectures in virtual reality (VR) under varying spatial conditions: close distance (personal space intrusion), optimal distance (user defined preferred proximity), and far distance (outside social interaction proximity). Attention was tracked via eye movement, and comprehension was evaluated through multiple-choice tests. Results show that close distance reduced comprehension despite increased gaze toward the professor, likely due to discomfort and attentional overload. Far distance increased visual distraction, while the optimal distance supported comfort, engagement, and performance. Our findings offer actionable design guidelines for avatar placements in virtual learning systems to optimize user experience and cognitive effectiveness.
Mustafa-Tevfik Lafci, Birgit Nierula, Dilara Damar, Sebastian Bosse
SMC4
2024 Multi-View Gesture Recognition in Conflict Situations
abstract
Non-verbal cues play a crucial role in social interactions and can influence conflict dynamics. For law enforcement officers, recognizing these cues is essential for effective deescalation, yet traditional training may not fully address their complexity. This paper focuses on body gestures and presents an automated system for recognizing specific body gestures relevant to social conflict situations, aiming to foster awareness for unconsciously performed body gestures and thereby enabling the training of de-escalation strategies.
Karam Tomotaki-Dawoud, Birgit Nierula, Farelle Toumaleu Siewe, Daniel Johannes Meyer, Andreas Bock, Marianne Heinze, Daniela Knuth, Denis Martin, Julia Schander, Anna Hilsmann, Peter Eisert, Sebastian Bosse
ISM13
2023 But That's Not Why: Inference Adjustment by Interactive Prototype Revision
Michael Gerstenberger, Thomas Wiegand 0001, Peter Eisert, Sebastian Bosse
CIARP4
2023 A Differentiable Gaussian Prototype Layer for Explainable Fruit Segmentation
abstract
We introduce a Gaussian Prototype Layer for gradient-based prototype learning and demonstrate two novel network architectures for explainable segmentation one of which relies on region proposals. Both models are evaluated on agricultural datasets. While Gaussian Mixture Models (GMMs) have been used to model latent distributions of neural networks before, they are typically fitted using the EM algorithm. Instead, the proposed prototype layer relies on gradient-based optimization and hence allows for end-to-end training. This facilitates development and allows to use the full potential of a trainable deep feature extractor. We show that it can be used as a novel building block for explainable neural networks. We employ our Gaussian Prototype Layer in (1) a model where prototypes are detected in the latent grid and (2) a model inspired by Fast-RCNN with SLIC superpixels as region proposals. The earlier achieves a similar performance as compared to the state-of-the art while the latter has the benefit of a more precise prototype localization that comes at the cost of slightly lower accuracies. By introducing a gradient-based GMM layer we combine the benefits of end-to-end training with the simplicity and theoretical foundation of GMMs which will allow to adapt existing semi-supervised learning strategies for prototypical part models in future.
Michael Gerstenberger, Steffen Maaß, Peter Eisert, Sebastian Bosse
ICIP4
2023 Pre-Training with Fractal Images Facilitates Learned Image Quality Estimation
abstract
Today’s image quality estimation is widely dominated by learning-based approaches. The availability of annotated, i.e. rated, images is often a bottleneck in training data-driven visual quality models and hinders their generalization power. This paper proposed a novel pre-training scheme for learning-based quality estimation that does not rely on human-annotated datasets, but leverages synthetic fractal images. These images can be synthesized inexhaustibly and are inherently labeled during generation. We evaluate the pre-training strategy on a popular neural network-based quality model and show that the training effort can be reduced significantly, resulting in better final accuracy and faster convergence speed.
Malte Silbernagel, Thomas Wiegand 0001, Peter Eisert, Sebastian Bosse
ICIP4
2023 An evaluation of hand interaction metaphors for immersive environments
abstract
User interfaces are essential tools for commanding machines on classical 2D displays and in immersive virtual reality (VR) environments. However, the limitless design options of 3D user interfaces and naturalistic interactions present challenges. While user experience for interaction in 2D scenarios has been studied for decades and has led to standards and well-established best practices, little is known about the rational design and the resulting user experience of user interfaces in VR. In this paper, we evaluate two user interaction methods: menu-based interaction and gesture interaction. We implemented an exemplary VR application for watching video content to explore 2D and 3D interaction concepts in a VR environment and collected user ratings for our interaction methods. Our results indicate that the menu-based hand interactions are superior to gesture-based ones in terms of task effectiveness and user satisfaction on average. However, a detailed analysis shows that for a significant subset of interaction elements gesture-based interaction is on par or superior to menu-based interaction while maintaining the subjective notion of naturalness. This provides insights into how to design efficient and natural gesture-based user interaction.
Mustafa-Tevfik Lafci, Robert Strzebkowski, Paul Chojecki, Sebastian Bosse
QoMEX4
2022 NDNetGaming - development of a no-reference deep CNN for gaming video quality prediction
abstract
Abstract Gaming video streaming services are growing rapidly due to new services such as passive video streaming of gaming content, e.g. Twitch.tv, as well as cloud gaming, e.g. Nvidia GeForce NOW and Google Stadia. In contrast to traditional video content, gaming content has special characteristics such as extremely high and special motion patterns, synthetic content and repetitive content, which poses new opportunities for the design of machine learning-based models to outperform the state-of-the-art video and image quality approaches for this special computer generated content. In this paper, we train a Convolutional Neural Network (CNN) based on an objective quality model, VMAF, as ground truth and fine-tuned it based on subjective image quality ratings. In addition, we propose a new temporal pooling method to predict gaming video quality based on frame-level predictions. Finally, the paper also describes how an appropriate CNN architecture can be chosen and how well the model performs on different contents. Our result shows that among four popular network architectures that we investigated, DenseNet performs best for image quality assessment based on the training dataset. By training the last 57 convolutional layers of DenseNet based on VMAF values, we obtained a high performance model to predict VMAF of distorted frames of video games with a Spearman’s Rank correlation (SRCC) of 0.945 and Root Mean Score Error (RMSE) of 7.07 on the image level, while achieving a higher performance on the video level leading to a SRCC of 0.967 and RMSE of 5.47 for the KUGVD dataset. Furthermore, we fine-tuned the model based on subjective quality ratings of images from gaming content which resulted in a SRCC of 0.93 and RMSE of 0.46 using one-hold-out cross validation. Finally, on the video level, using the proposed pooling method, the model achieves a very good performance indicated by a SRCC of 0.968 and RMSE of 0.30 for the used gaming video dataset.
Markus Utke, Saman Zad Tootaghaj, Steven Schmidt 0001, Sebastian Bosse, Sebastian Möller 0001
Multim. Tools Appl.4
2021 Curiously Effective Features For Image Quality Prediction
abstract
The performance of visual quality prediction models is commonly assumed to be closely tied to their ability to capture perceptually relevant image aspects. Models are thus either based on sophisticated feature extractors carefully designed from extensive domain knowledge or optimized through feature learning. In contrast to this, we find feature extractors constructed from random noise to be sufficient to learn a linear regression model whose quality predictions reach high correlations with human visual quality ratings, on par with a model with learned features. We analyze this curious result and show that besides the quality of feature extractors also their quantity plays a crucial role - with top performances only being achieved in highly overparameterized models.
Sören Becker 0001, Thomas Wiegand 0001, Sebastian Bosse
ICIP3
2021 EEG-Based Analysis of the Impact of Familiarity in the Perception of Deepfake Videos
abstract
We investigate the brain’s subliminal response to fake portrait videos using electroencephalography, with a special emphasis on the viewer’s familiarity with the depicted individuals.Deepfake videos are increasingly becoming popular but, while they are entertaining, they can also pose a threat to society. These face-swapped videos merge physiognomy and behaviour of two different individuals, both strong cues used for recognizing a person. We show that this mismatch elicits different brain responses depending on the viewer’s familiarity with the merged individuals.Using EEG, we classify perceptual differences of familiar and unfamiliar people versus their face-swapped counterparts. Our results show that it is possible to discriminate fake videos from genuine ones when at least one face-swapped actor is known to the observer. Furthermore, we indicate a correlation of classification accuracy with level of personal engagement between participant and actor, as well as with the participant’s familiarity with the used dataset.
Jan-Philipp Tauscher, Susana Castillo 0001, Sebastian Bosse, Marcus A. Magnor
ICIP3
2021 Inverse kinematics for full-body self representation in VR-based cognitive rehabilitation
abstract
Being self-represented through an avatar increases embodiment and the feeling of presence in virtual reality. Nevertheless, currently users in VR are typically represented only by their hands, as not enough tracking data is available for full body self-representation. In our use case of VR-based diagnostics and cognitive rehabilitation of stroke patients, we aim for full body self-representation in order to increase therapeutic effectivity of the treatment. To solve this problem, Inverse Kinematics (IK) can be used for pose estimation, where no tracking data is available. IK allows to minimize the use of hardware and efforts of patients and clinical staff and, at the same time, provides a full-body representation based only on positions of the VR-HMD and users hands as input. In some use cases tracking data from additional, visual full body tracking sensors can be used to estimate the position of the lower body joints. In this study, we evaluate existing IK-based pose estimators; find that VRIK from Final IK is the most suitable approach for the given use case; integrate VRIK in our VR-rehabilitation system; adapt VRIK to meet the use case requirements and conduct subjective tests to validate a significantly increased notion of embodiment and presence through full-body over hands-only self-representation.
Larissa Wagnerberger, Detlef Runde, Mustafa-Tevfik Lafci, David Przewozny, Sebastian Bosse, Paul Chojecki
ISM5
2021 Reproducibility Companion Paper: Kalman Filter-Based Head Motion Prediction for Cloud-Based Mixed Reality
abstract
In our MM'20 paper,, we presented a Kalman filter-based approach for prediction of head motion in 6DoF. The proposed approach was employed in our cloud-based volumetric video streaming system to reduce the interaction latency experienced by the user. In this companion paper, we present the dataset collected for our experiments and our simulation framework that reproduces the obtained experimental results. Our implementation is freely available on Github to facilitate further research.
Serhan Gul, Sebastian Bosse, Dimitri Podborski, Thomas Schierl, Cornelius Hellge, Marc A. Kastner 0001, Jan Zahálka
ACM Multimedia2
2021 Can You Do Real-Time Gesture Recognition with 5 Watts?
abstract
Accurate and reliable gesture recognition is a central problem in human-computer interaction (HCI). Many applications that make use of gesture recognition call for mobile devices with reduced power consumption, weight and form factors. Recent advances in computer vision were particularly brought by deep neural networks and come at the cost of high computational complexity that hinders the employment on mobile devices. In this study, we evaluate the usability of a low-cost Raspberry Pi 4B amended by a Coral USB Accelerator, or a Neural Compute Stick 2, respectively, for low power real-time gesture recognition. To this end we evaluate the accuracy, inference time and power consumption for two different deep neural network-based recognition models and compare the results to other computer systems available as standard. Our experiments show that a combination of a Raspberry Pi 4B and Coral USB Accelerator allows for hand gesture recognition at frame rates of up to 30 frames per second at a power consumption of less than 5 Watts.
Azrin Rahman, Mykyta Kovalenko, David Przewozny, Karam Tomotaki-Dawoud, Paul Chojecki, Peter Eisert, Sebastian Bosse
SMC7
2020 Xpsnr: A Low-Complexity Extension of The Perceptually Weighted Peak Signal-To-Noise Ratio For High-Resolution Video Quality Assessment
abstract
The objective PSNR metric is known to correlate quite poorly with subjective assessments of video coding quality. Thus, a number of alternative VQA measures such as (MS-)SSIM and VMAF have been proposed. These, however, are often algorithmically complex and difficult to use for visually motivated encoder optimization tasks, especially subjectively optimized bit allocation. In this paper we show that, by way of low-complexity enhancements of our previous work on a perceptually weighted PSNR (WPSNR) metric, addressing shortcomings with video and ultra high-definition content, the prediction of human judgments of video coding quality by the WPSNR can be improved. In fact, the resulting XPSNR seems to match the performance of the aforementioned state-of-the-art methods.
Christian R. Helmrich, Mischa Siekmann, Sören Becker 0001, Sebastian Bosse, Detlev Marpe, Thomas Wiegand 0001
ICASSP4
2020 EEG-Based Assessment of Perceived Quality in Complex Natural Images
abstract
Psychophysiological methods gained a lot of interest in recent years as a potential remedy for the inherent flaws of overt psychophysical quality assessment methods. Among the psychophysiological monitoring methods, electroencephalography showed to be a promising choice. Specifically, the steady state visually evoked potential (SSVEP) was shown to provide a reliable neural correlate of perceived visual quality for degraded texture image patches. This paper evaluates the feasibility of the SSVEP-based quality assessment approach for more realistic and practically relevant images. To this end, we collected overt, psychophysical responses and neural, psychophysiological responses of 14 participants to 6 HD images each compressed at 4 distortion levels. The psychophysical part followed the Degradation Category Rating procedure. In the subsequent psychophysiological part, the subjects were presented with distorted and reference images alternating at a fixed rate of $f_{stim}\,=5$ Hz to elicit the SSVEP. We show that the amplitude of the $1 ^{st}$ harmonic of the SSVEP correlates significantly with the psychophysical responses $(\vert \rho \vert = 0.85, p \lt 0.05)$ in a single channel analysis at the Oz electrode.
Tamer Ajaj, Klaus-Robert Müller, Gabriel Curio, Thomas Wiegand 0001, Sebastian Bosse
ICIP5
2020 Kalman Filter-based Head Motion Prediction for Cloud-based Mixed Reality
abstract
Volumetric video allows viewers to experience highly-realistic 3D content with six degrees of freedom in mixed reality (MR) environments. Rendering complex volumetric videos can require a prohibitively high amount of computational power for mobile devices. A promising technique to reduce the computational burden on mobile devices is to perform the rendering at a cloud server. However, cloud-based rendering systems suffer from an increased interaction (motion-to-photon) latency that may cause registration errors in MR environments. One way of reducing the effective latency is to predict the viewer's head pose and render the corresponding view from the volumetric video in advance.
Serhan Gul, Sebastian Bosse, Dimitri Podborski, Thomas Schierl, Cornelius Hellge
ACM Multimedia2
2020 EEG-Based Assessment of Perceived Realness in Stylized Face Images
abstract
In this paper, we investigate the perception of realness in rendered face images experimentally using electroencephalography. To this end, we presented ten subjects with 36 character images based on six different faces (varying in gender and emotional expression) rendered at six different levels of realness ranging from abstract, cartoon-like renderings to real photographs. In the first psychophysical part of our study, we asked participants to rate perceived realness, appeal, familiarity, reassurance, and attractiveness for the presented characters. In the second part, we recorded the electroencephalogram when presenting the character images at a stimulation frequency of fstim= 5 Hz. We show that the amplitudes of the odd harmonics of the elicited steady-state visual evoked potential correlate with the psychophysical responses (|ρ| = 0.83, p < 0.05).
Milena T. Bagdasarian, Anna Hilsmann, Peter Eisert, Gabriel Curio, Klaus-Robert Müller, Thomas Wiegand 0001, Sebastian Bosse
QoMEX7
2020 Exploring neural and peripheral physiological correlates of simulator sickness
abstract
Abstract This article investigates neural and physiological correlates of simulator sickness (SS) through a controlled experiment conducted within a fully immersive dome projection system. Our goal is to establish a reliable, objective, and in situ measurable predictive indicator of SS. SS is a problem common to all types of visual simulators consisting of motion sickness‐like symptoms that may be experienced while and after being exposed to a dynamic, immersive visualization. It leads to ethical concerns and impaired validity of simulator‐based research. Due to the popularity of virtual reality devices, the number of people exposed to this problem is increasing and, therefore, it is crucial to find reliable predictors of this condition before any symptoms appear. Despite its relevance and the several theories about its origins, SS cannot yet be quantitatively modeled and predicted. Our results indicate that, while neural correlates did not materialize, physiological measures may be a solid early indicator of oncoming SS.
Jan-Philipp Tauscher, Alexandra Witt, Sebastian Bosse, Fabian Wolf Schottky, Steve Grogorick, Susana Castillo 0001, Marcus A. Magnor
Comput. Animat. Virtual Worlds3
2019 Perceptually Optimized Bit-Allocation and Associated Distortion Measure for Block-Based Image or Video Coding
abstract
It is well known that input-invariant quantization in perceptual image or video coding often leads to visually suboptimal results and that quantization parameter adaptation (QPA) based on a model of the human visual system can improve subjective coding quality. This paper introduces a simple low-complexity QPA algorithm, controlled using a block-wise perceptually weighted distortion measure representing a generalization of the PSNR metric. The weighting scheme of this WPSNR metric is based on a psychovisual model. It directly leads to a perceptually adapted scaling of the block-wise Lagrange parameter used in the bit-allocation process in the encoder and, consequently, to a block-wise QPA. Unlike prior QPA approaches, the proposal avoids classifications of picture regions and easily extends from still-image or grayscale to video or chromatic coding. The WPSNR metric also uses fewer algorithmic operations than e. g. the multiscale structural similarity measure (MS-SSIM). Due to the results of two formal subjective tests indicating its visual benefit, the QPA proposal has been adopted into VTM, the currently developed Versatile Video Coding (VVC) reference software.
Christian R. Helmrich, Sebastian Bosse, Mischa Siekmann, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
DCC2
2019 Neural Network Guided Perceptually Optimized Bit-Allocation for Block-Based Image and Video Compression
abstract
Bit-allocation based on the MSE is computationally convenient in image and video compression, but leads to perceptually suboptimal compression results. Distortion sensitivity, modeled as a reference specific property, can be used to improve the accuracy of perceptual quality prediction based on the MSE. This paper shows how distortion sensitivity directly leads to computationally beneficial perceptual optimization of irrelevance reduction and, thereby, of bit-allocation in image and video compression. To this end distortion sensitivity is estimated using a deep convolutional neural network. The proposed method of distortion sensitive bit-allocation is evaluated experimentally using HEVC and on our testset shows average bit-rate reductions with regard to the MOS of 15.9% compared to constant QP-based bit-allocation and 7.3% compared to state-of-the-art perceptual bit-allocation schemes.
Sebastian Bosse, Michael Dietzel, Sören Becker 0001, Christian R. Helmrich, Mischa Siekmann, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
ICIP1
2019 A Study of the Perceptually Weighted Peak Signal-To-Noise Ratio (WPSNR) for Image Compression
abstract
The peak signal-to-noise ratio (PSNR) is the most used objective measure for assessing perceptual image quality when it comes to image and video compression tasks, despite the fact that it exhibits weak performance in reflecting human perception. To address this problem, many image quality assessment (IQA) methods were proposed, e. g. the structural similarity quality measure (SSIM) and its extension, the multi-scale SSIM (MS-SSIM). In this paper we revisit and evaluate a block-based perceptually weighted PSNR (WPSNR) which calculates weighting factors to capture visual sensitivity of local image regions. We further introduce a sample-based version of WPSNR which determines those sensitivity weights with higher spatial accuracy. These methods are computationally inexpensive compared to other similarity measures and are shown to outperform PSNR, SSIM and similar perceptual quality measures when it comes to approximate subjective ratings of JPEG or JPEG2000 compressed images.
Johannes Erfurt, Christian R. Helmrich, Sebastian Bosse, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
ICIP3
2019 Data-driven Optimization of Row-Column Transforms for Block-Based Hybrid Video Compression
abstract
In state-of-the-art video compression residual coding is done by transforming the prediction error signals into a less correlated representation and performing the quantization and entropy coding in the transform domain. For complexity reasons usually separable transforms are used. A more flexible transform structure is given by row-column transforms, which apply a separate transform to each row and each column of a signal block. This paper describes a method for training such structured transforms by maximizing the data likelihood under a parameterized probabilistic model with a compelled structure. An explicit model is derived for the case of row-column transforms and its efficiency is demonstrated in the application of video compression. It is shown that trained row-column transforms achieve almost the same coding gain as unconstrained KLTs when applied as secondary transforms, while the encoder and decoder runtime are the same as in the separable transform case.
Mischa Siekmann, Sebastian Bosse, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
PCS2
2018 Neural Network-Based Estimation of Distortion Sensitivity for Image Quality Prediction
abstract
Due to its computational simplicity, the PSNR is a popular and widely used image quality measure, although it correlates poorly with perceived visual quality. Distortion sensitivity, a reference image specific property, can be used to compensate for the lack of perceptual relevance of the PSNR. Based on the functional mapping between perceptual and computational quality a deep convolutional neural network is used to estimate patchwise distortion sensitivity. The local estimates are used for an imagewise perceptual adaptation of the PSNR. The performance of the proposed estimation approach is evaluated on the LIVE and TID2013 databases and shows comparable or superior performance as compared to benchmark image quality measures.
Sebastian Bosse, Sören Becker 0001, Zacharias V. Fisches, Wojciech Samek, Thomas Wiegand 0001
ICIP1
2018 On the Stimulation Frequency in SSVEP-based Image Quality Assessment
abstract
Steady-state visual evoked potentials (SSVEP) are brain responses elicited by periodic visual stimuli. Recently it was shown that the use of SSVEP in quality studies allows for accurate psychophysiological assessment of perceived visual quality, but the influence of the stimulation frequency is still unclear. This paper studies experimentally the relation between the SNR of the neural signal and the stimulation frequency in an psychophysiological quality assessment setup. For various source images tested at different distortion magnitudes over the range of 6 different stimulation frequencies, we show physiologically plausible results that provide insights into the temporal dynamics of neural distortion processing. Our findings inform a rational choice of stimulation frequency in SSVEP-based image quality assessment studies. This potentially improves the experimental setup of future image quality assessment studies exploiting the SSVEP paradigm.
Sebastian Bosse, Milena T. Bagdasarian, Wojciech Samek, Gabriel Curio, Thomas Wiegand 0001
QoMEX1
2018 A Haar wavelet-based perceptual similarity index for image quality assessment
Rafael Reisenhofer, Sebastian Bosse, Gitta Kutyniok, Thomas Wiegand 0001
Signal Process. Image Commun.2
2018 Assessing Perceived Image Quality Using Steady-State Visual Evoked Potentials and Spatio-Spectral Decomposition
abstract
Steady-state visual evoked potentials (SSVEPs) are neural responses, measurable using electroencephalography (EEG), that are directly linked to sensory processing of visual stimuli. In this paper, SSVEP is used to assess the perceived quality of texture images. The EEG-based assessment method is compared with conventional methods, and recorded EEG data are correlated to obtained mean opinion scores (MOSs). A dimensionality reduction technique for EEG data called spatio-spectral decomposition (SSD) is adapted for the SSVEP framework and used to extract physiologically meaningful and plausible neural components from the EEG recordings. It is shown that the use of SSD not only increases the correlation between neural features and MOS to r = -0.93, but also solves the problem of channel selection in an EEG-based image-quality assessment.
Sebastian Bosse, Laura Acqualagna, Wojciech Samek, Anne Porbadnigk, Gabriel Curio, Benjamin Blankertz, Klaus-Robert Müller, Thomas Wiegand 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Deep Neural Networks for No-Reference and Full-Reference Image Quality Assessment
abstract
We present a deep neural network-based approach to image quality assessment (IQA). The network is trained end-to-end and comprises ten convolutional layers and five pooling layers for feature extraction, and two fully connected layers for regression, which makes it significantly deeper than related IQA models. Unique features of the proposed architecture are that: 1) with slight adaptations it can be used in a no-reference (NR) as well as in a full-reference (FR) IQA setting and 2) it allows for joint learning of local quality and local weights, i.e., relative importance of local quality to the global quality estimate, in an unified framework. Our approach is purely data-driven and does not rely on hand-crafted features or other types of prior domain knowledge about the human visual system or image statistics. We evaluate the proposed approach on the LIVE, CISQ, and TID2013 databases as well as the LIVE In the wild image quality challenge database and show superior performance to state-of-the-art NR and FR IQA methods. Finally, cross-database evaluation shows a high ability to generalize between different databases, indicating a high robustness of the learned features.
Sebastian Bosse, Dominique Maniry, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek
IEEE Trans. Image Process.1
2017 A perceptually relevant shearlet-based adaptation of the PSNR
abstract
Although being one of the simplest and most widely used image quality metrics (IQMs) the peak signal-to-noise ratio (PSNR) correlates only poorly with visual quality as perceived by humans. Based on an analysis of the non-linear mapping from PSNR to mean opinion scores (MOS) we identify a functional mapping parameter to adapt the PSNR perceptually meaningful. Neurophysiologically motivated, a shearlet-based correction is proposed for controlling this perceptual PSNR adaption. The performance of the proposed perceptually adapted PSNR is evaluated on the LIVE and TID2013 databases and shows to be superior or comparable to benchmark IQMs.
Sebastian Bosse, Mischa Siekmann, Wojciech Samek, Thomas Wiegand 0001
ICIP1
2016 Shearlet-based reduced reference image quality assessment
abstract
This paper proposes a reduced reference image quality assessment method using only a low number of features. It involves a shearlet decomposition, directional pooling of the obtained coefficient and extracts the scalewise statistical location parameter as a feature. The proposed method is tested and compared to similar approaches on the LIVE image database. On this database it outperforms the compared methods on five of seven distortion types and on the full testset with a linear correlation of = 0.89.
Sebastian Bosse, Qiaobo Chen, Mischa Siekmann, Wojciech Samek, Thomas Wiegand 0001
ICIP1
2016 A deep neural network for image quality assessment
abstract
This paper presents a no reference image (NR) quality assessment (IQA) method based on a deep convolutional neural network (CNN). The CNN takes unpreprocessed image patches as an input and estimates the quality without employing any domain knowledge. By that, features and natural scene statistics are learnt purely data driven and combined with pooling and regression in one framework. We evaluate the network on the LIVE database and achieve a linear Pearson correlation superior to state-of-the-art NR IQA methods. We also apply the network to the image forensics task of decoder-sided quantization parameter estimation and also here achieve correlations of r = 0.989.
Sebastian Bosse, Dominique Maniry, Thomas Wiegand 0001, Wojciech Samek
ICIP1
2016 Quality assessment of image patches distorted by image compression using crowdsourcing
abstract
Three experiments addressing the assessment of perceived image quality in a patch-based manner are compared for HEVC compression artifacts. It is shown that image patches of a size small as 128×128 pixel are large enough to evaluate the perceived image quality in a Degradation Category Rating (DCR) setting. Ratings obtained with 128×128 pixel sized images patches and 512×512 pixel sized images of the same spatial statistics show a correlation of r=0.99. Based on this finding, image quality assessment of 128×128 pixel sized image patches degraded by HEVC compression is compared for controlled lab environment and uncontrolled crowdsourcing settings. Although we find high overall correlation between the quality ratings obtained in the two environments, observers tend to give worse ratings in the crowdsourcing setting and for conditions of higher quality a reduction of correlation is observed. These findings have implications for choosing controlled vs. uncontrolled viewing conditions for image quality assessment for real-life applications.
Sebastian Bosse, Mischa Siekmann, Jennifer Rasch, Thomas Wiegand 0001, Wojciech Samek
ICME1
2016 Neural network-based full-reference image quality assessment
abstract
S.334-338
Sebastian Bosse, Dominique Maniry, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek
PCS1
2016 Brain-Computer Interfacing for multimedia quality assessment
abstract
The assessment of perceived multimedia quality is a central research field in information and media technology. Conventionally, psychophysical techniques are used for determining the quality of multimedia signals. Recently, Brain-Computer Interfacing (BCI)-based methods have been proposed for the assessment of perceived multimedia signal quality. In this paper we give an overview over the shortcomings of conventional approaches, present the state-of-the art of BCI-based methods and discuss open questions and challenges relevant to the BCI community.
Sebastian Bosse, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek
SMC1
2014 Neurally informed assessment of perceived natural texture image quality
abstract
Conventionally, the quality of images and related codecs are assessed using subjective tests, such as Degradation Category Rating. These quality assessments consider the behavioral level only. Recently, it has been proposed to complement this approach by investigating how quality is processed in the brain of a user (using electroencephalography, EEG), potentially leading to results that are less biased by subjective factors. In this paper, a novel method is presented for assessing how image quality is processed on a neural level, using Steady-State Visual Evoked Potentials (SSVEPs) as EEG features. We tested our approach in an EEG study with 16 participants who were presented with distorted images of natural textures. Subsequently, we compared our approach analogously to the standardized Degradation Category Rating quality assessment. Remarkably, our novel method yields a correlation of |r| = 0.93 to MOS on the recorded dataset.
Sebastian Bosse, Laura Acqualagna, Anne Porbadnigk, Benjamin Blankertz, Gabriel Curio, Klaus-Robert Müller, Thomas Wiegand 0001
ICIP1
2013 3D High-Efficiency Video Coding for Multi-View Video and Depth Data
abstract
This paper describes an extension of the high efficiency video coding (HEVC) standard for coding of multi-view video and depth data. In addition to the known concept of disparity-compensated prediction, inter-view motion parameter, and inter-view residual prediction for coding of the dependent video views are developed and integrated. Furthermore, for depth coding, new intra coding modes, a modified motion compensation and motion vector coding as well as the concept of motion parameter inheritance are part of the HEVC extension. A novel encoder control uses view synthesis optimization, which guarantees that high quality intermediate views can be generated based on the decoded data. The bitstream format supports the extraction of partial bitstreams, so that conventional 2D video, stereo video, and the full multi-view video plus depth format can be decoded from a single bitstream. Objective and subjective results are presented, demonstrating that the proposed approach provides 50% bit rate savings in comparison with HEVC simulcast and 20% in comparison with a straightforward multi-view extension of HEVC without the newly developed coding tools.
Karsten Müller 0001, Heiko Schwarz, Detlev Marpe, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Philipp Merkle, Hunn Rhee, Gerhard Tech, Martin Winken, Thomas Wiegand 0001
IEEE Trans. Image Process.5
2012 Extension of High Efficiency Video Coding (HEVC) for multiview video and depth data
abstract
This paper presents an approach for 3D video coding that uses a format in which a small number of views as well as associated depth maps are coded and transmitted. At the receiver side, additional views required for displaying the 3D video on an autostereoscopic display can be generated based on the corresponding decoded signals by using depth image based rendering (DIBR) techniques. In terms of coding technology, the proposed coding scheme represents an extension of High Efficiency Video Coding (HEVC), similar to the Multiview Coding (MVC) extension of H.264/AVC. Besides the well-known disparity-compensated prediction, advanced techniques for inter-view and inter-component prediction, the representation of depth blocks, and the encoder control for depth signals have been developed and integrated. In comparison to simulcasting the different signals using HEVC, the proposed approach provides about 40% and 50% average bit rate savings for a whole test set when configured to comply with a 2- and 3-view scenario, respectively. The proposed codec was submitted as response to a Call for Proposals on 3D Video Technology issued by the ISO/IEC Moving Picture Experts Group (MPEG) and it was ranked as the overall best performing HEVC-based proposal in the related subjective tests.
Heiko Schwarz, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Philipp Merkle, Karsten Müller 0001, Hunn Rhee, Gerhard Tech, Martin Winken, Detlev Marpe, Thomas Wiegand 0001
ICIP3
2012 Encoder control for renderable regions in high efficiency multiview video plus depth coding
abstract
This paper describes a new encoder control method for multiview video plus depth coding. Since large parts of a multiview scenery are present in more than one of the captured video sequences, a depth-aware encoder control is introduced, which identifies those regions based on given depth maps and omits the coding of the residual signal for those regions. Experimental results indicate that bit rate reductions of about 5-9 %, depending on the bit rate, can be achieved for the 2-view case at a constant subjective quality.
Sebastian Bosse, Heiko Schwarz, Tobias Hinz, Thomas Wiegand 0001
PCS1
2012 3D video coding using advanced prediction, depth modeling, and encoder control methods
abstract
The presented approach for 3D video coding uses the multiview video plus depth format, in which a small number of video views as well as associated depth maps are coded. Based on the coded signals, additional views required for displaying the 3D video on an autostereoscopic display can be generated by depth image based rendering techniques. The developed coding scheme represents an extension of HEVC, similar to the MVC extension of H.264/AVC. However, in addition to the well-known disparity-compensated prediction advanced techniques for inter-view and inter-component prediction, the representation of depth blocks, and the encoder control for depth signals have been integrated. In comparison to simulcasting the different signals using HEVC, the proposed approach provides about 40% and 50% bit rate savings for the tested configurations with 2 and 3 views, respectively. Bit rate reductions of about 20% have been obtained in comparison to a straightforward multiview extension of HEVC without the newly developed coding tools.
Heiko Schwarz, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Detlev Marpe, Philipp Merkle, Karsten Müller 0001, Hunn Rhee, Gerhard Tech, Martin Winken, Thomas Wiegand 0001
PCS3
2012 Toward a Direct Measure of Video Quality Perception Using EEG
abstract
An approach to the direct measurement of perception of video quality change using electroencephalography (EEG) is presented. Subjects viewed 8-s video clips while their brain activity was registered using EEG. The video signal was either uncompressed at full length or changed from uncompressed to a lower quality level at a random time point. The distortions were introduced by a hybrid video codec. Subjects had to indicate whether they had perceived a quality change. In response to a quality change, a positive voltage change in EEG (the so-called P3 component) was observed at latency of about 400-600 ms for all subjects. The voltage change positively correlated with the magnitude of the video quality change, substantiating the P3 component as a graded neural index of the perception of video quality change within the presented paradigm. By applying machine learning techniques, we could classify on a single-trial basis whether a subject perceived a quality change. Interestingly, some video clips wherein changes were missed (i.e., not reported) by the subject were classified as quality changes, suggesting that the brain detected a change, although the subject did not press a button. In conclusion, abrupt changes of video quality give rise to specific components in the EEG that can be detected on a single-trial basis. Potentially, a neurotechnological approach to video assessment could lead to a more objective quantification of quality change detection, overcoming the limitations of subjective approaches (such as subjective bias and the requirement of an overt response). Furthermore, it allows for real-time applications wherein the brain response to a video clip is monitored while it is being viewed.
Simon Scholler, Sebastian Bosse, Matthias Treder, Benjamin Blankertz, Gabriel Curio, Klaus-Robert Müller, Thomas Wiegand 0001
IEEE Trans. Image Process.2
2010 Highly efficient video compression using quadtree structures and improved techniques for motion representation and entropy coding
abstract
This paper describes a novel video coding scheme that can be considered as a generalization of the block-based hybrid video coding approach of H.264/AVC. While the individual building blocks of our approach are kept simple similarly as in H.264/AVC, the flexibility of the block partitioning for prediction and transform coding has been substantially increased. This is achieved by the use of nested and pre-configurable quadtree structures, such that the block partitioning for temporal and spatial prediction as well as the space-frequency resolution of the corresponding prediction residual can be adapted to the given video signal in a highly flexible way. In addition, techniques for an improved motion representation as well as a novel entropy coding concept are included. The presented video codec was submitted to a Call for Proposals of ITU-T VCEG and ISO/IEC MPEG and was ranked among the five best performing proposals, both in terms of subjective and objective quality.
Detlev Marpe, Heiko Schwarz, Sebastian Bosse, Benjamin Bross, Philipp Helle, Tobias Hinz, Heiner Kirchhoffer, Haricharan Lakshman, Tung Nguyen 0001, Simon Oudin, Mischa Siekmann, Karsten Sühring, Martin Winken, Thomas Wiegand 0001
PCS3
2010 Separable Wiener filter based adaptive in-loop filter for video coding
abstract
Recent investigations have shown that a non-separable Wiener filter, that is applied inside the motion-compensation loop, can improve the coding efficiency of hybrid video coding designs. In this paper, we study the application of separable Wiener filters. Our design includes the possibility to adaptively choose between the application of the vertical, horizontal, or combined filter. The simulation results verify that a separable in-loop Wiener filter is capable of providing virtually the same increase in coding efficiency as a non-separable Wiener filter, but at a significantly reduced decoder complexity.
Mischa Siekmann, Sebastian Bosse, Heiko Schwarz, Thomas Wiegand 0001
PCS2
2010 Video Compression Using Nested Quadtree Structures, Leaf Merging, and Improved Techniques for Motion Representation and Entropy Coding
abstract
Abstract-A video coding architecture is described that is based on nested and pre-configurable quadtree structures for flexible and signal-adaptive picture partitioning. The primary goal of this partitioning concept is to provide a high degree of adaptability for both temporal and spatial prediction as well as for the purpose of space-frequency representation of prediction residuals. At the same time, a leaf merging mechanism is included in order to prevent excessive partitioning of a picture into prediction blocks and to reduce the amount of bits for signaling the prediction signal. For fractional-sample motion-compensated prediction, a fixed-point implementation of the maximal-order minimum-support algorithm is presented that uses a combination of infinite impulse response and FIR filtering. Entropy coding utilizes the concept of probability interval partitioning entropy codes that offers new ways for parallelization and enhanced throughput. The presented video coding scheme was submitted to a joint call for proposals of ITU-T Visual Coding Experts Group and ISO/IEC Moving Picture Experts Group and was ranked among the five best performing proposals, both in terms of subjective and objective quality.
Detlev Marpe, Heiko Schwarz, Sebastian Bosse, Benjamin Bross, Philipp Helle, Tobias Hinz, Heiner Kirchhoffer, Haricharan Lakshman, Tung Nguyen 0001, Simon Oudin, Mischa Siekmann, Karsten Sühring, Martin Winken, Thomas Wiegand 0001
IEEE Trans. Circuits Syst. Video Technol.3