VLDB 2026 Research / reviewers in the wild / expert
Sebastian Stober
dblp:73/650
· DBLP profile ↗
22ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-1717-4133ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | INAM: Image-Scale Neural Additive ModelsabstractNeural Additive Models are inherently interpretable models that can be applied to tabular data.However, when applying these models to images the value of a given pixel is not a meaningful feature for understanding the model.For that reason, we propose INAM -Image scale Neural Additive Model -a combination of trainable feature extractors and NAMs.We show INAMs can be successfully applied to image data sets with low variability while allowing global explanations of the models and data point-specific explanations. Jana Hüls, Jan-Ole Perschewski, Sebastian Stober |
ESANN | 3 |
| 2025 | Intersectional Bias Quantification in Facial Image Processing with Pre-Trained ImageNet ClassifiersabstractDeep Learning models have achieved significant success, often facilitated by transfer learning. This involves using pre-trained models as a basis for new tasks. However, this practice carries the risk of propagating biases that are present in the original training data. In this study, we examine biases related to the protected attributes of "race", "age", and "gender" in several pre-trained classifiers that were trained on the widely used ImageNet dataset. Our analysis emphasizes intersectionality, exploring how interactions between these attributes influence biases. We introduce and employ a novel, model-agnostic approach to analyze biases in the representations of pre-trained deep neural networks through activation similarity-based clustering, with a focus on intersectionality. Our results suggest that, regardless of the specific model, ImageNet classifiers representations strongly reflect age information, cluster certain ethnic groups, and differentiate genders in middle-aged individuals. Valerie Krug, Florian Röhrbein, Sebastian Stober |
IJCNN | 3 |
| 2025 | StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency SegmentationabstractDetecting and segmenting dysfluencies is crucial for effective speech therapy and real-time feedback. However, most methods only classify dysfluencies at the utterance level. We introduce StutterCut, a semi-supervised framework that formulates dysfluency segmentation as a graph partitioning problem, where speech embeddings from overlapping windows are represented as graph nodes. We refine the connections between nodes using a pseudo-oracle classifier trained on weak (utterance-level) labels, with its influence controlled by an uncertainty measure from Monte Carlo dropout. Additionally, we extend the weakly labelled FluencyBank dataset by incorporating frame-level dysfluency boundaries for four dysfluency types. This provides a more realistic benchmark compared to synthetic datasets. Experiments on real and synthetic datasets show that StutterCut outperforms existing methods, achieving higher F1 scores and more precise stuttering onset detection. Suhita Ghosh, Mélanie Jouaiti, Jan-Ole Perschewski, Sebastian Stober |
INTERSPEECH | 4 |
| 2024 | T-DVAE: A Transformer-Based Dynamical Variational Autoencoder for SpeechabstractAbstract In contrast to Variational Autoencoders, Dynamical Variational Autoencoders (DVAEs) learn a sequence of latent states for a time series. Initially, they were implemented using recurrent neural networks (RNNs) known for challenging training dynamics and problems with long-term dependencies. This led to the recent adoption of Transformers close to the RNN-based implementation. These implementations still use RNNs as part of the architecture even though the Transformer can solve the task as the sole building block. Hence, we improve the LigHT-DVAE architecture by removing the dependence on RNNs and Cross-Attention. Furthermore, we show that a trained LigHT-DVAE ignores output-to-hidden connections, which allows us to simplify the overall architecture by removing output-to-hidden connections. We demonstrate the capability of the resulting T-DVAE on librispeech and voice bank with an improvement in training time, memory consumption, and generative performance. Jan-Ole Perschewski, Sebastian Stober |
ICANN (7) | 2 |
| 2024 | Tilt your Head: Activating the Hidden Spatial-Invariance of ClassifiersabstractDeep neural networks are applied in more and more areas of everyday life. However, they still lack essential abilities, such as robustly dealing with spatially transformed input signals. Approaches to mitigate this severe robustness issue are limited to two pathways: Either models are implicitly regularised by increased sample variability (data augmentation) or explicitly constrained by hard-coded inductive biases. The limiting factor of the former is the size of the data space, which renders sufficient sample coverage intractable. The latter is limited by the engineering effort required to develop such inductive biases for every possible scenario. Instead, we take inspiration from human behaviour, where percepts are modified by mental or physical actions during inference. We propose a novel technique to emulate such an inference process for neural nets. This is achieved by traversing a sparsified inverse transformation tree during inference using parallel energy-based evaluations. Our proposed inference algorithm, called Inverse Transformation Search (ITS), is model-agnostic and equips the model with zero-shot pseudo-invariance to spatially transformed inputs. We evaluated our method on several benchmark datasets, including a synthesised ImageNet test set. ITS outperforms the utilised baselines on all zero-shot test scenarios. Johann Schmidt, Sebastian Stober |
ICML | 2 |
| 2024 | Anonymising Elderly and Pathological Speech: Voice Conversion Using DDSP and Query-by-ExampleabstractSpeech anonymisation aims to protect speaker identity by changing personal identifiers in speech while retaining linguistic content. Current methods fail to retain prosody and unique speech patterns found in elderly and pathological speech domains, which is essential for remote health monitoring. To address this gap, we propose a voice conversion-based method (DDSP-QbE) using differentiable digital signal processing and query-by-example. The proposed method, trained with novel losses, aids in disentangling linguistic, prosodic, and domain representations, enabling the model to adapt to uncommon speech patterns. Objective and subjective evaluations show that DDSP-QbE significantly outperforms the voice conversion state-of-the-art concerning intelligibility, prosody, and domain preservation across diverse datasets, pathologies, and speakers while maintaining quality and speaker anonymity. Experts validate domain preservation by analysing twelve clinically pertinent domain attributes. Suhita Ghosh, Mélanie Jouaiti, Yamini Sinha, Tim Polzehl, Ingo Siegert, Sebastian Stober |
INTERSPEECH | 7 |
| 2023 | Trustworthy Academic Risk Prediction with Explainable Boosting Machines
Vegenshanti Dsilva, Johannes Schleiss, Sebastian Stober |
AIED | 3 |
| 2023 | Emo-StarGAN: A Semi-Supervised Any-to-Many Non-Parallel Emotion-Preserving Voice ConversionabstractSpeech anonymisation prevents misuse of spoken data by removing any personal identifier while preserving at least linguistic content. However, emotion preservation is crucial for natural human-computer interaction. The well-known voice conversion technique StarGANv2-VC achieves anonymisation but fails to preserve emotion. This work presents an any-to-many semi-supervised StarGANv2-VC variant trained on partially emotion-labelled non-parallel data. We propose emotion-aware losses computed on the emotion embeddings and acoustic features correlated to emotion. Additionally, we use an emotion classifier to provide direct emotion supervision. Objective and subjective evaluations show that the proposed approach significantly improves emotion preservation over the vanilla StarGANv2-VC. This considerable improvement is seen over diverse datasets, emotions, target speakers, and inter-group conversions without compromising intelligibility and anonymisation. Suhita Ghosh, Yamini Sinha, Ingo Siegert, Tim Polzehl, Sebastian Stober |
INTERSPEECH | 6 |
| 2022 | Neural-Gas VAE
Jan-Ole Perschewski, Sebastian Stober |
ICANN (1) | 2 |
| 2021 | Uncertainty-aware temporal self-learning (UATS): Semi-supervised learning for segmentation of prostate zones and beyond
Anneke Meyer, Suhita Ghosh, Daniel Schindele, Martin Schostak, Sebastian Stober, Christian Hansen 0001, Marko Rak |
Artif. Intell. Medicine | 5 |
| 2020 | Analyzing Regions of Safety for Handling Shared Data in Cooperative SystemsabstractCooperative Systems promise increased performance by enriching environmental perception through shared data. Conversely, the entailed openness of the individual system architectures threatens their safety. Recent works focus on a runtime safety assessment to address this threat and thereby aim for high abstractions to provide general interfaces. On the other hand, uncertainty models of shared data, which are necessary inputs to such approaches, aim for low abstractions to provide detailed representations. The present work addresses the resulting incompatibilities by proposing a Lyapunov-based method to estimate so-called Regions of Safety. We show that these enable analyzing low-level uncertainty models to interface with state-of-the-art run-time safety assessment methods and thereby facilitate self-adaptivity and guaranteed safety of cooperative systems. The approach is evaluated in the simulated scenario of Cooperative Adaptive Cruise Control. Georg Jäger, Johannes Schleiss, Sasiporn Usanavasin, Sebastian Stober, Sebastian Zug |
ETFA | 4 |
| 2020 | PredNet and Predictive Coding: A Critical ReviewabstractPredNet, a deep predictive coding network developed by Lotter et al., combines a biologically inspired architecture based on the propagation of prediction error with self-supervised representation learning in video. While the architecture has drawn a lot of attention and various extensions of the model exist, there is a lack of a critical analysis. We fill in the gap by evaluating PredNet both as an implementation of the predictive coding theory and as a self-supervised video prediction model using a challenging video action classification dataset. We design an extended model to test if conditioning future frame predictions on the action class of the video improves the model performance. We show that PredNet does not yet completely follow the principles of predictive coding. The proposed top-down conditioning leads to a performance gain on synthetic data, but does not scale up to the more complex real-world action classification dataset. Our analysis is aimed at guiding future research on similar architectures based on the predictive coding theory. Roshan Prakash Rane, Edit Szügyi, Vageesh Saxena, André Ofner, Sebastian Stober |
ICMR | 5 |
| 2020 | Balancing Active Inference and Active Learning with Deep Variational Predictive Coding for EEGabstractThis paper discusses representation learning from electroencephalographic (EEG) signal with deep variational predictive coding networks. We introduce a hierarchical probabilistic network that minimises prediction error at multiple levels of spatio-temporal abstraction. While the lowest layer predicts brain activity directly, higher layers abstract away from the data and predict sequences of the hidden states in lower layers. The network captures both expected and actual uncertainty by relating predicted state posteriors. Each layer minimises (expected) surprise either with or without sampling new evidence from the layer below. This structure motivates both active learning and active inference as means to learn representations. Active learning refers to model parameter exploration which allows to learn regularities, especially when they are stable between trials. Active inference refers to hidden state exploration, a process that enables dynamic inference of the current context using the learned generative model. We train the model on EEG data recorded during free reading and evaluate adaptive EEG prediction in the context of Fixation Related Potentials (FRPs). André Ofner, Sebastian Stober |
SMC | 2 |
| 2019 | Window-Based Neural Tagging for Shallow Discourse Argument LabelingabstractThis paper describes a novel approach for the task of end-to-end argument labeling in shallow discourse parsing.Our method describes a decomposition of the overall labeling task into subtasks and a general distance-based aggregation procedure.For learning these subtasks, we train a recurrent neural network and gradually replace existing components of our baseline by our model.The model is trained and evaluated on the Penn Discourse Treebank 2 corpus.While it is not as good as knowledge-intensive approaches, it clearly outperforms other models that are also trained without additional linguistic features. René Knaebel, Manfred Stede, Sebastian Stober |
CoNLL | 3 |
| 2017 | Learning discriminative features from electroencephalography recordings by encoding similarity constraintsabstractThis paper introduces a pre-training technique for learning discriminative features from electroencephalography (EEG) recordings using deep neural networks. EEG data are generally only available in small quantities, they are high-dimensional with a poor signal-to-noise ratio, and there is considerable variability between individual subjects and recording sessions. Similarity-constraint encoders as introduced in this paper specifically address these challenges for feature learning. They learn features that allow to distinguish between classes by demanding that encodings of two trials from the same class are more similar to each other than to encoded trials from other classes. This tuple-based training approach is especially suitable for small datasets. The proposed technique is evaluated using the publicly available OpenMIIR dataset of EEG recordings taken while participants listened to and imagined music. For this dataset, a simple convolutional filter can be learned that significantly improves the signal-to-noise ratio while aggregating the 64 EEG channels into a single waveform. Sebastian Stober |
ICASSP | 1 |
| 2017 | Exploring Large Movie Collections: Comparing Visual Berrypicking and Traditional Browsing
Thomas Low, Christian Hentschel, Sebastian Stober, Harald Sack, Andreas Nürnberger |
MMM (2) | 3 |
| 2014 | Search result visualization with characters for childrenabstractIn this paper, we explore alternative ways to visualize search results for children. We propose a novel search result visualization using characters. The main idea is to represent each web document as a character where a character visually provides clues about the webpage's content. We focused on children between six and twelve as a target user group. Following the usercentered development approach, we conducted a preliminary user study to determine how children would represent a webpage as a sketch based on a given template of a character. Using the study results the first prototype of a search engine was developed. We evaluated the search interface on a touchpad and a touch table in a second user study and analyzed user's satisfaction and preferences. Tatiana Gossen, René Müller 0002, Sebastian Stober, Andreas Nürnberger |
IDC | 3 |
| 2014 | Using Convolutional Neural Networks to Recognize Rhythm Stimuli from Electroencephalography Recordings
Sebastian Stober, Daniel J. Cameron, Jessica A. Grahn |
NIPS | 1 |
| 2013 | Adaptive music retrieval-a state of the art
Sebastian Stober, Andreas Nürnberger |
Multim. Tools Appl. | 1 |
| 2011 | Analyzing the impact of data vectorization on distance relationsabstractSome popular algorithms used in Music Information Retrieval (MIR) such as Self-Organizing Maps (SOMs) require the objects they process to be represented as vectors, i.e. elements of a vector space. This is a rather severe restriction and if the data does not adhere to it, some means of vectorization is required. As a common practice, the full distance matrix is computed and each row of the matrix interpreted as an artificial feature vector. This paper empirically investigates the impact of this transformation. Further, an alternative approach for vectorization based on Multidimensional Scaling is pro posed that is able to better preserve the actual distance relations of the objects which is essential for obtaining a good retrieval performance. Sebastian Stober, Andreas Nürnberger |
ICME | 1 |
| 2010 | Multi-facet exploration of image collections with an adaptive multi-focus zoomable interfaceabstractSometimes it is not possible for a user to state a retrieval goal explicitly a priori. One common way to support such exploratory retrieval scenarios is to give an overview using a neighborhood-preserving projection of the collection onto two dimensions. However, neighborhood cannot always be preserved in the projection because of the dimensionality reduction. Further, there is usually more than one way to look at a collection of images - and diversity grows with the number of features that can be extracted. We describe an adaptive zoomable interface for exploration that addresses both problems: It makes use of a complex non-linear multi-focal zoom lens that exploits the distorted neighborhood relations introduced by the projection. We further introduce the concept of facet distances representing different aspects of image similarity. Given user-specific weightings of these aspects, the system can adapt to the user's way of exploring the collection by manipulation of the neighborhoods as well as the projection. Sebastian Stober, Christian Hentschel, Andreas Nürnberger |
IJCNN | 1 |
| 2006 | DAWN - A System for Context-Based Link Recommendation in Web Navigation
Sebastian Stober, Andreas Nürnberger |
KES (1) | 1 |