Fabian Berns

dblp:236/7409 · DBLP profile ↗
← Back
13ranked-venue papers in the field
9as first author
8since 2021 · last 2023
0000-0002-7033-3789ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 4 (3 first)Big Data, Cloud & Distributed Data Systems · 4 (2 first)Information Retrieval & Web Search · 3 (2 first)Database Systems & Data Management · 2 (2 first)
YearPublicationVenuePosition
2023 Trustworthy Medical Operational AI: Marrying AI and Regulatory Requirements
abstract
Despite recent advancements in AI and Data Science, the vast Big Data sources available to medical and health care providers are far from living up to their potential. Addressing the underlying transparency and data protection concerns helps to unleash these advancements in the medical domain and thus benefits research and patient care. Subsequently, we propose a system for trustworthy medical operational AI in this paper. We present our vision to align medical operational AI with regulatory demands of the medical domain. We propose guiding principles to marry data-driven diagnostic recommendations with legal frameworks, clinical protocols, and expert-driven reasoning. Through this research, we aim for AI systems in medicine that not only provide accurate predictions but also empower users to comprehend and trust the underlying decision-making processes.
Fabian Berns, Georg Zimmermann, Christian Borgelt, Niclas Heilig, Jan Kirchhoff, Florian Stumpe
IEEE Big Data1
2022 Tracing Patterns in Electrophysiological Time Series Data
abstract
When multiple sensors record spatially proximate areas of activity, spreading activity patterns appear as temporally shifted signals in multiple time series. This is particularly prominent in the domains of medical and health analysis, where multi-sensory data is the object of time-elastic investigation. Tracing the spread of these patterns still remains a challenge in time series analysis. In this paper, we propose Motif Tracking for Spatially Ordered Time Series (MoTrack), an algorithm to efficiently track the propagation of individual patterns of activity throughout spatially ordered time series. Additionally, we present the concept of propagation trees to represent this propagation for a given point of origin. We investigate our proposal by applying MoTrack to high-frequency recordings of the electrical activity of β-cells located inside the pancreatic islet. The results confirm MoTrack’s capability to trace dynamically evolving signals in such recordings and indicate that future work using this approach can address current challenges in diabetes research.
Jan David Hüwel, Anne Gresch, Fabian Berns, Ruben Koch, Martina Düfer, Christian Beecks
DSAA3
2022 A Comparative Performance Analysis of Fast K-Means Clustering Algorithms
Christian Beecks, Fabian Berns, Jan David Hüwel, Andrea Linxen, Georg Stefan Schlake, Tim Düsterhus
iiWAS2
2021 Automated Kernel Search for Gaussian Processes on Data Streams
abstract
Gaussian Processes offer non-parametric, probabilistic models that can be used in numerous fields of data analysis. One major drawback is their lack of adjustability in case of drifting and evolving streaming data, where inherent kernels need to be adapted in an efficient manner. To counteract this issue, we propose a novel automated kernel search method that allows us to incrementally adapt Gaussian Process models to evolving IoT data streams. Our approach, denoted as Adjusting Kernel Search (AKS), offers an efficient alternative to searching for suitable kernels from scratch. We evaluate the AKS algorithm on several IoT datasets and show that our approach is able to achieve higher accuracy with lower run-times compared to previous approaches.
Jan David Hüwel, Fabian Berns, Christian Beecks
IEEE BigData2
2021 LOGIC: Probabilistic Machine Learning for Time Series Classification
abstract
Time series data is one of the complex data types commonly encountered in many application areas ranging from automotive, finance, medicine to industry. A prominent task is time series classification, which entails identifying expressive features in oder to predict class labels of time series data. In this paper, we propose a novel approach for time series classification called Local Gaussian Process Model Inference Classification (LOGIC). Our concept consists in (i) learning latent characteristics of given time series data by means of Gaussian processes, (ii) using these characteristics to embed time series into a more expressive feature space and (iii) classifying time series data based on these features via existing classification methods. By making use of various general-purpose classification methods, we show that LOGIC is able to compete with state-of-the-art approaches in terms of accuracy and efficiency.
Fabian Berns, Jan David Hüwel, Christian Beecks
ICDM1
2021 Stochastic Time Series Representation for Interval Pattern Mining via Gaussian Processes
abstract
Trends, periodicities and local variations are among the main recognizable patterns in time series data. While humans are able to quickly explore the superimposition of such patterns, mining algorithms are often faced with the challenge of finding (i) a suitable time series representation model and (ii) an expressive query model which adapt to diverse application domains and information needs. In this paper, we propose a supervised stochastic approach which facilitates interval-based pattern analysis of time series data. Our proposal is based on non-parametric Gaussian Processes and is able to interrelate interesting patterns within single and across multiple time series. Our performance evaluation in different real-world application domains indicates that our approach is able to expose interesting patterns and knowledge.
Fabian Berns, Christian Beecks
SDM1
2021 Complexity-Adaptive Gaussian Process Model Inference for Large-Scale Data
abstract
A flexible, domain-agnostic function approximator, which is robust towards unreliable, noisy and partially missing data, would be an ideal tool for pattern mining in large-scale data. Although Gaussian Process Models (GPMs), which are widely regarded as a probabilistic tool for capturing inherent data characteristics, satisfy those requirements, full Gaussian Process inference and training is limited to a few thousand data records. Moreover, a process of automatic GPM inference is required to find an optimal model for a given dataset, despite prevailing default instantiations and existing prior knowledge in some scenarios, which both shortcut the way to an optimal GPM. Since non-approximate Gaussian Processes only allow for processing small datasets with low local statistical versatility, we propose a new approach that enables to automatically infer GPMs of adaptive local complexity on large scale multivariate data. The resulting model is composed of independent statistical representations for disjoint partitions varying in statistical versatility. Our performance evaluation indicates an improvement in inference runtime, while maintaining high model quality with regards to state-of-the-art GPM inference algorithms.
Fabian Berns, Christian Beecks
SDM1
2021 Local Gaussian Process Model Inference Classification for Time Series Data
abstract
One of the prominent types of time series analytics is classification, which entails identifying expressive class-wise features for determining class labels of time series data. In this paper, we propose a novel approach for time series classification called Local Gaussian Process Model Inference Classification (LOGIC). Our idea consists in (i) approximating the latent, class-wise characteristics of given time series data by means of Gaussian processes and (ii) aggregating these characteristics into a feature representation to (iii) provide a model-agnostic interface for state-of-the-art feature classification mechanisms. By making use of a fully-connected neural network as classification model, we show that the LOGIC model is able to compete with state-of-the-art approaches.
Fabian Berns, Joschka Strüber, Christian Beecks
SSDBM1
2020 Automatic Gaussian Process Model Retrieval for Big Data
abstract
Gaussian Process Models (GPMs) are widely regarded as a prominent tool for capturing the inherent characteristics of data. These bayesian machine learning models allow for data analysis tasks such as regression and classification. Usually a process of automatic GPM retrieval is needed to find an optimal model for a given dataset, despite prevailing default instantiations and existing prior knowledge in some scenarios, which both shortcut the way to an optimal GPM. Since non-approximative Gaussian Processes only allow for processing small datasets with low statistical versatility, we propose a new approach that allows to efficiently and automatically retrieve GPMs for large-scale data. The resulting model is composed of independent statistical representations for non-overlapping segments of the given data. Our performance evaluation of the new approach demonstrates the quality of resulting models, which clearly outperform default GPM instantiations, while maintaining reasonable model training time.
Fabian Berns, Christian Beecks
CIKM1
2020 Towards Large-scale Gaussian Process Models for Efficient Bayesian Machine Learning
abstract
275
Fabian Berns, Christian Beecks
DATA1
2019 Ptolemaic Indexing for Managing and Querying Internet of Things (IoT) Data
abstract
One of the key-enabling technologies in our digital information era is the Internet of Things (IoT). It provides a multitude of data technologies and analysis methodologies in order to connect internet-enabled devices and to collect, exchange, and analyze data snippets and information assets. It is forecasted that the majority of real-time data will be generated from devices interconnected within the Internet of Things. In order to be able to manage and access IoT data efficiently, we propose to make use of ptolemaic indexing. This domain-agnostic approach is applicable to many data-intensive IoT environments including a wide range of data types and (dis)similarity measures. In this paper, we provide an introduction to metric and ptolemaic indexing and evaluate their performance on different large-scale IoT datasets. The results of our performance evaluation indicate the high potential of our proposal for indexing and searching large-scale IoT data efficiently.
Christian Beecks, Fabian Berns, Kjeld Schmidt
IEEE BigData2
2019 A New Approach for Efficient Structure Discovery in IoT
abstract
Complex, multivariate data streams frequently comprise subjacent behavioral patterns, which are subsumable by a process of statistical structure discovery. Revealing these hidden patterns from raw data is a major challenge in abstracting information and thus for new opportunities of efficient data analysis at scale. State-of-the-art approaches, such as CKS and ABCD, leverage statistical data models and Gaussian Processes in order to abstract from raw data and to describe their major data characteristics by means of kernel-decomposed covariance functions. The process of identifying the most appropriate covariance function is a performance bottleneck due to its super-quadratic computation time complexity for model selection and evaluation. In this paper, we thus propose a new approach for the computation of large-scale statistical data models. To this end, we propose to bound the complexity of the statistical data model and develop a sequential agglomerative approach to reduce the computational load of the required evaluative calculations. Our performance analysis indicates that our proposal is able to outperform state-of-the-art kernel search algorithms such as CKS and ABCD with respect to the qualities of efficiency and accuracy.
Fabian Berns, Kjeld Schmidt, Alexander Graß, Christian Beecks
IEEE BigData1
2019 V3C1 Dataset: An Evaluation of Content Characteristics
abstract
In this work we analyze content statistics of the V3C1 dataset, which is the first partition of theVimeo Creative Commons Collection (V3C). The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, and will serve as evaluation basis for the Video Browser Showdown 2019-2021 and TREC Video Retrieval (TRECVID) Ad-Hoc Video Search tasks 2019-2021. The dataset comes with a shot segmentation (around 1 million shots) for which we analyze content specifics and statistics. Our research shows that the content of V3C1 is very diverse, has no predominant characteristics and provides a low self-similarity. Thus it is very well suited for video retrieval evaluations as well as for participants of TRECVID AVS or the VBS.
Fabian Berns, Luca Rossetto, Klaus Schöffmann, Christian Beecks, George Awad
ICMR1