Lukas Rauch

dblp:332/7293 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-6552-3270ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Uncertainty Calibration of Multi-Label Bird Sound Classifiers
abstract
4302
Raphael Schwinger, Ben McEwen, Vincent S. Kather, René Heinrich, Lukas Rauch, Sven Tomforde
ICAART (5)5
2025 BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics
abstract
Deep learning (DL) has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While AudioSet is a pivotal step to bridge this gap as a universal-domain dataset, its restricted accessibility and limited range of evaluation use cases challenge its role as the sole resource. Therefore, we introduce BirdSet, a large-scale benchmark data set for audio classification focusing on avian bioacoustics. BirdSet surpasses AudioSet with over 6,800 recording hours ($\uparrow17\%$) from nearly 10,000 classes ($\uparrow18\times$) for training and more than 400 hours ($\uparrow7\times$) across eight strongly labeled evaluation datasets. It serves as a versatile resource for use cases such as multi-label classification, covariate shift or self-supervised learning. We benchmark six well-known DL models in multi-label classification across three distinct training scenarios and outline further evaluation use cases in audio classification. We host our dataset on Hugging Face for easy accessibility and offer an extensive codebase to reproduce our results.
Lukas Rauch, Raphael Schwinger, Moritz Wirth, René Heinrich, Denis Huseljic, Marek Herde, Jonas Lange, Stefan Kahl, Bernhard Sick, Sven Tomforde, Christoph Scholz 0001
ICLR1
2025 Efficient Bayesian Updates for Deep Active Learning via Laplace Approximations
Denis Huseljic, Marek Herde, Lukas Rauch, Paul Hahn, Daniel Kottke, Stephan Vogt, Bernhard Sick
ECML/PKDD (2)3
2024 dopanim: A Dataset of Doppelganger Animals with Noisy Annotations from Multiple Humans
abstract
Human annotators typically provide annotated data for training machine learning models, such as neural networks. Yet, human annotations are subject to noise, impairing generalization performances. Methodological research on approaches counteracting noisy annotations requires corresponding datasets for a meaningful empirical evaluation. Consequently, we introduce a novel benchmark dataset, dopanim, consisting of about 15,750 animal images of 15 classes with ground truth labels. For approximately 10,500 of these images, 20 humans provided over 52,000 annotations with an accuracy of circa 67%. Its key attributes include (1) the challenging task of classifying doppelganger animals, (2) human-estimated likelihoods as annotations, and (3) annotator metadata. We benchmark well-known multi-annotator learning approaches using seven variants of this dataset and outline further evaluation use cases such as learning beyond hard class labels and active learning. Our dataset and a comprehensive codebase are publicly available to emulate the data collection process and to reproduce all empirical results.
Marek Herde, Denis Huseljic, Lukas Rauch, Bernhard Sick
NeurIPS3
2024 Fast Fishing: Approximating Bait for Efficient and Scalable Deep Active Image Classification
Denis Huseljic, Paul Hahn, Marek Herde, Lukas Rauch, Bernhard Sick
ECML/PKDD (7)4
2023 DADO - Low-Cost Query Strategies for Deep Active Design Optimization
abstract
In this work, we apply deep active learning to the field of design optimization to reduce the number of computationally expensive numerical simulations widely used in industry and engineering. We are interested in optimizing the design of structural components, where a set of parameters describes the shape. If we can predict the performance based on these parameters and consider only the promising candidates for simulation, there is an enormous potential for saving computing power. We present two query strategies for self-optimization to reduce the computational cost in multi-objective design optimization problems. Our proposed methodology provides an intuitive approach that is easy to apply, offers significant improvements over random sampling, and circumvents the need for uncertainty estimation. We evaluate our strategies on a large dataset from the domain of fluid dynamics and introduce two new evaluation metrics to determine the model's performance. Findings from our evaluation highlights the effectiveness of our query strategies in accelerating design optimization. Furthermore, the introduced method is easily transferable to other self-optimization problems in industry and engineering.
Jens Decke, Christian Gruhl, Lukas Rauch, Bernhard Sick
ICMLA3
2023 ActiveGLAE: A Benchmark for Deep Active Learning with Transformers
Lukas Rauch, Matthias Aßenmacher, Denis Huseljic, Moritz Wirth, Bernd Bischl, Bernhard Sick
ECML/PKDD (1)1