EDBT 2026 Demo / reviewers in the wild / expert
Sophie Fellenz
dblp:165/3605 · also Sophie Burkhardt
· DBLP profile ↗
13ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-5385-3926ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 39% Information extraction and text analysis · 20% Generative modeling · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 87% Information retrieval · 13% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
topic model |
1.6 | 3 | 2024 | Putting Back the Stops: Integrating Syntax with Neural Topic Models · IJCAI 2024 Evaluating Dynamic Topic Models · ACL (1) 2024 Decoupling Sparsity and Smoothness in the Dirichlet Variational Autoencoder Topic Model · J. Mach. Learn. Res. 2019 |
Machine learning › Trustworthy machine learning › interpretability › explainable AI
anomaly explanation |
1.0 | 1 | 2026 | Reimagining Anomalies: What If Anomalies Were Normal? · AAAI 2026 |
Machine learning › Generative modeling › diffusion model
diffusion planning |
1.0 | 1 | 2026 | TORA: Train Once, Realign Anytime for Offline Multi-Objective Reinforcement Learning · AAAI 2026 |
Machine learning › Trustworthy machine learning
interpretability |
1.0 | 1 | 2026 | Reimagining Anomalies: What If Anomalies Were Normal? · AAAI 2026 |
Machine learning › Reinforcement learning
multi-objective reinforcement learning |
1.0 | 1 | 2026 | TORA: Train Once, Realign Anytime for Offline Multi-Objective Reinforcement Learning · AAAI 2026 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.9 | 1 | 2025 | Mitigating Spurious Features in Contrastive Learning with Spectral Regularization · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Mitigating Spurious Features in Contrastive Learning with Spectral Regularization · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › robustness
spurious correlation |
0.9 | 1 | 2025 | Mitigating Spurious Features in Contrastive Learning with Spectral Regularization · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › debiasing
spurious feature mitigation |
0.9 | 1 | 2025 | Mitigating Spurious Features in Contrastive Learning with Spectral Regularization · NeurIPS 2025 |
Data mining
anomaly detection |
0.9 | 1 | 2025 | NoBOOM: Chemical Process Datasets for Industrial Anomaly Detection · NeurIPS 2025 |
Data mining › anomaly detection
time series anomaly detection |
0.9 | 1 | 2025 | NoBOOM: Chemical Process Datasets for Industrial Anomaly Detection · NeurIPS 2025 |
Natural language and speech › Information extraction and text analysis › topic model
neural topic model |
0.8 | 1 | 2024 | Putting Back the Stops: Integrating Syntax with Neural Topic Models · IJCAI 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › prior modeling
dirichlet prior |
0.4 | 1 | 2019 | Decoupling Sparsity and Smoothness in the Dirichlet Variational Autoencoder Topic Model · J. Mach. Learn. Res. 2019 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.3 | 1 | 2026 | TORA: Train Once, Realign Anytime for Offline Multi-Objective Reinforcement Learning · AAAI 2026 |
Machine learning › Time series and sequential data › anomaly detection
visual anomaly detection |
0.3 | 1 | 2026 | Reimagining Anomalies: What If Anomalies Were Normal? · AAAI 2026 |
Information retrieval › evaluation
benchmark dataset |
0.3 | 1 | 2025 | NoBOOM: Chemical Process Datasets for Industrial Anomaly Detection · NeurIPS 2025 |
Natural language and speech › Language models and text generation
text representation |
0.2 | 1 | 2024 | Putting Back the Stops: Integrating Syntax with Neural Topic Models · IJCAI 2024 |
Methods — techniques the papers use, named apart from their topics
preference integration · 1.0diffusion planning · 1.0deep anomaly detection · 1.0counterfactual generation · 1.0time series analysis · 0.9spectral regularization · 0.9neural topic model · 0.8variational inference · 0.4reparameterization · 0.4rejection sampling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TORA: Train Once, Realign Anytime for Offline Multi-Objective Reinforcement LearningabstractIntelligent agents in real-world applications must adapt their behavior to changing contexts and user preferences. For example, planning a road trip requires considering both travel time and cost. Multi-objective reinforcement learning (MORL) provides a principled approach to navigate such trade-offs. However, most existing approaches require predefined preference weights during training and jointly optimize the model for all objectives. In this paper, we introduce TORA (Train Once, Realign Anytime), a novel framework that defers preference integration to inference time, enabling flexible adaptation to user preferences without retraining. TORA independently trains diffusion planning models for each objective and combines them at inference time using user-specified preferences to generate behavior aligned with desired trade-offs. Furthermore, new objectives can be added seamlessly by training additional models without modifying existing ones. Empirical evaluations on standard offline MORL benchmarks demonstrate that TORA achieves competitive and consistent performance compared to methods that require fixed preference weights. Waleed Mustafa, Marcio Monteiro, Puyu Wang, Marius Kloft, Sophie Fellenz |
AAAI | 6 |
| 2026 | Reimagining Anomalies: What If Anomalies Were Normal?abstractDeep learning-based methods have achieved a breakthrough in image anomaly detection, but their complexity introduces a considerable challenge to understanding why an instance is predicted to be anomalous. We introduce a novel explanation method that generates multiple alternative modifications for each anomaly, capturing diverse concepts of anomalousness. Each modification is trained to be perceived as normal by the anomaly detector. The method provides a semantic explanation of the mechanism that triggered the detector, allowing users to explore ``what-if scenarios.'' Qualitative and quantitative analyses across various image datasets demonstrate that applying this method to state-of-the-art detectors provides high-quality semantic explanations. Philipp Liznerski, Saurabh Varshneya, Ece Calikus, Puyu Wang, Alexander Bartscher, Sebastian J. Vollmer, Sophie Fellenz, Marius Kloft |
AAAI | 7 |
| 2025 | Mitigating Spurious Features in Contrastive Learning with Spectral RegularizationabstractNeural networks generally prefer simple and easy-to-learn features. When these features are spuriously correlated with the labels, the network's performance can suffer, particularly for underrepresented classes or concepts. Self-supervised representation learning methods, such as contrastive learning, are especially prone to this issue, often resulting in worse performance on downstream tasks.
We identify a key spectral signature of this failure: early reliance on dominant singular modes of the learned feature matrix. To mitigate this, we propose a novel framework that promotes a uniform eigenspectrum of the feature covariance matrix, encouraging diverse and semantically rich representations. Our method operates in a fully self-supervised setting, without relying on ground-truth labels or any additional information. Empirical results on SimCLR and SimSiam demonstrate consistent gains in robustness and transfer performance, suggesting broad applicability across self-supervised learning paradigms. Code: https://github.com/NaghmehGh/SpuriousCorrelation_SSRL Naghmeh Ghanooni, Waleed Mustafa, Dennis Wagner, Sophie Fellenz, Anthony Widjaja Lin, Marius Kloft |
NeurIPS | 4 |
| 2025 | NoBOOM: Chemical Process Datasets for Industrial Anomaly DetectionabstractMonitoring chemical processes is essential to prevent catastrophic failures, optimize costs and profits, and ensure the safety of employees and the environment. A key component of modern monitoring systems is the automated detection of anomalies in sensor data over time, called time series, enabling partial automation of plant operation and adding additional layers of supervision to crucial components. The development of anomaly detection methods in this domain is challenging, since real chemical process data are usually proprietary, and simulated data are generally not a sufficient replacement. In this paper, we present NoBOOM, the first collection of datasets for anomaly detection in real-world chemical process data, including labeled data from a running process at our industry partner BASF SE — one of the world’s leading chemical companies — and several chemical processes run in laboratory‑scale and pilot‑scale plants. While we are not able to share every detail about the industrial process, for the laboratory‑ and pilot‑scale plants, we provide comprehensive information on plant configuration, process operation, and, in particular, anomaly events, enabling a differentiated analysis of anomaly detection methods. To demonstrate the complexity of the benchmark, we analyze the data with regard to common issues of time-series anomaly detection (TSAD) benchmarks, including potential triviality and bias. Dennis Wagner, Fabian Hartung, Justus Arweiler, Aparna Muraleedharan, Indra Jungjohann, Arjun Nair, Steffen Reithermann, Ralf Schulz, Michael Bortz, Daniel Neider, Heike Leitte, Joachim Pfeffinger, Stephan Mandt, Sophie Fellenz, Torsten Katz, Fabian Jirasek, Jakob Burger, Hans Hasse, Marius Kloft |
NeurIPS | 14 |
| 2024 | Evaluating Dynamic Topic ModelsabstractCharu Karakkaparambil James, Mayank Nagda, Nooshin Haji Ghassemi, Marius Kloft, Sophie Fellenz. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Charu James, Mayank Nagda, Nooshin Haji Ghassemi, Marius Kloft, Sophie Fellenz |
ACL (1) | 5 |
| 2024 | Ethics in Action: Training Reinforcement Learning Agents for Moral Decision-making In Text-based Adventure GamesabstractReinforcement Learning (RL) has demonstrated its potential in solving goal-oriented sequential tasks. However, with the increasing capabilities of RL agents, ensuring morally responsible agent behavior is becoming a pressing concern. Previous approaches have included moral considerations by statically assigning a moral score to each action at runtime. However, these methods do not account for the potential moral value of future states when evaluating immoral actions. This limits the ability to find trade-offs between different aspects of moral behavior and the utility of the action. In this paper, we aim to factor in moral scores by adding a constraint to the RL objective that is incorporated during training, thereby dynamically adapting the policy function. By combining Lagrangian optimization and meta-gradient learning, we develop an RL method that is able to find a trade-off between immoral behavior and performance in the decision-making process. Rati Devidze, Waleed Mustafa, Sophie Fellenz |
AISTATS | 4 |
| 2024 | Text Style Transfer Evaluation Using Large Language ModelsabstractEvaluating Text Style Transfer (TST) is a complex task due to its multi-faceted nature. The quality of the generated text is measured based on challenging factors, such as style transfer accuracy, content preservation, and overall fluency. While human evaluation is considered to be the gold standard in TST assessment, it is costly and often hard to reproduce. Therefore, automated metrics are prevalent in these domains. Nonetheless, it is uncertain whether and to what extent these automated metrics correlate with human evaluations. Recent strides in Large Language Models (LLMs) have showcased their capacity to match and even exceed average human performance across diverse, unseen tasks. This suggests that LLMs could be a viable alternative to human evaluation and other automated metrics in TST evaluation. We compare the results of different LLMs in TST evaluation using multiple input prompts. Our findings highlight a strong correlation between (even zero-shot) prompting and human evaluation, showing that LLMs often outperform traditional automated metrics. Furthermore, we introduce the concept of prompt ensembling, demonstrating its ability to enhance the robustness of TST evaluation. This research contributes to the ongoing efforts for more robust and diverse evaluation methods by standardizing and validating TST evaluation with LLMs. Phil Ostheimer, Mayank Nagda, Marius Kloft, Sophie Fellenz |
LREC/COLING | 4 |
| 2024 | Putting Back the Stops: Integrating Syntax with Neural Topic Models
Mayank Nagda, Sophie Fellenz |
IJCAI | 2 |
| 2023 | Learning to Play Text-Based Adventure Games with Maximum Entropy Reinforcement Learning
Rati Devidze, Sophie Fellenz |
ECML/PKDD (4) | 3 |
| 2019 | Decoupling Sparsity and Smoothness in the Dirichlet Variational Autoencoder Topic ModelabstractRecent work on variational autoencoders (VAEs) has enabled the development of generative topic models using neural networks. Topic models based on latent Dirichlet allocation (LDA) successfully use the Dirichlet distribution as a prior for the topic and word distributions to enforce sparseness. However, there is a trade-off between sparsity and smoothness in Dirichlet distributions. Sparsity is important for a low reconstruction error during training of the autoencoder, whereas smoothness enables generalization and leads to a better log-likelihood of the test data. Both of these properties are encoded in the Dirichlet parameter vector. By rewriting this parameter vector into a product of a sparse binary vector and a smoothness vector, we decouple the two properties, leading to a model that features both a competitive topic coherence and a high log-likelihood. Efficient training is enabled using rejection sampling variational inference for the reparameterization of the Dirichlet distribution. Our experiments show that our method is competitive with other recent VAE topic models. Sophie Fellenz, Stefan Kramer 0001 |
J. Mach. Learn. Res. | 1 |
| 2019 | Multi-label classification using stacked hierarchical Dirichlet processes with reduced sampling complexity
Sophie Fellenz, Stefan Kramer 0001 |
Knowl. Inf. Syst. | 1 |
| 2018 | Online multi-label dependency topic models for text classification
Sophie Fellenz, Stefan Kramer 0001 |
Mach. Learn. | 1 |
| 2017 | Online Sparse Collapsed Hybrid Variational-Gibbs Algorithm for Hierarchical Dirichlet Process Topic Models
Sophie Fellenz, Stefan Kramer 0001 |
ECML/PKDD (2) | 1 |