VLDB 2026 Research / reviewers in the wild / expert
David Berghaus
dblp:331/8339
· DBLP profile ↗
8ranked-venue papers in the field
1as first author
8since 2021 · last 2025
0000-0002-0740-154XORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 6 (1 first)Data Mining & Knowledge Discovery · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processingabstract2519 David Berghaus, Armin Berger, Lars Patrick Hillebrand, Kostadin Cvejoski, Rafet Sifa |
IEEE Big Data | 1 |
| 2025 | History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting
Sarthak Khanna, Armin Berger, Muskaan Chopra, David Berghaus, Rafet Sifa |
IEEE Big Data | 4 |
| 2025 | Reasoning LLMs in the Medical Domain: A Literature SurveyabstractThe emergence of advanced reasoning capabilities in Large Language Models (LLMs) marks a transformative development in healthcare applications. Beyond merely expanding functional capabilities, these reasoning mechanisms enhance decision transparency and explainability-critical requirements in medical contexts. This survey examines the transformation of medical LLMs from basic information retrieval tools to sophisticated clinical reasoning systems capable of supporting complex healthcare decisions. We provide a thorough analysis of the enabling technological foundations, with a particular focus on specialized prompting techniques like Chain-of-Thought and recent breakthroughs in Reinforcement Learning exemplified by DeepSeek-R1. Our investigation evaluates purpose-built medical frameworks while also examining emerging paradigms such as multi-agent collaborative systems and innovative prompting architectures. The survey critically assesses current evaluation methodologies for medical validation and addresses persistent challenges in field interpretation limitations, bias mitigation strategies, patient safety frameworks, and integration of multimodal clinical data. Through this survey, we seek to establish a roadmap for developing reliable LLMs that can serve as effective partners in clinical practice and medical research. Armin Berger, Sarthak Khanna, Lorenz Sparrenberg, Tobias Deußer, David Berghaus, Rafet Sifa |
DSAA | 5 |
| 2025 | Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal AttentionabstractWe propose STONK (Stock Optimization using News Knowledge), a multimodal framework integrating numerical market indicators with sentiment-enriched news embeddings to improve daily stock-movement prediction. By combining numerical & textual embeddings via feature concatenation and cross-modal attention, our unified pipeline addresses limitations of isolated analyses. Backtesting shows STONK outperforms numeric-only baselines. A comprehensive evaluation of fusion strategies and model configurations offers evidence-based guidance for scalable multimodal financial forecasting. Source code is available on GitHub11https://github.com/sarthak-12/thesis-dsaa/. Sarthak Khanna, Armin Berger, David Berghaus, Tobias Deußer, Lorenz Sparrenberg, Rafet Sifa |
DSAA | 3 |
| 2024 | Fine-Tuning Large Language Models for Compliance ChecksabstractThe auditing of financial documents, traditionally a labor-intensive task, is a promising field of application for Artificial Intelligence. Recommendation systems are capable of suggesting the most relevant passages from financial reports that meet accounting standards’ legal requirements. However, testing if the compliance requirements are satisfied is a non-trivial task. In this work, we tackle this problem from two directions. Our first approach leverages Large Language Models which we fine-tune specifically f or compliance checks. Our results show an improvement in performance over the generic baseline LLMs. A disadvantage of LLMs is that they result in high inference costs. For this reason, we explore a second approach in which we use smaller models that come with reduced running costs. Despite their smaller size, these models also show promising predictive performance. Thiago Bell, David Leonhard, Ali Hamza Bashir, Tim Dilmaghani Khameneh, Mohamed Khaled, Ulrich Warning, Rüdiger Loitz, Sandra Halscheidt, Jana Birr, Armin Berger, Rafet Sifa, David Berghaus |
IEEE Big Data | 12 |
| 2024 | Advancing Personalized Medicine: A Scalable LLM-based Recommender System for Patient MatchingabstractThis study explores efficient algorithms to enhance user matching in Unrare.me, a novel social networking platform designed to connect individuals affected by rare diseases. Our primary objective is to develop a recommender system that identifies and suggests users with similar medical conditions, facilitating meaningful connections within these unique communities. Utilizing textual user profile data, we train sentence embedder models to generate similar embeddings for users that have rated each other high. We investigate various fine-tuning strategies, as well as a hybrid approach between a dense embedder and sparse SPLADE embeddings. Furthermore, we investigate the efficacy of various clustering algorithms, such as TopicBERT for thematic analysis, K-Means for centroid-based grouping, and Latent Dirichlet Allocation (LDA) for probabilistic topic modeling, to reduce the matching complexity and enable better scalability of the platform. Armin Berger, David Berghaus, Ali Hamza Bashir, Lorenz Grigull, Lara Fendrich, Tom Anglim Lagones, Henriette Högl, Gundula Ernst, David Bascom, Tobias Deußer, Thiago Bell, Max Lübbering, Rafet Sifa |
IEEE Big Data | 2 |
| 2024 | Optimizing Rare Disease Patient Matching with Large Language ModelsabstractWe present RepLLaMA, a neural ranking model for optimizing patient matching in rare disease communities. Using data from Unrare.me consisting of over two thousand profiles and over ten thousand ratings, our bi-encoder architecture maps profiles to 4096-dimensional vectors, enabling efficient similarity computations. The system processes unstructured symptom descriptions and structured responses, incorporating expert-guided LLM enhancements. Results show Top-10 Recall of 49.36%$(\pm 2.03)$, surpassing baselines while maintaining generalization. The implementation provides a scalable solution for rare disease patient matching, addressing computational complexity challenges. Armin Berger, Ali Hamza Bashir, David Berghaus, Mowmita, Nazia Afsan, Lorenz Grigull, Lara Fendrich, Henriette Högl, Gundula Ernst, David Bascom, Tom Anglim Lagones, Tobias Deußer, Thiago Bell, Max Lübbering, Rafet Sifa |
IEEE Big Data | 3 |
| 2024 | Advancing Risk and Quality Assurance: A RAG Chatbot for Improved Regulatory ComplianceabstractRisk and Quality (R&Q) assurance in highly regulated industries requires constant navigation of complex regulatory frameworks, with employees handling numerous daily queries demanding accurate policy interpretation. Traditional methods relying on specialized experts create operational bottlenecks and limit scalability. We present a novel Retrieval Augmented Generation (RAG) system leveraging Large Language Models (LLMs), hybrid search and relevance boosting to enhance R&Q query processing. Evaluated on 124 expert-annotated real-world queries, our actively deployed system demonstrates substantial improvements over traditional RAG approaches. Additionally, we perform an extensive hyperparameter analysis to compare and evaluate multiple configuration setups, delivering valuable insights to practitioners. Lars Patrick Hillebrand, Armin Berger, Daniel Uedelhoven, David Berghaus, Ulrich Warning, Tim Dilmaghani Khameneh, Bernd Kliem, Rüdiger Loitz, Rafet Sifa |
IEEE Big Data | 4 |