VLDB 2026 Research / reviewers in the wild / expert
Herbert Woisetschlaeger
dblp:346/0243 · also Herbert Woisetschläger
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-9729-2895ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 24% Language models and text generation · 22% Multi-agent systems · 20% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems › normative multi-agent systems › social norms
convention formation |
1.0 | 1 | 2026 | SIGN: Schema Induced Games for Naming (Student Abstract) · AAAI 2026 |
Natural language and speech › Language models and text generation
LLM agents |
1.0 | 1 | 2026 | SIGN: Schema Induced Games for Naming (Student Abstract) · AAAI 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent communication |
1.0 | 1 | 2026 | SIGN: Schema Induced Games for Naming (Student Abstract) · AAAI 2026 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination |
1.0 | 1 | 2026 | SIGN: Schema Induced Games for Naming (Student Abstract) · AAAI 2026 |
Machine learning › Efficient and distributed learning
data reweighting |
0.9 | 1 | 2025 | Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.9 | 1 | 2025 | Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining · ICLR 2025 |
Machine learning › Deep learning architectures and training › loss function design
loss weighting |
0.9 | 1 | 2025 | Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining · ICLR 2025 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.9 | 1 | 2025 | MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees · NeurIPS 2025 |
Cloud and datacenter computing › inference serving
LLM serving |
0.9 | 1 | 2025 | MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees · NeurIPS 2025 |
Cloud and datacenter computing › datacenter services › online service systems
request routing |
0.9 | 1 | 2025 | MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees · NeurIPS 2025 |
Cloud and datacenter computing
serverless computing |
0.9 | 1 | 2025 | MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
federated learning |
0.8 | 1 | 2024 | A Survey on Efficient Federated Learning Methods for Foundation Model Training · IJCAI 2024 |
Machine learning › Deep learning architectures and training › foundation model
foundation model training |
0.8 | 1 | 2024 | A Survey on Efficient Federated Learning Methods for Foundation Model Training · IJCAI 2024 |
Machine learning › Efficient and distributed learning
dataset distillation |
0.7 | 1 | 2023 | A Survey on Dataset Distillation: Approaches, Applications and Future Directions · IJCAI 2023 |
Machine learning › Optimization for machine learning
convergence analysis |
0.3 | 1 | 2025 | Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
foundation model |
0.2 | 1 | 2024 | A Survey on Efficient Federated Learning Methods for Foundation Model Training · IJCAI 2024 |
Machine learning › Trustworthy machine learning
privacy and data protection |
0.2 | 1 | 2023 | A Survey on Dataset Distillation: Approaches, Applications and Future Directions · IJCAI 2023 |
Methods — techniques the papers use, named apart from their topics
stochastic optimization · 1.7request satisfaction prediction · 1.7schema-induced communication · 1.0virtual queues · 0.9virtual queue · 0.9gradient-based optimization · 0.9dynamic instance-level reweighting · 0.9federated learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SIGN: Schema Induced Games for Naming (Student Abstract)abstractReal-world AI systems are tackling increasingly complex problems, often through interactions among Large Language Model (LLM) agents. When these agents develop inconsistent conventions, coordination can break down. Applications such as collaborative coding and distributed planning therefore require reliable, consistent communication, and scalability is a central concern as systems grow. We introduce Schema-Induced Games for Naming (SIGN), a naming game that examines how lightweight structure can steer convention formation. We compare schema-induced communication to unconstrained natural language and find faster convergence with up to 5.8× higher agreement. These results suggest that minimal structure can act as a simple control knob for efficient multi-agent coordination, pointing toward broader applications beyond the naming game. Ryan Zhang, Herbert Woisetschlaeger |
AAAI | 2 |
| 2025 | Dynamic Loss-Based Sample Reweighting for Improved Large Language Model PretrainingabstractPretraining large language models (LLMs) on vast and heterogeneous datasets is crucial for achieving state-of-the-art performance across diverse downstream tasks. However, current training paradigms treat all samples equally, overlooking the importance or relevance of individual samples throughout the training process. Existing reweighting strategies, which primarily focus on group-level data importance, fail to leverage fine-grained instance-level information and do not adapt dynamically to individual sample importance as training progresses. In this paper, we introduce novel algorithms for dynamic, instance-level data reweighting aimed at improving both the efficiency and effectiveness of LLM pretraining. Our methods adjust the weight of each training sample based on its loss value in an online fashion, allowing the model to dynamically focus on more informative or important samples at the current training stage. In particular, our framework allows us to systematically devise reweighting strategies deprioritizing redundant or uninformative data, which we find tend to work best.
Furthermore, we develop a new theoretical framework for analyzing the impact of loss-based reweighting on the convergence of gradient-based optimization, providing the first formal characterization of how these strategies affect convergence bounds. We empirically validate our approach across a spectrum of tasks, from pretraining 7B and 1.4B parameter LLMs to smaller-scale language models and linear regression problems, demonstrating that our loss-based reweighting approach can lead to faster convergence and significantly improved performance. Daouda Sow, Herbert Woisetschlaeger, Saikiran Bulusu, Shiqiang Wang 0001, Hans-Arno Jacobsen, Yingbin Liang |
ICLR | 2 |
| 2025 | MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level GuaranteesabstractOpen-weight large language model (LLM) zoos provide access to numerous high-quality models, but selecting the appropriate model for specific tasks remains challenging and requires technical expertise. Most users simply want factually correct, safe, and satisfying responses without concerning themselves with model technicalities, while inference service providers prioritize minimizing operating costs. These competing interests are typically mediated through service level agreements (SLAs) that guarantee minimum service quality.
We introduce MESS+, a stochastic optimization algorithm for cost-optimal LLM request routing while providing rigorous SLA compliance guarantees. MESS+ learns request satisfaction probabilities of LLMs in real-time as users interact with the system, based on which model selection decisions are made by solving a per-request optimization problem. Our algorithm includes a novel combination of virtual queues and request satisfaction prediction, along with a theoretical analysis of cost optimality and constraint satisfaction.
Across a wide range of state-of-the-art LLM benchmarks, MESS+ achieves an average of $2\times$ cost savings compared to existing LLM routing techniques. Herbert Woisetschlaeger, Ryan Zhang, Shiqiang Wang 0001, Hans-Arno Jacobsen |
NeurIPS | 1 |
| 2024 | A Survey on Efficient Federated Learning Methods for Foundation Model Training
Herbert Woisetschlaeger, Alexander Erben, Shiqiang Wang 0001, Ruben Mayer, Hans-Arno Jacobsen |
IJCAI | 1 |
| 2024 | FLEdge: Benchmarking Federated Learning Applications in Edge Computing SystemsabstractFederated Learning (FL) has become a viable technique for realizing privacy-enhancing distributed deep learning on the network edge. Heterogeneous hardware, unreliable client devices, and energy constraints often characterize edge computing systems. In this paper, we propose FLEdge, which complements existing FL benchmarks by enabling a systematic evaluation of client capabilities. We focus on computational and communication bottlenecks, client behavior, and data security implications. Our experiments with models varying from 14K to 80M trainable parameters are carried out on dedicated hardware with emulated network characteristics and client behavior. We find that state-of-the-art embedded hardware has significant memory bottlenecks, leading to 4× longer processing times than on modern data center GPUs. Herbert Woisetschlaeger, Alexander Erben, Ruben Mayer, Shiqiang Wang 0001, Hans-Arno Jacobsen |
Middleware | 1 |
| 2023 | A Survey on Dataset Distillation: Approaches, Applications and Future DirectionsabstractDataset distillation is attracting more attention in machine learning as training sets continue to grow and the cost of training state-of-the-art models becomes increasingly high. By synthesizing datasets with high information density, dataset distillation offers a range of potential applications, including support for continual learning, neural architecture search, and privacy protection. Despite recent advances, we lack a holistic understanding of the approaches and applications. Our survey aims to bridge this gap by first proposing a taxonomy of dataset distillation, characterizing existing approaches, and then systematically reviewing the data modalities, and related applications. In addition, we summarize the challenges and discuss future directions for this field of research. Jiahui Geng, Zongxiong Chen, Yuandou Wang, Herbert Woisetschlaeger, Sonja Schimmler, Ruben Mayer, Zhiming Zhao, Chunming Rong |
IJCAI | 4 |