EDBT 2026 Demo / reviewers in the wild / expert
Matthias Langer
dblp:63/5644
· DBLP profile ↗
8ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 57% Cloud and datacenter computing · 43% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
0.8 | 2 | 2020 | Distributed Training of Deep Learning Models: A Taxonomic Perspective · IEEE Trans. Parallel Distributed Syst. 2020 MPCA SGD - A Method for Distributed Training of Deep Learning Models on Spark · IEEE Trans. Parallel Distributed Syst. 2018 |
Machine learning › Efficient and distributed learning › distributed training › asynchronous training
asynchronous stochastic gradient descent |
0.3 | 1 | 2018 | MPCA SGD - A Method for Distributed Training of Deep Learning Models on Spark · IEEE Trans. Parallel Distributed Syst. 2018 |
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training |
0.3 | 1 | 2018 | MPCA SGD - A Method for Distributed Training of Deep Learning Models on Spark · IEEE Trans. Parallel Distributed Syst. 2018 |
Parallel and multicore computing
parallel programming models and runtimes |
0.1 | 1 | 2020 | Distributed Training of Deep Learning Models: A Taxonomic Perspective · IEEE Trans. Parallel Distributed Syst. 2020 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.1 | 1 | 2018 | MPCA SGD - A Method for Distributed Training of Deep Learning Models on Spark · IEEE Trans. Parallel Distributed Syst. 2018 |
Methods — techniques the papers use, named apart from their topics
taxonomy analysis · 0.9bulk-synchronous parallel training · 0.7MPCA SGD · 0.7EASGD · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Embedding Optimization for Training Large-scale Deep Learning Recommendation Systems with EMBarkabstractTraining large-scale deep learning recommendation models (DLRMs) with embedding tables stretching across multiple GPUs in a cluster presents a unique challenge, demanding the efficient scaling of embedding operations that require substantial memory and network bandwidth within a hierarchical network of GPUs. To tackle this bottleneck, we introduce EMBark—a comprehensive solution aimed at enhancing embedding performance and overall DLRM training throughput at scale. EMBark empowers users to create and customize sharding strategies, and features a highly-automated sharding planner, to accelerate diverse model architectures on different cluster configurations. EMBark groups embedding tables, considering their preferred communication compression method to reduce communication overheads effectively. It embraces efficient data-parallel category distribution, combined with topology-aware hierarchical communication, and pipelining support to maximize the DLRM training throughput. Across four representative DLRM variants (DLRM-DCNv2, T180, T200, and T510), EMBark achieves an average end-to-end training throughput speedup of 1.5 × and up to 1.77 × over traditional table-row-wise sharding approaches. Xavier Simmons, Matthias Langer, Minseok Lee, Zehuan Wang 0001 |
RecSys | 6 |
| 2022 | Merlin HugeCTR: GPU-accelerated Recommender System Training and InferenceabstractIn this talk, we introduce Merlin HugeCTR. Merlin HugeCTR is an open source, GPU-accelerated integration framework for click-through rate estimation. It optimizes both training and inference, whilst enabling model training at scale with model-parallel embeddings and data-parallel neural networks. In particular, Merlin HugeCTR combines a high-performance GPU embedding cache with an hierarchical storage architecture, to realize low-latency retrieval of embeddings for online model inference tasks. In the MLPerf v1.0 DLRM model training benchmark, Merlin HugeCTR achieves a speedup of up to 24.6x on a single DGX A100 (8x A100) over PyTorch on 4x4-socket CPU nodes (4x4x28 cores). Merlin HugeCTR can also take advantage of multi-node environments to accelerate training even further. Since late 2021, Merlin HugeCTR additionally features a hierarchical parameter server (HPS) and supports deployment via the NVIDIA Triton server framework, to leverage the computational capabilities of GPUs for high-speed recommendation model inference. Using this HPS, Merlin HugeCTR users can achieve a 5~62x speedup (batch size dependent) for popular recommendation models over CPU baseline implementations, and dramatically reduce their end-to-end inference latency. Zehuan Wang 0001, Yingcan Wei, Minseok Lee, Matthias Langer, Daniel G. Abel, Jianbing Dong, Kunlun Li |
RecSys | 4 |
| 2022 | A GPU-specialized Inference Parameter Server for Large-Scale Deep Recommendation ModelsabstractRecommendation systems are of crucial importance for a variety of modern apps and web services, such as news feeds, social networks, e-commerce, search, etc. To achieve peak prediction accuracy, modern recommendation models combine deep learning with terabyte-scale embedding tables to obtain a fine-grained representation of the underlying data. Traditional inference serving architectures require deploying the whole model to standalone servers, which is infeasible at such massive scale. Yingcan Wei, Matthias Langer, Minseok Lee, Zehuan Wang 0001 |
RecSys | 2 |
| 2021 | The detection, tracking, and temporal action localisation of swimmers for automated analysis
Ashley Hall, Brandon Victor, Zhen He 0002, Matthias Langer, Marc Elipot, Aiden Nibali, Stuart Morgan |
Neural Comput. Appl. | 4 |
| 2020 | Distributed Training of Deep Learning Models: A Taxonomic PerspectiveabstractDistributed deep learning systems (DDLS) train deep neural network models by utilizing the distributed resources of a cluster. Developers of DDLS are required to make many decisions to process their particular workloads in their chosen environment efficiently. The advent of GPU-based deep learning, the ever-increasing size of datasets, and deep neural network models, in combination with the bandwidth constraints that exist in cluster environments require developers of DDLS to be innovative in order to train high-quality models quickly. Comparing DDLS side-by-side is difficult due to their extensive feature lists and architectural deviations. We aim to shine some light on the fundamental principles that are at work when training deep neural networks in a cluster of independent machines by analyzing the general properties associated with training deep learning models and how such workloads can be distributed in a cluster to achieve collaborative model training. Thereby we provide an overview of the different techniques that are used by contemporary DDLS and discuss their influence and implications on the training process. To conceptualize and compare DDLS, we group different techniques into categories, thus establishing a taxonomy of distributed deep learning systems. Matthias Langer, Zhen He 0002, Wenny Rahayu, Yanbo Xue |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | MPCA SGD - A Method for Distributed Training of Deep Learning Models on SparkabstractMany distributed deep learning systems have been published over the past few years, often accompanied by impressive performance claims. In practice these figures are often achieved in high performance computing (HPC) environments with fast InfiniBand network connections. For average deep learning practitioners this is usually an unrealistic scenario, since they cannot afford access to these facilities. Simple re-implementations of algorithms such as EASGD [1] for standard Ethernet environments often fail to replicate the scalability and performance of the original works [2] . In this paper, we explore this particular problem domain and present MPCA SGD, a method for distributed training of deep neural networks that is specifically designed to run in low-budget environments. MPCA SGD tries to make the best possible use of available resources, and can operate well if network bandwidth is constrained. Furthermore, MPCA SGD runs on top of the popular Apache Spark [3] framework. Thus, it can easily be deployed in existing data centers and office environments where Spark is already used. When training large deep learning models in a gigabit Ethernet cluster, MPCA SGD achieves significantly faster convergence rates than many popular alternatives. For example, MPCA SGD can train ResNet-152 [4] up to 5.3x faster than state-of-the-art systems like MXNet [5] , up to 5.3x faster than bulk-synchronous systems like SparkNet [6] and up to 5.3x faster than decentral asynchronous systems like EASGD [1] . Matthias Langer, Ashley Hall, Zhen He 0002, Wenny Rahayu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2009 | 3D object recognition and localization employing analysis by synthesis
Matthias Langer, Lars Kuhnert, Markus Ax, Duong Nguyen Van, H. Schmick, Klaus-Dieter Kuhnert |
IADIS AC (1) | 1 |
| 2008 | A new hierarchical approach in robust real-time image feature detection and matchingabstractObject recognition forms a ubiquitous problem in digital image processing. The detection of robust image features of high distinctiveness is one important key in this regard. We present a new hierarchical approach in object recognition targeting at high robustness, yet trying to fulfill hard real-time constraints. The former will be achieved using SIFT and SURF operators, while the latter is done by employing a fast pre-processing step exploiting decision-trees. Matthias Langer, Klaus-Dieter Kuhnert |
ICPR | 1 |