EDBT 2026 Demo / reviewers in the wild / expert
Thomas Lippert
dblp:13/5159
· DBLP profile ↗
24ranked-venue papers
6as first author
4since 2021 · last 2026
0000-0002-9407-6043ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Universal quantum computer simulation of 50 qubits on Europe's first exascale supercomputer harnessing its heterogeneous CPU-GPU architectureabstractWe have developed a new version of the high-performance Jülich universal quantum computer simulator (JUQCS-50) that leverages key features of the GH200 superchips as used in the JUPITER supercomputer, enabling simulations of a 50-qubit universal quantum computer for the first time. JUQCS-50 achieves this through three key innovations: (1) extending usable memory beyond GPU limits via high-bandwidth CPU–GPU interconnects and LPDDR5 memory; (2) adaptive data encoding to reduce memory footprint with acceptable trade-offs in precision and compute effort; and (3) an on-the-fly network traffic optimizer. These advances result in an 16.6-fold speedup over the previous 48-qubit record on the K computer. Hans De Raedt, Jiri Kraus, Andreas Herten, Vrinda Mehta, Mathis Bode, Markus Hrywniak, Kristel Michielsen, Thomas Lippert |
Future Gener. Comput. Syst. | 8 |
| 2024 | 3D DFT by block tensor-matrix multiplication via a modified Cannon's algorithm: Implementation and scaling on distributed-memory clusters with fat tree networksabstractA known scalability bottleneck of the parallel 3D FFT is its use of all-to-all communications. Here, we present S3DFT, a library that circumvents this by using point-to-point communication – albeit at a higher arithmetic complexity. This approach exploits three variants of Cannon's algorithm with adaptations for block tensor-matrix multiplications. We demonstrate S3DFT's efficient use of hardware resources, and its scaling using up to 16,464 cores of the JUWELS Cluster. However, in a comparison with well-established 3D FFT libraries, its parallel efficiency and performance were found to fall behind. A detailed analysis identifies the cause in two of its component algorithms, which scale poorly owing to how their communication patterns are mapped in subsets of the fat tree topology. This result exposes a potential drawback of running block-wise parallel algorithms on systems with fat tree networks caused by increased communication latencies along specific directions of the mesh of processing elements. Nitin Malapally, Viacheslav Bolnykh, Estela Suarez, Paolo Carloni, Thomas Lippert, Davide Mandelli |
J. Parallel Distributed Comput. | 5 |
| 2022 | A Comprehensive I/O Knowledge Cycle for Modular and Automated HPC Workload AnalysisabstractOn the way to the exascale era, millions of parallel processing elements are required. Accordingly, one major chal-lenge is the ever-widening gap between computational power and underlying I/O systems. To bridge this gap, I/O resources must be used efficiently, thus a profound I/O knowledge is required. In this work, we analyze state-of-the-art approaches that can be applied to improve the general I/O understanding and performance. Based on our analysis, we present an automated, modular, tool-agnostic I/Oanalysis workflow and a prototype implementation that can be used to generate, extract, store, analyze, and use I/O knowledge in a structured and reproducible way. Zhaobin Zhu, Sarah Neuwirth, Thomas Lippert |
CLUSTER | 3 |
| 2022 | Hybrid Quantum-Classical Workflows in Modular Supercomputing Architectures with the Julich Unified Infrastructure for Quantum ComputingabstractThe implementation of scalable processing workflows is essential to improve the access to and analysis of the vast amount of high-resolution and multi-source Remote Sensing (RS) data and to provide decision-makers with timely and valuable information. The Modular Supercomputing Architecture (MSA) systems that are operated by the Jülich Supercomputing Centre (JSC) are a concrete solution for data-intensive RS applications that rely on big data storage and processing capabilities. To meet the requirements of applications with more complex computational tasks, JSC plans to connect the High Performance Computing (HPC) systems of its MSA environment to different quantum computers via the Jülich UNified Infrastructure for Quantum computing (JUNIQ). The paper describes this unique computing environment and highlights its potential to address real RS application scenarios through high-performance and hybrid quantum-classical processing workflows. Gabriele Cavallaro, Morris Riedel, Thomas Lippert, Kristel Michielsen |
IGARSS | 3 |
| 2016 | The DEEP Project An alternative approach to heterogeneous cluster-computing in the many-core eraabstractSummary Homogeneous cluster architectures, which used to dominate high‐performance computing (HPC), are challenged today by heterogeneous approaches utilizing accelerator or co‐processor devices. The DEEP (Dynamical Exascale Entry Platform) project is implementing a novel architecture for HPC, in which a standard HPC cluster is directly connected to a so‐called ‘Booster’: a cluster of many‐core processors. By these means heterogeneity is organized differently as in today's standard approach, where accelerators are added to each node of the cluster. In order to adapt application codes to this Cluster‐Booster architecture as seamless as possible, DEEP has developed a complete programming environment. It integrates the offloading functionality given by the Message Passing Interface standard with an abstraction layer based on the task‐based OmpSs programming paradigm. This paper presents the DEEP project with an emphasis on the DEEP programming environment. Copyright © 2015 John Wiley & Sons, Ltd. Norbert Eicker, Thomas Lippert, Thomas Moschny, Estela Suarez |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | HPC for the human brain projectabstractThe Human Brain Project, one of two European flagship projects, is a collaborative effort to reconstruct the brain, piece by piece, in multi-scale models and their supercomputer-based simulation, integrating and federating giant amounts of existing information and creating new information and knowledge about the human brain. A fundamental impact on our understanding of the human brain and its diseases as well as on novel brain-inspired computing technologies is expected. Thomas Lippert |
ICS | 1 |
| 2014 | Smart data analytics methods for remote sensing applicationsabstractThe big data analytics approach emerged that can be interpreted as extracting information from large quantities of scientific data in a systematic way. In order to have a more concrete understanding of this term we refer to its refinement as smart data analytics in order to examine large quantities of scientific data to uncover hidden patterns, unknown correlations, or to extract information in cases where there is no exact formula (e.g. known physical laws). Our concrete big data problem is the classification of classes of land cover types in image-based datasets that have been created using remote sensing technologies, because the resolution can be high (i.e. large volumes) and there are various types such as panchromatic or different used bands like red, green, blue, and nearly infrared (i.e. large variety). We investigate various smart data analytics methods that take advantage of machine learning algorithms (i.e. support vector machines) and state-of-the-art parallelization approaches in order to overcome limitations of big data processing using non-scalable serial approaches. Gabriele Cavallaro, Morris Riedel, Jón Atli Benediktsson, Markus Götz, Tomas Runarsson, Kristjan Jonasson, Thomas Lippert |
IGARSS | 7 |
| 2014 | Advancements of the UltraScan scientific gateway for open standards-based cyberinfrastructuresabstractSUMMARY The UltraScan data analysis application is a software package that is able to take advantage of computational resources in order to support the interpretation of analytical ultracentrifugation experiments. Since 2006, the UltraScan scientific gateway has been used with Web browsers in TeraGrid by scientists studying the solution properties of biological and synthetic molecules. UltraScan supports its users with a scientific gateway in order to leverage the power of supercomputing. In this contribution, we will focus on several advancements of the UltraScan scientific gateway architecture with a standardized job management while retaining its lightweight design and end user interaction experience. This paper also presents insights into a production deployment of UltraScan in Europe. The approach is based on open standards with respect to job management and submissions to the Extreme Science and Engineering Discovery Environment in the USA and to similar infrastructures in Europe such as the European Grid Infrastructure or the Partnership for Advanced Computing in Europe (PRACE). Our implementation takes advantage of the Apache Airavata framework for scientific gateways that lays the foundation for easy integration into several other scientific gateways. Copyright © 2014 John Wiley & Sons, Ltd. M. Shahbaz Memon, Morris Riedel, Florian Janetzko, Borries Demeler, Gary Gorbet, Suresh Marru, Andrew S. Grimshaw, Lahiru Gunathilake, Raminderjeet Singh, Norbert Attig, Thomas Lippert |
Concurr. Comput. Pract. Exp. | 11 |
| 2013 | The DEEP Project - Pursuing Cluster-Computing in the Many-Core EraabstractHomogeneous cluster architectures dominating high-performance computing (HPC) today are challenged, in particular when thinking about reaching Exascale by the end of the decade, by heterogeneous approaches utilizing accelerator elements. The DEEP (Dynamical Exascale Entry Platform) project aims for implementing a novel architecture for high-performance computing consisting of two components - a standard HPC Cluster and a cluster of many-core processors called Booster. In order to make the adaptation of application codes to this Cluster-Booster architecture as seamless as possible, DEEP provides a complete programming environment. It integrates the offloading functionality given by the MPI standard with an abstraction layer based on the task-based OmpSs programming paradigm. This paper presents the DEEP project with an emphasis on the DEEP programming environment. Norbert Eicker, Thomas Lippert, Thomas Moschny, Estela Suarez |
ICPP | 2 |
| 2010 | Exploring the Potential of Using Multiple E-science Infrastructures with Emerging Open Standards-Based E-health Research ToolsabstractE-health makes use of information and communication methods and the latest e-research tools to support the understanding of body functions. E-scientists in this field take already advantage of one single infrastructure to perform computationally-intensive investigations of the human body that tend to consider each of the constituent parts separately without taking into account the multiple important interactions between them. But these important interactions imply an increasing complexity of applications that embrace multiple physical models (i.e. multi-physics) and consider a larger range of scales (i.e. multi-scale) thus creating a steadily growing demand for interoperable infrastructures that allow for new innovative application types of jointly using different infrastructures for one application. But interoperable infrastructures are still not seamlessly provided and we argue that this is due to the absence of a realistically implementable infrastructure interoperability reference model that is based on lessons learned from e-science usage. Therefore, the goal of this paper is to explore the potential of using multiple infrastructures for one scientific goal with a particular focus on e-health. Since e-scientists gain more interest in using multiple infrastructures there is a clear demand for interoperability between them to enable a use with one e-research tool. The paper highlights work in the context of an e-Health blood flow application while the reference model is applicable to other e-science applications as well. Morris Riedel, Bernd Schuller, Michael Rambadt, M. Shahbaz Memon, Ahmed Shiraz Memon, Achim Streit, Thomas Lippert, Stefan J. Zasada, Steven Manos, Peter V. Coveney, Felix Wolf 0001, Dieter Kranzlmüller |
CCGRID | 7 |
| 2009 | Interoperation of world-wide production e-Science infrastructuresabstractAbstract Many production Grid and e‐Science infrastructures have begun to offer services to end‐users during the past several years with an increasing number of scientific applications that require access to a wide variety of resources and services in multiple Grids. Therefore, the Grid Interoperation Now—Community Group of the Open Grid Forum—organizes and manages interoperation efforts among those production Grid infrastructures to reach the goal of a world‐wide Grid vision on a technical level in the near future. This contribution highlights fundamental approaches of the group and discusses open standards in the context of production e‐Science infrastructures. Copyright © 2009 John Wiley & Sons, Ltd. Morris Riedel, Erwin Laure, Thomas Soddemann, Laurence Field, John-Paul Navarro, James Casey, Maarten Litmaath, Jean-Philippe Baud, Birger Koblitz, Charles E. Catlett, Dane Skow, Cindy Zheng, Philip M. Papadopoulos, Mason J. Katz, Neha Sharma 0001, Oxana Smirnova, Balázs Kónya, Peter W. Arzberger, Frank Würthwein, Abhishek Singh Rana, Terrence Martin, M. Wan, Von Welch, Tony Rimovsky, Steven J. Newhouse, Andrea Vanni, Yoshio Tanaka, Yusuke Tanimura, Tsutomu Ikegami, David Abramson 0001, Colin Enticott, Graham Jenkins, Ruth Pordes, Steven Timm, Gidon Moont, Mona Aggarwal, Dave Colling, Olivier van der Aa, Alex Sim, Vijaya Natarajan, Arie Shoshani, Junmin Gu, Gerson Galang, Riccardo Zappi, Luca Magnoni, Vincenzo Ciaschini, Michele Pace, Valerio Venturi, Moreno Marzolla, Paolo Andreetto, Robert Cowles, Shaowen Wang 0001, Yuji Saeki, Hitoshi Sato, Satoshi Matsuoka, Putchong Uthayopas, Somsak Sriprayoonsakul, Oscar Koeroo, Matthew Viljoen, Laura Pearlman, Stephen Pickles, David Wallom, Glenn Moloney, Jerome Lauret, Jim Marsteller, Paul Sheldon, Surya Pathak, Shaun De Witt, Jirí Mencák, Jens Jensen, Matt Hodges, Derek Ross, Sugree Phatanapherom, Gilbert Netzer, Anders Rhod Gregersen, Mike Jones 0002, Péter Kacsuk, Achim Streit, Daniel Mallmann, Felix Wolf 0001, Thomas Lippert, Thierry Delaitre, Eduardo Huedo, Neil Geddes |
Concurr. Comput. Pract. Exp. | 83 |
| 2008 | Classification of Different Approaches for e-Science Applications in Next Generation Computing InfrastructuresabstractSimulation and thus scientific computing is the third pillar alongside theory and experiment in todays science and engineering. The term e-science evolved as a new research field that focuses on collaboration in key areas of science using next generation infrastructures to extend the powers of scientific computing. This paper contributes to the field of e-science as a study of how scientists actually work within currently existing Grid and e-science infrastructures. Alongside numerous different scientific applications, we identified several common approaches with similar characteristics in different domains. These approaches are described together with a classification on how to perform e-science in next generation infrastructures. The paper is thus a survey paper which provides an overview of the e-science research domain. Morris Riedel, Achim Streit, Felix Wolf 0001, Thomas Lippert, Dieter Kranzlmüller |
eScience | 4 |
| 2007 | Computational Steering and Online Visualization of Scientific Applications on Large-Scale HPC Systems within e-Science InfrastructuresabstractIn the past several years, many scientific applications from various domains have taken advantage of e-science infrastructures that share storage or computational resources such as supercomputers, clusters or PC server farms across multiple organizations. Especially within e-science infrastructures driven by high-performance computing (HPC) such as DEISA, online visualization and computational steering (COVS) has become an important technique to save compute time on shared resources by dynamically steering the parameters of a parallel simulation. This paper argues that future supercomputers in the Petaflop/s performance range with up to 1 million CPUs will create an even stronger demand for seamless computational steering technologies. We discuss upcoming challenges for the development of scalable HPC applications and limits of future storage/IO technologies in the context of next generation e- science infrastructures and outline potential solutions. Morris Riedel, Thomas Eickermann, Sonja Habbinga, Wolfgang Frings, Paul Gibbon, Daniel Mallmann, Felix Wolf 0001, Achim Streit, Thomas Lippert, Wolfram Schiffmann, Andreas Ernst, Rainer Spurzem, Wolfgang E. Nagel |
eScience | 9 |
| 2006 | Topic 16: Applications of High-Performance and Grid Computing
Thomas Lippert, Giovanni Erbacci, Denis Trystram |
Euro-Par | 2 |
| 2006 | Scalable Ethernet Clos-Switches
Norbert Eicker, Thomas Lippert |
Euro-Par | 2 |
| 2001 | Hyper-systolic matrix multiplication
Thomas Lippert, Nicolai Petkov, Paolo Palazzari, Klaus Schilling 0002 |
Parallel Comput. | 1 |
| 1999 | A Preconditioner for Improved Fermion Actions
Wolfgang Bietenholz, Norbert Eicker, Andreas Frommer, Thomas Lippert, Björn Medeke, Klaus Schilling 0002 |
Euro-Par | 4 |
| 1999 | Lattice QCD with two dynamical Wilson fermions on APE100 parallel systems
Stephan Güsken, Thomas Lippert, Klaus Schilling 0002 |
Parallel Comput. | 2 |
| 1999 | Parallel SSOR preconditioning for lattice QCD
Thomas Lippert |
Parallel Comput. | 1 |
| 1999 | Hyper-systolic algorithms for N-body computations and parallel level-3 BLAS libraries
Thomas Lippert |
Parallel Comput. | 1 |
| 1998 | Hyper-Systolic Parallel ComputingabstractWe introduce a new class of parallel algorithms for the exact computation of systems with pairwise mutual interactions of n elements, so called n/sup 2/-problems. Hitherto, practical conventional parallelization strategies could achieve a complexity of O(np) with respect to the inter-processor communication, p being the number of processors. Our new approach can reduce the inter-processor communication complexity to a number O(np). In the framework of Additive Number Theory, the determination of the optimal communication pattern can be formulated as h-range minimization problem that can be solved numerically. Based on a complexity model, the scaling behavior of the new algorithm is numerically tested on the connection machine CM5. As a real life example, we have implemented a fast code for globular cluster n-body simulations, a generic n/sup 2/-problem, on the CRAY T3D, with striking success. Our parallel method promises to be useful in various scientific and engineering fields like polymer chain computations, protein folding, signal processing, and, in particular, for parallel level-3 BLAS. Thomas Lippert, Armin Seyfried, Achim Bode, Klaus Schilling 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1997 | Scalable Parallel SSOR Preconditioning for Lattice Computations in Gauce Theories
Andreas Frommer, Thomas Lippert, Klaus Schilling 0002 |
Euro-Par | 2 |
| 1992 | Statistical analysis of simulation-generated time series: Systolic vs. semi-systolic correlation on the Connection Machine
T. Dontje, Thomas Lippert, Nicolai Petkov, Klaus Schilling 0002 |
Parallel Comput. | 2 |
| 1992 | Quark propagator on the Connection Machine
Thomas Lippert, Klaus Schilling 0002, Nicolai Petkov |
Parallel Comput. | 1 |