Cédric Lichtenau

dblp:19/1024 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Theory of computation · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enterprise Class On-Chip Accelerator Integration
abstract
The IBM Z®platform and the underlying processor chip designs supporting it are optimized for processing vast amounts of data and transactions, while delivering consistent system performance, throughput, and response latencies with a sustained processor utilization of over 90 % under all workload conditions in a highly virtualized and secured computing environment. The IBM Telum® series of processor chip designs that support the platform introduced the industry to a novel modular scalable heterogeneous processor compute framework with an integrated multi-tier unified cache hierarchy all within one chip. The processor chip design leverages a unique approach to ensure all elements work in unison to continuously deliver performance to the evolving needs of mission critical workloads running on the platform while responding to those changing demands at processor clock speeds. This paper will detail the varying compute, accelerator, and cache units within the IBM Telum II processor chip design, how they adaptively work in unison without generating a cacophony of agents competing for scarce hardware resources, and how this forms the backbone of the scalable multi-processor system that our modern economy is built upon.
Deanna Postles Dunn Berger, Alper Buyuktosunoglu, Craig R. Walters, Robert J. Sonnelitter, Hailey Nicholson, Ashraf ElSharif, Yamil Rivera, Avery Francois, Cédric Lichtenau, Jason Kohl
HPCA9
2023 Acceleration of Decision-Tree Ensemble Models on the IBM Telum Processor
abstract
This paper presents a tensor-based algorithm that leverages a hardware accelerator for inferencing decision-tree-based machine learning models. The algorithm has been integrated in a public software library and is demonstrated on an IBM z16 server, using the Telum processor with the Integrated Accelerator for AI. We describe the architecture and implementation of the algorithm and present experimental results that demonstrate its superior runtime performance compared with popular CPU-based machine learning inference implementations.
Nikolaos Papandreou, Jan van Lunteren, Andreea Anghel, Thomas P. Parnell, Martin Petermann, Milos Stanisavljevic, Cédric Lichtenau, Andrew Sica, Dominic Röhm, Elpida Tzortzatos, Haralampos Pozidis
ISCAS7
2022 AI accelerator on IBM telum processor: industrial product
abstract
IBM Telum is the next generation processor chip for IBM Z and LinuxONE systems. The Telum design is focused on enterprise class workloads and it achieves over 40% per socket performance growth compared to IBM z15. The IBM Telum is the first server-class chip with a dedicated on-chip AI accelerator that enables clients to gain real time insights from their data as it is getting processed.
Cédric Lichtenau, Alper Buyuktosunoglu, Ramon Bertran Monfort, Peter Figuli, Christian Jacobi 0002, Nikolaos Papandreou, Haralampos Pozidis, Anthony Saporito, Andrew Sica, Elpida Tzortzatos
ISCA1
2020 SIMD Multi Format Floating-Point Unit on the IBM z15(TM)
abstract
The IBM z Systems(TM) is the backbone of the insurance, banking, and retail industry. Innovation in these markets is driving the demand for new and additional applications to better serve the customers. These workloads like machine learning, data analytics, AI, etc. require a rapidly increasing number of computations in smaller precision formats. With IBM z15(TM) we completely redesigned the binary and hexadecimal floating-point unit to efficiently implement SIMD operations at 5.2GHz while maintaining the industry leading reliability, availability and serviceability standard. This paper describes the new design and special techniques used to achieve these goals like reusing the existing double precision unit pipeline for lower precision parallel SIMD, new approaches to formally verify the design, and improving error detection for the 14nm technology node.
Stefan Payer, Cédric Lichtenau, Michael Klein, Kerstin Schelm, Petra Leber, Nicol Hofmann, Tina Babinsky
ARITH2
2016 Quad Precision Floating Point on the IBM z13
abstract
When operating on a rapidly increasing amount of data, business analytics applications become sensitive to rounding errors, and profit from the higher stability and faster convergence of quad precision floating-point (FP-QP) arithmetic. The IBM z13TMsupports this emerging trend around Big Data with an outstanding FP-QP performance. The paper details the vector and floating-point unit of IBM z13TM, with special focus on binary FP-QP. Except for divide and square root, these instructions are executed in the decimal engine. To operate such an 8-cycle decimal and quad precision pipeline at 5GHz required innovation around exponent handling, normalization, and rounding.
Cédric Lichtenau, Steven R. Carlough, Silvia M. Müller
ARITH1
2002 Real PRAM Programming
Wolfgang J. Paul, Peter Bach, Michael Bosch, Cédric Lichtenau, Jochen Röhrig
Euro-Par5
1999 Highly Concurrent Locking in Shared Memory Database Systems
Christian Jacobi 0002, Cédric Lichtenau
Euro-Par2