EDBT 2026 Demo / reviewers in the wild / expert
Kenneth Czechowski
dblp:66/10046 · also Kent Czechowski
· DBLP profile ↗
6ranked-venue papers
3as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 32% Parallel and multicore computing · 18% Distributed systems · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › parallel numerical algorithms
communication-avoiding algorithms |
0.3 | 1 | 2017 | Design and Implementation of a Communication-Optimal Classifier for Distributed Kernel Support Vector Machines · IEEE Trans. Parallel Distributed Syst. 2017 |
Distributed systems › distributed machine learning
distributed SVM training |
0.3 | 1 | 2017 | Design and Implementation of a Communication-Optimal Classifier for Distributed Kernel Support Vector Machines · IEEE Trans. Parallel Distributed Syst. 2017 |
Parallel and multicore computing › parallel computing
parallel machine learning |
0.3 | 1 | 2017 | Design and Implementation of a Communication-Optimal Classifier for Distributed Kernel Support Vector Machines · IEEE Trans. Parallel Distributed Syst. 2017 |
Energy-efficient computing › energy-efficient architecture
energy-efficient microarchitecture |
0.2 | 1 | 2014 | Improving the energy efficiency of Big Cores · ISCA 2014 |
High-performance computing
scientific computing systems |
0.1 | 1 | 2012 | Optimizing the computation of n-point correlations on large-scale astronomical data · SC 2012 |
Performance modeling and evaluation › parallel system performance › speedup modeling
isoefficiency analysis |
0.1 | 1 | 2017 | Design and Implementation of a Communication-Optimal Classifier for Distributed Kernel Support Vector Machines · IEEE Trans. Parallel Distributed Syst. 2017 |
High-performance computing › performance optimization at scale
parallel scalability |
0.1 | 1 | 2017 | Design and Implementation of a Communication-Optimal Classifier for Distributed Kernel Support Vector Machines · IEEE Trans. Parallel Distributed Syst. 2017 |
Energy-efficient computing
power-performance modeling |
0.1 | 1 | 2014 | Improving the energy efficiency of Big Cores · ISCA 2014 |
Methods — techniques the papers use, named apart from their topics
parallel computation · 0.3kernel SVM · 0.3communication-avoiding algorithm · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Design and Implementation of a Communication-Optimal Classifier for Distributed Kernel Support Vector MachinesabstractWe consider the problem of how to design and implement communication-efficient versions of parallel kernel support vector machines, a widely used classifier in statistical machine learning, for distributed memory clusters and supercomputers. The main computational bottleneck is the training phase, in which a statistical model is built from an input data set. Prior to our study, the parallel isoefficiency of a state-of-the-art implementation scaled as W = Ω(P3), where W is the problem size and P the number of processors; this scaling is worse than even a one-dimensional block row dense matrix vector multiplication, which has W = Ω(P2). This study considers a series of algorithmic refinements, leading ultimately to a Communication-Avoiding SVM method that improves the isoefficiency to nearly W = Ω(P). We evaluate these methods on 96 to 1,536 processors, and show average speedups of 3 - 16x (7× on average) over Dis-SMO, and a 95 percent weak-scaling efficiency on six real-world datasets, with only modest losses in overall classification accuracy. The source code can be downloaded at [1]. Yang You 0001, James Demmel, Kenneth Czechowski, Richard W. Vuduc |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | CA-SVM: Communication-Avoiding Support Vector Machines on Distributed SystemsabstractWe consider the problem of how to design and implement communication-efficient versions of parallel support vector machines, a widely used classifier in statistical machine learning, for distributed memory clusters and supercomputers. The main computational bottleneck is the training phase, in which a statistical model is built from an input data set. Prior to our study, the parallel is efficiency of a state-of-the-art implementation scaled as W = Omega(P3), where W is the problem size and P the number of processors, this scaling is worse than even a one-dimensional block row dense matrix vector multiplication, which has W = Omega(P2). This study considers a series of algorithmic refinements, leading ultimately to a Communication-Avoiding SVM (CASVM) method that improves the is efficiency to nearly W = Omega(P). We evaluate these methods on 96 to 1536 processors, and show average speedups of 3 - 16× (7× on average) over Dis-SMO, and a 95% weak-scaling efficiency on six real world datasets, with only modest losses in overall classification accuracy. The source code can be downloaded at https://github.com/fastalgo/casvm. Yang You 0001, James Demmel, Kenneth Czechowski, Richard W. Vuduc |
IPDPS | 3 |
| 2014 | Improving the energy efficiency of Big CoresabstractTraditionally, architectural innovations designed to boost single-threaded performance incur overhead costs which significantly increase power consumption. In many cases the increase in power exceeds the improvement in performance, resulting in a net increase in energy consumption. Thus, it is reasonable to assume that modern attempts to improve single-threaded performance will have a negative impact on energy efficiency. This has led to the belief that “Big Cores” are inherently inefficient. To the contrary, we present a study which finds that the increased complexity of the core microarchitecture in recent generations of the Intel®Core™ processor have reduced both the time and energy required to run various workloads. Moreover, taking out the impact of process technology changes, our study still finds the architecture and microarchitecture changes -such as the increase in SIMD width, addition of the frontend caches, and the enhancement to the out-of-order execution engine- account for 1.2x improvement in energy efficiency for these processors. This paper provides real-world examples of how architectural innovations can mitigate inefficiencies associated with “Big Cores” -for example, micro-op caches obviate the costly decode of complex x86 instructions- resulting in a core architecture that is both high performance and energy efficient. It also contributes to the understanding of how microarchitecture affects performance, power and energy efficiency by modeling the relationship between them. Kenneth Czechowski, Victor W. Lee, Ed Grochowski, Ronny Ronen, Ronak Singhal, Richard W. Vuduc, Pradeep Dubey |
ISCA | 1 |
| 2013 | A Theoretical Framework for Algorithm-Architecture Co-designabstractWe consider the problem of how to enable computer architects and algorithm designers to reason directly and analytically about the relationship between high-level architectural features and algorithm characteristics. We propose a modeling framework designed to help understand the long-term and high-level impacts of algorithmic and technology trends. This model connects abstract communication complexity analysis-with respect to both the inter-core and inter-processor networks and the memory hierarchy-with current technology proposals and projections. We illustrate how one might use the framework by instantiating a particular model for a class of architectures and sample algorithms (three-dimensional fast Fourier transforms, matrix multiply, and three-dimensional stencil). Then, as a suggestive demonstration, we analyze a number of what-if scenarios within the model in light of these trends to suggest broader statements and alternative futures for power-constrained architectures and algorithms. Kenneth Czechowski, Richard W. Vuduc |
IPDPS | 1 |
| 2012 | On the communication complexity of 3D FFTs and its implications for ExascaleabstractThis paper revisits the communication complexity of large-scale 3D fast Fourier transforms (FFTs) and asks what impact trends in current architectures will have on FFT performance at exascale. We analyze both memory hierarchy traffic and network communication to derive suitable analytical models, which we calibrate against current software implementations; we then evaluate models to make predictions about potential scaling outcomes at exascale, based on extrapolating current technology trends. Of particular interest is the performance impact of choosing high-density processors, typified today by graphics co-processors (GPUs), as the base processor for an exascale system. Among various observations, a key prediction is that although inter-node all-to-all communication is expected to be the bottleneck of distributed FFTs, intra-node communication---expressed precisely in terms of the relative balance among compute capacity, memory bandwidth, and network bandwidth---will play a critical role. Kenneth Czechowski, Casey Battaglino, Chris McClanahan, Kartik Iyer, P.-K. Yeung, Richard W. Vuduc |
ICS | 1 |
| 2012 | Optimizing the computation of n-point correlations on large-scale astronomical dataabstractThe n-point correlation functions (npcf) are powerful statistics that are widely used for data analyses in astronomy and other fields. These statistics have played a crucial role in fundamental physical breakthroughs, including the discovery of dark energy. Unfortunately, directly computing the npcf at a single value requires O(Nn) time for N points and values of n of 2, 3, 4, or even larger. Astronomical data sets can contain billions of points, and the next generation of surveys will generate terabytes of data per night. To meet these computational demands, we present a highly-tuned npcf computation code that show an order-of-magnitude speedup over current state-of-the-art. This enables a much larger 3-point correlation computation on the galaxy distribution than was previously possible. We show a detailed performance evaluation on many different architectures. William B. March, Kenneth Czechowski, Marat Dukhan, Thomas Benson, Dongryeol Lee, Andrew J. Connolly, Richard W. Vuduc, Edmond Chow, Alexander G. Gray |
SC | 2 |