Kenneth S. McElvain

dblp:03/10927 · also Kenneth McElvain · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2018
0000-0002-1405-7935ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
High-performance computing · 72% Emerging computing paradigms · 9% Parallel and multicore computing · 9%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.522018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
Parallel implementation and performance optimization of the configuration-interaction method · SC 2015
High-performance computing
performance optimization at scale
0.422018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
Parallel implementation and performance optimization of the configuration-interaction method · SC 2015
High-performance computing › supercomputing
exascale computing
0.312018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
High-performance computing › scientific computing systems
lattice quantum chromodynamics
0.312018
Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing · SC 2018
Parallel and multicore computing
load balancing
0.212015
Parallel implementation and performance optimization of the configuration-interaction method · SC 2015
Emerging computing paradigms › quantum computing › quantum simulation
quantum many-body simulation
0.212015
Parallel implementation and performance optimization of the configuration-interaction method · SC 2015
Electronic design automation › physical design › placement › circuit placement
FPGA placement
0.112012
A fast discrete placement algorithm for FPGAs · FPGA 2012
High-performance computing › sparse linear algebra
sparse matrix computation
0.112015
Parallel implementation and performance optimization of the configuration-interaction method · SC 2015
Electronic design automation › physical design › placement
global placement
0.012012
A fast discrete placement algorithm for FPGAs · FPGA 2012
Electronic design automation › logic synthesis
circuit optimization
0.011987
An Intelligent Compiler Subsystem for a Silicon Compiler · DAC 1987
Electronic design automation
logic synthesis
0.011987
An Intelligent Compiler Subsystem for a Silicon Compiler · DAC 1987
Electronic design automation › physical design
module generation
0.011987
An Intelligent Compiler Subsystem for a Silicon Compiler · DAC 1987
Electronic design automation
physical design
0.011987
An Intelligent Compiler Subsystem for a Silicon Compiler · DAC 1987
Integrated circuit design
digital circuit design
0.011987
An Intelligent Compiler Subsystem for a Silicon Compiler · DAC 1987

Methods — techniques the papers use, named apart from their topics

monte carlo simulation · 0.3lattice QCD · 0.3matrix-vector multiplication · 0.2lanczos reorthogonalization · 0.2simulated annealing · 0.1acceleration techniques · 0.1constraint-based optimization · 0.0
YearPublicationVenuePosition
2018 Simulating the weak death of the Neutron in a femtoscale universe with near-exascale computing
Evan Berkowitz, Michael A. Clark, Arjun Singh Gambhir, Kenneth S. McElvain, Amy N. Nicholson, Enrico Rinaldi, Pavlos Vranas, André Walker-Loud, Chia-Cheng Chang, Bálint Joó, Thorsten Kurth, Konstantinos Orginos
SC4
2015 Parallel implementation and performance optimization of the configuration-interaction method
abstract
The configuration-interaction (CI) method, long a popular approach to describe quantum many-body systems, is cast as a very large sparse matrix eigenpair problem with matrices whose dimension can exceed one billion. Such formulations place high demands on memory capacity and memory bandwidth --- two quantities at a premium today. In this paper, we describe an efficient, scalable implementation, BIGSTICK, which, by factorizing both the basis and the interaction into two levels, can reconstruct the nonzero matrix elements on the fly, reduce the memory requirements by one or two orders of magnitude, and enable researchers to trade reduced resources for increased computational time. We optimize BIGSTICK on two leading HPC platforms --- the Cray XC30 and the IBM Blue Gene/Q. Specifically, we not only develop an empirically-driven load balancing strategy that can evenly distribute the matrix-vector multiplication across 256K threads, we also developed techniques that improve the performance of the Lanczos reorthogonalization. Combined, these optimizations improved performance by 1.3-8× depending on platform and configuration.
Hongzhang Shan, Samuel Williams 0001, Calvin W. Johnson, Kenneth S. McElvain, W. Erich Ormand
SC4
2012 A fast discrete placement algorithm for FPGAs
abstract
Good FPGA placement is crucial to obtain the best Quality of Results (QoR) from FPGA hardware. Although many published global placement techniques place objects in a continuous ASIC-like environment, FPGAs are discrete in nature, and a continuous algorithm cannot always achieve superior QoR by itself. Therefore, discrete FPGA-specific detail placement algorithms are used to improve the global placement results. Unfortunately, most of these detail placement algorithms do not have a global view. This paper presents a discrete "middle" placer that fills the gap between the two placement steps. It works like simulated annealing, but leverages various acceleration techniques. It does not pay the runtime penalty typical of simulated annealing solutions. Experiments show that with this placer, final QoR is significantly better than with the global-detail placer approach.
Qinghong Wu, Kenneth S. McElvain
FPGA2
1987 An Intelligent Compiler Subsystem for a Silicon Compiler
abstract
This paper presents a module generator which automatically generates and optimizes circuitry to satisfy constraints of speed, area and power. The user has complete control over the clock timing driving the circuitry and the area, width, or height of the resulting module. Unlike other programs that have been optimized for area and speed, this program supports more degrees of freedom and a broad range of circuit constructs, permitting a complete integrated circuit to be designed to meet the overall IC project objectives.
D. L. Johannsen, S. K. Tsubota, Kenneth S. McElvain
DAC3