Norman H. Christ

dblp:01/6988 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2013
0000-0003-3856-9641ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
High-performance computing · 66% Performance modeling and evaluation · 19% Processor architecture and microarchitecture · 8%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.222013
The origin of mass · SC 2013
A Very Fast Parallel Processor · IEEE Trans. Computers 1984
High-performance computing › scientific computing systems
lattice quantum chromodynamics
0.212013
The origin of mass · SC 2013
High-performance computing
performance optimization at scale
0.212013
The origin of mass · SC 2013
Performance modeling and evaluation › parallel system performance
weak scaling
0.212013
The origin of mass · SC 2013
Processor architecture and microarchitecture
instruction set architecture
0.012004
QCDOC: A 10 Teraflops Computer for Tightly-Coupled Calculations · SC 2004
High-performance computing
supercomputing
0.012004
QCDOC: A 10 Teraflops Computer for Tightly-Coupled Calculations · SC 2004
High-performance computing
supercomputer architecture
0.011997
QCDSP: A Teraflop Scale Massively Parallel Supercomputer · SC 1997
Interconnection networks and networks-on-chip
network topology
0.011984
A Very Fast Parallel Processor · IEEE Trans. Computers 1984
Processor architecture and microarchitecture
SIMD
0.011984
A Very Fast Parallel Processor · IEEE Trans. Computers 1984

Methods — techniques the papers use, named apart from their topics

multigrid · 0.2domain decomposition · 0.2conjugate gradient solver · 0.0custom hardware design · 0.0
YearPublicationVenuePosition
2013 The origin of mass
abstract
The origin of mass is one of the deepest mysteries in science. Neutrons and protons, which account for almost all visible mass in the Universe, emerged from a primordial plasma through a cataclysmic phase transition microseconds after the Big Bang. However, most mass in the Universe is invisible. The existence of dark matter, which interacts with our world so weakly that it is essentially undetectable, has been established from its galactic-scale gravitational effects. Here we describe results from the first truly physical calculations of the cosmic phase transition and a groundbreaking first-principles investigation into composite dark matter, studies impossible with previous state-of-the-art methods and resources. By inventing a powerful new algorithm, "DSDR," and implementing it effectively for contemporary supercomputers, we attain excellent strong scaling, perfect weak scaling to the LLNL BlueGene/Q two million cores, sustained speed of 7.2 petaflops, and time-to-solution speedup of more than 200 over the previous state-of-the-art.
Peter A. Boyle, Michael I. Buchoff, Norman H. Christ, Taku Izubuchi, Chulwoo Jung, Thomas C. Luu, Robert D. Mawhinney, Chris Schroeder, Ron Soltz, Pavlos Vranas, Joseph Wasem
SC3
2004 QCDOC: A 10 Teraflops Computer for Tightly-Coupled Calculations
abstract
Numerical simulations of the strong nuclear force, known as quantum chromodynamics or QCD, have proven to be a demanding, forefront problem in high-performance computing. In this report, we describe a new computer, QCDOC (QCD On a Chip), designed for optimal price/performance in the study of QCD. QCDOC uses a six-dimensional, low-latency mesh network to connect processing nodes, each of which includes a single custom ASIC, designed by our collaboration and built by IBM, plus DDR SDRAM. Each node has a peak speed of 1Gigaflops and two 12,288node, 10+ Teraflops machines are to be completed in the fall of 2004. Currently, a 512 node machine is running, delivering efficiencies as high as 45% of peak on the conjugate gradient solvers that dominate our calculations and a 4096-node machine with a cost of $1.6M is under construction. This should give us a price/performance less than $1per sustained Megaflops.
Peter A. Boyle, Dong Chen 0005, Norman H. Christ, Michael A. Clark, Saul D. Cohen, Zhihua Dong, Alan Gara, Bálint Joó, Chulwoo Jung, Ludmila A. Levkova, Xiaodong Liao, Guofeng Liu, Robert D. Mawhinney, Shigemi Ohta, Konstantin Petrov, Tilo Wettig, Azusa Yamaguchi, Calin Cristian
SC3
1997 QCDSP: A Teraflop Scale Massively Parallel Supercomputer
abstract
We discuss the work of the QCDSP collaboration to build an inexpensive Teraflop scale massively parallel computer suitable for computations in Quantum Chromodynamics (QCD). The computer is a collection of nodes connected in a four dimensional toroidial grid with nearest neighbor bit serial communications. A node is composed of a Texas Instruments Digital Signal Processor (DSP), memory, and a custom made communications and memory controller chip. An 8192 node computer with a peak speed of 0.4 Teraflops is being constructed at Columbia University for a cost of $1.8 Million. A 12,288-node machine with a peak speed of 0.6 Teraflops is being constructed for the RIKEN Brookhaven Research Center. Other computers have been built including a 50 Gigaflop version for Florida State University.
Dong Chen 0005, Norman H. Christ, Robert G. Edwards, George Fleming, Alan Gara, Sten Hansen, Chulwoo Jung, Adrian Kahler, Stephen Kasow, Anthony D. Kennedy, Greg Kilcup, Yu Bing Luo, Catalin Malureanu, Robert D. Mawhinney, John Parsons, Jim Sexton, ChengZhong Sui, Pavlos Vranas
SC3
1984 A Very Fast Parallel Processor
abstract
A parallel processor specially designed for an important problem in theoretical physics is described. The final device will contain 256 nodes running in lock-step in a SIMD mode with a computational power of 4 billion 22-bit floating point operations per second. Each node is controlled by an Intel 80286/287 microprocessor, contains 160K bits of memory and has a pipelined, microprogrammable arithmetic unit which performs floating point multiplications and additions. The nodes are interconnected in a two-dimensional rectangular grid so that vectorized arithmetic can be performed using data stored on any two adjacent nodes.
Norman H. Christ, Anthony E. Terrano
IEEE Trans. Computers1