Chuck Haymes

dblp:78/8707 · also Charles L. Haymes · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0002-5056-3528ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 since 2021Computer networks · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Electronic design automation · 34% Cloud and datacenter computing · 18% Memory systems · 17%
Computer networks
1 paper
Software-defined and programmable networks · 100%
Network and information security
1 paper
Network security · 100%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software-defined and programmable networks
programmable data plane
0.712023
Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation · FPGA 2023
Cloud and datacenter computing
cloud infrastructure
0.712023
Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation · FPGA 2023
Electronic design automation
hardware/software co-design
0.712023
Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation · FPGA 2023
Reconfigurable computing and FPGAs
FPGA prototyping
0.632017
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017
Efficient in-system RTL verification and debugging using FPGAs (abstract only) · FPGA 2012
A cycle-accurate, cycle-reproducible multi-FPGA system for accelerating multi-core processor simulation · FPGA 2012
Electronic design automation › hardware verification and test
hardware verification
0.322012
Efficient in-system RTL verification and debugging using FPGAs (abstract only) · FPGA 2012
A cycle-accurate, cycle-reproducible multi-FPGA system for accelerating multi-core processor simulation · FPGA 2012
Emerging computing paradigms
neuromorphic computing
0.212014
Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014
Electronic design automation › hardware verification and test › functional verification
logic verification
0.112012
A cycle-accurate, cycle-reproducible multi-FPGA system for accelerating multi-core processor simulation · FPGA 2012
Electronic design automation › hardware verification and test › design validation
processor design validation
0.112012
Efficient in-system RTL verification and debugging using FPGAs (abstract only) · FPGA 2012
Performance modeling and evaluation › simulation › monte carlo simulation
importance sampling
0.112009
Low BER performance estimation of LDPC codes via application of importance sampling to trapping sets · IEEE Trans. Commun. 2009
Performance modeling and evaluation
simulation
0.112009
Low BER performance estimation of LDPC codes via application of importance sampling to trapping sets · IEEE Trans. Commun. 2009
Coding theory › error-correcting codes › error probability analysis
error floor estimation
0.112009
Low BER performance estimation of LDPC codes via application of importance sampling to trapping sets · IEEE Trans. Commun. 2009
Coding theory › error-correcting codes
LDPC codes
0.112009
Low BER performance estimation of LDPC codes via application of importance sampling to trapping sets · IEEE Trans. Commun. 2009
Memory systems
emerging memory technologies
0.112017
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017
Memory systems
non-volatile memory
0.112017
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017
Memory systems › non-volatile memory › non-volatile main memory
NVDIMM
0.112017
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM
0.112017
Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor · MICRO 2017
Hardware accelerators and domain-specific architectures › neural network hardware
brain-inspired computing accelerator
0.112014
Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014
Energy-efficient computing
power management
0.112014
Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014
Processor architecture and microarchitecture
chip multiprocessor
0.012012
A cycle-accurate, cycle-reproducible multi-FPGA system for accelerating multi-core processor simulation · FPGA 2012

Methods — techniques the papers use, named apart from their topics

p4 · 2.0hardware-enforced isolation · 2.0hardware offloading · 2.0FPGA prototyping · 0.3event-driven kernel · 0.2chip tiling · 0.2in-system RTL debugging · 0.1design partitioning · 0.1clock synchronization · 0.1FPGA emulation · 0.1trapping-set analysis · 0.1importance sampling · 0.1
YearPublicationVenuePosition
2023 Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation
abstract
Disparate deployment models of cloud computing pose varying requirements on cloud infrastructure components such as networking, storage, provisioning, and security. Infrastructure providers need to study these and often create custom infrastructure components to satisfy these requirements. A major challenge in the research and development of these cloud infrastructure solutions, however, is the availability of customizable platforms for experimentation and trade-off analysis of the various hardware and software components. Most platforms are either general purpose or bespoke solutions created to assist a particular task, too rigid to allow meaningful customization. In this work, we present a 100G reconfigurable smartNIC prototyping platform called Janus that enables cloud infrastructure research and hardware-software co-design of infrastructure components such as hypervisor, secure boot, software defined networking and distributed storage. The platform provides a path to optimize the stack by offloading the functionalities from the host x86 to the embedded processor on the smartNIC and optimize performance by moving pieces to hardware using P4. Further, our platform provides hardware-enforced isolation of cloud network control plane, thereby securing the control plane from the tenants even for bare-metal deployments.
Bharat Sukhwani, Mohit Kapur, Alda Ohmacht, Liran Schour, Martin Ohmacht, Chris Ward, Chuck Haymes, Sameh W. Asaad
FPGA7
2017 Contutto: a novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor
abstract
We demonstrate the use of an FPGA as a memory buffer in a POWER8® system, creating a novel prototyping platform that enables innovation in the memory subsystem of POWER-based servers. Our platform, called ConTutto, is pin-compatible with POWER8 buffered memory DIMMs and plugs into a memory slot of a standard POWER8 processor system, running at aggregate memory channel speeds of 35 GB/s per link. ConTutto, which means "with everything", is a platform to experiment with different memory technologies, such as STT-MRAM and NAND Flash, in an end-to-end system context. Enablement of STT-MRAM and NVDIMM using ConTutto shows up to 12.5x lower latency and 7.5x higher bandwidth compared to the respective technologies when attached to the PCIe bus. Moreover, due to the unique attach-point of the FPGA between the processor and system memory, ConTutto provides a means for in-line acceleration of certain computations on-route to memory, and enables sensitivity analysis for memory latency while running real applications. To the best of our knowledge, ConTutto is the first ever FPGA platform on the memory bus of a server class processor.
Bharat Sukhwani, Chuck Haymes, Kyu-Hyoun Kim, Adam J. McPadden, Daniel M. Dreps, Dean Sanner, Jan van Lunteren, Sameh W. Asaad
MICRO3
2014 Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution
abstract
Drawing on neuroscience, we have developed a parallel, event-driven kernel for neurosynaptic computation, that is efficient with respect to computation, memory, and communication. Building on the previously demonstrated highly optimized software expression of the kernel, here, we demonstrate True North, a co-designed silicon expression of the kernel. True North achieves five orders of magnitude reduction in energy to-solution and two orders of magnitude speedup in time-to solution, when running computer vision applications and complex recurrent neural network simulations. Breaking path with the von Neumann architecture, True North is a 4,096 core, 1 million neuron, and 256 million synapse brain-inspired neurosynaptic processor, that consumes 65mW of power running at real-time and delivers performance of 46 Giga-Synaptic OPS/Watt. We demonstrate seamless tiling of True North chips into arrays, forming a foundation for cortex-like scalability. True North's unprecedented time-to-solution, energy-to-solution, size, scalability, and performance combined with the underlying flexibility of the kernel enable a broad range of cognitive applications.
Andrew S. Cassidy, Rodrigo Alvarez-Icaza, Filipp Akopyan, Jun Sawada, John V. Arthur, Paul Merolla, Pallab Datta, Marc González 0001, Brian Taba, Alexander Andreopoulos, Arnon Amir, Steven K. Esser, Jeffrey A. Kusnitz, Rathinakumar Appuswamy, Chuck Haymes, Bernard Brezzo, Roger Moussalli, Ralph Bellofatto, Christian W. Baks, Michael Mastro, Kai Schleupen, Charles E. Cox, Ken Inoue, Steven E. Millman, Nabil Imam, Emmett McQuinn, Yutaka Y. Nakamura, Ivan Vo, Chen Guok, Don Nguyen, Scott Lekuch, Sameh W. Asaad, Daniel J. Friedman, Bryan L. Jackson, Myron Flickner, William P. Risk, Rajit Manohar, Dharmendra S. Modha
SC15
2012 A cycle-accurate, cycle-reproducible multi-FPGA system for accelerating multi-core processor simulation
abstract
Software based tools for simulation are not keeping up with the demands for increased chip and system design complexity. In this paper, we describe a cycle-accurate and cycle-reproducible large-scale FPGA platform that is designed from the ground up to accelerate logic verification of the Bluegene/Q compute node ASIC, a multi-processor SOC implemented in IBM's 45 nm SOI CMOS technology. This paper discusses the challenges for constructing such large-scale FPGA platforms, including design partitioning, clocking & synchronization, and debugging support, as well as our approach for addressing these challenges without sacrificing cycle accuracy and cycle reproducibility. The resulting fullchip simulation of the Bluegene/Q compute node ASIC runs at a simulated processor clock speed of 4 MHz, over 100,000 times faster than the logic level software simulation of the same design. The vast increase in simulation speed provides a new capability in the design cycle that proved to be instrumental in logic verification as well as early software development and performance validation for Bluegene/Q.
Sameh W. Asaad, Ralph Bellofatto, Bernard Brezzo, Chuck Haymes, Mohit Kapur, Benjamin D. Parker, Proshanta Saha, Todd Takken, José A. Tierno
FPGA4
2012 Efficient in-system RTL verification and debugging using FPGAs (abstract only)
abstract
FPGAs have become indispensible in processor design, bring-up and debug. Traditionally FPGAs have been used in prototyping, allowing end-users to emulate functionality of a specific component of a processor. However, as the complexity of processors grows, another aspect of processor design, RTL verification, has become a prime target for acceleration using FPGAs. Software-only RTL simulation and verification tools are no longer sufficient for many verification tasks as they often incur long execution time penalties. Software simulation time for a basic Linux kernel bring-up on a BlueGene/Q [1] processor, with 16 user PowerPC A2 cores, for example, could easily exceed several years.
Proshanta Saha, Chuck Haymes, Ralph Bellofatto, Bernard Brezzo, Mohit Kapur, Sameh W. Asaad
FPGA2
2009 Low BER performance estimation of LDPC codes via application of importance sampling to trapping sets
abstract
We introduce an importance sampling (IS) method that successfully simulates the performance of Low density Parity Check (LDPC) Codes in an AWGN channel at very low bit error rates (BERs). By effectively finding and biasing bit node combinations that are the dominant sources of error events, called trapping sets, the developed technique provokes more frequent decoder failures. Consequently, fewer simulation runs and higher simulation gains are achieved.
Enver Cavus, Chuck Haymes, Babak Daneshrad
IEEE Trans. Commun.2
2007 2-Gbps Uncompressed HDTV Transmission over 60-GHz SiGe Radio Link
abstract
We report a proof-of-concept demonstration of 2- Gbps uncompressed HDTV transmission using a 60-GHz SiGe radio chipset. We took a single-carrier approach with a usual DQPSK modulation scheme, assuming an LOS environment, and implemented the system with FPGAs. At the same time, in order to take care of more frequent sync/burst errors in high-data-rate single-carrier approaches, we equipped the baseband with effi- cient random/packet error recovery and symbol-timing recovery with an effective interpolation method. As a result, a clear and crisp image was obtained in the end-to-end transmission. I. INTRODUCTION
Yasunao Katayama, Chuck Haymes, Daiju Nakano, Troy J. Beukema, Brian A. Floyd, Scott K. Reynolds, Ullrich R. Pfeiffer, Brian P. Gaucher, Kai Schleupen
CCNC2
2006 An IS Simulation Technique for Very Low BER Performance Evaluation of LDPC Codes
abstract
We introduce an Importance Sampling (IS) method that successfully simulates the performance of Low density Parity Check (LDPC) Codes in an AWGN channel at very low bit error rates (BERs). By effectively finding and biasing bit node combinations that are the dominant sources of error events, called trapping sets, the developed technique provokes more frequent decoder failures. Consequently, fewer simulation runs and higher simulation gains are achieved. Regardless of the block size of an LDPC code, only a few dominant trapping set classes cause decoder failures at low BER regions. Therefore, the proposed technique allows the performance evaluation for any size LDPC code at very low BER regions with remarkable simulation gains. For BERs of 10-20, we observed simulation gains on the order of 1014.
Enver Cavus, Chuck Haymes, Babak Daneshrad
ICC2