Kai Schleupen

dblp:62/6287 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 89% Hardware accelerators and domain-specific architectures · 5% Energy-efficient computing · 5%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
neuromorphic computing
0.422016
Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016
Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014
Emerging computing paradigms › neuromorphic computing
brain-inspired computing
0.212016
Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016
Emerging computing paradigms
neuromorphic hardware
0.212016
Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016
Hardware accelerators and domain-specific architectures › neural network hardware
brain-inspired computing accelerator
0.112014
Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014
Energy-efficient computing
power management
0.112014
Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution · SC 2014

Methods — techniques the papers use, named apart from their topics

software ecosystem · 0.2scalable systems · 0.2event-driven kernel · 0.2chip tiling · 0.2
YearPublicationVenuePosition
2023 IBM NorthPole Neural Inference Machine
Dharmendra S. Modha, Filipp Akopyan, Alexander Andreopoulos, Rathinakumar Appuswamy, John V. Arthur, Andrew S. Cassidy, Pallab Datta, Michael DeBole, Steven K. Esser, Carlos Ortega Otero, Jun Sawada, Brian Taba, Arnon Amir, Deepika Bablani, Peter J. Carlson, Myron Flickner, Rajamohan Gandhasri, Guillaume Garreau, Megumi Ito, Jennifer L. Klamo, Jeffrey A. Kusnitz, Nathaniel J. McClatchey, Jeffrey L. McKinstry, Yutaka Y. Nakamura, Tapan K. Nayak, William P. Risk, Kai Schleupen, Ben Shaw 0001, Jay Sivagnaname, Daniel F. Smith, Ignacio G. Terrizzano, Takanori Ueda
HCS27
2016 Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications
abstract
Abstract not provided
Jun Sawada, Filipp Akopyan, Andrew S. Cassidy, Brian Taba, Michael DeBole, Pallab Datta, Rodrigo Alvarez-Icaza, Arnon Amir, John V. Arthur, Alexander Andreopoulos, Rathinakumar Appuswamy, Heinz Baier, Davis Barch, David J. Berg, Carmelo di Nolfo, Steven K. Esser, Myron Flickner, Thomas A. Horvath, Bryan L. Jackson, Jeffrey A. Kusnitz, Scott Lekuch, Michael Mastro, Timothy Melano, Paul Merolla, Steven E. Millman, Tapan K. Nayak, Norm Pass, Hartmut Penner, William P. Risk, Kai Schleupen, Ben Shaw 0001, Hayley Wu, Brian Giera, Adam Moody, T. Nathan Mundhenk, Brian Van Essen, Eric X. Wang, David P. Widemann, William E. Murphy, Jamie K. Infantolino, James A. Ross, Dale R. Shires, Manuel M. Vindiola, Raju Namburu, Dharmendra S. Modha
SC30
2014 Real-Time Scalable Cortical Computing at 46 Giga-Synaptic OPS/Watt with ~100× Speedup in Time-to-Solution and ~100, 000× Reduction in Energy-to-Solution
abstract
Drawing on neuroscience, we have developed a parallel, event-driven kernel for neurosynaptic computation, that is efficient with respect to computation, memory, and communication. Building on the previously demonstrated highly optimized software expression of the kernel, here, we demonstrate True North, a co-designed silicon expression of the kernel. True North achieves five orders of magnitude reduction in energy to-solution and two orders of magnitude speedup in time-to solution, when running computer vision applications and complex recurrent neural network simulations. Breaking path with the von Neumann architecture, True North is a 4,096 core, 1 million neuron, and 256 million synapse brain-inspired neurosynaptic processor, that consumes 65mW of power running at real-time and delivers performance of 46 Giga-Synaptic OPS/Watt. We demonstrate seamless tiling of True North chips into arrays, forming a foundation for cortex-like scalability. True North's unprecedented time-to-solution, energy-to-solution, size, scalability, and performance combined with the underlying flexibility of the kernel enable a broad range of cognitive applications.
Andrew S. Cassidy, Rodrigo Alvarez-Icaza, Filipp Akopyan, Jun Sawada, John V. Arthur, Paul Merolla, Pallab Datta, Marc González 0001, Brian Taba, Alexander Andreopoulos, Arnon Amir, Steven K. Esser, Jeffrey A. Kusnitz, Rathinakumar Appuswamy, Chuck Haymes, Bernard Brezzo, Roger Moussalli, Ralph Bellofatto, Christian W. Baks, Michael Mastro, Kai Schleupen, Charles E. Cox, Ken Inoue, Steven E. Millman, Nabil Imam, Emmett McQuinn, Yutaka Y. Nakamura, Ivan Vo, Chen Guok, Don Nguyen, Scott Lekuch, Sameh W. Asaad, Daniel J. Friedman, Bryan L. Jackson, Myron Flickner, William P. Risk, Rajit Manohar, Dharmendra S. Modha
SC21
2007 2-Gbps Uncompressed HDTV Transmission over 60-GHz SiGe Radio Link
abstract
We report a proof-of-concept demonstration of 2- Gbps uncompressed HDTV transmission using a 60-GHz SiGe radio chipset. We took a single-carrier approach with a usual DQPSK modulation scheme, assuming an LOS environment, and implemented the system with FPGAs. At the same time, in order to take care of more frequent sync/burst errors in high-data-rate single-carrier approaches, we equipped the baseband with effi- cient random/packet error recovery and symbol-timing recovery with an effective interpolation method. As a result, a clear and crisp image was obtained in the end-to-end transmission. I. INTRODUCTION
Yasunao Katayama, Chuck Haymes, Daiju Nakano, Troy J. Beukema, Brian A. Floyd, Scott K. Reynolds, Ullrich R. Pfeiffer, Brian P. Gaucher, Kai Schleupen
CCNC9
2007 Dynamic Partial FPGA Reconfiguration in a Prototype Microprocessor System
abstract
Modern FPGAs' parallel computing capability and their ability to be reconfigured make them an ideal platform to build accelerators for supercomputing systems. As a multi-core processor, the recently announced Cell Broadband EngineTM1 offers tremendous computing power. In this paper, we introduce a prototype system that combines these two types of computing devices together in a reconfigurable blade and we describe its architecture, memory system and abundant interfaces. On the reconfigurable blade it is desirable that the FPGA devices can be partially reconfigured at run-time. This paper presents the dynamic partial reconfiguration (DPR) technique and its design flow for the reconfigurable blade. We report our experimental results of the blade doing partial reconfiguration. DPR allows the reconfigurable blade to be a powerful, run-time changeable computing engine. A sample application is presented that was both simulated for the Cell processor and dynamically loaded to run on the FPGA.
Kai Schleupen, Scott Lekuch, Ryan Mannion, Zhi Guo, Walid A. Najjar, Frank Vahid
FPL1