EDBT 2026 Demo / reviewers in the wild / expert
Nicholas P. Carter
dblp:09/254
· DBLP profile ↗
18ranked-venue papers
4as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Processor architecture and microarchitecture · 26% Memory systems · 23% Energy-efficient computing · 15% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
many-core architecture |
0.2 | 1 | 2013 | Runnemede: An architecture for Ubiquitous High-Performance Computing · HPCA 2013 |
Energy-efficient computing › voltage scaling
near-threshold voltage operation |
0.2 | 1 | 2013 | Runnemede: An architecture for Ubiquitous High-Performance Computing · HPCA 2013 |
Memory systems
on-chip memory |
0.2 | 1 | 2013 | Runnemede: An architecture for Ubiquitous High-Performance Computing · HPCA 2013 |
Distributed systems › fault tolerance
checkpointing |
0.1 | 1 | 2007 | Architecture of a Self-Checkpointing Microprocessor that Incorporates Nanomagnetic Devices · IEEE Trans. Computers 2007 |
Memory systems
non-volatile memory |
0.1 | 1 | 2007 | Architecture of a Self-Checkpointing Microprocessor that Incorporates Nanomagnetic Devices · IEEE Trans. Computers 2007 |
Processor architecture and microarchitecture
clustered architecture |
0.0 | 1 | 2004 | A reconfigurable unit for a clustered programmable-reconfigurable processor · FPGA 2004 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable arrays |
0.0 | 1 | 2004 | A reconfigurable unit for a clustered programmable-reconfigurable processor · FPGA 2004 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable logic |
0.0 | 1 | 2004 | A magnetoelectronic macrocell employing reconfigurable threshold logic · FPGA 2004 |
Integrated circuit design › digital circuit design › threshold logic
threshold logic circuits |
0.0 | 1 | 2004 | A magnetoelectronic macrocell employing reconfigurable threshold logic · FPGA 2004 |
Processor architecture and microarchitecture
multithreading |
0.0 | 3 | 1998 | Exploiting Fine-grain Thread Level Parallelism on the MIT Multi-ALU Processor · ISCA 1998 The M-Machine multicomputer · MICRO 1995 Hardware Support for Fast Capability-based Addressing · ASPLOS 1994 |
Hardware reliability and fault tolerance
power failure resilience |
0.0 | 1 | 2007 | Architecture of a Self-Checkpointing Microprocessor that Incorporates Nanomagnetic Devices · IEEE Trans. Computers 2007 |
Parallel and multicore computing › synchronization
synchronization mechanisms |
0.0 | 1 | 1998 | Exploiting Fine-grain Thread Level Parallelism on the MIT Multi-ALU Processor · ISCA 1998 |
Processor architecture and microarchitecture
pipelining |
0.0 | 1 | 2004 | A reconfigurable unit for a clustered programmable-reconfigurable processor · FPGA 2004 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable fabric |
0.0 | 1 | 2004 | A magnetoelectronic macrocell employing reconfigurable threshold logic · FPGA 2004 |
Interconnection networks and networks-on-chip
network topology |
0.0 | 1 | 1995 | The M-Machine multicomputer · MICRO 1995 |
Operating systems › resource management › memory management
memory protection |
0.0 | 1 | 1994 | Hardware Support for Fast Capability-based Addressing · ASPLOS 1994 |
Processor architecture and microarchitecture › instruction set architecture
capability-based addressing |
0.0 | 1 | 1994 | Hardware Support for Fast Capability-based Addressing · ASPLOS 1994 |
Memory systems
memory protection |
0.0 | 1 | 1994 | Hardware Support for Fast Capability-based Addressing · ASPLOS 1994 |
Parallel and multicore computing › parallel programming models
message passing |
0.0 | 1 | 1995 | The M-Machine multicomputer · MICRO 1995 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.3hardware-software co-design · 0.2threshold logic · 0.0pipelining scheme · 0.0interleaved reconfigurable array · 0.0hybrid hall effect · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Runnemede: An architecture for Ubiquitous High-Performance ComputingabstractDARPA's Ubiquitous High-Performance Computing (UHPC) program asked researchers to develop computing systems capable of achieving energy efficiencies of 50 GOPS/Watt, assuming 2018-era fabrication technologies. This paper describes Runnemede, the research architecture developed by the Intel-led UHPC team. Runnemede is being developed through a co-design process that considers the hardware, the runtime/OS, and applications simultaneously. Near-threshold voltage operation, fine-grained power and clock management, and separate execution units for runtime and application code are used to reduce energy consumption. Memory energy is minimized through application-managed on-chip memory and direct physical addressing. A hierarchical on-chip network reduces communication energy, and a codelet-based execution model supports extreme parallelism and fine-grained tasks. We present an initial evaluation of Runnemede that shows the design process for our on-chip network, demonstrates 2-4x improvements in memory energy from explicit control of on-chip memory, and illustrates the impact of hardware-software co-design on the energy consumption of a synthetic aperture radar algorithm on our architecture. Nicholas P. Carter, Aditya Agrawal, Shekhar Borkar, Romain Cledat, Howard David, Dave Dunning, Joshua B. Fryman, Ivan Ganev, Roger A. Golliver, Rob C. Knauerhase, Richard A. Lethin, Benoît Meister, Asit K. Mishra, Wilfred R. Pinfold, Justin Teller, Josep Torrellas, Nicolas Vasilache, Ganesh Venkatesh |
HPCA | 1 |
| 2011 | DeNovo: Rethinking the Memory Hierarchy for Disciplined ParallelismabstractFor parallelism to become tractable for mass programmers, shared-memory languages and environments must evolve to enforce disciplined practices that ban "wild shared-memory behaviors;'' e.g., unstructured parallelism, arbitrary data races, and ubiquitous non-determinism. This software evolution is a rare opportunity for hardware designers to rethink hardware from the ground up to exploit opportunities exposed by such disciplined software models. Such a co-designed effort is more likely to achieve many-core scalability than a software-oblivious hardware evolution. This paper presents DeNovo, a hardware architecture motivated by these observations. We show how a disciplined parallel programming model greatly simplifies cache coherence and consistency, while enabling a more efficient communication and cache architecture. The DeNovo coherence protocol is simple because it eliminates transient states - verification using model checking shows 15X fewer reachable states than a state-of-the-art implementation of the conventional MESI protocol. The DeNovo protocol is also more extensible. Adding two sophisticated optimizations, flexible communication granularity and direct cache-to-cache transfers, did not introduce additional protocol states (unlike MESI). Finally, DeNovo shows better cache hit rates and network traffic, translating to better performance and energy. Overall, a disciplined shared-memory programming model allows DeNovo to seamlessly integrate message passing-like interactions within a global address space for improved design complexity, performance, and efficiency. Byn Choi, Rakesh Komuravelli, Hyojin Sung, Robert Smolinski, Nima Honarmand, Sarita V. Adve, Vikram S. Adve, Nicholas P. Carter, Ching-Tsun Chou |
PACT | 8 |
| 2010 | Design techniques for cross-layer resilienceabstractCurrent electronic systems implement reliability using only a few layers of the system stack, which simplifies the design of other layers but is becoming increasingly expensive over time. In contrast, cross-layer resilient systems, which distribute the responsibility for tolerating errors, device variation, and aging across the system stack, have the potential to provide the resilience required to implement reliable, high-performance, low-power systems in future fabrication processes at significantly lower cost. These systems can implement less-frequent resilience tasks in software to save power and chip area, can tune their reliability guarantees to the needs of applications, and can use the information available at each level in the system stack to optimize performance and power consumption. In this paper, we outline an approach to cross-layer system design that describes resilience as a set of tasks that systems must perform in order to detect and tolerate errors and variation. We then present strawman examples of how this task-based design process could be used to implement general-purpose computing and SoC systems, drawing on previous work and identifying key areas for future research. Nicholas P. Carter, Helia Naeimi, Donald S. Gardner |
DATE | 1 |
| 2010 | Vision for cross-layer optimization to address the dual challenges of energy and reliabilityabstractWe are rapidly approaching an inflection point where the conventional target of producing perfect, identical transistors that operate without upset can no longer be maintained while continuing to reduce the energy per operation. With power requirements already limiting chip performance, continuing to demand perfect, upset-free transistors would mean the end of scaling benefits. The big challenges in device variability and reliability are driven by uncommon tails in distributions, infrequent upsets, one-size-fits-all technology requirements, and a lack of information about the context of each operation. Solutions co-designed across traditional layer boundaries in our system stack can change the game, allowing architecture and software (a) to compensate for uncommon variation, environments, and events, (b) to pass down invariants and requirements for the computation, and (c) to monitor the health of collections of devices. Cross-layer codesign provides a path to continue extracting benefits from further scaled technologies despite the fact that they may be less predictable and more variable. While some limited multi-layer mitigation strategies do exist, to move forward redefining traditional layer abstractions and developing a framework that facilitates cross-layer collaboration is necessary. André DeHon, Heather M. Quinn, Nicholas P. Carter |
DATE | 3 |
| 2007 | Architecture of a Self-Checkpointing Microprocessor that Incorporates Nanomagnetic DevicesabstractMemory and latch circuits in CMOS systems rely on capacitatively-stored charge to hold state information. When power is removed from a chip, this charge quickly drains off, destroying any information that was contained in the chip. This causes a number of problems for computer systems, including data loss from power failures, the need to load operating systems from nonvolatile storage each time the system is powered on, and high "idle" power consumption due to leakage currents in memory arrays. Magnetoelectronic devices that combine ferromagnetic elements with semiconductor structures have the potential to overcome this limitation by providing high-performance nonvolatile storage that can be tightly integrated with logic. In this paper, we present the architecture of a microprocessor that uses magnetoelectronic devices to "snapshot" the state of the currently executing program at regular intervals. If its power supply is interrupted, this self-checkpointing microprocessor can near instantly restore its state from the last checkpoint, allowing it to resume execution with little loss of progress. Simulations of a self-checkpointing version of the Pentium 4 microprocessor show that the magnetoelectronic memories increase power consumption by only 62 mW, with little to no cost in system performance Love Kothari, Nicholas P. Carter |
IEEE Trans. Computers | 2 |
| 2005 | Exploiting Pipelining to Tolerate Wire Delays in a Programmable-Reconfigurable ProcessorabstractAs fabrication technologies advance, increasing wire delays in semiconductor systems are leading to larger and larger gaps between the clock rates of circuits implemented in reconfigurable logic and those of conventional microprocessors. In this paper, we present a pipelining scheme for the Amalgam programmable-reconfigurable processor that divides long wire delays into multi-cycle operations and supports overlapping of independent computations. On streaming benchmark programs, this pipelining scheme increases the clock rates of Amalgam's reconfigurable clusters by up to 72%, allowing the pipelined Amalgam to maintain a 2.6/spl times/ performance advantage over a purely-programmable processor in a wide range of fabrication processes. Chi-Wei Wang, Nicholas P. Carter, Richard B. Kujoth, Jeffrey J. Cook, Derek B. Gottlieb |
FPL | 2 |
| 2004 | A magnetoelectronic macrocell employing reconfigurable threshold logicabstractIn this paper, we introduce a reconfigurable fabric based around a new class of circuit element: the hybrid Hall effect (HHE) magnetoelectronic device. Because they incorporate a ferromagnetic element, HHE devices are inherently non-volatile, retaining their state without a power supply. In addition, HHE devices are extremely well-suited to implementing threshold logic circuits, which allows many complex logic functions to be implemented in fewer gates than are required in systems based on AND-OR logic. We present the design of an HHE-based reconfigurable macrocell based on two-level threshold logic that can be configured on a cycle-by-cycle basis while internally storing non-volatile configuration data and computation state. The performance of this macrocell is characterized, and compared to that of competing technologies, showing that it has a significantly better power-delay product when implementing complex functions of many inputs. Steve Ferrera, Nicholas P. Carter |
FPGA | 2 |
| 2004 | A reconfigurable unit for a clustered programmable-reconfigurable processorabstractIn a clustered programmable-reconfigurable processor, multiple programmable processors and blocks of reconfigurable logic communicate through a register-based communication mechanism, which reduces the impact of wire delay on clock cycle time. In this paper, we present a circuit-level design for the reconfigurable clusters used on the Amalgam programmable-reconfigurable processor. We outline our interleaved reconfigurable array design, which provides high bandwidth to and from the register file without requiring large amounts of register control logic. We characterize the latency of operations in our array, and present results that show the impact that this latency has on overall system performance in a range of fabrication processes. Finally, we present a pipelining scheme that enables the array to operate at clock rates closer to those of programmable processors and allows for better scaling in future technologies. Richard B. Kujoth, Chi-Wei Wang, Derek B. Gottlieb, Jeffrey J. Cook, Nicholas P. Carter |
FPGA | 5 |
| 2003 | Reconfigurable Circuits Using Hybrid Hall Effect Devices
Steve Ferrera, Nicholas P. Carter |
FPL | 2 |
| 2003 | Mapping computation kernels to clustered programmable-reconfigurable processorsabstractReconfigurable computing systems have shown the potential to surpass conventional processor architectures in performance for a growing range of applications. That performance, however, must be attained without significantly changing the design effort on the programmer's part, and without drastically increasing compilation time. In this paper, we present our compiler framework for mapping computation kernels to the reconfigurable clusters of Amalgam, a clustered programmable-reconfigurable processor. We first promote the use of the gated singular-assignment program dependence graph, a parallel intermediate program representation, to represent computation kernels. We then present an algorithm for mapping a computation kernel into the control FSM and datapath for a reconfigurable cluster. Finally, we describe our fast datapath synthesis tool-flow which preserves regularity and reduces the problem size by not flattening the datapath to gates. Jeffrey J. Cook, Lee Baugh, Derek B. Gottlieb, Nicholas P. Carter |
FPT | 4 |
| 2002 | Mapping Algorithms to the Amalgam Programmable-Reconfigurable ProcessorabstractThe Amalgam programmable-reconfigurable processor is designed to provide the computational power required by upcoming embedded applications without requiring the design of application-specific hardware. It integrates multiple programmable processors and blocks of reconfigurable logic onto a single chip, using a clustered architecture, similar to the one used on the M-Machine to reduce wire length and delay and allow implementation at high clock rates. The clustered architecture provides tremendous flexibility, allowing applications to exploit parallelism at whatever granularity is best-suited to the application, while the combination of reconfigurable logic and programmable processors delivers much higher performance than could be achieved through programmable processors alone. This abstract presents the results of our initial experiments in hand-mapping applications onto Amalgam. Five applications (IDCT, Rijndael encryption, nQueens, DNA sequence comparison, and image dithering) have been implemented, achieving speedups ranging from 8.7/spl times/ to 23.2/spl times/ over the performance of a single programmable cluster by using the complete resources of an Amalgam chip. Jeffrey J. Cook, Derek B. Gottlieb, Joshua D. Walstrom, Steve Ferrera, Chi-Wei Wang, Nicholas P. Carter |
FCCM | 6 |
| 2002 | The Design of the Amalgam Reconfigurable ClusterabstractAmalgam is a novel architecture for multifunction embedded systems. It integrates multiple reconfigurable and programmable processing resources (known as clusters) to achieve high-performance with low design effort on a variety of multimedia applications. The reconfigurable cluster (RClust) enables Amalgam to exploit the natural parallelism and operator granularities of a target application. The RClust contains a ring of reconfigurable logic interleaved with a banked register file to support Amalgam's register-based inter-cluster communication mechanism. This low-latency mechanism allows the RClust to coordinate with a programmable cluster (PClust) as a special purpose junctional unit implementing small custom operations. The relatively large size of the cluster, however, allows it to implement larger, more independent computational kernels. In this extended abstract, we describe the initial design of the RClust and present results from mapping several benchmarks to Amalgam architectures with and without RClust elements. Joshua D. Walstrom, Jeffrey J. Cook, Derek B. Gottlieb, Steve Ferrera, Chi-Wei Wang, Nicholas P. Carter |
FCCM | 6 |
| 2002 | Clustered programmable-reconfigurable processorsabstractIn order to pose a successful challenge to conventional processor architectures, reconfigurable computing systems must achieve significantly better performance than conventional programmable processors by both greatly reducing the number of clock cycles required to execute a wide range of applications and achieving high clock rates when implemented in deep-submicron fabrication technologies. In this paper, we describe the architecture of Amalgam, a clustered programmable-reconfigurable processor that integrates multiple conventional processors and blocks of reconfigurable logic onto a single chip. Amalgam's distributed architecture allows implementation at high clock rates by limiting the impact of wire delay on cycle time and delivers an average of 13.7/spl times/ speedup on our benchmark applications when compared to an equivalent architecture that contains only a single programmable processor. Derek B. Gottlieb, Jeffrey J. Cook, Joshua D. Walstrom, Steve Ferrera, Chi-Wei Wang, Nicholas P. Carter |
FPT | 6 |
| 1998 | The effects of explicitly parallel mechanisms on the multi-ALU processor cluster pipelineabstractContinuing reductions in on-chip geometries yield increasing numbers of transistors per chip and fundamentally faster devices but also result in effectively slower wires. This combination presents significant challenges for new microprocessor architectures. The disparity in performance between on-chip arithmetic units and memory creates longer effectively latencies. The changing balance between gate delay and wire delay penalizes global interactions. The MIT Multi-ALUP processor (IMRP) architecture incorporates three explicitly parallel mechanisms to address these challenges. Efficient intercluster interactions enable instruction scheduling across clustered arithmetic units. Deferred exceptions based on ERRVAL's facilitate aggressive instruction reordering and speculation. Zero-cycle multithreading provides latency tolerance without sacrificing single threaded performance. In this paper; we describe each of these mechanisms and quantify their impact on the area and routing of the cluster pipeline in the 5 Million transistor MAP chip. Zero-cycle multithreading accounts for over 44% of the total cluster area. Support for ERRVAL's requires very little area (less than 4%). The intercluster interaction mechanisms require minimal cluster area and less than 5% of the available global routing resources, but enable fully general access across clusters and between all arithmetic units. Andrew Chang 0001, William J. Dally, Stephen W. Keckler, Nicholas P. Carter, Whay Sing Lee |
ICCD | 4 |
| 1998 | Exploiting Fine-grain Thread Level Parallelism on the MIT Multi-ALU ProcessorabstractMuch of the improvement in computer performance over the last twenty years has come from faster transistors and architectural advances that increase parallelism. Historically, parallelism has been exploited either at the instruction level with a grain-size of a single instruction or by partitioning applications into coarse threads with grain-sizes of thousands of instructions. Fine-grain threads fill the parallelism gap between these extremes by enabling tasks with run lengths as small as 20 cycles. As this fine-grain parallelism is orthogonal to ILP and coarse threads, it complements both methods and provides an opportunity for greater speedup. This paper describes the efficient communication and synchronization mechanisms implemented in the Multi-ALU Processor (MAP) chip, including a thread creation instruction, register communication, and a hardware barrier. These register-based mechanisms provide 10 times faster communication and 60 times faster synchronization than mechanisms that operate via a shared on-chip cache. With a three-processor implementation of the MAP: fine-grain speedups of 1.2-2.1 are demonstrated on a suite of applications. Stephen W. Keckler, William J. Dally, Daniel Maskit, Nicholas P. Carter, Andrew Chang 0001, Whay Sing Lee |
ISCA | 4 |
| 1995 | The M-Machine multicomputerabstractThe M-Machine is an experimental multicomputer being developed to test architectural concepts motivated by the constraints of modern semiconductor technology and the demands of programming systems. The M-Machine computing nodes are connected with a 3-D mesh network; each node is a multithreaded processor incorporating 12 function units, on-chip cache, and local memory. The multiple function units are used to exploit both instruction-level and thread-level parallelism. A user accessible message passing system yields fast communication and synchronization between nodes. Rapid access to remote memory is provided transparently to the user with a combination of hardware and software mechanisms. This paper presents the architecture of the M-Machine and describes how its mechanisms attempt to maximize both single thread performance and overall system throughput. The architecture is complete and the MAP chip, which will serve as the M-Machine processing node, is currently being implemented. Marco Fillo, Stephen W. Keckler, William J. Dally, Nicholas P. Carter, Andrew Chang 0001, Yevgeny Gurevich, Whay Sing Lee |
MICRO | 4 |
| 1994 | Hardware Support for Fast Capability-based AddressingabstractTraditional methods of providing protection in memory systems do so at the cost of increased context switch time and/or increased storage to record access permissions for processes. With the advent of computers that supported cycle-by-cycle multithreading, protection schemes that increase the time to perform a context switch are unacceptable, but protecting unrelated processes from each other is still necessary if such machines are to be used in non-trusting environments. Nicholas P. Carter, Stephen W. Keckler, William J. Dally |
ASPLOS | 1 |
| 1992 | Segmentation and preliminary recognition of madrigals notated in white mensural notation
Nicholas P. Carter |
Mach. Vis. Appl. | 1 |