EDBT 2026 Demo / reviewers in the wild / expert
Ethan Schuchman
dblp:29/45
· DBLP profile ↗
7ranked-venue papers
3as first author
0since 2021 · last 2010
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorSecurity and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Reconfigurable computing and FPGAs · 43% Processor architecture and microarchitecture · 26% Energy-efficient computing · 12% | |
| Theoretical computer science
1 paper |
Quantum computing and quantum information · 100% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA-based emulation |
0.1 | 1 | 2010 | Intel nehalem processor core made FPGA synthesizable · FPGA 2010 |
Reconfigurable computing and FPGAs
FPGA prototyping |
0.1 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Reconfigurable computing and FPGAs
processor emulation |
0.1 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Emerging computing paradigms
quantum computer architecture |
0.1 | 1 | 2006 | A program transformation and architecture support for quantum uncomputation · ASPLOS 2006 |
Processor architecture and microarchitecture › pipelining
pipelined processor |
0.1 | 1 | 2005 | Balancing Resource Utilization to Mitigate Power Density in Processor Pipelines · MICRO 2005 |
Energy-efficient computing
power density |
0.1 | 1 | 2005 | Balancing Resource Utilization to Mitigate Power Density in Processor Pipelines · MICRO 2005 |
Reconfigurable computing and FPGAs › FPGA resource optimization
resource utilization balancing |
0.1 | 1 | 2005 | Balancing Resource Utilization to Mitigate Power Density in Processor Pipelines · MICRO 2005 |
Energy-efficient computing
thermal management |
0.1 | 1 | 2005 | Balancing Resource Utilization to Mitigate Power Density in Processor Pipelines · MICRO 2005 |
Reconfigurable computing and FPGAs › FPGA-based emulation
multi-FPGA emulation |
0.0 | 1 | 2010 | Intel nehalem processor core made FPGA synthesizable · FPGA 2010 |
Electronic design automation › hardware verification and test
hardware verification |
0.0 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Electronic design automation › hardware verification and test › functional verification
pre-silicon verification |
0.0 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Electronic design automation
hardware verification and test |
0.0 | 1 | 2005 | Rescue: A Microarchitecture for Testability and Defect Tolerance · ISCA 2005 |
Processor architecture and microarchitecture › out-of-order execution
issue queue |
0.0 | 1 | 2005 | Balancing Resource Utilization to Mitigate Power Density in Processor Pipelines · MICRO 2005 |
Hardware reliability and fault tolerance
permanent fault tolerance |
0.0 | 1 | 2005 | Rescue: A Microarchitecture for Testability and Defect Tolerance · ISCA 2005 |
Processor architecture and microarchitecture
register file |
0.0 | 1 | 2005 | Balancing Resource Utilization to Mitigate Power Density in Processor Pipelines · MICRO 2005 |
Electronic design automation › hardware verification and test › design for testability
scan-based testing |
0.0 | 1 | 2005 | Rescue: A Microarchitecture for Testability and Defect Tolerance · ISCA 2005 |
Methods — techniques the papers use, named apart from their topics
scheduling policy · 0.2program transformation · 0.1latch mapping · 0.1clock gating conversion · 0.1FPGA synthesis · 0.1logic transformation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2010 | Intel nehalem processor core made FPGA synthesizableabstractWe present a FPGA-synthesizable version of the Intel Nehalem processor core, synthesized, partitioned and mapped to a multi-FPGA emulation system consisting of Xilinx Virtex-4 and Virtex-5 FPGAs. To our knowledge, this is the first time a modern state-of-the-art x86 design with the out-of-order micro-architecture is made FPGA synthesizable and capable of high-speed cycle-accurate emulation. Unlike the Intel Atom core which was made FPGA synthesizable on a single Xilinx Virtex-5 in a previous endeavor, the Nehalem core is a more complex design with aggressive clock-gating, double phase latch RAMs, and RTL constructs that have no true equivalent in FPGA architectures. Despite these challenges, we are successful in making the RTL synthesizable with only 5% RTL code modifications, partitioning the design across five FPGAs, and emulating the core at 520 KHz. The synthesizable Nehalem core is able to boot Linux and execute standard x86 workloads with all architectural features enabled. Graham Schelle, Jamison D. Collins, Ethan Schuchman, Perry H. Wang, Gautham N. Chinya, Ralf Plate, Thorsten Mattner, Franz Olbrich, Per Hammarlund, Ronak Singhal, Jim Brayton, Sebastian Steibl, Hong Wang 0003 |
FPGA | 3 |
| 2009 | Intel® atomTM processor core made FPGA-synthesizableabstractWe present an FPGA-synthesizable version of the Intel Atom processor core, synthesized to a Virtex-5 based FPGA emulation system. To make the production Atom design in SystemVerilog synthesizable through industry standard EDA tool flow, we transformed and mapped latches in the design, converted clock gating, and replaced nonsynthesizable constructs with FPGA-synthesizable counterparts. Additionally, as the target FPGA emulator is hosted on a PC platform with the Pentium-based CPU socket that supports a significantly different front side bus (FSB) protocol from that of the Atom processor, we replaced the existing bus control logic in the Atom core with an alternate FSB protocol to communicate with the rest of the PC platform. With these efforts, we succeeded in synthesizing the entire Atom processor core to fit within a single Virtex-5 LX330 FPGA. The synthesizable Atom core runs at 50Mhz on the Pentium PC motherboard with fully functional I/O peripherals. It is capable of booting off-the-shelf MS-DOS, Windows XP and Linux operating systems, and executing standard x86 workloads. Perry H. Wang, Jamison D. Collins, Christopher T. Weaver, Belliappa Kuttanna, Shahram Salamian, Gautham N. Chinya, Ethan Schuchman, Oliver Schilling, Thorsten Doil, Sebastian Steibl, Hong Wang 0003 |
FPGA | 7 |
| 2008 | Pangaea: a tightly-coupled IA32 heterogeneous chip multiprocessorabstractMoore's Law and the drive towards performance efficiency have led to the on-chip integration of general-purpose cores with special-purpose accelerators. Pangaea is a heterogeneous CMP design for non-rendering workloads that integrates IA32 CPU cores with non-IA32 GPU-class multi-cores, extending the current state-of-the-art CPU-GPU integration that physically "fuses" existing CPU and GPU designs. Pangaea introduces (1) a resource repartitioning of the GPU, where the hardware budget dedicated for 3D-specific graphics processing is used to build more general-purpose GPU cores, and (2) a 3-instruction extension to the IA32 ISA that supports tighter architectural integration and fine-grain shared memory collaborative multithreading between the IA32 CPU cores and the non-IA32 GPU cores. We implement Pangaea and the current CPU-GPU designs in fully-functional synthesizable RTL based on the production quality RTL of an IA32 CPU and an Intel GMA X4500 GPU. On a 65 nm ASIC process technology, the legacy graphics-specific fixed-function hardware has the area of 9 GPU cores and total power consumption of 5 GPU cores. With the ISA extensions, the latency from the time an IA32 core spawns a GPU thread to the time the thread begins execution is reduced from thousands of cycles to fewer than 30 cycles. Pangaea is synthesized on a FPGA-based prototype and runs off-the-shelf IA32 OSes. A set of general-purpose non-graphics workloads demonstrate speedups of up to 8.8x. Henry Wong, Anne Bracy, Ethan Schuchman, Tor M. Aamodt, Jamison D. Collins, Perry H. Wang, Gautham N. Chinya, Ankur Khandelwal Groen, Hong Wang 0003 |
PACT | 3 |
| 2007 | BlackJack: Hard Error Detection with Redundant Threads on SMTabstractTesting is a difficult process that becomes more difficult with scaling. With smaller and faster devices, tolerance for errors shrinks and devices may act correctly under certain condition and not under others. As such, hard errors may exist but are only exercised by very specific machine state and signal pathways. Targeting these errors is difficult, and creating test cases that cover all machine states and pathways is not possible. In addition, new complications during burn-in may mean latent hard errors are not exposed in the fab and reach the customer before becoming active. To address this problem, we propose an architecture we call BlackJack that allows hard errors to be detected using redundant threads running on a single SMT core. This technique provides a safety-net that catches hard errors that were either latent during test or just not covered by the test cases at all. Like SRT, our technique works by executing redundant copies and verifying that their resulting machine states agree. Unlike SRT, BlackJack is able to achieve high hard error instruction coverage by executing redundant threads on different front and backend resources in the pipeline. We show that for a 15% performance penalty over SRT, BlackJack achieves 97% hard error instruction coverage compared to SRT's 35%. Ethan Schuchman, T. N. Vijaykumar |
DSN | 1 |
| 2006 | A program transformation and architecture support for quantum uncomputationabstractQuantum computing's power comes from new algorithms that exploit quantum mechanical phenomena for computation. Quantum algorithms are different from their classical counterparts in that quantum algorithms rely on algorithmic structures that are simply not present in classical computing. Just as classical program transformations and architectures have been designed for common classical algorithm structures, quantum program transformations and quantum architectures should be designed with quantum algorithms in mind. Because quantum algorithms come with these new algorithmic structures, resultant quantum program transformations and architectures may look very different from their classical counterparts.This paper focuses on uncomputation, a critical and prevalent structure in quantum algorithms, and considers how program transformations, and architecture support should be designed to accommodate uncomputation. In this paper,we show a simple quantum program transformation that exposes independence between uncomputation and later computation. We then propose a multicore architecture tailored to this exposed parallelism and propose a scheduling policy that efficiently maps such parallelism to the multicore architecture. Our policy achieves parallelism between uncomputation and later computation while reducing cumulative communication distance. Our scheduling and architecture allows significant speedup of quantum programs (between 1.8x and 2.8x speedup in Shor's factoring algorithm), while reducing cumulative communication distance 26%. Ethan Schuchman, T. N. Vijaykumar |
ASPLOS | 1 |
| 2005 | Rescue: A Microarchitecture for Testability and Defect ToleranceabstractScaling feature size improves processor performance but increases each device's susceptibility to defects (i.e., hard errors). As a result, fabrication technology must improve significantly to maintain yields. Redundancy techniques in memory have been successful at improving yield in the presence of defects. Apart from core sparing which disables faulty cores in a chip multiprocessor, little has been done to target the core logic. While previous work has proposed that either inherent or added redundancy in the core logic can be used to tolerate defects, the key issues of realistic testing and fault isolation have been ignored. This paper is the first to consider testability and fault isolation in designing modern high-performance, defect-tolerant microarchitectures. We define intra-cycle logic independence (ICI) as the condition needed for conventional scan test to isolate faults quickly to the microarchitectural-block granularity. We propose logic transformations to redesign conventional superscalar microarchitecture to comply with ICI. We call our novel, testable, and defect-tolerant microarchitecture Rescue. Ethan Schuchman, T. N. Vijaykumar |
ISCA | 1 |
| 2005 | Balancing Resource Utilization to Mitigate Power Density in Processor PipelinesabstractPower density is a growing problem in high-performance processors in which small, high-activity resources overheat. Two categories of techniques, temporal and spatial, can address power density in a processor. Temporal solutions slow computation and heating either through frequency and voltage scaling or through stopping computation long enough to allow the processor to cool; both degrade performance. Spatial solutions reduce heat by moving computation from a hot resource to an alternate resource (e.g., a spare ALU) to allow cooling. Spatial solutions are appealing because they have negligible impact on performance, but they require availability of spatial slack in the form of spare or underutilized resource copies. Previous work focusing on spatial slack within a pipeline has proposed adding extra resource copies to the pipeline, which adds substantial complexity because the resources that overheat, issue logic, register files, and ALUs, are the resources in some of the tightest critical paths in the pipeline. Previous work has not considered exploiting the spatial slack already existing within pipeline resource copies. Utilization can be quite asymmetric across resource copies, leaving some copies substantially cooler than others. We observe that asymmetric utilization within copies of three key back-end resources, the issue queue, register files, and ALUs, creates spatial slack opportunities. By balancing asymmetry in their utilization, we can reduce power density. Scheduling policies for these resources were designed for maximum simplicity before power density was a concern; our challenge is to address asymmetric heating while keeping the pipeline simple. Balancing asymmetric utilization reduces the need for other performance-degrading temporal power-density techniques. While our techniques do not obviate temporal techniques in high-resource-utilization applications, we greatly reduce their use, improving overall performance. Michael D. Powell, Ethan Schuchman, T. N. Vijaykumar |
MICRO | 2 |