EDBT 2026 Demo / reviewers in the wild / expert
Perry H. Wang
dblp:49/4683
· DBLP profile ↗
13ranked-venue papers
6as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Processor architecture and microarchitecture · 40% Reconfigurable computing and FPGAs · 28% Memory systems · 11% | |
| Software engineering, system software, and programming languages
5 papers |
Programming languages and type systems · 54% Compilers and program optimization · 24% Operating systems · 14% |
Topics — the 25 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
multithreading |
0.1 | 3 | 2006 | Multiple Instruction Stream Processor · ISCA 2006 Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004 Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002 |
Reconfigurable computing and FPGAs
FPGA-based emulation |
0.1 | 1 | 2010 | Intel nehalem processor core made FPGA synthesizable · FPGA 2010 |
Reconfigurable computing and FPGAs
FPGA prototyping |
0.1 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Reconfigurable computing and FPGAs
processor emulation |
0.1 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
GPUs and heterogeneous computing › heterogeneous programming models
heterogeneous multicore programming |
0.1 | 1 | 2007 | EXOCHI: architecture and programming environment for a heterogeneous multi-core multithreaded system · PLDI 2007 |
Processor architecture and microarchitecture
multiple instruction stream |
0.1 | 1 | 2006 | Multiple Instruction Stream Processor · ISCA 2006 |
Memory systems
cache |
0.0 | 1 | 2004 | Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004 |
Processor architecture and microarchitecture › multithreading
helper threads |
0.0 | 1 | 2004 | Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004 |
Memory systems › cache
prefetching |
0.0 | 1 | 2004 | Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004 |
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution |
0.0 | 2 | 2002 | Register Renaming and Scheduling for Dynamic Execution of Predicated Code · HPCA 2001 Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002 |
Processor architecture and microarchitecture
memory latency tolerance |
0.0 | 1 | 2002 | Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002 |
Memory systems › memory access optimization
memory prefetching |
0.0 | 1 | 2002 | Post-Pass Binary Adaptation for Software-Based Speculative Precomputation · PLDI 2002 |
Processor architecture and microarchitecture
out-of-order execution |
0.0 | 1 | 2002 | Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002 |
Reconfigurable computing and FPGAs › FPGA-based emulation
multi-FPGA emulation |
0.0 | 1 | 2010 | Intel nehalem processor core made FPGA synthesizable · FPGA 2010 |
Parallel and multicore computing › task scheduling
dynamic scheduling |
0.0 | 1 | 2001 | Register Renaming and Scheduling for Dynamic Execution of Predicated Code · HPCA 2001 |
Electronic design automation › hardware verification and test
hardware verification |
0.0 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
Electronic design automation › hardware verification and test › functional verification
pre-silicon verification |
0.0 | 1 | 2009 | Intel® atomTM processor core made FPGA-synthesizable · FPGA 2009 |
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous multicore architecture |
0.0 | 1 | 2007 | EXOCHI: architecture and programming environment for a heterogeneous multi-core multithreaded system · PLDI 2007 |
Operating systems › resource management › process management
thread management |
0.0 | 1 | 2006 | Multiple Instruction Stream Processor · ISCA 2006 |
Compilers and program optimization
register allocation |
0.0 | 1 | 2004 | Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platform · ASPLOS 2004 |
Program analysis › dynamic analysis › instrumentation
binary instrumentation |
0.0 | 1 | 2002 | Post-Pass Binary Adaptation for Software-Based Speculative Precomputation · PLDI 2002 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 1 | 2002 | Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.0 | 1 | 2002 | Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative Precomputation · HPCA 2002 |
Compilers and program optimization › instruction scheduling
instruction-level parallelism |
0.0 | 1 | 2001 | Register Renaming and Scheduling for Dynamic Execution of Predicated Code · HPCA 2001 |
Compilers and program optimization
predicated compilation |
0.0 | 1 | 2001 | Register Renaming and Scheduling for Dynamic Execution of Predicated Code · HPCA 2001 |
Methods — techniques the papers use, named apart from their topics
fat binary compilation · 0.1OpenMP pragma extension · 0.1sequencer architecture · 0.1cache-coherent shared memory · 0.1asynchronous control transfer · 0.1latch mapping · 0.1clock gating conversion · 0.1FPGA synthesis · 0.1prefetching · 0.1performance evaluation · 0.1user-level multithreading · 0.0speculative prefetch threads · 0.0post-pass compilation · 0.0microarchitecture simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL InferenceabstractGeneral Matrix Multiply (GEMM) units, consisting of multiply-accumulate (MAC) arrays, perform bulk of the computation in deep learning (DL). Recent work has proposed a novel MAC design, Bit-Pragmatic (PRA), capable of dynamically exploiting bit sparsity. This work presents OzMAC (Omit-zero-MAC), a modified re-implementation of PRA, but extends beyond earlier works by performing rigorous post-synthesis evaluation against binary MAC design across multiple bitwidths and clock frequencies using TSMC N5 process node to assess commercial implementation potential. We demonstrate the existence of high bit sparsity in eight pretrained INT8 DL workloads and show that 8-bit OzMAC improves all three metrics of area, power, and energy significantly by 21%, 70%, and 28%, respectively. Similar improvements are achieved when scaling data precisions (4, 8, 16 bits) and clock frequencies (0.5 GHz, 1 GHz, 1.5 GHz). For the 8-bit OzMAC, scaling its frequency to normalize the throughput, it still achieves 30% improvement on both power and energy. Harideep Nair, Prabhu Vellaisamy, Tsung-Han Lin, Perry H. Wang, R. D. (Shawn) Blanton, John Paul Shen |
VLSI-SoC | 4 |
| 2010 | Intel nehalem processor core made FPGA synthesizableabstractWe present a FPGA-synthesizable version of the Intel Nehalem processor core, synthesized, partitioned and mapped to a multi-FPGA emulation system consisting of Xilinx Virtex-4 and Virtex-5 FPGAs. To our knowledge, this is the first time a modern state-of-the-art x86 design with the out-of-order micro-architecture is made FPGA synthesizable and capable of high-speed cycle-accurate emulation. Unlike the Intel Atom core which was made FPGA synthesizable on a single Xilinx Virtex-5 in a previous endeavor, the Nehalem core is a more complex design with aggressive clock-gating, double phase latch RAMs, and RTL constructs that have no true equivalent in FPGA architectures. Despite these challenges, we are successful in making the RTL synthesizable with only 5% RTL code modifications, partitioning the design across five FPGAs, and emulating the core at 520 KHz. The synthesizable Nehalem core is able to boot Linux and execute standard x86 workloads with all architectural features enabled. Graham Schelle, Jamison D. Collins, Ethan Schuchman, Perry H. Wang, Gautham N. Chinya, Ralf Plate, Thorsten Mattner, Franz Olbrich, Per Hammarlund, Ronak Singhal, Jim Brayton, Sebastian Steibl, Hong Wang 0003 |
FPGA | 4 |
| 2009 | Intel® atomTM processor core made FPGA-synthesizableabstractWe present an FPGA-synthesizable version of the Intel Atom processor core, synthesized to a Virtex-5 based FPGA emulation system. To make the production Atom design in SystemVerilog synthesizable through industry standard EDA tool flow, we transformed and mapped latches in the design, converted clock gating, and replaced nonsynthesizable constructs with FPGA-synthesizable counterparts. Additionally, as the target FPGA emulator is hosted on a PC platform with the Pentium-based CPU socket that supports a significantly different front side bus (FSB) protocol from that of the Atom processor, we replaced the existing bus control logic in the Atom core with an alternate FSB protocol to communicate with the rest of the PC platform. With these efforts, we succeeded in synthesizing the entire Atom processor core to fit within a single Virtex-5 LX330 FPGA. The synthesizable Atom core runs at 50Mhz on the Pentium PC motherboard with fully functional I/O peripherals. It is capable of booting off-the-shelf MS-DOS, Windows XP and Linux operating systems, and executing standard x86 workloads. Perry H. Wang, Jamison D. Collins, Christopher T. Weaver, Belliappa Kuttanna, Shahram Salamian, Gautham N. Chinya, Ethan Schuchman, Oliver Schilling, Thorsten Doil, Sebastian Steibl, Hong Wang 0003 |
FPGA | 1 |
| 2008 | Pangaea: a tightly-coupled IA32 heterogeneous chip multiprocessorabstractMoore's Law and the drive towards performance efficiency have led to the on-chip integration of general-purpose cores with special-purpose accelerators. Pangaea is a heterogeneous CMP design for non-rendering workloads that integrates IA32 CPU cores with non-IA32 GPU-class multi-cores, extending the current state-of-the-art CPU-GPU integration that physically "fuses" existing CPU and GPU designs. Pangaea introduces (1) a resource repartitioning of the GPU, where the hardware budget dedicated for 3D-specific graphics processing is used to build more general-purpose GPU cores, and (2) a 3-instruction extension to the IA32 ISA that supports tighter architectural integration and fine-grain shared memory collaborative multithreading between the IA32 CPU cores and the non-IA32 GPU cores. We implement Pangaea and the current CPU-GPU designs in fully-functional synthesizable RTL based on the production quality RTL of an IA32 CPU and an Intel GMA X4500 GPU. On a 65 nm ASIC process technology, the legacy graphics-specific fixed-function hardware has the area of 9 GPU cores and total power consumption of 5 GPU cores. With the ISA extensions, the latency from the time an IA32 core spawns a GPU thread to the time the thread begins execution is reduced from thousands of cycles to fewer than 30 cycles. Pangaea is synthesized on a FPGA-based prototype and runs off-the-shelf IA32 OSes. A set of general-purpose non-graphics workloads demonstrate speedups of up to 8.8x. Henry Wong, Anne Bracy, Ethan Schuchman, Tor M. Aamodt, Jamison D. Collins, Perry H. Wang, Gautham N. Chinya, Ankur Khandelwal Groen, Hong Wang 0003 |
PACT | 6 |
| 2007 | Sequencer virtualizationabstractThe Multiple Instruction Stream Processor (MISP) architecture introduces the sequencer as a new class of architectural resource, and provides a minimalist user-level MIMD instruction set extension for application programs to directly control execution of concurrent instruction streams on these sequencers. As with classic architectural resources, namely, registers and memory, the sequencer architectural resource can be subject to virtualization. This paper details the idea of Sequencer Virtualization (SV), a foundational architectural support to decouple architectural virtual sequencers from physical sequencers. SV enables more efficient utilization of sequencer resources at the microarchitectural level while maintaining a consistent programming interface at the architectural level. To evaluate the key tradeoffs for SV, we conduct extensive experiments by implementing a prototype SV system using a custom firmware on a large-scale multiprocessor system. Using the prototype SV system, we demonstrate that SV improves efficiency in sequencer utilization while incurring little performance overhead. In particular, for a set of real multithreaded workloads, SV can significantly improve sequencer utilization, achieving an average of 32% better wall-clock performance than MISP without SV support in a multi-programming environment. Perry H. Wang, Jamison D. Collins, Gautham N. Chinya, Bernard Lint, Asit Mallick, Koichi Yamada, Hong Wang 0003 |
ICS | 1 |
| 2007 | EXOCHI: architecture and programming environment for a heterogeneous multi-core multithreaded systemabstractFuture mainstream microprocessors will likely integrate specialized accelerators, such as GPUs, onto a single die to achieve better performance and power efficiency. However, it remains a keen challenge to program such a heterogeneous multicore platform, since these specialized accelerators feature ISAs and functionality that are significantly different from the general purpose CPU cores. In this paper, we present EXOCHI: (1) Exoskeleton Sequencer(EXO), an architecture to represent heterogeneous acceleratorsas ISA-based MIMD architecture resources, and a shared virtual memory heterogeneous multithreaded program execution model that tightly couples specialized accelerator cores with generalpurpose CPU cores, and (2) C for Heterogeneous Integration(CHI), an integrated C/C++ programming environment that supports accelerator-specific inline assembly and domain-specific languages. The CHI compiler extends the OpenMP pragma for heterogeneous multithreading programming, and produces a single fat binary with code sections corresponding to different instruction sets. The runtime can judiciously spread parallel computation across the heterogeneous cores to optimize performance and power. Perry H. Wang, Jamison D. Collins, Gautham N. Chinya, Xinmin Tian, Milind Girkar, Nick Y. Yang, Guei-Yuan Lueh, Hong Wang 0003 |
PLDI | 1 |
| 2006 | Multiple Instruction Stream ProcessorabstractMicroprocessor design is undergoing a major paradigm shift towards multi-core designs, in anticipation that future performance gains will come from exploiting threadlevel parallelism in the software. To support this trend, we present a novel processor architecture called the Multiple Instruction Stream Processing (MISP) architecture. MISP introduces the sequencer as a new category of architectural resource, and defines a canonical set of instructions to support user-level inter-sequencer signaling and asynchronous control transfer. MISP allows an application program to directly manage user-level threads without OS intervention. By supporting the classic cache-coherent shared-memory programming model, MISP does not require a radical shift in the multithreaded programming paradigm. This paper describes the design and evaluation of the MISP architecture for the IA-32 family of microprocessors. Using a research prototype MISP processor built on an IA-32-based multiprocessor system equipped with special firmware, we demonstrate the feasibility of implementing the MISP architecture. We then examine the utility of MISP by (1) assessing the key architectural tradeoffs of the MISP architecture design and (2) showing how legacy multithreaded applications can be migrated to MISP with relative ease. Richard A. Hankins, Gautham N. Chinya, Jamison D. Collins, Perry H. Wang, Ryan N. Rakvic, Hong Wang 0003, John Paul Shen |
ISCA | 4 |
| 2004 | Helper threads via virtual multithreading on an experimental itanium® 2 processor-based platformabstractHelper threading is a technology to accelerate a program by exploiting a processor's multithreading capability to run ``assist'' threads. Previous experiments on hyper-threaded processors have demonstrated significant speedups by using helper threads to prefetch hard-to-predict delinquent data accesses. In order to apply this technique to processors that do not have built-in hardware support for multithreading, we introduce virtual multithreading (VMT), a novel form of switch-on-event user-level multithreading, capable of fly-weight multiplexing of event-driven thread executions on a single processor without additional operating system support. The compiler plays a key role in minimizing synchronization cost by judiciously partitioning register usage among the user-level threads. The VMT approach makes it possible to launch dynamic helper thread instances in response to long-latency cache miss events, and to run helper threads in the shadow of cache misses when the main thread would be otherwise stalled.The concept of VMT is prototyped on an Itanium ® 2 processor using features provided by the Processor Abstraction Layer (PAL) firmware mechanism already present in currently shipping processors. On a 4-way MP physical system equipped with VMT-enabled Itanium 2 processors, helper threading via the VMT mechanism can achieve significant performance gains for a diverse set of real-world workloads, ranging from single-threaded workstation benchmarks to heavily multithreaded large scale decision support systems (DSS) using the IBM DB2 Universal Database. We measure a wall-clock speedup of 5.8% to 38.5% for the workstation benchmarks, and 5.0% to 12.7% on various queries in the DSS workload. Perry H. Wang, Jamison D. Collins, Hong Wang 0003, Dongkeun Kim, Bill Greene, Kai-Ming Chan, Aamir B. Yunus, Terry Sych, Stephen F. Moore, John Paul Shen |
ASPLOS | 1 |
| 2004 | Physical Experimentation with Prefetching Helper Threads on Intel's Hyper-Threaded ProcessorsabstractPre-execution techniques have received much attention as an effective way of prefetching cache blocks to tolerate the ever-increasing memory latency. A number of pre-execution techniques based on hardware, compiler, or both have been proposed and studied extensively by researchers. They report promising results on simulators that model a simultaneous multithreading (SMT) processor. We apply the helper threading idea on a real multithreaded machine, i.e., Intel Pentium 4 processor with hyper-threading technology, and show that indeed it can provide wall-clock speedup on real silicon. To achieve further performance improvements via helper threads, we investigate three helper threading scenarios that are driven by automated compiler infrastructure, and identify several key challenges and opportunities for novel hardware and software optimizations. Our study shows a program behavior changes dynamically during execution. In addition, the organizations of certain critical hardware structures in the hyper-threaded processors are either shared or partitioned in the multithreading mode and thus, the tradeoffs regarding resource contention can be intricate. Therefore, it is essential to judiciously invoke helper threads by adapting to the dynamic program behavior so that we can alleviate potential performance degradation due to resource contention. Moreover, since adapting to the dynamic behavior requires frequent thread synchronization, having light-weight thread synchronization mechanisms is important. Dongkeun Kim, Shih-Wei Liao, Perry H. Wang, Juan del Cuvillo, Xinmin Tian, Hong Wang 0003, Donald Yeung, Milind Girkar, John Paul Shen |
CGO | 3 |
| 2003 | Inferno: a functional simulation infrastructure for modeling microarchitectural data speculationsabstractThis paper presents key insights and design rationales behind Inferno, a functional simulation construction framework developed at Intel to support execution-driven cycle-accurate performance modeling and simulation of advanced microarchitectural data speculation techniques for future processor designs and explorations. Inferno divides the task of functional simulation into three essential components, namely: (1) context manager of in-flight speculatively executed instructions, (2) stateless emulator of instruction semantics, and (3) a high speed functional simulator capable of booting OS and running large-scale system workloads. These building block components work together in concert via a set of well-architected functional model convergence APIs. With a novel abstraction called speculative domain, the context manager serves effectively as a relational database about the in-flight instructions and on-going speculations. Through a set of functional model usage APIs, the context manager enables performance models to express arbitrary microarchitectural data speculation scenarios. The contribution of this paper is to demonstrate the importance of providing a functional model with modular construction, proper abstraction and expressive APIs for speculative state management. Inferno is such a functional model construction framework and has significantly improved productivity in modeling a variety of sophisticated data speculation microarchitecture techniques. Hong Wang 0003, Shiri Manor, Dave LaFollette, Nadav Nesher, Ku-jei King, Perry H. Wang, Shay Levy, Shai Satt, Gal Carmeli, Arjun Kapur, Ioannis Schoinas, Ed Rubinstein, Rahul Bhatt |
ISPASS | 6 |
| 2002 | Memory Latency-Tolerance Approaches for Itanium Processors: Out-of-Order Execution vs. Speculative PrecomputationabstractThe performance of in-order execution Itanium/sup TM/ processors can suffer significantly due to cache misses. Two memory latency tolerance approaches can be applied for the Itanium processors. One uses an out-of-order (OOO) execution core; the other assumes multithreading support and exploits cache prefetching via speculative precomputation (SP). This paper evaluates and contrasts these two approaches. In addition, this paper assesses the effectiveness of combining the two approaches. For a select set of memory-intensive programs, an in-order SMT Itanium processor using speculative precomputation can achieve performance improvement (92%) comparable to that of an out-of-order design (87%). Applying both 000 and SP yields a total performance improvement of 141% over the baseline in-order machine. OOO tends to be effective in prefetching-for L1 misses; whereas SP is primarily good at covering L2 and L3 misses. Our analysis indicates that the two approaches can be redundant or complementary depending on the type of delinquent loads that each targets. Both approaches are effective on delinquent loads in the loop body; however only SP is effective on delinquent loads found in loop control code. Perry H. Wang, Hong Wang 0003, Jamison D. Collins, Ed Grochowski, Ralph-Michael Kling, John Paul Shen |
HPCA | 1 |
| 2002 | Post-Pass Binary Adaptation for Software-Based Speculative PrecomputationabstractRecently, a number of thread-based prefetching techniques have been proposed. These techniques aim at improving the latency of single-threaded applications by leveraging multithreading resources to perform memory prefetching via speculative prefetch threads. Software-based speculative precomputation (SSP) is one such technique, proposed for multithreaded Itanium models. SSP does not require expensive hardware support-instead it relies on the compiler to adapt binaries to perform prefetching on otherwise idle hardware thread contexts at run time. This paper presents a post-pass compilation tool for generating SSP-enhanced binaries. The tool is able to: (1) analyze a single-threaded application to generate prefetch threads; (2) identify and embed trigger points in the original binary; and (3) produce a new binary that has the prefetch threads attached. The execution of the new binary spawns the speculative prefetch threads, which are executed concurrently with the main thread. Our results indicate that for a set of pointer-intensive benchmarks, the prefetching performed by the speculative threads achieves an average of 87% speedup on an in-order processor and 5% speedup on an out-of-order processor. Shih-Wei Liao, Perry H. Wang, Hong Wang 0003, John Paul Shen, Gerolf Hoflehner, Daniel M. Lavery |
PLDI | 2 |
| 2001 | Register Renaming and Scheduling for Dynamic Execution of Predicated CodeabstractTo achieve higher processor performance requires greater synergy between advanced hardware features and innovative compiler techniques. Recent advancement in compilation techniques for predicated execution has provided significant opportunity in exploiting instruction level parallelism. However, little research has been done on how to efficiently execute predicated code in a dynamic microarchitecture. In this paper, we evaluate hardware optimizations for executing predicated code on a dynamically scheduled microarchitecture. We provide two novel ideas to improve the efficiency of executing predicated code. On a generic Intel Itanium processor pipeline model, we demonstrate that, with some microarchitecture enhancements, a dynamic execution processor can achieve about 16% performance improvement over an equivalent static execution processor. Perry H. Wang, Hong Wang 0003, Ralph-Michael Kling, Kalpana Ramakrishnan, John Paul Shen |
HPCA | 1 |