EDBT 2026 Demo / reviewers in the wild / expert
H. Peter Hofstee
dblp:67/1067
· DBLP profile ↗
36ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0001-9649-7338ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Theory of computation · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenMPI: Cluster Scalable Variant Calling for Short/Long Reads Sequencing DataabstractRapid technological advancements in sequencing technologies allow producing cost effective and high volume sequencing data. Processing this data for real-time clinical diagnosis is potentially time-consuming if done on a single computing node. This work presents a complete variant calling workflow, implemented using the Message Passing Interface (MPI) to leverage the benefits of high bandwidth interconnects. This solution (GenMPI) is portable and flexible, meaning it can be deployed to any private or public cluster/cloud infrastructure. Any alignment or variant calling application can be used with minimal adaptation. To achieve high performance, compressed input data can be streamed in parallel to alignment applications while uncompressed data can use internal file seek functionality to eliminate the bottleneck of streaming input data from a single node. Alignment output can be directly stored in multiple chromosome-specific SAM files or a single SAM file. After alignment, a distributed queue using MPI RMA (Remote Memory Access) atomic operations is created for sorting, indexing, marking of duplicates (if necessary) and variant calling applications. We ensure the accuracy of variants as compared to the original single node methods. We also show that for 300x coverage data, alignment scales almost linearly up to 64 nodes (8192 CPU cores). Overall, this work outperforms existing Big Data based workflows by a factor of two and is almost 20% faster than other MPI-based implementations for alignment without any extra memory overheads. Sorting, indexing, duplicate removal and variant calling is also scalable up to 8 nodes cluster. For pair-end short-reads (Illumina) data, we integrated the BWA-MEM aligner and three variant callers (GATK HaplotypeCaller, DeepVariant and Octopus), while for long-reads data, we integrated the Minimap2 aligner and three different variant callers (DeepVariant, DeepVariant with WhatsHap for phasing (PacBio) and Clair3 (ONT)). Joseph Schuchart, Zaid Al-Ars, Christoph Niethammer, José Gracia, H. Peter Hofstee |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | Open-Vocabulary Object Detection with Driving-Aware Multi-Scale Feature Fusion for Autonomous DrivingabstractOpen-vocabulary object detection (OVD) is crucial for handling dynamic real-world driving scenarios. Inspired by YOLO-World, we propose OpenVocab-Auto, an open-vocabulary object detection framework with driving-aware multi-scale feature fusion for autonomous driving scenarios. Our system introduces three key innovations: (1) a context-adaptive prompt engine that significantly reduces computational overhead compared to global prompt strategies, (2) hierarchical vision-language alignment for improved small object detection, and (3) real-time optimization achieving 27 FPS on NVIDIA Jetson AGX Orin through TensorRT acceleration. On RTX 3080 (FP16 full model), the framework achieves 0.923 F1 for parking space detection and 0.847 F1 for zero-shot obstacle recognition. On Jetson Orin (TensorRT INT8 model), the corresponding scores are 0.811 and 0.333, respectively, under the same evaluation protocol. Yongtao Yao, H. Peter Hofstee, Weisong Shi |
SEC | 3 |
| 2024 | Hardware-Accelerator Design by Composition: Dataflow Component Interfaces With Tydi-ChiselabstractAs dedicated hardware is becoming more prevalent in accelerating complex applications, methods are needed to enable easy integration of multiple hardware components into a single accelerator system. However, this vision of composable hardware is hindered by the lack of standards for interfaces that allow such components to communicate. To address this challenge, the Tydi standard was proposed to facilitate the representation of streaming data in digital circuits, notably providing interface specifications of composite and variable-length data structures. At the same time, constructing hardware in a Scala embedded language (Chisel) provides a suitable environment for deploying Tydi-centric components due to its abstraction level and customizability. This article introduces Tydi-Chisel, a library that integrates the Tydi standard within Chisel, along with a toolchain and methodology for designing data-streaming accelerators. This toolchain reduces the effort needed to design streaming hardware accelerators by raising the abstraction level for streams and module interfaces, hereby avoiding writing boilerplate code, and allows for easy integration of accelerator components from different designers. This is demonstrated through an example project incorporating various scenarios where the interface-related declaration is reduced by 6–14 times. Tydi-Chisel project repository is available athttps://github.com/abs-tudelft/Tydi-Chisel. Casper Cromjongh, Yongding Tian, H. Peter Hofstee, Zaid Al-Ars |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Improving Gradient Paths for Binary Convolutional Neural Networks
Baozhou Zhu, H. Peter Hofstee, Jinho Lee 0001, Zaid Al-Ars |
BMVC | 2 |
| 2022 | SALoBa: Maximizing Data Locality and Workload Balance for Fast Sequence Alignment on GPUsabstractSequence alignment forms an important backbone in many sequencing applications. A commonly used strategy for sequence alignment is an approximate string matching with a two-dimensional dynamic programming approach. Although some prior work has been conducted on GPU acceleration of a sequence alignment, we identify several shortcomings that limit exploiting the full computational capability of modern GPUs. This paper presents SALoBa, a GPU-accelerated sequence alignment library focused on seed extension. Based on the analysis of previous work with real-world sequencing data, we propose techniques to exploit the data locality and improve work-load balancing. The experimental results reveal that SALoBa significantly improves the seed extension kernel compared to state-of-the-art GPU-based methods. Seongyeon Park, Hajin Kim, Nauman Ahmed, Zaid Al-Ars, H. Peter Hofstee, Youngsok Kim, Jinho Lee 0001 |
IPDPS | 6 |
| 2022 | Communication-Efficient Cluster Scalable Genomics Data Processing Using Apache Arrow FlightabstractCurrent cluster scaled genomics data processing solutions rely on big data frameworks like Apache Spark, Hadoop and HDFS for data scheduling, processing and storage. These frameworks come with additional computation and memory overheads by default. It has been observed that scaling genomics dataset processing beyond 32 nodes is not efficient on such frameworks.To overcome the inefficiencies of big data frameworks for processing genomics data on clusters, we introduce a low-overhead and highly scalable solution on a SLURM based HPC batch system. This solution uses Apache Arrow as in-memory columnar data format to store genomics data efficiently and Arrow Flight as a network protocol to move and schedule this data across the HPC nodes with low communication overhead.As a use case, we use NGS short reads DNA sequencing data for pre-processing and variant calling applications. This solution outperforms existing Apache Spark based big data solutions in term of both computation time (2x) and lower communication overhead (more than 20-60% depending on cluster size). Our solution has similar performance to MPI-based HPC solutions, with the added advantage of easy programmability and transparent big data scalability. The whole solution is Python and shell script based, which makes it flexible to update and integrate alternative variant callers. Our solution is publicly available on GitHub at https://github.com/abs-tudelft/time-to-fly-high/tree/main/genomics Chengxin Ma, Zaid Al-Ars, H. Peter Hofstee |
ISPDC | 4 |
| 2022 | The Future of FPGA Acceleration in Datacenters and the CloudabstractIn this article, we survey existing academic and commercial efforts to provide Field-Programmable Gate Array (FPGA) acceleration in datacenters and the cloud. The goal is a critical review of existing systems and a discussion of their evolution from single workstations with PCI-attached FPGAs in the early days of reconfigurable computing to the integration of FPGA farms in large-scale computing infrastructures. From the lessons learned, we discuss the future of FPGAs in datacenters and the cloud and assess the challenges likely to be encountered along the way. The article explores current architectures and discusses scalability and abstractions supported by operating systems, middleware, and virtualization. Hardware and software security becomes critical when infrastructure is shared among tenants with disparate backgrounds. We review the vulnerabilities of current systems and possible attack scenarios and discuss mitigation strategies, some of which impact FPGA architecture and technology. The viability of these architectures for popular applications is reviewed, with a particular focus on deep learning and scientific computing. This work draws from workshop discussions, panel sessions including the participation of experts in the reconfigurable computing field, and private discussions among these experts. These interactions have harmonized the terminology, taxonomy, and the important topics covered in this manuscript. Christophe Bobda, Joel Mandebi, Paul Chow, Mohammad Ewais, Naif Tarafdar, Juan Camilo Vega, Kenneth Eguro, Dirk Koch, Suranga Handagala, Miriam Leeser, Martin C. Herbordt, Hafsah Shahzad, H. Peter Hofstee, Burkhard Ringlein, Jakub Szefer, Ahmed Sanaullah, Russell Tessier |
ACM Trans. Reconfigurable Technol. Syst. | 13 |
| 2021 | An Attention Module for Convolutional Neural Networks
Baozhou Zhu, H. Peter Hofstee, Jinho Lee 0001, Zaid Al-Ars |
ICANN (1) | 2 |
| 2021 | AutoReCon: Neural Architecture Search-based Reconstruction for Data-free CompressionabstractData-free compression raises a new challenge because the original training dataset for a pre-trained model to be compressed is not available due to privacy or transmission issues. Thus, a common approach is to compute a reconstructed training dataset before compression. The current reconstruction methods compute the reconstructed training dataset with a generator by exploiting information from the pre-trained model. However, current reconstruction methods focus on extracting more information from the pre-trained model but do not leverage network engineering. This work is the first to consider network engineering as an approach to design the reconstruction method. Specifically, we propose the AutoReCon method, which is a neural architecture search-based reconstruction method. In the proposed AutoReCon method, the generator architecture is designed automatically given the pre-trained model for reconstruction. Experimental results show that using generators discovered by the AutoRecon method always improve the performance of data-free compression. Baozhou Zhu, H. Peter Hofstee, Johan Peltenburg, Jinho Lee 0001, Zaid Al-Ars |
IJCAI | 2 |
| 2021 | Low Latency and High Throughput Write-Ahead Logging Using CAPI-FlashabstractHigh-velocity data imposes high durability overheads on Big Data technology components such as NoSQL data stores. In Apache Cassandra and MongoDB, widely used NoSQL solutions with high scalability and availability, write-ahead logging is used to provide durability. However, current write-ahead logging techniques are limited by the excessive overhead in the I/O subsystem. To address this performance gap, we have designed a novel CAPI-Flash based high performance durable logging mechanism for Apache Cassandra and MongoDB. We take advantage of the high throughput, low latency path to flash storage provided by the Coherent Accelerator Processor Interface (CAPI) on IBM POWER8 Systems. Our experimental results show that for insert-only workloads, CAPI-Flash logging provides up to 70 and 514 percent improvement in throughput compared to Cassandra and MongoDB’s durable alternatives, respectively. It also provides average of 45 percent increase in throughput with Cassandra and average of 115 percent increase in throughput with MongoDB for update-mostly and update-only workloads. Bedri Sendir, Madhusudhan Govindaraju, Rei Odaira, H. Peter Hofstee |
IEEE Trans. Cloud Comput. | 4 |
| 2020 | NASB: Neural Architecture Search for Binary Convolutional Neural NetworksabstractBinary Convolutional Neural Networks (CNNs) have significantly reduced the number of arithmetic operations and the size of memory storage needed for CNNs, which makes their deployment on mobile and embedded systems more feasible. However, after binarization, the CNN architecture has to be redesigned and refined significantly due to two reasons: 1. the large accumulation error of binarization in the forward propagation, and 2. the severe gradient mismatch problem of binarization in the backward propagation. Even though substantial effort has been invested in designing architectures for single and multiple binary CNNs, it is still difficult to find an optimized architecture for binary CNNs. In this paper, we propose a strategy, named NASB, which adapts Neural Architecture Search (NAS) to find an optimized architecture for the binarization of CNNs. In the NASB strategy, the operations and their connections define a unique searching space and the training and binarization of the network progress in the three-stage training algorithm.1Due to the flexibility of this automated strategy, the obtained architecture is not only suitable for binarization but also has low overhead, achieving a better trade-off between the accuracy and computational complexity compared to hand-optimized binary CNNs. The implementation of the NASB strategy is evaluated on the ImageNet dataset and demonstrated as a better solution compared to existing quantized CNNs. With insignificant overhead increase, NASB outperforms existing single and multiple binary CNNs by up to 4.0% and 1.0% Top-1 accuracy respectively, bringing them closer to the precision of their full precision counterpart. Baozhou Zhu, Zaid Al-Ars, H. Peter Hofstee |
IJCNN | 3 |
| 2020 | ThymesisFlow: A Software-Defined, HW/SW co-Designed Interconnect Stack for Rack-Scale Memory DisaggregationabstractWith cloud providers constantly seeking the best infrastructure trade-off between performance delivered to customers and overall energy/utilization efficiency of their data-centres, hardware disaggregation comes in as a new paradigm for dynamically adapting the data-centre infrastructure to the characteristics of the running workloads. Such an adaptation enables an unprecedented level of efficiency both from the standpoint of energy and the utilization of system resources. In this paper, we present - ThymesisFlow - the first, to our knowledge, full-stack prototype of the holy-grail of disaggregation of compute resources: pooling of remote system memory. Thymesis-Flow implements a HW/SW co-designed memory disaggregation interconnect on top of the POWER9 architecture, by directly interfacing the memory bus via the OpenCAPI port. We use ThymesisFlow to evaluate how disaggregated memory impacts a set of cloud workloads, and we show that for many of them the performance degradation is negligible. For those cases that are severely impacted, we offer insights on the underlying causes and viable cross-stack mitigation paths. Christian Pinto, Dimitris Syrivelis, Michele Gazzetti, Panos K. Koutsovasilis, Andrea Reale, Kostas Katrinis, H. Peter Hofstee |
MICRO | 7 |
| 2020 | In-memory database acceleration on FPGAs: a surveyabstractAbstract While FPGAs have seen prior use in database systems, in recent years interest in using FPGA to accelerate databases has declined in both industry and academia for the following three reasons. First, specifically for in-memory databases, FPGAs integrated with conventional I/O provide insufficient bandwidth, limiting performance. Second, GPUs, which can also provide high throughput, and are easier to program, have emerged as a strong accelerator alternative. Third, programming FPGAs required developers to have full-stack skills, from high-level algorithm design to low-level circuit implementations. The good news is that these challenges are being addressed. New interface technologies connect FPGAs into the system at main-memory bandwidth and the latest FPGAs provide local memory competitive in capacity and bandwidth with GPUs. Ease of programming is improving through support of shared coherent virtual memory between the host and the accelerator, support for higher-level languages, and domain-specific tools to generate FPGA designs automatically. Therefore, this paper surveys using FPGAs to accelerate in-memory database systems targeting designs that can operate at the speed of main memory. Jian Fang 0004, Yvo T. B. Mulder, Jan Hidders, Jinho Lee 0001, H. Peter Hofstee |
VLDB J. | 5 |
| 2019 | Refine and Recycle: A Method to Increase Decompression ParallelismabstractRapid increases in storage bandwidth, combined with a desire for operating on large datasets interactively, drives the need for improvements in high-bandwidth decompression. Existing designs either process only one token per cycle or process multiple tokens per cycle with low area efficiency and/or low clock frequency. We propose two techniques to achieve high single-decoder throughput at improved efficiency by keeping only a single copy of the history data across multiple BRAMs and operating on each BRAM independently. A first stage efficiently refines the tokens into commands that operate on a single BRAM and steers the commands to the appropriate one. In the second stage, a relaxed execution model is used where each BRAM command executes immediately and those with invalid data are recycled to avoid stalls caused by the read-after-write dependency. We apply these techniques to Snappy decompression and implement a Snappy decompression accelerator on a CAPI2-attached FPGA platform equipped with a Xilinx VU3P FPGA. Experimental results show that our proposed method achieves up to 7.2 GB/s output throughput per decompressor, with each decompressor using 14.2% of the logic and 7% of the BRAM resources of the device. Therefore, a single decompressor can easily keep pace with an NVMe device (PCIe Gen3 x4) on a small FPGA, while a larger device, integrated on a host bridge adapter and instantiating multiple decompressors, can keep pace with the full OpenCAPI 3.0 bandwidth of 25 GB/s. Jian Fang 0004, Jinho Lee 0001, Zaid Al-Ars, H. Peter Hofstee |
ASAP | 5 |
| 2019 | A Fine-Grained Parallel Snappy Decompressor for FPGAs Using a Relaxed Execution ModelabstractSnappy is a widely used (de) compression algorithm in many big data applications. Such a data compression technique has been proven to be successful to save storage space and to reduce the amount of data transmission from/to storage devices. In this paper, we present a fine-grained parallel Snappy decompressor on FPGAs running under a relaxed execution model that addresses the following main challenges in existing solutions. First, existing designs either can only process one token per cycle or can process multiple tokens per cycle with low area efficiency and/or low clock frequency. Second, the high read-after-write data dependency during decompression introduces stalls which pull down the throughput. Jian Fang 0004, Jinho Lee 0001, Zaid Al-Ars, H. Peter Hofstee |
FCCM | 5 |
| 2019 | Fletcher: A Framework to Efficiently Integrate FPGA Accelerators with Apache ArrowabstractModern big data systems are highly heterogeneous. The components found in their many layers of abstraction are often implemented in a wide variety of programming languages and frameworks. Due to language implementation differences, interfaces between these components, including hardware accelerated components, are often burdened by serialization overhead. Serialization bandwidth of many high-level language frameworks is an order of magnitude lower than contemporary FPGA accelerator interface bandwidth, especially when objects are small but numerous. Therefore, serialization bounds the effective end-to-end performance of FPGA-accelerated solutions integrated with applications written in high-level languages. The Apache Arrow project defines a language agnostic columnar in-memory format optimized for big data applications, preventing the need to serialize or even make copies during communication between components. To enable FPGA accelerators to benefit from the approach of Arrow, we first investigate the properties of its format in relation to hardware interfaces and establish that the format is usable. Second, we present the Fletcher framework, that automatically generates highly efficient hardware interfaces to access data of potentially complex, nested Arrow data types. Our approach allows 11 of the languages supported by Apache Arrow libraries to efficiently communicate large data sets with FPGA accelerators at system bandwidth. Furthermore, on the hardware side, the generated interfaces deliver any data type that Arrow can represent as groups of streams, providing a better starting point for data-flow-oriented kernel development, compared to manually creating custom interfaces to address issues related to pointer arithmetic, bus word misalignment and latency. For example applications, as measured on an AWS EC2 F1 and CAPI2-enabled POWER9 system, accelerated end-to-end application performance improves by 1.3x - 49x compared to a hardware accelerated solution that still requires serialization. Johan Peltenburg, Jeroen van Straten, Lars Wijtemans, Lars van Leeuwen, Zaid Al-Ars, H. Peter Hofstee |
FPL | 6 |
| 2018 | CAPI-Flash Accelerated Persistent Read Cache for Apache CassandraabstractIn real-world NoSQL deployments, users have to trade off CPU, memory, I/O bandwidth and storage space to achieve the required performance and efficiency goals. Data compression is a vital component to improve storage space efficiency, but reading compressed data increases response time. Therefore, compressed data stores rely heavily on using the memory as a cache to speed up read operations. However, as large DRAM capacity is expensive, NoSQL databases have become costly to deploy and hard to scale. In our work, we present a persistent caching mechanism for Apache Cassandra on a high-throughput, low-latency FPGA-based NVMe Flash accelerator (CAPI-Flash), replacing Cassandra's in-memory cache. Because flash is dramatically less expensive per byte than DRAM, our caching mechanism provides Apache Cassandra with access to a large caching layer at lower cost. The experimental results show that for read-intensive workloads, our caching layer provides up to 85% improved throughput and also reduces CPU usage by 25% compared to default Cassandra. Bedri Sendir, Madhusudhan Govindaraju, Rei Odaira, H. Peter Hofstee |
IEEE CLOUD | 4 |
| 2018 | A hardware compilation framework for text analytics queries
Raphael Polig, Kubilay Atasu, Heiner Giefers, Christoph Hagleitner, Laura Chiticariu, Frederick Reiss 0001, Huaiyu Zhu 0001, H. Peter Hofstee |
J. Parallel Distributed Comput. | 8 |
| 2017 | ExtraV: Boosting Graph Processing Near Storage with a Coherent AcceleratorabstractIn this paper, we propose ExtraV, a framework for near-storage graph processing. It is based on the novel concept of graph virtualization , which efficiently utilizes a cache-coherent hardware accelerator at the storage side to achieve performance and flexibility at the same time. ExtraV consists of four main components: 1) host processor, 2) main memory, 3) AFU (Accelerator Function Unit) and 4) storage. The AFU, a hardware accelerator, sits between the host processor and storage. Using a coherent interface that allows main memory accesses, it performs graph traversal functions that are common to various algorithms while the program running on the host processor (called the host program) manages the overall execution along with more application-specific tasks. Graph virtualization is a high-level programming model of graph processing that allows designers to focus on algorithm-specific functions. Realized by the accelerator, graph virtualization gives the host programs an illusion that the graph data reside on the main memory in a layout that fits with the memory access behavior of host programs even though the graph data are actually stored in a multi-level, compressed form in storage. We prototyped ExtraV on a Power8 machine with a CAPI-enabled FPGA. Our experiments on a real system prototype offer significant speedup compared to state-of-the-art software only implementations. Jinho Lee 0001, Heesu Kim, Sungjoo Yoo, Kiyoung Choi, H. Peter Hofstee, Gi-Joon Nam, Mark Nutter, Damir A. Jamsek |
Proc. VLDB Endow. | 5 |
| 2016 | Optimized Durable Commitlog for Apache Cassandra Using CAPI-FlashabstractHigh-velocity data imposes high durability overheads on Big Data technology components such as NoSQL data stores. In Apache Cassandra, a widely used NoSQL solution with high scalability and availability, write-ahead logging is used to support Commitlog operations, which in turn provides fault tolerance to applications. However, current write-ahead logging techniques are limited by the excessive overhead in the I/O subsystem. To address this performance gap, we have designed a novel CAPI-Flash based high performance durable Commitlog for Apache Cassandra. We take advantage of the high throughput, low latency path to flash storage provided by the Coherent Accelerator Processor Interface (CAPI) on IBM POWER8 Systems. Our experimental results show that for write-intensive workloads CAPI-Flash logging provides up to 107% improvement in throughput compared to Cassandra's durable alternative. We also provide 77% better throughput in update-mostly workloads. Bedri Sendir, Madhusudhan Govindaraju, Rei Odaira, H. Peter Hofstee |
CLOUD | 4 |
| 2016 | Auto-tuning Spark Big Data Workloads on POWER8: Prediction-Based Dynamic SMT ThreadingabstractMuch research work devotes to tuning big data analytics in modern data centers, since %the truth that even a small percentage of performance improvement immediately translates to huge cost savings because of the large scale. Simultaneous multithreading (SMT) receives great interest from data center communities, as it has the potential to boost performance of big data analytics by increasing the processor resources utilization. For example, the emerging processor architectures like POWER8 support up to 8-way multithreading. However, as different big data workloads have disparate architectural characteristics, how to identify the most efficient SMT configuration to achieve the best performance is challenging in terms of both complex application behaviors and processor architectures. In this paper, we specifically focus on auto-tuning SMT configuration for Spark-based big data workloads on POWE-R8. However, our methodology could be generalized and extended to other programming software stacks and other architectures. Zhen Jia 0001, Guancheng Chen, Jianfeng Zhan, Lixin Zhang 0002, Yonghua Lin, H. Peter Hofstee |
PACT | 7 |
| 2014 | Hardware-accelerated text analyticsabstractPresents a collection of slides covering the following topics: SystemT text analytics software; hardware-accelerated SystemT; Big text Data; and field programmable gate array. Raphael Polig, Kubilay Atasu, Christoph Hagleitner, Laura Chiticariu, Frederick Reiss 0001, Huaiyu Zhu 0001, H. Peter Hofstee |
Hot Chips Symposium | 7 |
| 2007 | The future of multi-core technologies
Michael T. Clark, H. Peter Hofstee, Edward J. Barragy, Ian Buck, Stephen W. Keckler |
CLUSTER | 2 |
| 2006 | Key features of the design methodology enabling a multi-core SoC implementation of a first-generation CELL processorabstractThis paper reviews the design challenges that current and future processors must face, with stringent power limits and high frequency targets, and the design methods required to overcome the above challenges and address the continuing Giga-scale system integration trend. This paper then describes the details behind the design methodology that was used to successfully implement a first-generation CELL processor - a multi-core SoC. Key features of this methodology are broad optimization with fast rule-based analysis engines using macro-level abstraction for constraints propagation up/down the design hierarchy, coupled with accurate transistor level simulation for detailed analysis. The methodology fostered the modular design concept that is inherent to the CELL architecture, enabling a high frequency design by maximizing custom circuit content through re-use, and balanced power, frequency, and die size targets through global convergence capabilities. The design has roughly 241 million transistors implemented in 90 nm SOI technology with 8 levels of copper interconnects and one local interconnect layer. The chip has been tested at various temperatures, voltages, and frequencies. Correct operation has been observed in the lab on first pass silicon at frequencies well over 4GHz. Dac C. Pham, Hans-Werner Anderson, Erwin Behnen, Mark Bolliger, H. Peter Hofstee, Paul E. Harvey, Charles R. Johns, James A. Kahle, Atsushi Kameyama, John M. Keaty, Bob Le, Sang Lee, Tuyen V. Nguyen, John G. Petrovick, Mydung Pham, Juergen Pille, Stephen D. Posluszny, Mack W. Riley, Joseph Verock, James D. Warnock, Steve Weitzel, Dieter F. Wendel |
ASP-DAC | 6 |
| 2006 | Invited speakers II - Real-time supercomputing and technology for games and entertainmentabstractThis talk will explore the influence of supercomputing technology on games and the influence of game technology on supercomputing. H. Peter Hofstee |
SC | 1 |
| 2005 | Power Efficient Processor Architecture and The Cell ProcessorabstractThis paper provides a background and rationale for some of the architecture and design decisions in the cell processor, a processor optimized for compute-intensive and broadband rich media applications, jointly developed by Sony Group, Toshiba, and IBM. The paper discusses some of the challenges microprocessor designers face and provides motivation for performance per transistor as a reasonable first-order metric for design efficiency. Common microarchitectural enhancements relative to this metric are provided. Also alternate architectural choices and some of its limitations are discussed and non-homogeneous SMP as a means to overcome these limitations is proposed. H. Peter Hofstee |
HPCA | 1 |
| 2002 | Power-Constrained Microprocessor DesignabstractPower dissipation and power density have become first-order design constraints, even for high-performance systems. For future designs it will be the dominant constraint. In this paper we suggest a systematic approach to optimizing a processor design under (only) a power constraint. The approach uses the energy-performance ratio (EPR) of the various design parameters as the key to identifying opportunities for improving energy-efficiency. H. Peter Hofstee |
ICCD | 1 |
| 2001 | Derivation of a rotator circuit with homogeneous interconnect
H. Peter Hofstee, Jun Sawada |
Inf. Process. Lett. | 1 |
| 2001 | Timed circuit verification using TEL structuresabstractRecent design examples have shown that significant performance gains are realized when circuit designers are allowed to make aggressive timing assumptions. Circuit correctness in these aggressive styles is highly timing dependent and, in industry, they are typically designed by hand. In order to automate the process of designing and verifying timed circuits, algorithms for their synthesis and verification are necessary. This paper presents timed event/level (TEL) structures, a specification formalism for timed circuits that corresponds directly to gate-level circuits. It also presents an algorithm based on partially ordered sets to make the state-space exploration of TEL structures more tractable. The combination of the new specification method and algorithm significantly improves efficiency for gate-level timing verification. Results on a number of circuits, including many from the recently published gigahertz unit Test Site (guTS) processor from IBM indicate that modules of significant size can be verified using a level of abstraction that preserves the interesting timing properties of the circuit. Accurate circuit level verification allows the designer to include less margin in the design, which can lead to increased performance. Wendy Belluomini, Chris J. Myers, H. Peter Hofstee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | "Timing closure by design, " a high frequency microprocessor design methodologyabstractThis paper presents a design methodology emphasizing early and quick timing closure for high frequency microprocessor designs. This methodology was used to design a Gigahertz class PowerPC microprocessor with 19 million transistors. Characteristics of “Timing Closure by Design are 1) logic partitioned on timing boundaries, 2) predictable control structures (PLAs), 3) static interfaces for dynamic circuits, 4) low skew clock distribution, 5) deterministic method of macro placement, 6) simplified timing analysis, and 7) refinement method of chip integration with early timing analysis. Stephen D. Posluszny, Naoaki Aoki, David W. Boerstler, Paula K. Coulman, Sang H. Dhong, Brian K. Flachs, H. Peter Hofstee, Nobuo Kojima, Ohsang Kwon, Kyung Tek Lee, David Meltzer, Kevin J. Nowka, J. Peter, Joel Silberman, Osamu Takahashi, Paul G. Villarrubia |
DAC | 7 |
| 1998 | Design methodology for a 1.0 GHz microprocessorabstractThis paper describes the design methodology used to build an experimental 1.0 GigaHertz PowerPC integer microprocessor at IBM's Austin Research Laboratory. The high frequency requirements dictated the chip composition to be almost entirely custom macros using dynamic circuit techniques. The methodology presented will cover design and verification tools as well as circuit constraints and microarchitecture philosophy. The microarchitecture, circuits and tools were defined by the high frequency requirements of the processor as well as the aggressive design schedule and size of the design team. Stephen D. Posluszny, Naoaki Aoki, David W. Boerstler, Jeffrey L. Burns, Sang H. Dhong, Uttam Ghoshal, H. Peter Hofstee, David P. LaPotin, Kyung Tek Lee, David Meltzer, Hung C. Ngo, Kevin J. Nowka, Joel Silberman, Osamu Takahashi, Ivan Vo |
ICCD | 7 |
| 1998 | A 690 ps read-access latency register file for a GHz integer microprocessorabstractThis paper describes a 690 ps read-access latency, 32 entry by 64 bit, 3 read-port, 2 write-port, register file with internal bypass. The register file has been fabricated as a pan of 1.0 GHz single-issue 64-bit PowerPC integer processor. Fabrication technology was IBM CMOS6X: 0.25-/spl mu/m drawn channel length, six-metal-layer (Al), 1.8 V nom. V/sub DD/. Self-resetting custom dynamic circuits are used exclusively. Read operation is accomplished by sensing the differential voltage of dual rail bit-lines. Read operation is followed by write operation in the same cycle. Whenever a read address is identical to a write address, the write data is forwarded by an output multiplexer. The register file has been tested and cycle by cycle operation in the processor environment verified at frequencies up to 1.0 GHz (1.8 V, 25/spl deg/C). Osamu Takahashi, Joel Silberman, Sang H. Dhong, H. Peter Hofstee, Naoaki Aoki |
ICCD | 4 |
| 1998 | High-Speed Serializing/De-Serializing Design-For-Test Method for Evaluating a 1 GHz MicroprocessorabstractAs microprocessor speeds approach 1 GHz and beyond the difficulties of at-speed testing continue to increase. In particular, automated test equipment which operates at these frequencies is very limited. This paper discusses a design-for-test method which serializes parallel circuit inputs and de-serializes circuit outputs to achieve 1 GHz operation on test equipment operating at frequencies below 100 MHz. This method has been used to successfully characterize the operation of a 1 GHz microprocessor chip. David F. Heidel, Sang H. Dhong, H. Peter Hofstee, Michael Immediato, Kevin J. Nowka, Joel Silberman, Kevin Stawiasz |
VTS | 3 |
| 1994 | Distributing a Class of Sequential Programs
H. Peter Hofstee |
Sci. Comput. Program. | 1 |
| 1992 | Distributing a Class of Sequential Programs
H. Peter Hofstee |
MPC | 1 |
| 1990 | Distributed Sorting
H. Peter Hofstee, Alain J. Martin, Jan L. A. van de Snepscheut |
Sci. Comput. Program. | 1 |