EDBT 2026 Demo / reviewers in the wild / expert
Derek Chiou
dblp:06/5295
· DBLP profile ↗
46ranked-venue papers
10as first author
6since 2021 · last 2026
0009-0008-6762-4527ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 8 first-author · 3 since 2021Computer networks · 5 · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorTheory of computation · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SmartNIC-Enabled Live Migration for Storage-Optimized VMs with PYROCUMULUS
Jiechen Zhao 0002, Ran Shu 0001, Ziyue Yang 0002, Rui Ma 0021, Derek Chiou, Natalie D. Enright Jerger, Peng Cheng 0005, Yongqiang Xiong |
NSDI | 6 |
| 2026 | Rules Offload Engine (ROE): Accelerating Host SDN Policy Evaluation
Anshuman Verma, Tian Tan 0007, Ahmed Abdelsalam, Milan Dasgupta, Jonathan Hunter, Zach Libby, Narayanan Ravichandran, Harish Srinivasan, Matt Reat, Nadeen Gebara, Vishal Gondaliya, Ezz Hamed, Rahul Garlapati, Lok Chand Koppaka, Abdullah Mughrabi, Dev Desai, Alexander Malysh, Shwetha Bhat, Rohan Kandi, Megan Sng, Tushar Garg, Muluken Hailesellasie, Andrew Putnam, Derek Chiou, Osman Ertugay, Alireza Dabagh, Vivek Bhanu, Daniel Firestone |
SIGCOMM | 24 |
| 2025 | Miniature: Fast AI Supercomputer Networks Simulation on FPGAs
Yicheng Qian, Ran Shu 0001, Rui Ma 0021, Yang Wang 0053, Derek Chiou, Nadeen Gebara, Luca Piccolboni, Miriam Leeser, Yongqiang Xiong |
APNet | 5 |
| 2024 | Accelerating Fault Injection for Validating Processor RTL ImplementationsabstractWith the increasing probability of software errors being caused by hardware transient faults, hardware engineers require better methods to evaluate the impact of faults and determine how to harden their products against them. This paper presents novel techniques to accurately and efficiently simulate faults on processors by dynamically switching, with little designer effort, a simulation from RTL-level-or-lower to instruction-level. When combined with known techniques, the proposed techniques make fault modeling orders-of-magnitude faster than RTL simulation while achieving the same level of accuracy, enabling fast identification of the portions of the design that would benefit most from hardening, and providing statistical guarantees of a design's resistance to faults. Derek Chiou |
ICCAD | 2 |
| 2021 | DO-GPU: Domain Optimizable Soft GPUsabstract”Soft” GPUs are overlays that implement GPGPU-like data parallel processor architectures in FPGA logic to make FPGAs as software-programmable as ”hard” GPGPUs. Unlike hard GPUs, soft GPU architectures can be specialized to further improve efficiency by leveraging FPGA’s flexibility. Prior work has shown the software programmability potential for soft GPUs but only studied general-purpose soft GPUs with minor specializations (e.g., FPGU, FlexGrip, MIAOW, and SCRATCH) or only domain-optimized for a particular application domain (e.g., PDL-FGPU for the persistent deep learning domain.) This paper proposes a soft GPU development framework to automate the creation of soft GPU instances with aggressive application-domain optimizations (i.e., domain-optimized GPUs, or DOGPUs) that consists of a baseline general soft GPU architecture ”template” with an improved architecture over prior general purpose soft GPUs, along with a customizable partition that enables a custom datapath (macro unit) to be inserted to optimize for a target application domain. Unlike the prior PDL-FGPU which targets the persistent deep learning domain, the proposed framework can be used to target optimization for any application domain. Our evaluation on a set of data parallel workloads shows that (i) the proposed general soft GPU architecture offers average speedup of 1.8x versus the best prior soft GPUs we know of (i.e., FGPU, PDL-FGPU), (ii) DO-GPUs with domain-optimizations provide an average of 218x speedup over general soft GPUs, (iii) the proposed framework enabled building six new domain-optimized soft GPU instances in a matter of days, and (iv) enables quick GPU-like development effort (hours), where code is concise (low 100s of lines) and can be compiled in seconds without FPGA EDA tools in the loop, assuming an appropriate soft DO-GPU bitstream for the application domain is already built. Rui Ma 0021, Jia-Ching Hsu, Tian Tan 0007, Eriko Nurvitadhi, Rajesh Vivekanandham, Aravind Dasu, Martin Langhammer, Derek Chiou |
FPL | 8 |
| 2021 | Specializing FGPU for Persistent Deep LearningabstractOverlay architectures are a good way to enable fast development and debug on FPGAs at the expense of potentially limited performance compared to fully customized FPGA designs. When used in concert with hand-tuned FPGA solutions, performant overlay architectures can improve time-to-solution and thus overall productivity of FPGA solutions. This work tunes and specializes FGPU, an open source OpenCL-programmable GPU overlay for FPGAs. We demonstrate that our persistent deep learning (PDL )-FGPU architecture maintains the ease-of-programming and generality of GPU programming while achieving high performance from specialization for the persistent deep learning domain. We also propose an easy method to specialize for other domains. PDL-FGPU includes new instructions, along with micro-architecture and compiler enhancements. We evaluate both the FGPU baseline and the proposed PDL-FGPU on a modern high-end Intel Stratix 10 2800 FPGA in simulation running persistent DL applications (RNN, GRU, LSTM), and non-DL applications to demonstrate generality. PDL-FGPU requires 1.4–3× more ALMs, 4.4–6.4× more M20ks, and 1–9.5× more DSPs than baseline, but improves performance by 56–693× for PDL applications with an average 23.1% degradation on non-PDL applications. We integrated the PDL-FGPU overlay into Intel OPAE to measure real-world performance/power and demonstrate that PDL-FGPU is only 4.0–10.4× slower than the Nvidia V100. Rui Ma 0021, Jia-Ching Hsu, Tian Tan 0007, Eriko Nurvitadhi, David Sheffield, Rob Pelt, Martin Langhammer, Jaewoong Sim, Aravind Dasu, Derek Chiou |
ACM Trans. Reconfigurable Technol. Syst. | 10 |
| 2019 | Specializing FGPU for Persistent Deep LearningabstractOverlay architectures are a good way to enable fast development and debug on FPGAs at the expense of potentially limited performance when compared to fully customized FPGA designs. When used in concert with a hand-tuned FPGA solution, a performant overlay architecture can improve the time-to-solution and thus overall productivity of FPGA solutions. In this work, we tune and specialize FGPU, an open source OpenCL-programmable GPU overlay for FPGAs. We demonstrate that our PDL-FGPU architecture is able to maintain the ease-of-programming and generality of a software programmable soft GPU while achieving high performance due to specialization in the persistent deep learning domain. We also propose a easy method to specialize for different domains. PDL-FGPU includes new instructions, along with micro-architecture and compiler enhancements. We evaluate both the FGPU baseline and the proposed PDL-FGPU on a modern high-end Intel Stratix 10 2800 FPGA running a set of persistent DL applications (RNN, GRU, LSTM), as well as general non-DL applications to demonstrate generality. PDL-FGPU requires 1.5-3x more ALMs, 4.4-6.4x more M20ks, and 4.6-10x more DSPs than the FGPU baseline, but improves performance by 55-727x for persistent DL applications with an average 15% degradation on general non-PDL applications. We also demonstrate that the PDL-FGPU is only 4-7x slower than the Nvidia Volta V100 GPU. Rui Ma 0021, Derek Chiou, Jia-Ching Hsu, Tian Tan 0007, Eriko Nurvitadhi, David Sheffield, Rob Pelt, Martin Langhammer, Jaewoong Sim, Aravind Dasu |
FPL | 2 |
| 2019 | Direct Universal Access: Making Data Center Resources Available to FPGA
Ran Shu 0001, Peng Cheng 0005, Guo Chen 0001, Yongqiang Xiong, Derek Chiou, Thomas Moscibroda |
NSDI | 7 |
| 2019 | FlexSaaS: A Reconfigurable Accelerator for Web Search SelectionabstractWeb search engines deploy large-scale selection services on CPUs to identify a set of web pages that match user queries. An FPGA-based accelerator can exploit various levels of parallelism and provide a lower latency, higher throughput, more energy-efficient solution than commodity CPUs. However, maintaining such a customized accelerator in a commercial search engine is challenging because selection services are changed often. This article presents our design for FlexSaaS (Flexible Selection as a Service), an FPGA-based accelerator for web search selection. To address efficiency and flexibility challenges, FlexSaaS abstracts computing models and separates memory access from computation. Specifically, FlexSaaS (i) contains a reconfigurable number of matching processors that can handle various possible query plans, (ii) decouples index stream reading from matching computation to fetch and decode index files, and (iii) includes a universal memory accessor that hides the complex memory hierarchy and reduces host data access latency. Evaluated on FPGAs in the selection service of a commercial web search--the Bing web search engine—FlexSaaS can be evolved quickly to adapt to new updates. Compared to the software baseline, FlexSaaS on Arria 10 reduces average latency by 30% and increases throughput by 1.5×. Shijie Cao, Lanshun Nie, Dechen Zhan, Ningyi Xu, Ramashis Das, Ming Wu 0007, Derek Chiou |
ACM Trans. Reconfigurable Technol. Syst. | 9 |
| 2018 | Minnow: Lightweight Offload Engines for Worklist Management and Worklist-Directed PrefetchingabstractThe importance of irregular applications such as graph analytics is rapidly growing with the rise of Big Data. However, parallel graph workloads tend to perform poorly on general-purpose chip multiprocessors (CMPs) due to poor cache locality, low compute intensity, frequent synchronization, uneven task sizes, and dynamic task generation. At high thread counts, execution time is dominated by worklist synchronization overhead and cache misses. Researchers have proposed hardware worklist accelerators to address scheduling costs, but these proposals often harden a specific scheduling policy and do not address high cache miss rates. We address this with Minnow, a technique that augments each core in a CMP with a lightweight Minnow accelerator. Minnow engines offload worklist scheduling from worker threads to improve scalability. The engines also perform worklist-directed prefetching, a technique that exploits knowledge of upcoming tasks to issue nearly perfectly accurate and timely prefetch operations. On a simulated 64-core CMP running a parallel graph benchmark suite, Minnow improves scalability and reduces L2 cache misses from 29 to 1.2 MPKI on average, resulting in 6.01x average speedup over an optimized software baseline for only 1% area overhead. Dan Zhang 0004, Michael Thomson, Derek Chiou |
ASPLOS | 4 |
| 2018 | Evaluating The Highly-Pipelined Intel Stratix 10 FPGA Architecture Using Open-Source BenchmarksabstractIntel® Stratix® 10 FPGAs offer a novel architectural feature called HyperFlex that enables an extreme degree of pipelining resulting in up to 1GHz clock frequencies. Prior work has evaluated HyperFlex on pre-production Stratix 10 FPGAs using internal designs not accessible to the general public. This paper presents an updated evaluation of HyperFlex on the latest publicly-available production Stratix 10 FPGA using open-source benchmarks. In particular, our evaluation started with seven RTL designs from existing open-source projects, carefully chosen to capture a variety of architectures (simple pipeline to pipeline with loop/M20Ks/DSPs) implementing well-known functions such as crypto, math, and image processing. An FPGA developer who was not an expert in HyperFlex then spent around 250 engineering hours to develop 24 optimized versions of these designs, following the Intel Stratix 10 FPGA HyperFlex optimization guide. Those optimized designs run at 400MHz to 850MHz. In this paper, we describe the optimizations, efforts, and results. Upon publication, those optimized designs will be open-sourced and published in GitHub. Tian Tan 0007, Eriko Nurvitadhi, David Shih, Derek Chiou |
FPT | 4 |
| 2018 | Azure Accelerated Networking: SmartNICs in the Public Cloud
Daniel Firestone, Andrew Putnam, Sambrama Mundkur, Derek Chiou, Alireza Dabagh, Mike Andrewartha, Hari Angepat, Vivek Bhanu, Adrian M. Caulfield, Eric S. Chung, Harish Kumar Chandrappa, Somesh Chaturmohta, Matt Humphrey, Jack Lavier, Norman Lam, Fengfen Liu, Kalin Ovtcharov, Jitendra Padhye, Gautham Popuri, Shachar Raindel, Tejas Sapre, Mark Shaw 0001, Gabriel Silva, Madhan Sivakumar, Nisheeth Srivastava, Anshuman Verma, Qasim Zuhair, Deepak Bansal, Doug Burger, Kushagra Vaid, David A. Maltz, Albert G. Greenberg |
NSDI | 4 |
| 2017 | FPGA-Accelerated Transactional Execution of Graph Workloads
Dan Zhang 0004, Derek Chiou |
FPGA | 3 |
| 2016 | Agile Co-Design for a Reconfigurable DatacenterabstractIn 2015, a team of software and hardware developers at Microsoft shipped the world?s first commercial search engine accelerated using FPGAs in the datacenter. During the sprint to production, new algorithms in the Bing ranking service were ported into FPGAs and deployed to a production bed within several weeks of conception, leading to significant gains in latency and throughput. The fast turnaround time of new features demanded by an agile software culture would not have been possible without a disciplined and effective approach to co-design in the datacenter. This talk will describe some of the learnings and best practices developed from this unique experience. Shlomi Alkalay, Hari Angepat, Adrian M. Caulfield, Eric S. Chung, Oren Firestein, Michael Haselman, Stephen Heil, Kyle Holohan, Matt Humphrey, Tamás Juhász, Puneet Kaur, Sitaram Lanka, Daniel Lo, Todd Massengill, Kalin Ovtcharov, Michael Papamichael, Andrew Putnam, Raja Seera, Rimon Tadros, Jason Thong, Lisa Woods, Derek Chiou, Doug Burger |
FPGA | 22 |
| 2016 | Intel Acquires Altera: How Will the World of FPGAs be Affected?abstractIntel's purchase of Altera is very likely to be the biggest single event in FPGA history and, therefore, have a profound impact on the FPGA world. This panel intends to explore the business and research opportunities that are potentially enabled and potentially squashed by the acquisition. Questions that will be explored by the panel include: Derek Chiou |
FPGA | 1 |
| 2016 | HGum: Messaging Framework for Hardware Accelerators (Abstact Only)abstractSoftware messaging frameworks help avoid errors and reduce engineering effort in building distributed systems by (i) providing an interface definition language (IDL) to precisely specify the structure of the message (the message schema) and (ii) automatically generating the serialization and deserialization functions that transform user data structures into binary data for sending across the network and vice versa. Similarly, a hardware-accelerated system that consists of host software and multiple FPGAs, could also benefit from a messaging framework to handle messages both between software and FPGA and also between different FPGAs. The key challenge for a hardware messaging framework is that it must be able to support large messages with complex schema while meeting critical constraints such as clock frequency, area, and throughput. We present HGum, a messaging framework for hardware accelerators that meets all the above requirements. HGum is able to generate high-performance and low-cost hardware logic by employing a novel design that algorithmically parses the message schema to perform serialization and deserialization. Our evaluation of HGum shows that it not only significantly reduces engineering effort but also generates hardware with comparable quality to manual implementation. Sizhuo Zhang, Hari Angepat, Derek Chiou |
FPGA | 3 |
| 2016 | Heterogeneous Computing and Infrastructure for Energy Efficiency in Microsoft Data Centers: Extended AbstractabstractThis paper describes some of the ways heterogeneous computing, specifically FPGAs, is used in Microsoft data centers to improve energy efficiency and performance. Improved energy efficiency can have a significant effect on the viability of the data center in question. We discuss the Bing ranking application along with Azure networking infrastructure production projects. Derek Chiou |
ISLPED | 1 |
| 2016 | A cloud-scale acceleration architectureabstractHyperscale datacenter providers have struggled to balance the growing need for specialized hardware (efficiency) with the economic benefits of homogeneity (manageability). In this paper we propose a new cloud architecture that uses reconfigurable logic to accelerate both network plane functions and applications. This Configurable Cloud architecture places a layer of reconfigurable logic (FPGAs) between the network switches and the servers, enabling network flows to be programmably transformed at line rate, enabling acceleration of local applications running on the server, and enabling the FPGAs to communicate directly, at datacenter scale, to harvest remote FPGAs unused by their local servers. We deployed this design over a production server bed, and show how it can be used for both service acceleration (Web search ranking) and network acceleration (encryption of data in transit at high-speeds). This architecture is much more scalable than prior work which used secondary rack-scale networks for inter-FPGA communication. By coupling to the network plane, direct FPGA-to-FPGA messages can be achieved at comparable latency to previous work, without the secondary network. Additionally, the scale of direct inter-FPGA messaging is much larger. The average round-trip latencies observed in our measurements among 24, 1000, and 250,000 machines are under 3, 9, and 20 microseconds, respectively. The Configurable Cloud architecture has been deployed at hyperscale in Microsoft's production datacenters worldwide. Adrian M. Caulfield, Eric S. Chung, Andrew Putnam, Hari Angepat, Jeremy Fowers, Michael Haselman, Stephen Heil, Matt Humphrey, Puneet Kaur, Joo-Young Kim 0001, Daniel Lo, Todd Massengill, Kalin Ovtcharov, Michael Papamichael, Lisa Woods, Sitaram Lanka, Derek Chiou, Doug Burger |
MICRO | 17 |
| 2016 | Introduction to Special Issue on Reconfigurable Components with Source CodeabstractNo abstract available. André DeHon, Derek Chiou |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2015 | Accelerating data centers with reconfigurable logicabstractData centers are a highly competitive environment that demands high performance and energy efficiency and, in many cases, low latency. Custom hardware can provide significant improvements over conventional microprocessors on those metrics. Microsoft has been investigating the use of reconfigurable logic, in the form of field programmable gate arrays, to accelerate its data centers. In this talk, I will describe some of our efforts in this area. Derek Chiou |
ASAP | 1 |
| 2015 | Growing a Healthy FPGA EcosystemabstractThe personal computer market grew exponentially in the 1980's for vendors such as Apple, Microsoft, and Intel when there was a healthy mix of software, tools, and microprocessor devices. John Lockwood, Michael Adler, Dan Mansur, Derek Chiou, Mike Strickland, Jason Cong, Steve Teig |
FPGA | 4 |
| 2015 | GPGPU performance and power estimation using machine learningabstractGraphics Processing Units (GPUs) have numerous configuration and design options, including core frequency, number of parallel compute units (CUs), and available memory bandwidth. At many stages of the design process, it is important to estimate how application performance and power are impacted by these options. This paper describes a GPU performance and power estimation model that uses machine learning techniques on measurements from real GPU hardware. The model is trained on a collection of applications that are run at numerous different hardware configurations. From the measured performance and power data, the model learns how applications scale as the GPU's configuration is changed. Hardware performance counter values are then gathered when running a new application on a single GPU configuration. These dynamic counter values are fed into a neural network that predicts which scaling curve from the training data best represents this kernel. This scaling curve is then used to estimate the performance and power of the new application at different GPU configurations. Over an 8× range of the number of CUs, a 3.3× range of core frequencies, and a 2.9× range of memory bandwidth, our model's performance and power estimates are accurate to within 15% and 10% of real hardware, respectively. This is comparable to the accuracy of cycle-level simulators. However, after an initial training phase, our model runs as fast as, or faster than the program running natively on real hardware. Gene Y. Wu, Joseph L. Greathouse, Alexander Lyashevsky, Nuwan Jayasena, Derek Chiou |
HPCA | 5 |
| 2015 | Keynote talk II: Accelerating data centers using reconfigurable logicabstractReconfigurable logic has the potential to provide hardware level performance with the flexibility of software. Such properties make it an interesting solution in data center environments that value high throughput, low latency, low power, and uniformity of hardware. Microsoft has been exploring the use of reconfigurable logic in its data centers. In this talk, I will describe some of our efforts in this area. Derek Chiou |
MEMOCODE | 1 |
| 2014 | Cryptoraptor: high throughput reconfigurable cryptographic processorabstractThis paper describes a high performance, low power, and highly flexible cryptographic processor, Cryptoraptor, which is designed to support both today's and tomorrow's symmetric-key cryptography algorithms and standards. To the best of our knowledge, the proposed cryptographic processor supports the widest range of cryptographic algorithms compared to other solutions in the literature and is the only crypto-specific processor targeting future standards as well. Our 1GHz design achieves a peak throughput of 128Gbps for AES-128 which is competitive with ASIC designs and has 25X and 160X higher throughputs per area than CPU and GPU solutions, respectively. Gokhan Sayilar, Derek Chiou |
ICCAD | 2 |
| 2014 | A reconfigurable fabric for accelerating large-scale datacenter servicesabstractDatacenter workloads demand high computational capabilities, flexibility, power efficiency, and low cost. It is challenging to improve all of these factors simultaneously. To advance datacenter capabilities beyond what commodity server designs can provide, we have designed and built a composable, reconfigurable fabric to accelerate portions of large-scale software services. Each instantiation of the fabric consists of a 6×8 2-D torus of high-end Stratix V FPGAs embedded into a half-rack of 48 machines. One FPGA is placed into each server, accessible through PCIe, and wired directly to other FPGAs with pairs of 10 Gb SAS cables. In this paper, we describe a medium-scale deployment of this fabric on a bed of 1,632 servers, and measure its efficacy in accelerating the Bing web search engine. We describe the requirements and architecture of the system, detail the critical engineering challenges and solutions needed to make the system robust in the presence of failures, and measure the performance, power, and resilience of the system when ranking candidate documents. Under high load, the largescale reconfigurable fabric improves the ranking throughput of each server by a factor of 95% for a fixed latency distribution—or, while maintaining equivalent throughput, reduces the tail latency by 29%. Andrew Putnam, Adrian M. Caulfield, Eric S. Chung, Derek Chiou, Kypros Constantinides, John Demme, Hadi Esmaeilzadeh, Jeremy Fowers, Gopi Prashanth Gopal, Jan Gray, Michael Haselman, Scott Hauck, Stephen Heil, Amir Hormati, Joo-Young Kim 0001, Sitaram Lanka, James R. Larus, Eric Peterson, Simon Pope, Aaron Smith, Jason Thong, Phillip Yi Xiao, Doug Burger |
ISCA | 4 |
| 2013 | Implementing microprocessors from simplified descriptionsabstractDespite the proliferation of high-level synthesis tools, hardware description of microprocessors remains complex. We argue that much of the incidental complexity can be relieved by untangling the description into separate functional and microarchitectural components. Such an untangling can be achieved using a high-level microcode compiler that can generate not only microcode, but also the micro-instruction format and the interpretations of each control bit. Simplifying hardware description will help the designer make better design-space trade-offs, and close the design and verification loop faster. This paper takes the reader through an implementation of a simple Y86 processor to qualitatively illustrate the complexity reduction from the untangling. Nikhil A. Patil, Derek Chiou |
ASP-DAC | 2 |
| 2013 | An FPGA-based in-line accelerator for Memcached
Maysam Lavasani, Hari Angepat, Derek Chiou |
Hot Chips Symposium | 3 |
| 2012 | On the asymptotic costs of multiplexer-based reconfigurabilityabstractExisting literature documents a number of techniques for combining a set of independent datapath designs into a single datapath that is run-time configurable to the functionality of any datapath in the set. This paper explores how delay, energy and area overhead attributable to reconfigurability scales with the number of configurable functionalities, independent of the design of specific datapaths. Distinct design space regions are identified based upon common scaling properties, with implications on the design and feasible efficiency bounds of reconfigurable devices. Johnathan York, Derek Chiou |
DAC | 2 |
| 2012 | Compiling high throughput network processorsabstractGorilla is a methodology for generating FPGA-based solutions especially well suited for data parallel applications with fine grain irregularity. Irregularity simultaneously destroys performance and increases power consumption on many data parallel processors such as General Purpose Graphical Processor Units (GPGPUs). Gorilla achieves high performance and low power through the use of FPGA-tailored parallelization techniques and application-specific hardwired accelerators, processing engines, and communication mechanisms. Automatic compilation from a stylized C language and templates that define the hardware structure coupled with the intrinsic flexibility of FPGAs provide high performance, low power, and programmability. Maysam Lavasani, Larry Dennison, Derek Chiou |
FPGA | 3 |
| 2011 | Enforcing architectural contracts in high-level synthesisabstractWe present a high-level synthesis technique that takes as input two orthogonal descriptions: (a) a behavioral architectural contract between the implementation and the user, and (b) a microarchitecture on which the architectural contract can be implemented. We describe a prototype compiler that generates control required to enforce the contract, and thus, synthesizes the pair of descriptions to hardware. Nikhil A. Patil, Ankit Bansal, Derek Chiou |
DAC | 3 |
| 2011 | MEMOCODE 2011 Hardware/Software CoDesign Contest: NoC simulatorabstractThe objective of the 2011 MEMOCODE Hardware Software CoDesign Contest was to improve the performance of a configurable Network-on-Chip (NoC) simulator. A C++ reference simulator was provided, whose output was defined to be correct. The simulator takes two inputs: (i) a specification of the NoC to be simulated (the target) and (ii) a traffic pattern that specifies the traffic to be simulated on that NoC. The simulator was designed to support virtually any NoC topology. Results were judged firstly on correctness (that the number of cycles is accurately modeled) and secondly on simulation speed (the amount of time it takes to complete the entire simulation). Derek Chiou |
MEMOCODE | 1 |
| 2010 | NIFD: Non-intrusive FPGA Debugger -- Debugging FPGA 'Threads' for Rapid HW/SW Systems PrototypingabstractDebugging hardware has always been difficult when compared to debugging software, in large part due to a lack of convenient visibility. This paper describes the open NIFD framework that provides software-like debugging facilities to both pure FPGA and hybrid FPGA/software platforms, allowing a designer to treat the hardware logic like a specialized remote software debug target. NIFD provides features such as single stepping, breakpoints, and examination of the full hardware state from a standard debug console such as GDB. The framework leverages built-in readback support to enable non-intrusive, transparent debugging with full observability and controllability. This technique is not only useful for debugging, but can also be used in production environments for infrequent events such as the slow sampling of counters. Hari Angepat, Gage Eads, Christopher Craik, Derek Chiou |
FPL | 4 |
| 2010 | PrEsto: An FPGA-accelerated Power Estimation Methodology for Complex SystemsabstractReduced or bounded power consumption has become a first-order requirement for modern hardware design. As a design progresses and more detailed information becomes available, more accurate power estimations become possible but at the cost of significantly slower simulation speeds. Power simulation that is both sufficiently-accurate and fast would have a positive impact on architecture and design. In this paper, we propose PrEsto, a power modeling methodology that improves the speed and accuracy of power estimation through FPGA-acceleration. PrEsto automatically generates FPGA-based power estimators consisting of linear models that are designed to be integrated into fast, accurate FPGA-based performance simulators of microprocessors. Our prototype implementation predicts the cycle-by-cycle power dissipation of the LEON3 core and the ARM Cortex-A8 core to within 6% of a commercial gate-level power estimation tool, while running several orders of magnitude faster. The combination of simulation speed and accuracy is not only useful to architects and designers, it is fast enough to be useful for power-sensitive operating system and application developers. Dam Sunwoo, Gene Y. Wu, Nikhil A. Patil, Derek Chiou |
FPL | 4 |
| 2009 | Soft connections: addressing the hardware-design modularity problemabstractHardware-design languages typically impose a rigid communication hierarchy that follows module instantiation. This leads to an undesirable side-effect where changes to a child's interface result in changes to the parents. Soft connections address this problem by allowing the user to specify connection endpoints that are automatically connected at compilation time, rather than by the user. Michael Pellauer, Michael Adler, Derek Chiou, Joel S. Emer |
DAC | 3 |
| 2009 | QUICK: A flexible full-system functional modelabstractIn this paper, we introduce the concept of full-system complete-and-rollback functional simulators that make efficient functional models in functional/timing partitioned simulators. Complete-and-rollback functional simulators can efficiently drive simulators of resolutions ranging from functional-only to cycle-accurate for a wide range of simulated machines. Complete-and-rollback functional models achieve their capabilities by executing instructions to completion, enabling their execution to be highly optimized, but providing rollback capabilities to enable on-the-fly modifications to the functional execution. We also introduce QUICK, an implementation of a full-system complete-and-rollback functional model that supports the x86 and PowerPC ISAs, boots unmodified Windows XP and Linux, and runs unmodified applications such as YouTube on Internet Explorer while fully supporting rollbacks, including across I/O operations. We present various case studies using QUICK and conduct performance analyses to demonstrate its simulation performance. Dam Sunwoo, Joonsoo Kim, Derek Chiou |
ISPASS | 3 |
| 2008 | Parallelizing computer system simulatorsabstractThis paper describes NSF-supported work in parallelized computer system simulators being done in the Electrical and Computer Engineering Department at the University of Texas at Austin. Our work is currently following two paths: (i) the FAST simulation methodology[9, 11, 10] that is capable of simulating complex systems accurately and quickly (currently about 1.2MIPS executing the x86 ISA, modeling an out-of-order superscalar processor and booting Windows XP and Linux) and (ii) the RAMP-White (White)[1, 22] platform that will soon be capable of simulating very large systems of around 1000 cores. We plan to combine the projects to provide fast and accurate simulation of multicore systems. Derek Chiou, Dam Sunwoo, Hari Angepat, Joonsoo Kim, Nikhil A. Patil, William H. Reinhart, Darrel Eric Johnson |
IPDPS | 1 |
| 2007 | The FAST methodology for high-speed SoC/computer simulationabstractThis paper describes the FAST methodology that enables a single FPGA to accelerate the performance of cycle-accurate computer system simulators modeling modem, realistic SoCs, embedded systems and standard desktop/laptop/server computer systems. The methodology partitions a simulator into (i) a functional model that simulates the functionality of the computer system and (ii) a predictive model that predicts performance and other metrics. The partitioning is crafted to map most of the parallel work onto a hardware-based predictive model, eliminating much of the complexity and difficulty of simulating parallel constructs on a sequential platform. FAST conventions and libraries have been designed to make creating, modifying, using and measuring such simulators straightforward. We describe a prototype FAST system: a full-system, RTL-level cycle-accurate-capable computer system simulator that executes the x86 ISA, boots unmodified Linux and executes unmodified x86 applications. The prototype runs two to three orders of magnitude faster than the fastest Intel and AMD RTL-level cycle-accurate x86 software-based simulators and about six to seven times faster than RTL simulation. Derek Chiou, Dam Sunwoo, Joonsoo Kim, Nikhil A. Patil, William H. Reinhart, Darrel Eric Johnson, Zheng Xu 0004 |
ICCAD | 1 |
| 2007 | FPGA-Accelerated Simulation Technologies (FAST): Fast, Full-System, Cycle-Accurate SimulatorsabstractThis paper describes FAST, a novel simulation methodology that can produce simulators that (i) are orders of magnitude faster than comparable simulators, (ii) are cycle- accurate, (Hi) model the entire system running unmodified applications and operating systems, (iv) provide visibility with minimal simulation performance impact and (v) are capable of running current instruction sets such as x86. It achieves its capabilities by partitioning simulators into a speculative functional model component that simulates the instruction set architecture and a timing model component that predicts performance. The speculative functional model enables the simulator to be parallelized, implementing the timing model in FPGA hardware for speed and the functional model using a modified full-system simulators. We currently achieve an average simulation speed of 1.2MIPS running x86 applications on x86 Linux and Windows XP and expect to achieve 10MIPS over time. Such simulators are useful to virtually all computer system simulator users ranging from architects, through RTL designers and verifiers to software developers. Sharing a common simulation/design infrastructure couldfoster better communication between these groups, potentially resulting in better system designs. Derek Chiou, Dam Sunwoo, Joonsoo Kim, Nikhil A. Patil, William H. Reinhart, Darrel Eric Johnson, Jebediah Keefe, Hari Angepat |
MICRO | 1 |
| 2006 | Research accelerator for multiple processors
David A. Patterson 0001, Arvind 0001, Krste Asanovic, Derek Chiou, James C. Hoe, Christoforos E. Kozyrakis, Shih-Lien Lu, Mark Oskin, Jan M. Rabaey, John Wawrzynek |
Hot Chips Symposium | 4 |
| 2000 | Application-specific memory management for embedded systems using software-controlled cachesabstractWe propose a way to improve the performance of embedded processors running data-intensive applications by allowing software to allocate on-chip memory on an application-specific basis. On-chip memory in the form of cache can be made to act like scratch-pad memory via a novel hardware mechanism, which we call column caching. Column caching enables dynamic cache partitioning in software, by mapping data regions to a specified sets of cache “columns” or “ways.” When a region of memory is exclusively mapped to an equivalent sized partition of cache, column caching provides the same functionality and predictability as a dedicated scratchpad memory for time-critical parts of a real-time application. The ratio between scratchpad size and cache size can be easily and quickly varied for each application, or each task within an application. Thus, software has much finer software control of on-chip memory, providing the ability to dynamically tradeoff performance for on-chip memory. Derek Chiou, Prabhat Jain, Larry Rudolph, Srini Devadas |
DAC | 1 |
| 2000 | Micro-Architectures of High Performance, Multi-User System Area Network Interface CardsabstractThis paper examines two Network Interface Card micro-architectures that support low latency, high bandwidth user level message passing in multi-user environments. The two are at different ends of a design spectrum-the Resident queues design relies completely on hardware, while the Non-resident queues design is heavily firmware driven. Through actual implementation of these designs and simulation-based micro-benchmark studies, we identify issues critical to the performance and functionality of the firmware-based approach. The firmware-based approach offers much flexibility at a moderate performance penalty, while the Resident design has superior performance for the functions it implements. This leads us to conclude that a hybrid design combining complete hardware support for common operations and a firmware implementation of less common functions achieves both high performance and flexibility. Boon Seong Ang, Derek Chiou, Larry Rudolph, Arvind 0001 |
IPDPS | 2 |
| 1998 | Message passing support on StarT-VoyagerabstractNo single message passing mechanism can efficiently support all types of communication that commonly occur in most parallel or distributed programs. MIT's StarT-Voyager, a hybrid message passing/shared memory parallel machine, provides four message passing mechanisms to achieve high performance over a wide spectrum of communication types and sizes. Hardware and address translation enforced protection allows direct user-level access to message passing facilities in a multiuser environment. StarT-Voyager's protection scheme improves upon past designs by not requiring strictly synchronized gang-scheduling, and by supporting non-monolithic protection domains. To minimize the development effort and cost, the machine is designed to use unmodified commercial PowerPC 604-based SMP systems as the building block. A Network End-point Subsystem (NES) card which plugs into one of each SMP's processor card slots provides the interface to Arctic, a low-latency, high-bandwidth network developed at MIT. This paper describes StarT-Voyager's message passing mechanisms and their predicted performance. Boon Seong Ang, Derek Chiou, Larry Rudolph, Arvind 0001 |
HiPC | 2 |
| 1998 | StarT-Voyager: A Flexible Platform for Exploring Scalable SMP IssuesabstractThis paper describes StarT-Voyager, a machine designed as an experimental platform for research in cluster system communication. The heart of StarT-Voyager is a network interface unit (NIU) that connects the memory bus of a PowerPC-based SMP to the MIT Arctic network. The NIU is highly flexible, with its set of functions easily modified by firmware or by programmable hardware, making it possible to compare different communication interfaces and implementation strategies on a common platform. Its flexibility comes from a fast embedded processor and large, fast FPGAs that surround a high-speed protected communication core. Its efficiency comes from a set of primitive operations that are implemented in hardware and are designed to reduce the firmware overhead. Our initial configuration of StarT-Voyager implements four forms of message passing along with S-COMA and NUMA shared memory support. With experimentation on the machine, it can be reconfigured to introduce new mechanisms improving usability and performance. Boon Seong Ang, Derek Chiou, Daniel L. Rosenband, Mike Ehrlich, Larry Rudolph, Arvind 0001 |
SC | 2 |
| 1995 | START-NG: Delivering Seamless Parallel Computing
Derek Chiou, Boon Seong Ang, Robert Greiner, Arvind 0001, James C. Hoe, Michael J. Beckerle, James E. Hicks, G. Andrew Boughton |
Euro-Par | 1 |
| 1993 | Performance Studies of Id on the Monsoon Dataflow System
James E. Hicks, Derek Chiou, Boon Seong Ang, Arvind 0001 |
J. Parallel Distributed Comput. | 2 |
| 1993 | Performance Visualization on Monsoon
Venkat Natarajan, Derek Chiou, Boon Seong Ang |
J. Parallel Distributed Comput. | 2 |