EDBT 2026 Demo / reviewers in the wild / expert
Chen Yang 0010
dblp:01/2478-10
· DBLP profile ↗
16ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0001-7262-2878ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | FPGA-Accelerated Range-Limited Molecular DynamicsabstractLong timescale Molecular Dynamics (MD) simulation of small molecules is crucial in drug design and basic science. To accelerate a small data set that is executed for a large number of iterations, high-efficiency is required. Recent work in this domain has demonstrated that among COTS devices only FPGA-centric clusters can scale beyond a few processors. The problem addressed here is that, as the number of on-chip processors has increased from fewer than 10 into the hundreds, previous intra-chip routing solutions are no longer viable. We find, however, that through various design innovations, high efficiency can be maintained. These include replacing the previous broadcast networks with ring-routing and then augmenting the rings with out-of-order and caching mechanisms. Others are adding a level of hierarchical filtering and memory recycling. Two novel optimized architectures emerge, together with a number of variations. These are validated, analyzed, and evaluated. We find that in the domain of interest speed-ups over GPUs are achieved. The potential impact is that this system promises to be the basis for scalable long timescale MD with commodity clusters. Chunshu Wu, Chen Yang 0010, Sahan Bandara, Tong Geng, Anqi Guo, Pouya Haghi, Ang Li 0006, Martin C. Herbordt |
IEEE Trans. Computers | 2 |
| 2022 | Reconfigurable switches for high performance and flexible MPI collectivesabstractAbstract There has been much effort in offloading MPI collective operations into hardware. But while NIC‐based collective acceleration is well‐studied, offloading their processing into the switching fabric, despite numerous advantages, has been much more limited. A major problem with fixed logic implementations is that either only a fraction of the possible collective communication is accelerated or that logic is wasted in the applications that do not need a particular capability. Using reconfigurable logic has numerous advantages: exactly the required operations can be implemented; the level of desired performance can be specified; and new, possibly complex, operations can be defined and implemented. We have designed an in‐switch collective accelerator,MPI‐FPGA, and demonstrated its use with seven MPI collectives and over a set of benchmarks and proxy applications (MiniApps). The accelerator uses a novel two‐level switch design containing fully pipelined vectorized aggregation logic units. Essential to this work is providing support for sub‐communicator collectives that enables communicators of arbitrary shape, and that is scalable to large systems. A streaming interface improves the performance for long messages. While this reconfigurable design is generally applicable, we prototype it with an FPGA‐centric cluster. A sampleMPI‐FPGAdesign in a direct network achieves considerable speedups over conventional clusters in the most likely scenarios. We also present results for indirect networks with reconfigurable high‐radix switches and show that this approach is competitive withSHArPtechnology for the subset of operations thatSHArPsupports.MPI‐FPGAis fully integrated into MPICH and is transparent to MPI applications. Pouya Haghi, Anqi Guo, Qingqing Xiong, Chen Yang 0010, Tong Geng, Justin T. Broaddus, Ryan J. Marshall, Derek Schafer, Anthony Skjellum, Martin C. Herbordt |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Upgrade of FPGA Range-Limited Molecular Dynamics to Handle Hundreds of ProcessorsabstractWith the current pandemic, the central role that Molecular Dynamics simulation (MD) plays in drug discovery makes advances in MD performance urgent. Recent work has demonstrated that among COTS devices only FPGA-centric clusters can scale beyond a few processors for relevant targets; other work has shown that single FPGA performance compares favorably to that of a GPU. In this study we demonstrate that an additional factor of 4× performance can be achieved which results in a factor of 5× speed up over a GPU. The problem addressed is that the designs of the last decade no longer scale when the number of processing pipelines grows from around ten to the hundreds. We begin by systematically evaluating existing work, exposing its flaws, and proposing a series of new design solutions. There are four major contributions. First, we address the massive routing problem by augmenting the design with three minimal networks in logic and latency. Second, we have developed a novel asynchronous out-of-order communication mechanism that removes nearly all bubbles from the routing networks. Third, we find that inverting the standard particle access algorithm results in improved locality and performance. Finally, we have created a custom numerical format that increases precision while saving space and logic. Chunshu Wu, Tong Geng, Sahan Bandara, Chen Yang 0010, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
FCCM | 4 |
| 2020 | Accelerating MPI Collectives with FPGAs in the Network and Novel Communicator SupportabstractMPI collective operations can often be performance killers in HPC applications; we seek to solve this bottleneck by offloading them to reconfigurable hardware within the switch itself, rather than, e.g., the NIC. We have designed a hardware accelerator MPI-FPGA to implement six MPI collectives in the network. Preliminary results show that MPI-FPGA achieves $10 \times $ speedup in the most likely scenarios over conventional clusters. We introduce a novel mechanism that enables the hardware to support a large number of communicators of arbitrary shape, and that is scalable to very large systems. MPI-FPGA is fully integrated into MPICH and so transparent to MPI applications. Qingqing Xiong, Chen Yang 0010, Pouya Haghi, Anthony Skjellum, Martin C. Herbordt |
FCCM | 2 |
| 2019 | LP-BNN: Ultra-low-Latency BNN Inference with Layer ParallelismabstractHigh inference latency seriously limits the deployment of DNNs in real-time domains such as autonomous driving, robotic control, and many others. To address this emerging challenge, researchers have proposed approximate DNNs with reduced precision, e.g., Binarized Neural Networks (BNNs). While BNNs can be built to have little loss in accuracy, latency reduction still has much room for improvement. In this paper, we propose a single-FPGA-based BNN accelerator that achieves microsecond-level ultra-low-latency inference of ImageNet, LP-BNN. We obtain this performance via several design optimizations. First, we optimize the network structure by removing Batch Normalization (BN) functions which leads to significant latency in BNNs without any loss on accuracy. Second, we propose a parameterized architecture which is based on layer parallelism and supports nearly perfect load balancing for various types of BNNs. Third, we fuse all the convolution layers and the first fully connected layer. We process them in parallel through fine-grained inter-layer pipelining. With our proposed accelerator, the inference of binarized AlexNet, VGGNet, and ResNet are completed within 21.5us, 335us, and 67.8us respectively, with no loss in accuracy as compared with other BNN implementations. Tong Geng, Chunshu Wu, Chen Yang 0010, Shuaiwen Song, Ang Li 0006, Martin C. Herbordt |
ASAP | 4 |
| 2019 | Molecular Dynamics Range-Limited Force Evaluation Optimized for FPGAsabstractFPGA Molecular Dynamics was much studied from 2004-2010. Due to limited chip resources of that era, and the inherent variety and complexity of tasks comprising Molecular Dynamics simulations (MD), those FPGA accelerators relied on host or embedded processors to organize and pre-process input and output data. This introduced long latency for data movement between simulation iterations and, as technology advanced, drastically limited performance. Current generation FPGAs are equipped not only with abundant on-chip resources, but also have hardware support for floating point operations; these advances provide an opportunity for creating self-contained MD simulation systems on a single device. In this paper, we demonstrate such a system based on the range-limited force, which comprises 90% of the flops in a typical MD simulation. It features online particle-pair generation, hundreds of force evaluation pipelines, motion update, and particle data migration. We integrate into OpenMM and find that, for a representative dataset (liquid argon with 20K atoms), we can achieve a simulation throughput of 1.4us/day with a single FPGA, more than twice the performance of a comparable generation GPU. The bulk of the work presented here explores the design of an independent MD range-limited force evaluation system tailored for modern FPGAs without data exchange with any off-chip devices. The primary contributions are the designs of the new features, the methods for coupling those features into an integrated system, and, especially, the analysis of the most likely mappings among particles/cells, on-chip memories (BRAMs), and on-chip compute units (pipelines). Chen Yang 0010, Tong Geng, Charles Lin, Jiayi Sheng, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
ASAP | 1 |
| 2019 | GhostSZ: A Transparent FPGA-Accelerated Lossy Compression FrameworkabstractHigh-performance computing (HPC) applications often generate enormous amounts of data that must be transferred for check-pointing, in situ processing, or post-execution analysis. To reduce the related network traffic and storage consumption, lossy compression schemes that target scientific data are often used. SZ compression emerged three years ago and has gained much attention because of its high compression ratio. However, performing SZ compression can take half a day per Terabyte of data; this could be a drawback to adoption. We propose GhostSZ an FPGA framework for accelerating tasks in SZ at line rate, and so transparently. The critical problem to be overcome is the tight data dependence central to SZ. GhostSZ solves this with a data transfer path having novel staged hardware. We test our implementation with both synthetic and real HPC application data and show 9.5×-80× core versus pipeline speedup over the optimized production version running on a state-of-the-art CPU and 8.2× per chip. Much of the variance in performance is due to the FPGA already running at line rate and so benefiting less from optimizations applicable to the CPU only on the most favorable data sets. The significance of this work is the possibility of a major reduction in required networking and storage in HPC installations. For example, using GhostSZ, fewer than 10 FPGAs would be sufficient to handle the entire I/O bandwidth of the top entry on the latest IO-500 list. Qingqing Xiong, Rushi Patel, Chen Yang 0010, Tong Geng, Anthony Skjellum, Martin C. Herbordt |
FCCM | 3 |
| 2019 | O3BNN: an out-of-order architecture for high-performance binarized neural network inference with fine-grained pruningabstractBinarized Neural Networks (BNN) have drawn tremendous attention due to significantly reduced computational complexity and memory demand. They have especially shown great potential in cost- and power-restricted domains, such as IoT and smart edge-devices, where reaching a certain accuracy bar is often sufficient, and real-time is highly desired. Tong Geng, Chunshu Wu, Chen Yang 0010, Wei Wu 0016, Ang Li 0006, Martin C. Herbordt |
ICS | 4 |
| 2019 | Fully integrated FPGA molecular dynamics simulationsabstractThe implementation of Molecular Dynamics (MD) on FPGAs has received substantial attention. Previous work, however, has consisted of either proof-of-concept implementations of components, usually the range-limited force; full systems, but with much of the work shared by the host CPU; or prototype demonstrations, e.g., using OpenCL, that neither implement a whole system nor have competitive performance. In this paper, we present what we believe to be the first full-scale FPGA-based simulation engine, and show that its performance is competitive with a GPU (running Amber in an industrial production environment). The system features on-chip particle data storage and management, short- and long-range force evaluation, as well as bonded forces, motion update, and particle migration. Other contributions of this work include exploring numerous architectural trade-offs and analysis of various mappings schemes among particles/cells and the various on-chip compute units. The potential impact is that this system promises to be the basis for long timescale Molecular Dynamics with a commodity cluster. Chen Yang 0010, Tong Geng, Rushi Patel, Qingqing Xiong, Ahmed Sanaullah, Chunshu Wu, Jiayi Sheng, Charles Lin, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
SC | 1 |
| 2018 | FPDeep: Acceleration and Load Balancing of CNN Training on FPGA ClustersabstractFPGA-based CNN accelerators have advantages in flexibility and power efficiency and so are being deployed by a number of cloud computing service providers, including Microsoft, Amazon, Tencent, and Alibaba. Given the increasing complexity of neural networks, however, it is becoming challenging to efficiently map CNNs to multi-FPGA platforms. In this work, we present a scalable framework, FPDeep, which helps engineers map a specific CNN's training logic to a multi-FPGA cluster or cloud and to build RTL implementations for the target network. With FPDeep, multi-FPGA accelerators work in a deeply-pipelined manner using a simple 1-D topology; this enables the accelerators to map directly onto many existing platforms, including Catapult, Catapult2, and almost any tightly-coupled FPGA cluster. FPDeep uses two mechanisms to facilitate high-performance and energy-efficiency. First, FPDeep provides a strategy to balance workload among FPGAs, leading to improved utilization. Second, training of CNNs is executed in a fine-grained inter- and intra-layer pipelined manner, minimizing the time that features need to remain available while waiting for back-propagation. This reduces the storage demand to where only on-chip memory is required for convolution layers. Experiments show that FPDeep has good scalability to a large number of FPGAs, with the limiting factor being the FPGA-to-FPGA bandwidth. Using six transceivers per FPGA, FPDeep shows linearity up to 60 FPGAs. We evaluate energy efficiency in GOPs/J and find that FPDeep provides up to 3.4 times higher energy efficiency than the Tesla K80 GPU. Tong Geng, Ahmed Sanaullah, Chen Yang 0010, Rushi Patel, Martin C. Herbordt |
FCCM | 4 |
| 2018 | High Performance Dynamic Communication on Reconfigurable ClustersabstractFPGA clusters with the FPGAs directly linked through their Multi-Gigabit Transceiver (MGT) ports have a proven advantage over other commodity architectures for communication-bound applications. We find that the standard wormhole routers need some modification to be appropriate for clusters with tightly coupled FPGAs, and create such a router. We generalize this router so that it is parameterized by several parameters including routing algorithm, arbitration policy, virtual channels, and buffers. We have evaluated these designs with respect to a number standard communication patterns and packet sizes. These results enable selection of the appropriate router for any resource budget. Finally, We find that the optimality of the router design varies significantly with workload. We observe that for a 512 FPGA cluster, connected in an 8^3 torus, compared with the router configuration with the best average performance, application-aware router configurations reduce average batch latency by 3%, improve the throughput by 6% on average, and improve area consumption by 50%. Jiayi Sheng, Chen Yang 0010, Martin C. Herbordt |
FCCM | 2 |
| 2018 | A Framework for Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters with Work and Weight Load BalancingabstractTo improve flexibility and energy efficiency of Convolutional Neural Networks, a number of cloud computing service providers-including Microsoft, Amazon, and Alibaba-are using FPGA-based CNN accelerators. However, the growing size and complexity of neural networks, coupled with communication and off-chip memory bottlenecks, make it increasingly difficult for multi-FPGA designs to achieve high resource utilization and performance, especially when training. In this work, we present new results for a scalable framework, FPDeep, which helps users efficiently map CNN training logic to multiple FPGAs and automatically generates the resulting RTL implementation. FPDeep is equipped with two mechanisms to facilitate high-performance and energy-efficient training. First, FPDeep improves DSP slice utilization across FPGAs by balancing workload using dedicated partition and mapping strategies. Second, only on-chip memory is used in the CONV layers: a) FPDeep balances CNN weight allocation among FPGAs to improve BRAM utilization; b) training of CNNs is executed in a fine-grained pipelined manner, minimizing the time features need to be cached while waiting for back-propagation leading to a reduced storage demand. We evaluate our framework by training AlexNet, VGG-16, and VGG-19. Experimental results show FPDeep has good scalability to a large number of FPGAs, with the limiting factor being the inter-FPGA bandwidth. With 6 transceivers per FPGA, FPDeep shows linearity up to 83 FPGAs. FPDeep provides, on average, 6.36x higher energy efficiency than GPU servers. Tong Geng, Ahmed Sanaullah, Chen Yang 0010, Rushi Patel, Martin C. Herbordt |
FPL | 4 |
| 2018 | High Performance Communication on Reconfigurable ClustersabstractFPGA clusters with the FPGAs directly linked through their Multi-Gigabit Transceivers (MGT) have a proven advantage over other commodity architectures for communication-bound applications. To date, however, communication infrastructure for such clusters has generally taken one of two approaches: nearest neighbor only, which is fast but has limited utility, and processor-based, which is general, but relatively slow. What is needed is for communication microarchitecture of these systems to be systematically explored, as has been done for HPC clusters and for Networks on Chip (NoC) on both FPGAs and ASICs. Our first contribution is finding that the properties of clusters of tightly coupled FPGAs substantially influence the router design space. We create a candidate router and generalize it so that it is parameterized by routing algorithm, arbitration policy, and virtual channels (VC). We have created a cycle-accurate simulator validated on a four-FPGA system. We evaluate the design space with respect to a number of standard communication patterns and packet sizes. These results enable selection of the appropriate router for any resource budget. We find that the optimality of the router design varies significantly with workloads. We present a framework that helps to determine appropriate parameters based on different applications and generate the HDL design. We observe that for a 512 FPGA cluster, compared with the router configuration with the best average performance, application-aware router selection can lead to substantial improvement in performance or reduction in area. Jiayi Sheng, Chen Yang 0010, Martin C. Herbordt |
FPL | 2 |
| 2018 | Real-time data analysis for medical diagnosis using FPGA-accelerated neural networksabstractBACKGROUND: Real-time analysis of patient data during medical procedures can provide vital diagnostic feedback that significantly improves chances of success. With sensors becoming increasingly fast, frameworks such as Deep Neural Networks are required to perform calculations within the strict timing constraints for real-time operation. However, traditional computing platforms responsible for running these algorithms incur a large overhead due to communication protocols, memory accesses, and static (often generic) architectures. In this work, we implement a low-latency Multi-Layer Perceptron (MLP) processor using Field Programmable Gate Arrays (FPGAs). Unlike CPUs and Graphics Processing Units (GPUs), our FPGA-based design can directly interface sensors, storage devices, display devices and even actuators, thus reducing the delays of data movement between ports and compute pipelines. Moreover, the compute pipelines themselves are tailored specifically to the application, improving resource utilization and reducing idle cycles. We demonstrate the effectiveness of our approach using mass-spectrometry data sets for real-time cancer detection. RESULTS: We demonstrate that correct parameter sizing, based on the application, can reduce latency by 20% on average. Furthermore, we show that in an application with tightly coupled data-path and latency constraints, having a large amount of computing resources can actually reduce performance. Using mass-spectrometry benchmarks, we show that our proposed FPGA design outperforms both CPU and GPU implementations, with an average speedup of 144x and 21x, respectively. CONCLUSION: In our work, we demonstrate the importance of application-specific optimizations in order to minimize latency and maximize resource utilization for MLP inference. By directly interfacing and processing sensor data with ultra-low latency, FPGAs can perform real-time analysis during procedures and provide diagnostic feedback that can be critical to achieving higher percentages of successful patient outcomes. Ahmed Sanaullah, Chen Yang 0010, Yuri Alexeev, Kazutomo Yoshii, Martin C. Herbordt |
BMC Bioinform. | 2 |
| 2017 | HPC on FPGA clouds: 3D FFTs and implications for molecular dynamicsabstractThe architecture of the Microsoft Catapult II cloud places the accelerator (FPGA) as a bump-in-the-wire on the way to the network and thus promises a dramatic reduction in latency as layers of hardware and software are avoided. We demonstrate this capability with an implementation of the 3D FFT. Next we examine phased application elasticity, i.e., the use of a reduced set of nodes for some phases of an HPC application. We find that, for the FFT phase within Molecular Dynamics, such contraction is beneficial with a 13%–14% performance improvement. Turning to MD, we show how this elasticity can be integrated into the existing data transformation to hide its communication overhead and increase the performance benefit to 16%–29%. Jiayi Sheng, Chen Yang 0010, Ahmed Sanaullah, Michael Papamichael, Adrian M. Caulfield, Martin C. Herbordt |
FPL | 2 |
| 2016 | Application-Aware Collective Communication (Extended Abstract)abstractPreliminary results are presented of hardware support for collective communication that takes advantage of a priori routing information. Jiayi Sheng, Qingqing Xiong, Chen Yang 0010, Martin C. Herbordt |
FCCM | 3 |