EDBT 2026 Demo / reviewers in the wild / expert
Jing Jane Li
dblp:181/2820-73 · also Jing Li 0073
· DBLP profile ↗
41ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0001-5139-938XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 9 first-author · 11 since 2021Software engineering, systems software and programming languages · 6 · 3 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiLFS: FPGA-Orchestrated File System for High-Level SynthesisabstractField Programmable Gate Arrays (FPGAs) deliver high performance, and High-Level Synthesis (HLS) simplifies computation description. However, modern FPGA systems cannot directly exploit the convenience and advanced features of contemporary file systems to enable performant, secure, and robust access to high-speed storage devices such as SSDs. This limitation significantly impedes the adoption of FPGAs in data- and I/O-intensive applications, including large language models (LLMs). Existing HLS storage solutions typically either rely on host CPUs to manage file systems via the operating system stack or provide only low-level block access, both of which introduce considerable performance and programmability overheads. Host-mediated access to storage incurs additional latency due to multiple round-trips through the OS kernel on the host CPU, while block-level management on the FPGA side requires substantial engineering efforts that often require recreating file system functionality, such as raw block management, security, and robustness guarantees. These challenges substantially complicate FPGA development and create a 100× gap in scalability compared to GPUs for deploying modern, large-scale machine learning models. To close this gap, we propose HiLFS, the first file system and storage stack for HLS that manages storage entirely within the FPGA. HiLFS exposes a POSIX-like file interface to HLS kernels to ease programming and maintains an on-chip cache of recently accessed file metadata to accelerate file access. It also provides rich file system features, including data integrity, crash consistency, durability guarantees, and efficient concurrent access. As such, HiLFS enables high-performance, secure, and reliable storage management, completely eliminating the need for host intervention. We prototype HiLFS on an AMD/Xilinx Alveo U200 FPGA with a Solidigm DC-P4610 SSD. On Mixtral 8×7B, HiLFS outperforms Nvidia Titan RTX by 1.1/1.3× in performance and 3.0/3.5× in energy efficiency with/without GPUDirect Storage. To the best of our knowledge, this represents the largest-scale LLM deployment on an FPGA to date. Moreover, HiLFS delivers 1.5/1.8× average latency and bandwidth improvements over state-of-the-art commercial CPU-centric storage platforms with/without PCIe P2P, while incurring 13% bandwidth and latency overhead to state-of-the-art HLS low-level block storage works. Furthermore, HiLFS reduces the LoC by 1.5× and 5.3× compared to CPU-centric and block-level storage platforms, respectively. YoungSeok Na, Linus Y. Wong, André DeHon, Jing Jane Li |
FPGA | 4 |
| 2026 | Hardware Accelerated FPGA Divide-and-Conquer Page Placement in MillisecondsabstractExcessive FPGA compilation times, often measured in hours, stifle rapid iterative development, design-space exploration, and runtime reconfiguration applications. Coarse-grain divide-and-conquer techniques, which break large applications into separately compiled pages, offer moderate speedups, potentially bringing compilation down to minutes, but leave significant fine-grain parallelism opportunities untapped. Systolic-array-based accelerators have previously offered orders of magnitude speedup for FPGA placement (a major bottleneck in compilation), by exploiting massive fine-grain parallelism, however poor scalability restricts them to small designs, and lack of support for modern heterogeneous netlists (CLBs, BRAMs, DSPs) prevents their use today. We introduce an enhanced, FPGA-based systolic placement accelerator, capable of placing divide-and-conquer page-sized netlists of CLBs, BRAMs, and DSPs onto VTR architectures, with 2-3 orders of magnitude speedup over VTR-9 running on a modern workstation-class processor. We demonstrate page-placement in milliseconds on realistic benchmarks, including HLS dataflow designs, run on an AMD Versal VCK190 implementation of our systolic placer, forging a path towards real-time, self-hosted FPGA compilation. Ezra Thomas, Jing Jane Li, André DeHon |
FPGA | 2 |
| 2024 | DONGLE 2.0: Direct FPGA-Orchestrated NVMe Storage for HLSabstractRapid growth in data size poses significant computational and memory challenges to data processing. FPGA accelerators and near-storage processing have emerged as compelling solutions for tackling the growing computational and memory requirements. Many FPGA-based accelerators have shown to be effective in processing large data sets by leveraging the storage capability of either host-attached or FPGA-attached storage devices. However, the current HLS development environment does not allow direct access to host-or FPGA-attached NVMe storage from the HLS code. As such, users must frequently hand off between HLS and host code to access data in storage, and such a process requires tedious programming to ensure functional correctness. Moreover, since the HLS code uses radically different methods to access storage compared to DRAM, the HLS codebase targeting DRAM-based platforms cannot be easily ported to NVMe-based platforms, resulting in limited code portability and reusability. Furthermore, frequent suspension of HLS kernel and synchronization between CPU and FPGA introduce significant latency overhead and require sophisticated scheduling mechanisms to hide latency. To address these challenges, we propose a new HLS storage interface named DONGLE 2.0 that enables direct FPGA-orchestrated NVMe storage access. By providing a unified interface for storage and memory access, DONGLE 2.0 allows a single-source HLS program to target multiple memory/storage devices, thus making the codebase cleaner, portable, and more efficient. DONGLE 2.0 is an extension to DONGLE 1.0 [ 1 ] but adds support for host-attached storage. While its primary focus is still on FPGA NVMe access in near-storage configurations, the added host storage support ensures its compatibility with platforms that lack native support for FPGA-attached NVMe storage. We implemented a prototype of DONGLE 2.0 using an AMD/Xilinx Alveo U200 FPGA and Solidigm DC-P4610 SSD. Our evaluation on various workloads showed a geometric mean speed-up of 2.3× and a reduction in lines of code (LoC) by 2.4× compared to the state-of-the-art commercial platform when using FPGA-attached NVMe storage. Moreover, DONGLE 2.0 demonstrated a geometric mean speed-up of 1.5× and a reduction in LoC by 2.4× compared to the state-of-the-art commercial platform when using host-attached NVMe storage. Linus Y. Wong, Jing Jane Li |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2023 | DONGLE: Direct FPGA-Orchestrated NVMe Storage for HLSabstractRapid growth in data size poses increasing computational and memory challenges to data processing. FPGA accelerators and near-storage processing are promising candidates for tackling computational and memory requirements, and many near-storage FPGA accelerators have been shown to be effective in processing large data. However, the current HLS development environment does not allow direct NVMe storage access from the HLS code. As such, users must frequently hand off between HLS and host code to access data in storage, and such a process requires tedious programming to ensure functional correctness. Moreover, since the HLS code uses radically different methods to access storage compared to DRAM, the HLS codebase targeting DRAM-based platforms cannot be easily ported to NVMe-based platforms, resulting in limited code portability and reusability. Furthermore, frequent suspension of HLS kernel and synchronization between CPU and FPGA introduce significant latency overhead and require sophisticated scheduling mechanisms to hide latency. Linus Y. Wong, Jing Jane Li |
FPGA | 3 |
| 2023 | Introduction to the Special Section on FCCM 2022abstractNo abstract available. Jing Jane Li, Martin C. Herbordt |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | Software-defined address mapping: a case on 3D memoryabstract3D-stacking memory such as High-Bandwidth Memory (HBM) and Hybrid Memory Cube (HMC) provides orders of magnitude more bandwidth and significantly increased channel-level parallelism (CLP) due to its new parallel memory architecture. However, it is challenging to fully exploit the abundant CLP for performance as the bandwidth utilization is highly dependent on address mapping in the memory controller. Unfortunately, CLP is very sensitive to a program’s data access pattern, which is not made available to OS/hardware by existing mechanisms. Michael M. Swift, Jing Jane Li |
ASPLOS | 3 |
| 2022 | Augmenting HLS with Zero-Overhead Application-Specific Address Mapping for Optane DCPMMabstractFPGAs have been introduced to datacenters as a mainstream computing device to accelerate a wide range of data-intensive applications when paired with heterogeneous memory. Leveraging High-Level Synthesis (HLS), application engineers not only can accelerate their applications but also the development time of designing, debugging and validating accelerators. However, existing HLS flows do not have effective support for emerging memory devices such as Intel’s Optane DC Persistent Memory Modules (Optane DCPMM) – a storage-class memory in a DIMM form factor. In fact, we observe that some HLS kernels can at best utilize only one-tenth of the total memory bandwidth of Optane DCPMM.To remedy the poor performance of HLS with Optane DCPMM, we augment the existing HLS external memory interface with zero-overhead, application-specific address mapping capabilities. The proposed scheme utilizes both fine-grained information from variable access patterns and coarse-grained variable-interleaving information to select an optimal hybrid address mapping for high memory bandwidth utilization, compared to a default fixed address mapping in existing HLS. Furthermore, our scheme is compatible with existing tool flows such as the Intel FPGA SDK for OpenCL and Vitis Application Flow to maintain a low adoption barrier. We observe that by using our proposed address mapping scheme and interface, we achieve 10× speedup on a diverse set of benchmarks including merge join, matrix multiplication and convolution without any additional hardware cost. Nicholas Beckwith, Jing Jane Li |
FCCM | 3 |
| 2022 | Revisiting PathFinder Routing AlgorithmabstractPathFinder, a popular routing algorithm widely used in the state-of-the-art FPGA compilation tools such as VTR 8, has been reported to have a non-negligible variation in the routing quality under routing resource constraints. Several workaround methods have been proposed to mitigate this undesired effect while keeping the algorithm unchanged. In this paper, we take a fresh look at the PathFinder algorithm itself and identify that its inappropriate definition of routing congestion problems causes an issue in its congestion coefficient updating strategy and eventually leads to the variation in the routing quality under different routing orders. This problem is aggravated in VTR~8 and inevitably degrades the routing quality as the predefined routing order is not optimal in most cases. Based on these findings, we propose an enhanced PathFinder algorithm with redefined routing congestion problems to address the issue in the congestion coefficient updating strategy and algorithmically resolve the routing quality variation issue. This enhanced algorithm can be easily integrated into VTR~8 to benefit every researcher in our community. Evaluation results show that it reduces the longest critical path delay and the variation by up to 49.4% and 96.2% under routing resource constraints, respectively. The variation can be fully eliminated when sufficient routing iterations are performed. Yue Zha, Jing Jane Li |
FPGA | 2 |
| 2021 | When application-specific ISA meets FPGAs: a multi-layer virtualization framework for heterogeneous cloud FPGAsabstractWhile field-programmable gate arrays (FPGAs) have been widely deployed into cloud platforms, the high programming complexity and the inability to manage FPGA resources in an elastic/scalable manner largely limits the adoption of FPGA acceleration. Existing FPGA virtualization mechanisms partially address these limitations. Application-specific (AS) ISA provides a nice abstraction to enable a simple software programming flow that makes FPGA acceleration accessible by the mainstream software application developers. Nevertheless, existing AS ISA-based approaches can only manage FPGA resources at a per-device granularity, leading to a low resource utilization. Alternatively, hardware-specific (HS) abstraction improves the resource utilization by spatially sharing one FPGA among multiple applications. But it cannot reduce the programming complexity due to the lack of a high-level programming model. Yue Zha, Jing Jane Li |
ASPLOS | 2 |
| 2021 | GORDON: Benchmarking Optane DC Persistent Memory Modules on FPGAsabstractScalable nonvolatile memory DIMMs become commercially available on FPGAs with the release of Intel's Optane DC Persistent Memory (DCPM) product. This new class of memory combines the benefits of DRAM-like solid-state memory (fast, byte addressable) and Flash-like persistent storage (cost- effective, non-volatile), making FPGA highly competitive in accelerating large-scale machine learning and data analytics applications. Despite of the great promise, the performance characteristics of Optane DCPM remains relatively alien to FPGA developers compared to conventional DDRx DRAM or SSD. Recent preliminary studies all use CPU-based systems running full OS stack that limits their ability to characterize the detailed performance characteristics of the Optane DCPM. To fully exploit the advantages of Optane DCPM in FPGA-based accelerator design, we present the first FPGA-based Optane DCPM profiling framework, named GORDON1on Stratix 10 DX FPGA. By leveraging the flexibility of the FPGA in building custom logic and the FPGA-specific features for Optane DCPM, GORDON addresses the fundamental limitations of prior CPU- based profiling. The detailed understanding on Optane DCPM may also benefit system design and optimization beyond FPGAs using CPUs and GPUs. Nicholas Beckwith, Jing Jane Li |
FCCM | 3 |
| 2021 | Hetero-ViTAL: A Virtualization Stack for Heterogeneous FPGA ClustersabstractWith field-programmable gate arrays (FPGAs) being widely deployed into data centers, an efficient virtualization support is required to fully unleash the potential of cloud FPGAs. Nevertheless, existing FPGA virtualization solutions only support a homogeneous FPGA cluster comprising identical FPGA devices. Representative work such as ViTAL provides sufficient system support for scale-out acceleration and improves the overall resource utilization through a fine-grained spatial sharing. While these existing solutions (including ViTAL) can efficiently virtualize a homogeneous cluster, it is hard to extend them to virtualizing a heterogeneous cluster which comprises multiple types of FPGAs. We expect the future cloud FPGAs are likely to be more heterogeneous due to hardware rolling upgrade.In this paper, we rethink FPGA virtualization from ground up and propose Hetero-ViTAL to virtualize heterogeneous FPGA clusters. We identify the conflicting requirements of runtime management and offline compilation when designing the abstraction for a heterogeneous cluster, which is also the fundamental reason why the single-level abstraction as proposed in ViTAL (and other prior works) cannot be trivially extended to the heterogeneous case. To decouple these conflicting requirements, we provide a two-level system abstraction in Hetero-ViTAL. Specifically, the high-level abstraction is FPGA-agnostic and provides a simple and homogeneous view of the FPGA resources to simplify the runtime management. On the contrary, the low-level abstraction is FPGA-specific and exposes sufficient spatial resource constraints to the compilation framework to ensure the mapping quality. Rather than simply adding a layer on top of the single-level abstraction as proposed in ViTAL and other prior work, we judiciously determine how much hardware details should be exposed at each level to balance the management complexity, mapping quality and compilation cost. We then develop a compilation framework to map applications onto this two-level abstraction with several optimization techniques to further improve the mapping quality. We also provide a runtime management policy to alleviate the fragmentation issue, which becomes more severe in a heterogeneous cluster due to the distinct resource capacities of diverse FPGAs.We evaluate Hetero-ViTAL on a custom-built FPGA cluster and demonstrate its effectiveness using machine learning and image processing applications. Results show that Hetero-ViTAL reduces the average response time (a critical metric for QoS) by 79.2% for a heterogeneous cluster compared to the non-virtualized baseline. When virtualizing a homogeneous cluster, Hetero-ViTAL also reduces the average response time by 42.0% compared with ViTAL due to a better system design. Yue Zha, Jing Jane Li |
ISCA | 2 |
| 2020 | Virtualizing FPGAs in the CloudabstractField-Programmable Gate Arrays (FPGAs) have been integrated into the cloud infrastructure to enhance its computing performance by supporting on-demand acceleration. However, system support for FPGAs in the context of the cloud environment is still in its infancy with two major limitations, i.e., the inefficient runtime management due to the tight coupling between compilation and resource allocation, and the high programming complexity when exploiting scale-out acceleration. The root cause is that FPGA resources are not virtualized. In this paper, we propose a full-stack solution, namely ViTAL, to address the aforementioned limitations by virtualizing FPGA resources. Specifically, ViTAL provides a homogeneous abstraction to decouple the compilation and resource allocation. Applications are offline compiled onto the abstraction, while the resource allocation is dynamically determined at runtime. Enabled by a latency-insensitive communication interface, applications can be mapped flexibly onto either one FPGA or multiple FPGAs to maximize the resource utilization and the aggregated system throughput. Meanwhile, ViTAL creates an illusion of a single and large FPGA to users, thereby reducing the programming complexity and supporting scale-out acceleration. Moreover, ViTAL also provides virtualization support for peripheral components (e.g., on-board DRAM and Ethernet), as well as protection and isolation support to ensure a secure execution in the multi-user cloud environment. We evaluate ViTAL on a real system - an FPGA cluster composed of the latest Xilinx UltraScale+ FPGAs (XCVU37P). The results show that, compared with the existing management method, ViTAL enables fine-grained resource sharing and reduces the response time by 82% on average (improving Quality-of-Service) with a marginal virtualization overhead. Moreover, ViTAL also reduces the response time by 25% compared to AmorphOS (operating in high-throughput mode), a recently proposed FPGA virtualization method. Yue Zha, Jing Jane Li |
ASPLOS | 2 |
| 2020 | Hyper-Ap: Enhancing Associative Processing Through A Full-Stack OptimizationabstractAssociative processing (AP) is a promising PIM paradigm that overcomes the von Neumann bottleneck (memory wall) by virtue of a radically different execution model. By decomposing arbitrary computations into a sequence of primitive memory operations (i.e., search and write), AP's execution model supports concurrent SIMD computations in-situ in the memory array to eliminate the need for data movement. This execution model also provides a native support for flexible data types and only requires a minimal modification on the existing memory design (low hardware complexity). Despite these advantages, the execution model of AP has two limitations that substantially increase the execution time, i.e., 1) it can only search a single pattern in one search operation and 2) it needs to perform a write operation after each search operation. In this paper, we propose the Highly Performant Associative Processor (Hyper- AP) to fully address the aforementioned limitations. The core of Hyper- AP is an enhanced execution model that reduces the number of search and write operations needed for computations, thereby reducing the execution time. This execution model is generic and improves the performance for both CMOS-based and RRAM-based AP, but it is more beneficial for the RRAMbased AP due to the substantially reduced write operations. We then provide complete architecture and micro-architecture with several optimizations to efficiently implement Hyper-AP. In order to reduce the programming complexity, we also develop a compilation framework so that users can write C-like programs with several constraints to run applications on Hyper- AP. Several optimizations have been applied in the compilation process to exploit the unique properties of Hyper- AP. Our experimental results show that, compared with the recent work IMP, Hyper- AP achieves up to 54×/4.4× better power-/area-efficiency for various representative arithmetic operations. For the evaluated benchmarks, Hyper-AP achieves 3.3× speedup and 23.8× energy reduction on average compared with IMP. Our evaluation also confirms that the proposed execution model is more beneficial for the RRAM-based AP than its CMOS-based counterpart. Yue Zha, Jing Jane Li |
ISCA | 2 |
| 2020 | MEG: A RISCV-based System Emulation Infrastructure for Near-data Processing Using FPGAs and High-bandwidth MemoryabstractEmerging three-dimensional (3D) memory technologies, such as the Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM), provide high-bandwidth and massive memory-level parallelism. With the growing heterogeneity and complexity of computer systems (CPU cores and accelerators, etc.), efficiently integrating emerging memories into existing systems poses new challenges and requires detailed evaluation in a realistic computing environment. In this article, we propose MEG, an open source, configurable, cycle-exact, and RISC-V-based full-system emulation infrastructure using FPGA and HBM. MEG provides a highly modular hardware design and includes a bootable Linux image for a realistic software flow, so that users can perform cross-layer software-hardware co-optimization in a full-system environment. To improve the observability and debuggability of the system, MEG also provides a flexible performance monitoring scheme to guide the performance optimization. The proposed MEG infrastructure can potentially benefit broad communities across computer architecture, system software, and application software. Leveraging MEG, we present two cross-layer system optimizations as illustrative cases to demonstrate the usability of MEG. In the first case study, we present a reconfigurable memory controller to improve the address mapping of standard memory controller. This reconfigurable memory controller along with its OS support allows us to optimize the address mapping scheme to fully exploit the massive parallelism provided by the emerging three-dimensional (3D) memories. In the second case study, we present a lightweight IOMMU design to tackle the unique challenges brought by 3D memory in providing virtual memory support for near-memory accelerators. We provide a prototype implementation of MEG on a Xilinx VU37P FPGA and demonstrate its capability, fidelity, and flexibility on real-world benchmark applications. We hope MEG fills a gap in the space of publicly available FPGA-based full-system emulation infrastructures, specifically targeting memory systems, and inspires further collaborative software/hardware innovations. Yue Zha, Nicholas Beckwith, Bangya Liu, Jing Jane Li |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2019 | MEG: A RISCV-Based System Simulation Infrastructure for Exploring Memory Optimization Using FPGAs and Hybrid Memory CubeabstractEmerging 3D memory technologies, such as the Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM), provide increased bandwidth and massive memory-level parallelism. Efficiently integrating emerging memories into existing system pose new challenges and require detailed evaluation in a real computing environment. In this paper, we propose MEG, an open-source, configurable, cycle-exact, and RISC-V based full system simulation infrastructure using FPGA and HMC. MEG has three highly configurable design components: (i) a HMC adaptation module that not only enables communication between the HMC device and the processor cores but also can be extended to fit other memories (e.g., HBM, nonvolatile memory) with minimal effort, (ii) a reconfigurable memory controller along with its OS support that can be effectively leveraged by system designers to perform software-hardware co-optimization, and (iii) a performance monitor module that effectively improves the observability and debuggability of the system to guide performance optimization. We provide a prototype implementation of MEG on Xilinx VCU110 board and demonstrate its capability, fidelity, and flexibility on real-world benchmark applications. We hope that our open-source release of MEG fills a gap in the space of publicly-available FPGA-based full system simulation infrastructures specifically targeting memory system and inspires further collaborative software/hardware innovations. Gaurav Jain, Yue Zha, Jonathan Ta, Jing Jane Li |
FCCM | 6 |
| 2019 | Unleashing the Power of Soft Logic for Convolutional Neural Network Acceleration via Product QuantizationabstractTo reduce the load of taxing CNN infrastructures, both industry and academia show great interest in building specialized hardware for CNN acceleration. Numerous FPGA-based accelerators have been proposed to better utilize hard blocks. However, the capability of soft logic has not been fully explored. Prior works either fail to utilize soft logic for computation or inefficiently use soft logic to mimic the function of hard blocks. In this work, to better utilize soft logic we propose to use the native function of soft logic, i.e. as distributed memory, which yields more than 20x throughput compared to implementing multiplication operations. To fully leverage this potential, we present a framework to more efficiently accelerate CNN-based inference. Firstly, we employ product quantization (PQ) to convert most of the multiplications to a reduced number of distributed memory accesses and additions. We then constructed an analytical model to reveal the complex relationship between the inference accuracy and the utilization of hard blocks and soft blocks. Based on the model, we select optimal PQ parameters for a balanced design. In the rest of this paper, we describe the complete system implementation and discuss the experimental results. According to our results on Xilinx VU9P FPGA, we can achieve a 140 Tops equivalent throughput and 475 Gops/W energy efficiency with less than 0.5% accuracy degradation. Jing Jane Li |
FPGA | 2 |
| 2018 | Liquid Silicon-Monona: A Reconfigurable Memory-Oriented Computing Fabric with Scalable Multi-Context SupportabstractWith the recent trend of promoting Field-Programmable Gate Arrays (FPGAs) to first-class citizens in accelerating compute-intensive applications in networking, cloud services and artificial intelligence, FPGAs face two major challenges in sustaining competitive advantages in performance and energy efficiency for diverse cloud workloads: 1) limited configuration capability for supporting light-weight computations/on-chip data storage to accelerate emerging search-/data-intensive applications. 2) lack of architectural support to hide reconfiguration overhead for assisting virtualization in a cloud computing environment. In this paper, we propose a reconfigurable memory-oriented computing fabric, namely Liquid Silicon-Monona (L-Si), enabled by emerging nonvolatile memory technology i.e. RRAM, to address these two challenges. Specifically, L-Si addresses the first challenge by virtue of a new architecture comprising a 2D array of physically identical but functionally-configurable building blocks. It, for the first time, extends the configuration capabilities of existing FPGAs from computation to the whole spectrum ranging from computation to data storage. It allows users to better customize hardware by flexibly partitioning hardware resources between computation and memory, greatly benefiting emerging search- and data-intensive applications. To address the second challenge, L-Si provides scalable multi-context architectural support to minimize reconfiguration overhead for assisting virtualization. In addition, we provide compiler support to facilitate the programming of applications written in high-level programming languages (e.g. OpenCL) and frameworks (e.g. TensorFlow, MapReduce) while fully exploiting the unique architectural capability of L-Si. Our evaluation results show L-Si achieves 99.6% area reduction, 1.43× throughput improvement and 94.0% power reduction on search-intensive benchmarks, as compared with the FPGA baseline. For neural network benchmarks, on average, L-Si achieves 52.3× speedup, 113.9× energy reduction and 81% area reduction over the FPGA baseline. In addition, the multi-context architecture of L-Si reduces the context switching time to - 10ns, compared with an off-the-shelf FPGA (∼100ms), greatly facilitating virtualization. Yue Zha, Jing Jane Li |
ASPLOS | 2 |
| 2018 | Efficient Large-Scale Approximate Nearest Neighbor Search on OpenCL FPGAabstractWe present a new method for Product Quantization (PQ) based approximated nearest neighbor search (ANN) in high dimensional spaces. Specifically, we first propose a quantization scheme for the codebook of coarse quantizer, product quantizer, and rotation matrix, to reduce the cost of accessing these codebooks. Our approach also combines a highly parallel k-selection method, which can be fused with the distance calculation to reduce the memory overhead. We implement the proposed method on Intel HARPv2 platform using OpenCL-FPGA. The proposed method significantly outperforms state-of-the-art methods on CPU and GPU for high dimensional nearest neighbor queries on billion-scale datasets in terms of query time and accuracy regardless of the batch size. To our best knowledge, this is the first work to demonstrate FPGA performance superior to CPU and GPU on high-dimensional, large-scale ANN datasets. Soroosh Khoram, Jing Jane Li |
CVPR | 3 |
| 2018 | PQ-CNN: Accelerating Product Quantized Convolutional Neural Network on FPGAabstractThis work presents an efficient CNN computation framework on FPGA, which utilizes Product Quantization (PQ). Compared to other compression methods, PQ has larger compression ratios and, furthermore, it alleviates the irregularity problem. However, its algorithmic benefits do not translate to system performance gains because of: 1) a large codebook that diminishes the compression ratio; 2) large numbers of look-up operations that are inefficient on CPU and GPU architectures. In this work, to address these problems, we first provide an analytical model to guide our design and find a dilemma for selecting PQ parameters. Then, we propose a software/hardware method to tackle these issues. We present a complete framework to optimally implement PQ-CNN on FPGA. According to our experimental results, we can achieve 140 Tops equivalent throughput, 475 Gops/w energy efficiency and with less than 0.5% accuracy degradation. Jing Jane Li |
FCCM | 2 |
| 2018 | Accelerating Graph Analytics by Co-Optimizing Storage and Access on an FPGA-HMC PlatformabstractGraph analytics, which explores the relationships among interconnected entities, is becoming increasingly important due to its broad applicability, from machine learning to social sciences. However, due to the irregular data access patterns in graph computations, one major challenge for graph processing systems is performance. The algorithms, softwares, and hardwares that have been tailored for mainstream parallel applications are generally not effective for massive, sparse graphs from the real-world problems, due to their complex and irregular structures. To address the performance issues in large-scale graph analytics, we leverage the exceptional random access performance of the emerging Hybrid Memory Cube (HMC) combined with the flexibility and efficiency of modern FPGAs. In particular, we develop a collaborative software/hardware technique to perform a level-synchronized Breadth First Search (BFS) on a FPGA-HMC platform. From the software perspective, we develop an architecture-aware graph clustering algorithm that exploits the FPGA-HMC platform»s capability to improve data locality and memory access efficiency. From the hardware perspective, we further improve the FPGA-HMC graph processor architecture by designing a memory request merging unit to take advantage of the increased data locality resulting from graph clustering. We evaluate the performance of our BFS implementation using the AC-510 development kit from Micron and achieve $2.8 \times$ average performance improvement compared to the latest FPGA-HMC based graph processing system over a set of benchmarks from a wide range of applications. Soroosh Khoram, Maxwell Strange, Jing Jane Li |
FPGA | 4 |
| 2018 | Liquid Silicon: A Data-Centric Reconfigurable Architecture Enabled by RRAM TechnologyabstractThis paper presents a data-centric reconfigurable architecture, namely Liquid Silicon, enabled by emerging non-volatile memory, i.e., RRAM. Compared to the heterogeneous architecture of commercial FPGAs, Liquid Silicon is inherently a homogeneous architecture comprising a two-dimensional (2D) array of identical 'tiles'. Each tile can be configured into one or a combination of four modes: TCAM, logic, interconnect, and memory. Such flexibility allows users to partition resources based on applications? needs, in contrast to the fixed hardware design using dedicated hard IP blocks in FPGAs. In addition to better resource usage, its 'memory friendly' architecture effectively addresses the limitations of commercial FPGAs i.e., scarce on-chip memory resources, making it an effective complement to FPGAs. Moreover, its coarse-grained logic implementation results in shallower logic depth, less inter-tile routing overhead, and thus smaller area and better performance, compared with its FPGA counterpart. Our study shows that, on average, for both traditional and emerging applications, we achieve 62% area reduction, 27% speedup and 31% improvement in energy efficiency when mapping applications onto Liquid Silicon instead of FPGAs. Yue Zha, Jing Jane Li |
FPGA | 2 |
| 2018 | Degree-aware Hybrid Graph Traversal on FPGA-HMC PlatformabstractGraph traversal is a core primitive for graph analytics and a basis for many higher-level graph analysis methods. However, irregularities in the structure of scale-free graphs (e.g., social network) limit our ability to analyze these important and growing datasets. A key challenge is the redundant graph computations caused by the presence of high-degree vertices which not only increase the total amount of computations but also incur unnecessary random data access. Jing Jane Li |
FPGA | 2 |
| 2018 | Adaptive Quantization of Neural Networks
Soroosh Khoram, Jing Jane Li |
ICLR (Poster) | 2 |
| 2017 | Accelerating Large-Scale Graph Analytics with FPGA and HMCabstractGraph analytics that explores the relationship among interconnected entities is becoming increasingly important due to its broad applicability from machine learning to social science. However, one major challenge for graph processing systems is the irregular data access pattern of graph computation which can significantly degrade the performance. The algorithms, software, and hardware that have been tailored for mainstream parallel applications are, as a result, generally not effective for massive-scale sparse graphs from the real world due to their complexity and irregularity. To address the performance issues in large-scale graph analytics, we combine the emerging Hybrid Memory Cube (HMC) with a modern FPGA in order to achieve exceptional random access performance without any loss of flexibility or efficiency in computation. In particular, we develop collaborative software/hardware techniques to perform a level-synchronized breadth first search (BFS) on the FPGA-HMC platform. From the software perspective, we develop an architecture-aware graph clustering algorithm that fully exploits the platform's capability to improve data locality and memory access efficiency. For each input graph, this algorithm provides an efficient data layout that allows the FPGA to coalesce memory requests into the largest possible HMC payload requests so that the number of memory requests, which is the primary factor in runtime, can be minimized. From the hardware perspective, we further improve the FPGA-HMC graph processor architecture by adding a merging unit. The merging unit takes the best advantage of the increased data locality resulting from graph clustering. We evaluated the performance of our BFS implementation using the AC-510 development kit from Micron over a set of benchmarks from a wide range of applications. We observed that the combination of the clustering algorithm and the merging hardware achieved 2.8 × average performance improvement compared to the latest FPGA-HMC based graph processing system. Soroosh Khoram, Maxwell Strange, Jing Jane Li |
FCCM | 4 |
| 2017 | A Mixed-Signal Data-Centric Reconfigurable Architecture enabled by RRAM Technology (Abstract Only)
Yue Zha, Jing Jane Li |
FPGA | 4 |
| 2017 | Boosting the Performance of FPGA-based Graph Processor using Hybrid Memory Cube: A Case for Breadth First Search
Soroosh Khoram, Jing Jane Li |
FPGA | 3 |
| 2017 | RRAM-based reconfigurable in-memory computing architecture with hybrid routing
Yue Zha, Jing Jane Li |
ICCAD | 2 |
| 2017 | Challenges and Opportunities: From Near-memory Computing to In-memory ComputingabstractThe confluence of the recent advances in technology and the ever-growing demand for large-scale data analytics created a renewed interest in a decades-old concept, processing-in-memory (PIM). PIM, in general, may cover a very wide spectrum of compute capabilities embedded in close proximity to or even inside the memory array. In this paper, we present an initial taxonomy for dividing PIM into two broad categories: 1) Near-memory processing and 2) In-memory processing. This paper highlights some interesting work in each category and provides insights into the challenges and possible future directions. Soroosh Khoram, Yue Zha, Jing Jane Li |
ISPD | 4 |
| 2016 | Reconfigurable in-memory computing with resistive memory crossbarabstractDriven by recent advances in resistive random-access memory (RRAM), there have been growing interests in exploring alternative computing concept, i.e., in-memory processing, to address the classical von Neumann bottlenecks. Despite of their great promise in improving performance and energy efficiency, most existing works are built on the inherent matrix-vector multiplication capability of RRAM crossbar structure, and thus lack the flexibility to adapt to future market/technology induced changes in data-intensive applications. To address these challenges, we propose an in-memory reconfigurable architecture based on RRAM crossbar structure. For the first time, it achieves a full programmability across computation and storage, and thereby provides more flexibilities of partitioning the hardware resources based on applications' needs. We further develop two complete CAD design flows to facilitate development of applications written in hardware description languages (HDLs) for our architecture, based on: 1) adaption from existing tool set developed for FPGA, 2) a custom tool design optimized towards the new architecture. Our experiments show that, both design flows are effective in exploiting flexible resources offered by our architecture and thus achieves better efficiency than state-of-art FPGAs (30% improvement in performance with 66% reduction in area). In addition, compared to adapted design flow, our custom design flow achieves speedup by 3.3×, and further improves mapping quality. Yue Zha, Jing Jane Li |
ICCAD | 2 |
| 2015 | Enabling phase-change memory for data-centric computing: Technology, circuitand systemabstractEmerging nonvolatile memory (NVM) technology i.e., phase-change memory (PCM) has been commonly employed as a drop-in replacement for either DRAM or Flash. However, the inherent nature of PCM technology does not align perfectly with either applications in terms of cost-per-bit, performance, power, endurance or retention. The missing killer applications for PCM have slowed down the technology development from becoming mainstream. From the systems perspective, the ever-growing big data problems call for a paradigm shift from traditional compute-centric system to data-centric system for better performance, efficiency and productivity. A data-centric system will demand new hardware features to support the efficient storage and manipulation of data in a wide range of data-intensive applications. Considering these factors, this paper will give an overview of recent research efforts to enable PCM for building a reliable and efficient ternary content addressable memory (TCAM). By leveraging multiple levels of computing stack - technology, circuit and architecture, we will show a holistic approach of tailoring PCM to meeting the new system requirements. It opens up opportunities of accelerating NVM development as the methodology can be generalized to other NVM technologies. Jing Jane Li |
ISCAS | 1 |
| 2012 | A case for small row buffers in non-volatile main memoriesabstractDRAM-based main memories have read operations that destroy the read data, and as a result, must buffer large amounts of data on each array access to keep chip costs low. Unfortunately, system-level trends such as increased memory contention in multi-core architectures and data mapping schemes that improve memory parallelism lead to only a small amount of the buffered data to be accessed. This makes buffering large amounts of data on every memory array access energy-inefficient; yet organizing DRAM chips to buffer small amounts of data is costly, as others have shown [11]. Emerging non-volatile memories (NVMs) such as PCM, STT-RAM, and RRAM, however, do not have destructive read operations, opening up opportunities for employing small row buffers without incurring additional area penalty and/or design complexity. In this work, we discuss and evaluate architectural changes to enable small row buffers at a low cost in NVMs. We find that on a multi-core system, reducing the row buffer size can greatly reduce main memory dynamic energy compared to a DRAM baseline with large row sizes, without greatly affecting endurance, and for some NVM technologies, leads to improved performance. Justin Meza, Jing Jane Li, Onur Mutlu |
ICCD | 2 |
| 2011 | Phase change memory
Jing Jane Li, Chung Lam |
Sci. China Inf. Sci. | 1 |
| 2010 | Variable-Latency Adder (VL-Adder) Designs for Low Power and NBTI ToleranceabstractIn this paper, we proposed a new adder design called variable-latency adder (VL-adder). This technique allows the adder to work at a lower supply voltage than that required by a conventional adder while maintaining the same throughput. The VL-adder design can be further modified to overcome the effects of negative bias temperature instability (NBTI) on circuit delay. By applying VL-adder concept to a 64-bit carry-select adder design, more than 40% energy saving is obtained when a similar throughput is maintained. Yiran Chen 0001, Hai Li 0001, Cheng-Kok Koh, Guangyu Sun 0003, Jing Jane Li, Yuan Xie 0001, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2010 | Design Paradigm for Robust Spin-Torque Transfer Magnetic RAM (STT MRAM) From Circuit/Architecture PerspectiveabstractSpin-torque transfer magnetic RAM (STT MRAM) is a promising candidate for future embedded applications. It combines the desirable attributes of current memory technologies such as SRAM, DRAM, and flash memories (fast access time, low cost, high density, and non-volatility). It also solves the critical drawbacks of conventional MRAM technology: poor scalability and high write current. However, variations in process parameters can lead to a large number of cells to fail, severely affecting the yield of the memory array. In this paper, we analyzed and modeled the failure probabilities of STT MRAM cells due to parameter variations. Based on the model, we performed a thorough analysis of the impact of design parameters on parametric failures due to process variations. To achieve high memory yield without incurring expensive technology modification, we developed an efficient design paradigm from circuit and/or architecture perspective-to improve the robustness and integration density. The proposed technique effectively relaxes or completely decouples the conflicting design requirements for read stability, writability and cell area. It can be used at an early stage of the design cycle for yield enhancement. Jing Jane Li, Patrick Ndai, Ashish Goel, Sayeef S. Salahuddin, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | An alternate design paradigm for robust spin-torque transfer magnetic RAM (STT MRAM) from circuit/architecture perspectiveabstractSpin-Torque Transfer Magnetic RAM (STT MRAM) is a promising candidate for future embedded applications. It provides desirable memory attributes such as fast access time, low cost, high density and non-volatility. However, variations in process parameters can lead to a large number of cells to fail, severely affecting the yield of the memory array. In this paper, we provide a thorough analysis of the impact of design parameters on parametric failures due to process variations. To achieve high memory yield without incurring expensive technology modification, we developed an alternate design paradigm -circuit/architecture co-design - to take advantage of different levels of design hierarchy (circuit and architecture) to improve the yield and memory density. The technique decouples the conflicting design requirements for read stability/writability and density. Consequently, the memory cell failure probability reduces by 48% and cell area reduces by 21% with negligible performance degradation (~0.4%). Jing Jane Li, Patrick Ndai, Ashish Goel, Kaushik Roy 0001 |
ASP-DAC | 1 |
| 2009 | Variation Estimation and Compensation Technique in Scaled LTPS TFT Circuits for Low-Power Low-Cost ApplicationsabstractLow-temperature polycrystalline-silicon thin-film transistor (LTPS TFT) has emerged as one of the promising candidates for low-power low-cost applications on flexible substrates. In this paper, we propose a statistical simulation methodology to estimate parametric variations in scaled LTPS TFT due to the inherent properties [such as the number, location, and orientation of grain boundaries (GBs)] of the polycrystalline material. Our simulation technique employs the response surface method (RSM) to consider multiple process parameters which affect the performance distribution of LTPS TFT devices/circuits. Simulation results show that inherent GB variations result in multimodal delay distributions in basic logic building blocks (inv, nand, and nor) in scaled LTPS TFT technology, contrary to unimodal distributions in conventional bulk CMOS technology. We also observed that with increasing logic depth, the multimodal distribution converges to a unimodal distribution. Hence, to ensure robust and stable functionality of TFT technology under inherent process variations, we propose a multifinger (MF) design technique to improve the reliability of TFT circuits and to reduce the impact of GB-induced variations on TFT performance. Simulation results obtained from a 20-stage inverter chain show that by applying the proposed MF-based design, one can achieve 28% and 61% reductions in delay variations using two- and four-finger structures, respectively. Jing Jane Li, Kunhyuk Kang, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | Modeling of failure probability and statistical design of spin-torque transfer magnetic random access memory (STT MRAM) array for yield enhancementabstractSpin-Torque Transfer Magnetic RAM (STT MRAM) is a promising candidate for future universal memory. It combines the desirable attributes of current memory technologies such as SRAM, DRAM and flash memories. It also solves the key drawbacks of conventional MRAM technology: poor scalability and high write current. In this paper, we analyzed and modeled the failure probabilities of STT MRAM cells due to parameter variations. Based on the model, we developed an efficient simulation tool to capture the coupled electro/magnetic dynamics of spintronic device, leading to effective prediction for memory yield. We also developed a statistical optimization methodology to minimize the memory failure probability. The proposed methodology can be used at an early stage of the design cycle to enhance memory yield. Jing Jane Li, Charles Augustine, Sayeef S. Salahuddin, Kaushik Roy 0001 |
DAC | 1 |
| 2008 | An alternate design paradigm for low-power, low-cost, testable hybrid systems using scaled LTPS TFTsabstractThis article presents a holistic hybrid design methodology for low-power, low-cost, testable digital designs using low-temperature polycrystalline-silicon thin-film transistors (LTPS TFTs). An alternate scaling rule under low thermal budget (due to flexible substrate) is developed to improve the performance of TFTs in the presence of process variation. We demonstrate that LTPS TFTs can be further optimized for ultralow-power subthreshold operation with performances comparable to contemporary single-crystal silicon-on-insulator (c-Si SOI) devices after process optimization. The optimized LTPS TFTs with high current drivability and less variability can comprise a promising low-cost option to augment Si CMOS technology, opening up a plethora of new hybrid 3D applications. We illustrate one such application: IC testing. Testing of complex VLSI systems is a prime concern due to design cost of DFT circuits, area/delay overheads, and poor test confidence. To harness the benefits of TFT technology, a novel low-power, process-tolerant, generic, and reconfigurable test structure designed using LTPS TFTs is proposed to reduce the test cost, as well as to improve diagnosability and verifiability, of complex VLSI systems. Due to proper optimization of TFT devices, the proposed test structure consumes low power but operates with reasonable performance. Furthermore, the test circuits do not consume any silicon area because they can be integrated on-chip using 3D technology. Since the test architecture is reconfigurable, this eliminates the need to redesign built-in-self-test (BIST) components that may vary from one processor generation to another. We have developed test structures using 200nm TFT devices and evaluated them on designs implemented in 130nm bulk CMOS. For circuit simulations, we have developed a SPICE-compatible model for TFT devices. The BIST components designed using the test structures operate at 0.8--4.3 GHz (compared to 8.2 GHz in bulk CMOS) with low power consumption. The enhanced scan cells partially implemented in TFT (3D hybrid design) consume ∼24% less power and ∼15--20% less area of Si die compared to conventional bulk-Si design (2D planar design), with minimal delay overhead. Jing Jane Li, Aditya Bansal, Swaroop Ghosh, Kaushik Roy 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2007 | High Performance and Low Power Electronics on Flexible SubstrateabstractWe propose a design and optimization methodology for high performance and ultra low power digital applications on flexible substrate using low temperature polycrystalline silicon thin film transistor (LTPS TFT). We show that by using ultra-thin bodies and minimizing the mid-gap trap density by hydrogenation, LTPS TFTs (in 200 nm technology) can achieve higher performance than standard TFTs. We also demonstrate that it can be a promising candidate for both sub-threshold and super-threshold operation with performances comparable to contemporary bulk silicon. However, due to grain boundaries (GBs), there can be large intrinsic variations in such devices. Hence, there is a need for GB-tolerant design. Integration of proposed digital electronics in conjunction with conventional display application of LTPS TFTs on flexible substrates (system-on-panel) will open up plethora of new and interesting applications. Jing Jane Li, Kunhyuk Kang, Aditya Bansal, Kaushik Roy 0001 |
DAC | 1 |
| 2007 | Variable-latency adder (VL-adder): new arithmetic circuit design practice to overcome NBTIabstractNegative bias temperature instability (NBTI) has become a dominant reliability concern for nanoscale PMOS transistors. In this paper, we propose variable-latency adder (VL-adder) technique for NBTI tolerance. By detecting the circuit failure on-the-fly, the proposed VL-adder can automatically shift data capturing clock edge to tolerate NBTI-induced delay degradation on critical timing paths. VL-adder operates with a fixed supply voltage and clock period, avoiding the high design and manufacturing costs incurred by existing NBTI-tolerant techniques. Compared to other related lower-power adder designs, VL-adder technique always provides better energy efficiency through the whole chip lifetime with very limited performance degradation (4.6% or less). Yiran Chen 0001, Hai Li 0001, Jing Jane Li, Cheng-Kok Koh |
ISLPED | 3 |
| 2007 | A generic and reconfigurable test paradigm using Low-cost integrated Poly-Si TFTsabstractIn this work, we propose a novel low power, process tolerant, generic and reconfigurable test structure to reduce the test cost, improve diagnosability and verifiability of complex VLSI systems. The test structure contains a variety of configurable design-for-test units designed with low cost Low Temperature Polycrystalline Silicon Thin Film Transistors (LTPS TFTs) that are fabricated on a separate substrate (e.g., polymer, glass etc). The proposed test circuits do not consume any silicon area because they can be integrated on the chip using 3-D technology. This reconfigurable test paradigm eliminates the need to re-design the BIST components that may vary from one processor generation to another. Jing Jane Li, Swaroop Ghosh, Kaushik Roy 0001 |
ITC | 1 |