VLDB 2026 Research / reviewers in the wild / expert
Sang Woo Jun
dblp:124/7201 · also Sang-Woo Jun
· DBLP profile ↗
31ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 8 first-author · 16 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lembas: Cost-Efficient Genome Alignment with External Memory and FPGA Acceleration
Seongyoung Kang, Se-Min Lim, Sang Woo Jun |
ISCA | 3 |
| 2025 | Bancroft: Genomics Acceleration Beyond On-Device MemoryabstractThis paper presents Bancroft, a computational genomics acceleration platform for processing datasets far exceeding accelerator memory capacity. Bancroft overcomes the capacity limitations of accelerator memory by storing genomic data in a novel genomic compression format on the host server, and decompressing it on-demand within the accelerator after fetching it over PCIe. The key innovation of Bancroft is algorithmic optimizations for reference-based compression and decompression of popular genomic file formats. The algorithm achieves high enough compression ratios to improve the effective bandwidth of the PCIe to DRAM-levels, while facilitating highly efficient hardware implementations. We evaluate a prototype implementation of Bancroft on an affordable Alveo U50 FPGA accelerator card equipped with 8 GB of High-Bandwidth Memory (HBM). Our evaluation demonstrates that Bancroft delivers speeds exceeding on-device DDR4 memory and over 30% of HBM performance, while incurring only 30% chip space overhead. This is an order of magnitude higher performance and efficiency compared to conventional PCIe-limited architectures. Using a real-world pre-alignment filtering application, Bancroft demonstrates over $7 \times$ performance improvement over conventional accelerators on scalable datasets. Se-Min Lim, Seongyoung Kang, Sang Woo Jun |
PACT | 3 |
| 2025 | Labidus: RISC-V Overlay with Streaming Asynchronous Custom InstructionsabstractHigh development complexity is one of the most critical issues preventing more widespread use of reconfigurable hardware accelerators such as FPGAs. While soft processor overlays allow productive development with high-level software tools, they suffer from low performance. We address this issue with Labidus, a parallel RISC-V soft processor overlay that addresses this performance gap while maintaining development simplicity. Labidus automatically generates custom instructions based on static analysis of user software. These custom instructions achieve high utilization through two key innovations: asynchronous semantics and sharing across four-core tiles. Our evaluation across four scientific computing applications shows that Labidus matches or exceeds the performance of even manually optimized FPGA accelerators for gigabyte-scale tasks, at a fraction of development effort. Gongjin Sun, Seongyoung Kang, Jane He, Se-Min Lim, Sang Woo Jun |
ASAP | 5 |
| 2025 | IceSpy: Reconfigurable Edge Accelerator for Scalable and Private Structural Health MonitoringabstractStructural Health Monitoring (SHM) uses pervasive sensors to monitor the health and integrity of buildings and civil infrastructure, significantly reducing maintenance costs and improving safety. Despite its potential, widespread adoption of SHM is hindered by high deployment costs and privacy concerns from building tenants. This work introduces IceSpy as one solution to both problems: reducing the cost and improving the privacy of SHM applications. IceSpy uses a programmable multi-chip systolic array of small, low-power Commercial off-the Shelf (COTS) FPGAs to implement data filtering and differential privacy. Data filtering reduces cost by reducing power-hungry wireless data transmission, leading to smaller batteries and power harvesters. Meanwhile, local differential privacy removes privacy-sensitive information before transmitting it to an untrusted server. Even with additional differential privacy measures, IceSpy achieves a 3× reduction of power consumption and a consequent reduction in battery costs due to decreased wireless communication. This low-power privacy-preserving filtering technology reduces deployment costs and mitigates privacy concerns from potential deployment sites. With many obstacles removed, we are working on deploying IceSpy in dams and commercial buildings to obtain deeper insight into the structural health of buildings and infrastructure. Alexandra Zhang Jiang, Jonathan Ta, Yuqiao Li, Zhou Li 0001, Nalini Venkatasubramanian, Monica D. Kohler, Sang Woo Jun |
FCCM | 7 |
| 2025 | MAPLE: Flexible-Precision Processing-In-Memory Architecture for Efficient On-Device ML
Jaewon Park, Quang Anh Hoang, Jonathan Ta, Shinhaeng Kang, Kyomin Sohn, Sang Woo Jun |
ACM Great Lakes Symposium on VLSI | 6 |
| 2024 | Durin: CPU-FPGA Heterogeneous Platform for Scalable Low-Dimensional Data ClusteringabstractReconfigurable hardware accelerators, known for their high performance and power efficiency, have yet to be fully leveraged for clustering low-dimensional data at realistic scales. In this work, we identify and address two hurdles for accelerating this important class of applications: the overhead of implementing indexing data structures in hardware, and the PCIe bottleneck when the data capacity spills over from accelerator memory to host storage. We overcome these hurdles for the first time with Durin, a CPU-FPGA heterogeneous system with a hardware-software codesigned index structure, which minimizes the hardware resource overhead of high-performance neighbor search. It also minimizes PCIe overhead by facilitating asynchronous acceleration of distance calculation, and block floating-point compression. We show that a desktop computer with Durin implemented on a mid-range Alveo U50 FPGA can outperform a 32-thread Xeon server by almost 20× , with an order of magnitude power efficiency improvements. Furthermore, Durin outperforms even the best-case projection of a conventional standalone accelerator design, which implements the entirety of the clustering algorithm in the FPGA and its High Bandwidth Memory (HBM), by 2×. Se-Min Lim, Esmerald Aliaj, Sang Woo Jun |
IEEE Big Data | 3 |
| 2024 | Sting: Near-storage accelerator framework for scalable triangle counting and beyondabstractOne of the most critical limitations to scalable graph mining is memory capacity, as graphs of interest continue to grow while the rate of DRAM scaling diminishes. While high-performance NVMe storage is cheap and dense enough to better support larger graphs, the relative performance limitations of secondary storage force a cost-performance trade-off. We present STING, which uses an asynchronous callback function to provide a general interface to in-storage graphs while allowing transparent near-storage acceleration. Using triangle counting, we show with transparent filtering and sorting acceleration, STING can improve state-of-the-art by 3x for cost and power efficiency. Seongyoung Kang, Sang Woo Jun |
DAC | 2 |
| 2024 | FlexForge: Efficient Reconfigurable Cloud Acceleration via Peripheral Resource DisaggregationabstractReconfigurable hardware acceleration in the cloud using Field-Programmable Gate Arrays (FPGAs) is an increasingly popular solution for scaling performance and cost-effectiveness. For efficient utilization of FPGA resources, cloud platforms typically support elastic FPGA resource allocation. However, FPGAs are usually allocated in a homogeneous unit consisting of logic, memory, and PCIe bandwidth. Because user kernels have a wide and varying combination of resource requirements, this can result in high internal fragmentation and underutilization of each resource. To address this issue, we present FlexForge, a platform facilitating high-performance disaggregation of peripheral resources over a network of potentially untrusted FPGAs, aided by a secondary inter-FPGA network. Evaluated on a mix of representative accelerator applications deployed on a prototype cluster, FlexForge improves the overall performance of the cloud by up to 70 % and 20 % on average across all possible combinations without significant additional hardware resource requirements. Se-Min Lim, Sang Woo Jun |
DATE | 2 |
| 2024 | Morbius: Platform-Adaptive Hardware Accelerator for Scalable Sequence Motif DiscoveryabstractMotif finding is one of the fundamental tools of computational biology. Unfortunately, the benefits of high-performance, low-power acceleration have not been available to this application due to the limited parallelism available within efficient classes of algorithms, such as probabilistic Gibbs sampling. In this work, we present Morbius, which demonstrates the benefits of reconfigurable application-specific hardware acceleration using FPGAs as a solution to the execution time and cost overhead of motif finding. Morbius employs a novel, hardware-optimized data structure called base pair matrix to minimize off-chip data movement, and implements a small number of deep hardware pipelines to achieve high sequential performance. Furthermore, we develop performance and chip space prediction models based on microarchitectural parameters of the accelerator, to facilitate optimal performance on a wide range of accelerators potential users may already have. We compare Morbius on a wide range of FPGA platforms spanning low-profile $250 M.2 FPGA cards to Amazon F1, and demonstrate up to orders of magnitude performance improvements compared to costly server machines, and even higher cost and power efficiency. Se-Min Lim, Esmerald Aliaj, Sang Woo Jun |
e-Science | 3 |
| 2024 | Xyloni: Very Low Power Neural Network Accelerator for Intermittent Remote Visual Detection of Wildfire and BeyondabstractWildfires are one of the most catastrophic natural disasters, causing increasingly severe ecological and economic damage. Early response is critically important for wildfire management, but also difficult due to the wide geographical area to monitor, often far from utility infrastructures such as stable power and high-bandwidth network. In this work, we present Xyloni, a very low-cost, low-power neural network accelerator for sensor nodes, which improves the cost-effectiveness and scalability of real-time wildfire detection by drastically reducing wireless data transmission and overall power consumption. Xyloni uses low-power flash and FeRAM memories to store a hardware co-optimized Neural Network model for fire and smoke detection, as well as intermediate activations during inference. It also time-shares a Field-Programmable Gate Array across different model layers for power-efficient computation. The detection model prevents benign images from consuming network traffic, allowing the use of low-bandwidth, low-power network fabrics such as a LoRa mesh network with enough range for the necessary geographical coverage. Compared to a wide range of edge and sensor platforms capable of real-time data collection, Xyloni demonstrated an order of magnitude reduction in power consumption for the network transmission reduction task, leading to a corresponding reduction in battery and deployment cost. Jeffrey Chen, Sang Woo Jun, Aditi Mundra, Jonathan Ta |
ISLPED | 2 |
| 2024 | Eciton: Very Low-power Recurrent Neural Network Accelerator for Real-time Inference at the EdgeabstractThis article presents Eciton, a very low-power recurrent neural network accelerator for time series data within low-power edge sensor nodes, achieving real-time inference with a power consumption of 17 mW under load. Eciton reduces memory and chip resource requirements via 8-bit quantization and hard sigmoid activation, allowing the accelerator as well as the recurrent neural network model parameters to fit in a low-cost, low-power Lattice iCE40 UP5K FPGA. We evaluate Eciton on multiple, established time-series classification applications including predictive maintenance of mechanical systems, sound classification, and intrusion detection for IoT nodes. Binary and multi-class classification edge models are explored, demonstrating that Eciton can adapt to a variety of deployable environments and remote use cases. Eciton demonstrates real-time processing at a very low power consumption with minimal loss of accuracy on multiple inference scenarios with differing characteristics, while achieving competitive power efficiency against the state-of-the-art of similar scale. We show that the addition of this accelerator actually reduces the power budget of the sensor node by reducing power-hungry wireless transmission. The resulting power budget of the sensor node is small enough to be powered by a power harvester, potentially allowing it to run indefinitely without a battery or periodic maintenance. Jeffrey Chen, Sang Woo Jun, Sehwan Hong, Warrick He, Jinyeong Moon |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2023 | Barad-dur: Near-Storage Accelerator for Training Large Graph Neural NetworksabstractGraph Neural Networks (GNNs) enable effective machine learning on graph-structured data, but their performance and scalability are often limited by the irregular structure and large size of real-world graphs. Conventional accelerators such as GPUs suffer a sharp performance loss when the target graph exceeds its fast random-access memory, because the overhead of partitioning graphs into memory-size chunks quickly dominates performance. We address this issue with Barad-dur, a near-storage GNN accelerator which processes the entire graph in cost-efficient solid-state storage (SSD) to remove the partitioning overhead. Barad-dur minimizes the performance impact of using relatively slow SSDs instead of DRAM, via storage-optimized graph encodings, hardware-accelerated access reordering, and the higher internal storage bandwidth available to near-storage accelerators. We demonstrate that Barad-dur can better maintain computational throughput on larger graphs compared to in-memory CPU and GPU systems. On even moderately large graphs such as the Twitter graph, Barad-dur achieves an order of magnitude higher performance compared to a highly optimized Pytorch Geometric implementation running on the NVIDIA V100 GPU, while requiring a fraction of capital cost and energy. Jiyoung An, Esmerald Aliaj, Sang Woo Jun |
PACT | 3 |
| 2023 | FarSlayer: Turnkey Acceleration of Legacy Software on Commodity FPGA CardsabstractApplication-specific hardware acceleration of computation-intensive kernels can often provide significant performance and power efficiency improvements over general-purpose software, but it is difficult and costly to incorporate them into existing software systems. Designing hardware accelerators and modifying legacy software to incorporate them are already complex tasks. Furthermore, identifying a suitable kernel for acceleration is complicated by the PCIe-attached architecture of commodity FPGA cards, meaning bandwidth and latency overhead must be considered for kernel selection. For example, small blocking kernels may not benefit from acceleration due to PCIe latency. As a remedy, we present FarSlayer, a high-level source-to-source compiler for end-to-end acceleration of legacy software. FarSlayer analyzes existing software code and emits an accelerated version of it, where the kernel is automatically selected considering data movement over PCIe. Specifically, FarSlayer identifies kernels which can be called asynchronously to hide the PCIe latency, while also having a high operational intensity for low bandwidth requirements. The entire process is automatic, meaning the programmer does not necessarily need to understand the existing code, or reason about hardware development. We demonstrate FarSlayer on multiple existing scientific computing software systems, and demonstrate it can automatically achieve significant performance improvements. Esmerald Aliaj, Alberto Krone-Martins, Joshua Garcia, Sang Woo Jun |
ASAP | 4 |
| 2023 | PreCog: Near-Storage Accelerator for Heterogeneous CNN InferenceabstractComputational Storage Devices (CSD) with near-storage acceleration is gaining popularity for data-intensive applications, by moving power-efficient hardware acceleration closer to data. However, because the power-constrained near-storage accelerator is often not powerful enough by itself to handle all computation requirements of an application, it must intelligently cooperate with other computation and acceleration units in the system. In this work, we explore how a near-storage accelerator can best fit into a larger computer system in the context of CNN inference. We demonstrate that an attractive configuration is using the near-storage accelerator to offload only the first convolution and pooling layers, where the accelerator can achieve almost an order of magnitude better performance compared to a general convolution accelerator. Targeting only the first layer allows some FPGA-specific floating-point computation optimizations such as pre-determine the range of output exponents, performing a costly floating point normalization task only once, as well as pack more input into the datapath to mitigate the performance impact of wide strides of the convolution filter. We package these optimizations into a flexible library we call Static Range Float (SRFloat), and construct a prototype system called PreCog. We evaluate PreCog implemented on the Samsung SmartSSD platform, and demonstrate over$\mathbf{3}\times$performance efficiency compared to conventional convolution accelerators on prominent CNN models without introducing a communications bottleneck. Jiyoung An, Esmerald Aliaj, Sang Woo Jun |
ASAP | 3 |
| 2022 | BunchBloomer: Cost-Effective Bloom Filter Accelerator for Genomics ApplicationsabstractBloom filters are a very important tool for many applications including genomics, where they are used as a compact data structure for counting k-mers, represent de Bruijn graphs, and more. Due to their random-access nature coupled with the large size required for genomics, Bloom filters for genomics can easily become bound by the random access performance of off-chip memory. This is especially true for accelerators such as FPGAs and GPUs, which can easily remove the computation overhead of the multiple hash functions. As a result, Bloom filter accelerators have typically focused either on small filters which can fit in fast on-chip memory, or require fast off-chip memory fabric such as Hybrid Memory Cubes. In this work, we present BunchBloomer, which improves the cost-effectiveness of FPGA Bloom filter accelerators by making better use of cheaper, lower-power DDR memory. BunchBloomer uses a multi-layer radix sorter to group table updates into bursts directed to the same 8 KiB memory region, which can be efficiently cached in on-chip memory. A single BunchBloomer device outperforms a costly 12-core server by over 2×, demonstrating an order of magnitude better power efficiency. It even achieves better power efficiency compared to published FPGA Bloom filter accelerators equipped with Hybrid Memory Cubes. Seongyoung Kang, Tarun Sai Ganesh Nerella, Shashank Uppoor, Sang Woo Jun |
FPL | 4 |
| 2022 | BurstZ+: Eliminating The Communication Bottleneck of Scientific Computing Accelerators via Accelerated CompressionabstractWe present BurstZ+, an accelerator platform that eliminates the communication bottleneck between PCIe-attached scientific computing accelerators and their host servers, via hardware-optimized compression. While accelerators such as GPUs and FPGAs provide enormous computing capabilities, their effectiveness quickly deteriorates once data is larger than its on-board memory capacity, and performance becomes limited by the communication bandwidth of moving data between the host memory and accelerator. Compression has not been very useful in solving this issue due to performance and efficiency issues of compressing floating point numbers, which scientific data often consists of. BurstZ+ is an FPGA-based prototype accelerator platform which addresses the bandwidth issue via a class of novel hardware-optimized floating point compression algorithm called ZFP-V. We demonstrate that BurstZ+ can completely remove the host-side communication bottleneck for accelerators, using multiple stencil kernels with a wide range of operational intensities. Evaluated against hand-optimized implementations of kernel accelerators of the same architecture, our single-pipeline BurstZ+ prototype outperforms an accelerator without compression by almost 4×, and even an accelerator with enough memory for the entire dataset by over 2×. Furthermore, the projected performance of BurstZ+ on a future, faster FPGA scales to almost 7× that of the same accelerator without compression, whose performance is still limited by the PCIe bandwidth. Gongjin Sun, Seongyoung Kang, Sang Woo Jun |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2021 | Eciton: Very Low-Power LSTM Neural Network Accelerator for Predictive Maintenance at the EdgeabstractThis paper presents Eciton, a very low-power LSTM neural network accelerator for low-power edge sensor nodes, demonstrating real-time processing on predictive maintenance applications with a power consumption of 17 mW under load. Eciton reduces memory and chip resource requirements via 8-bit quantization and hard sigmoid activation, allowing the accelerator as well as the LSTM model parameters to fit in a lowcost, low-power Lattice iCE40 UP5K FPGA. Eciton demonstrates real-time processing at a very low power consumption with minimal loss of accuracy on two predictive maintenance scenarios with differing characteristics, while achieving competitive power efficiency against the state-of-the-art of similar scale. We also show that the addition of this accelerator actually reduces the power budget of the sensor node by reducing power-hungry wireless transmission.The resulting power budget of the sensor node is small enough to be powered by a power harvester, potentially allowing it to run indefinitely without a battery or periodic maintenance. Jeffrey Chen, Sehwan Hong, Warrick He, Jinyeong Moon, Sang Woo Jun |
FPL | 5 |
| 2021 | : Near-Storage Accelerator for High-Performance Log AnalyticsabstractThis paper presents, a log analytics platform with near-storage accelerators for high-performance, cost- and power-efficient unstructured log processing. offloads log analytics queries to an efficient near-storage FPGA implementation of a token querying engine, which can take advantage of the high internal bandwidth of storage devices within the available chip resource limitations. This engine is flexible enough to handle complex queries including template search based on user-defined tree-based template libraries, as well as concurrent execution of multiple queries. also uses a log-optimized version of a simple, high-throughput compression algorithm in order to further improve the effective bandwidth of backing storage. Seongyoung Kang, Jiyoung An, Jinpyo Kim, Sang Woo Jun |
MICRO | 4 |
| 2020 | FPGA-Accelerated Time Series Mining on Low-Power IoT DevicesabstractWe present a case for FPGA-accelerated edge processing for low-power Internet-of-Things (IoT) devices, using time series similarity search as a driving application. As the data collection capabilities of low-power IoT device increase, the primary constraint on their capacity is becoming the resource requirements of wirelessly transferring collected data to a central repository. This work presents a solution to this limitation by augmenting the IoT device with a inexpensive, power-efficient FPGA accelerator, which can perform fairly complex edge mining operations and drastically reduce the wireless data transfer requirements. This approach reduces the total power consumption of the device despite the added FPGA component, while also reducing the computation requirements at the central server. We use the Dynamic Time Warping (DTW) algorithm as an example workload. Using a low-cost Lattice iCE40 UltraPlus FPGA, we demonstrate that the FPGA-augmented mining algorithm can both support significantly higher data collection rate while improving the computation power efficiency of the entire deployment by an order of magnitude. Seongyoung Kang, Jinyeong Moon, Sang Woo Jun |
ASAP | 3 |
| 2020 | BurstZ: a bandwidth-efficient scientific computing accelerator platform for large-scale dataabstractWe present BurstZ, a bandwidth-efficient accelerator platform for scientific computing. While accelerators such as GPUs and FPGAs provide enormous computing capabilities, their effectiveness quickly deteriorates once the working set becomes larger than the on-board memory capacity, causing the performance to become bottlenecked either by the communication bandwidth between the host and the accelerator. Compression has not been very useful in solving this issue due to the difficulty of efficiently compressing floating point numbers, which scientific data often consists of. Most compression algorithms are either ineffective with floating point numbers, or has a high performance overhead. Gongjin Sun, Seongyoung Kang, Sang Woo Jun |
ICS | 3 |
| 2019 | Wire-Speed Multirate Accelerator for Aggregation Operations on Sorted DataabstractWe present an accelerator architecture for wire-speed aggregation of sorted key-value pairs on a wide datapath, in a bump-in-the-wire fashion. The presented accelerator is capable of maintaining wire-speed regardless of data distribution, even when (1) the aggregation function has multiple-cycle latency, and (2) the input stream is multi-rate, i.e., multiple elements arrive every cycle. To the best of our knowledge, it is the first accelerator architecture that satisfies both properties. Sang Woo Jun, Arvind Arvind |
FCCM | 1 |
| 2018 | GraFBoost: Using Accelerated Flash Storage for External Graph AnalyticsabstractWe describe GraFBoost, a flash-based architecture with hardware acceleration for external analytics of multi-terabyte graphs. We compare the performance of GraFBoost with 1 GB of DRAM against various state-of-the-art graph analytics software including FlashGraph, running on a 32-thread Xeon server with 128 GB of DRAM. We demonstrate that despite the relatively small amount of DRAM, GraFBoost achieves high performance with very large graphs no other system can handle, and rivals the performance of the fastest software platforms on sizes of graphs that existing platforms can handle. Unlike in-memory and semi-external systems, GraFBoost uses a constant amount of memory for all problems, and its performance decreases very slowly as graph sizes increase, allowing GraFBoost to scale to much larger problems than possible with existing systems while using much less resources on a single-node system. The key component of GraFBoost is the sort-reduce accelerator, which implements a novel method to sequentialize fine-grained random accesses to flash storage. The sort-reduce accelerator logs random update requests and then uses hardware-accelerated external sorting with interleaved reduction functions. GraFBoost also stores newly updated vertex values generated in each superstep of the algorithm lazily with the old vertex values to further reduce I/O traffic. We evaluate the performance of GraFBoost for PageRank, breadth-first search and betweenness centrality on our FPGA-based prototype (Xilinx VC707 with 1 GB DRAM and 1 TB flash) and compare it to other graph processing systems including a pure software implementation of GrapFBoost. Sang Woo Jun, Andy Wright, Sizhuo Zhang, Shuotao Xu, Arvind 0001 |
ISCA | 1 |
| 2017 | Terabyte Sort on FPGA-Accelerated Flash StorageabstractSorting is one of the most fundamental and usefulapplications in computer science, and continues to be animportant tool in analyzing large datasets. An important andchallenging subclass of sorting problems involves sorting terabytescale datasets with hundreds of billions of records. Theconventional method of sorting such large amounts of datais to distribute the data and computation over a cluster ofmachines. Such solutions can be fast but are often expensiveand power-hungry. In this paper, we propose a solution basedon flash storage connected to a collection of FPGA-based sortingaccelerators that perform large-scale merge-sort in storage. Theaccelerators include highly efficient sorting networks and mergetrees that use bitonic sorting to emit multiple sorted valuesevery cycle. We show that by appropriate use of acceleratorswe can remove all the computation bottlenecks so that the endto-endsorting performance is limited only by the flash storagebandwidth. We demonstrate that our flash-based system matchesthe performance of existing distributed-cluster solutions of muchlarger scale. More importantly, our prototype is able to showalmost twice the power efficiency compared to the existingJoulesort record holder. An optimized system with less wastefulcomponents is projected to be four times more efficient comparedto the current record holder, sorting over 200,000 records perjoule of energy. Sang Woo Jun, Shuotao Xu, Arvind 0001 |
FCCM | 1 |
| 2016 | minFlash: A minimalistic clustered flash array
Sang Woo Jun, Sungjin Lee 0001, Jamey Hicks, Arvind 0001 |
DATE | 2 |
| 2016 | Application-Managed Flash
Sungjin Lee 0001, Sang Woo Jun, Shuotao Xu, Jihong Kim 0001, Arvind 0001 |
FAST | 3 |
| 2016 | BlueCache: A Scalable Distributed Flash-based Key-value StoreabstractA key-value store (KVS), such as memcached and Redis, is widely used as a caching layer to augment the slower persistent backend storage in data centers. DRAM-based KVS provides fast key-value access, but its scalability is limited by the cost, power and space needed by the machine cluster to support a large amount of DRAM. This paper offers a 10X to 100X cheaper solution based on flash storage and hardware accelerators. In BlueCache key-value pairs are stored in flash storage and all KVS operations, including the flash controller are directly implemented in hardware. Furthermore, BlueCache includes a fast interconnect between flash controllers to provide a scalable solution. We show that BlueCache has 4.18X higher throughput and consumes 25X less power than a flash-backed KVS software implementation on x86 servers. We further show that BlueCache can outperform DRAM-based KVS when the latter has more than 7.4% misses for a read-intensive aplication. BlueCache is an attractive solution for both rack-level appliances and data-center-scale key-value cache. Shuotao Xu, Sungjin Lee 0001, Sang Woo Jun, Jamey Hicks, Arvind 0001 |
Proc. VLDB Endow. | 3 |
| 2016 | BlueDBM: Distributed Flash Storage for Big Data AnalyticsabstractComplex data queries, because of their need for random accesses, have proven to be slow unless all the data can be accommodated in DRAM. There are many domains, such as genomics, geological data, and daily Twitter feeds, where the datasets of interest are 5TB to 20TB. For such a dataset, one would need a cluster with 100 servers, each with 128GB to 256GB of DRAM, to accommodate all the data in DRAM. On the other hand, such datasets could be stored easily in the flash memory of a rack-sized cluster. Flash storage has much better random access performance than hard disks, which makes it desirable for analytics workloads. However, currently available off-the-shelf flash storage packaged as SSDs does not make effective use of flash storage because it incurs a great amount of additional overhead during flash device management and network access. In this article, we present BlueDBM, a new system architecture that has flash-based storage with in-store processing capability and a low-latency high-throughput intercontroller network between storage devices. We show that BlueDBM outperforms a flash-based system without these features by a factor of 10 for some important applications. While the performance of a DRAM-centric system falls sharply even if only 5% to 10% of the references are to secondary storage, this sharp performance degradation is not an issue in BlueDBM. BlueDBM presents an attractive point in the cost/performance tradeoff for Big Data analytics. Sang Woo Jun, Sungjin Lee 0001, Jamey Hicks, John Ankcorn, Myron King, Shuotao Xu, Arvind 0001 |
ACM Trans. Comput. Syst. | 1 |
| 2015 | A transport-layer network for distributed FPGA platformsabstractWe present a transport-layer network that aids developers in building safe, high-performance distributed FPGA applications. Two essential features of such a network are virtual channels and end-to-end flow control. Our network implements these features, taking advantage of the low error characteristic of a rack level FPGA network to implement a low overhead credit based end-to-end flow control. Our design has many parameters in the source code which can be set at the time of FPGA synthesis, to provide flexibility in setting buffer size and flow control credits to make best use of scarce on-chip memory resources and match the traffic pattern of a virtual channel. Our prototype cluster, which is composed of 20 Xilinx VC707 boards, each with 4 20Gb/s serial links, achieves effective bandwidth of 85% of the maximum physical bandwidth, and a latency of 0.5us per hop. User feedback suggest that these features make distributed application development significantly easier. Sang Woo Jun, Shuotao Xu, Arvind 0001 |
FPL | 1 |
| 2015 | BlueDBM: an appliance for big data analyticsabstractComplex data queries, because of their need for random accesses, have proven to be slow unless all the data can be accommodated in DRAM. There are many domains, such as genomics, geological data and daily twitter feeds where the datasets of interest are 5TB to 20 TB. For such a dataset, one would need a cluster with 100 servers, each with 128GB to 256GBs of DRAM, to accommodate all the data in DRAM. On the other hand, such datasets could be stored easily in the flash memory of a rack-sized cluster. Flash storage has much better random access performance than hard disks, which makes it desirable for analytics workloads. In this paper we present BlueDBM, a new system architecture which has flash-based storage with in-store processing capability and a low-latency high-throughput inter-controller network. We show that BlueDBM outperforms a flash-based system without these features by a factor of 10 for some important applications. While the performance of a ram-cloud system falls sharply even if only 5%~10% of the references are to the secondary storage, this sharp performance degradation is not an issue in BlueDBM. BlueDBM presents an attractive point in the cost-performance trade-off for Big Data analytics. Sang Woo Jun, Sungjin Lee 0001, Jamey Hicks, John Ankcorn, Myron King, Shuotao Xu, Arvind 0001 |
ISCA | 1 |
| 2014 | Scalable multi-access flash store for big data analyticsabstractFor many "Big Data" applications, the limiting factor in performance is often the transportation of large amount of data from hard disks to where it can be processed, i.e. DRAM. In this paper we examine an architecture for a scalable distributed flash store which aims to overcome this limitation in two ways. First, the architecture provides a high-performance, high-capacity, scalable random-access storage. It achieves high-throughput by sharing large numbers of flash chips across a low-latency, chip-to-chip backplane network managed by the flash controllers. The additional latency for remote data access via this network is negligible as compared to flash access time. Second, it permits some computation near the data via a FPGA-based programmable flash controller. The controller is located in the datapath between the storage and the host, and provides hardware acceleration for applications without any additional latency. We have constructed a small-scale prototype whose network bandwidth scales directly with the number of nodes, and where average latency for user software to access flash store is less than 70mus, including 3.5mus of network overhead. Sang Woo Jun, Kermin Fleming, Arvind 0001 |
FPGA | 1 |
| 2012 | ZIP-IO: Architecture for application-specific compression of Big DataabstractWe have entered the “Big Data” age: scaling of networks and sensors has led to exponentially increasing amounts of data. Compression is an effective way to deal with many of these large data sets, and application-specific compression algorithms have become popular in problems with large working sets. Unfortunately, these compression algorithms are often computationally difficult and can result in application-level slow-down when implemented in software. To address this issue, we investigate ZIP-IO, a framework for FPGA-accelerated compression. Using this system we demonstrate that an unmodified industrial software workload can be accelerated 3x while simultaneously achieving more than 1000x compression in its data set. Sang Woo Jun, Kermin Fleming, Michael Adler, Joel S. Emer |
FPT | 1 |