Jaeyoung Kang 0004

dblp:123/5488-4 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0006-0023-743XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 UPP: Universal Predicate Pushdown to Smart Storage
abstract
In large-scale analytics, in-storage processing (ISP) can significantly boost query performance by letting ISP engines (e.g., FPGAs) preselect only the relevant data before sending them to databases.This reduces the amount of not only data transfer between storage and host, but also database computation, facilitating faster query processing.However, existing ISP solutions cannot effectively support a wide range of modern analytical queries because they only support simple combinations of frequently used operators (e.g., =, <), particularly on fixed-length columns.As modern databases allow filter predicates to include numerous operators/functions (e.g., dateadd) compatible with diverse data formats (and their complex combinations), it becomes more challenging for existing approaches to accelerate such queries efficiently.To address the limitations, we propose a new ISP approach, called Universal Predicate Pushdown (UPP), that can accelerate modern analytical databases, leveraging hardware/software co-design for a high level of flexibility.Our core insight is that instead of programming for individual filter operators/functions, we should devise a compact instruction set architecture (ISA) tailored explicitly for predicate pushdown.The software (i.e., database) layer recognizes and compiles various general filters (called a universal predicate) to a set of UPP-compliant instructions, which are then processed efficiently by FPGA using bitwise comparisons, leveraging lightweight metadata.In our experiments with a 100 GB TPC-H dataset, UPP running on SmartSSD could speed up Spark's end-to-end query performance by 1.2×-7.9×without changing input data formats.
Ipoom Jeong, Jinghan Huang 0001, Chuxuan Hu, Dohyun Park, Jaeyoung Kang 0004, Nam Sung Kim, Yongjoo Park
ISCA5
2025 Intel ® in-Memory Analytics Accelerator: Performance Characterization and Guidelines
abstract
Improvements in CPU performance have significantly slowed due to the demise of Dennard scaling, making it increasingly challenging to cost-effectively process the exponentially growing volumes of data. As a result, there is a growing trend of offloading frequently used functions to hardware accelerators to enhance application performance while reducing expensive CPU cycle consumption. In line with this trend, Intel introduced the InMemory Analytics Accelerator (IAA) as an on-chip accelerator in its$\mathbf{4}^{\text{th}}$-generation Xeon®Scalable CPUs (Sapphire Rapids). IAA is designed to offload common big data and in-memory analytics functions from CPUs, such as CRC64, expand, extract, scan, and select, in addition to (de)compression. As IAA is integrated with CPUs as an on-chip accelerator, it can directly access the CPU's cache and memory in a cache-coherent manner, offering low latency and power consumption with reduced programming complexity compared to off-chip accelerators. In this paper, we first introduce the hardware and software architectures of IAA, highlighting its latest features. We then evaluate IAA's performance using various microbenchmarks and widely used analytics applications, including Pandas, Citus, and ClickHouse. Finally, based on our findings, we provide guidelines for effectively utilizing and optimizing IAA for analytics applications.
Jaeyoung Kang 0004, Qirong Xia, Ipoom Jeong, Yongjoo Park, Nam Sung Kim
ISPASS1
2025 NetZIP: Algorithm/Hardware Co-design of In-network Lossless Compression for Distributed Large Model Training
abstract
In distributed large model training, the long communication time required to exchange large volumes of gradients and activations among GPUs dominates the training time.To reduce the communication times, lossy or lossless compression of gradients and/or activations can be employed.However, lossy compression of gradients and activations may demand more training iterations to achieve the same model accuracy and cause convergence failure, respectively.Lossless compression, on the other hand, may not reduce the volumes of gradients and activations enough to offset the significant latency associated with compression and decompression on current platforms.To address these challenges, we propose NetZIP, an algorithm/hardware co-design for in-network lossless compression of both gradients and activations.NetZIP consists of two components.(1) NetZIP-algorithm transforms gradients and activations at the bit and value levels to help lightweight standard lossless compression achieve more compression of the gradients and activations.(2) NetZIP-accelerator integrates Net-ZIP-algorithm with a lightweight lossless compression accelerator within a NIC in a bump-in-the-wire fashion to reduce the compression/decompression latency under the resource constraints.NetZIP-algorithm compresses gradients and activations 40-63 and 43-75 percentage points more, respectively, than heavy standard lossless compression for Llama-3 70B, GPT-3 175B, and Llama-3 405B.NetZIP-accelerator, implemented within FPGA-NICs and connected to commodity servers, provides orders of magnitude lower
Jinghan Huang 0001, Hyungyo Kim, Nachuan Wang, Jaeyoung Kang 0004, Hrishi Shah, Minjia Zhang, Fan Lai 0001, Nam Sung Kim
MICRO4