EDBT 2026 Demo / reviewers in the wild / expert
Jihoon Jang 0001
dblp:03/5066-1
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-8311-1328ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PriME: PIM-Aware Efficient Compression for Memory-Bound Embedding Layers in sLLMsabstractWith the growing demand for on-device AI, increasing efforts have been directed toward deploying lightweight small-scale large language models (sLLMs) on edge and mobile devices to enhance inference performance while minimizing computational cost and latency. As the number of decoder layers in sLLMs decreases, the embedding layer constitutes a substantial portion of the model's overall parameters and memory consumption. Consequently, efficient data compression is crucial; however, existing methods, such as quantization and pruning, drastically degrade accuracy when applied to the error-sensitive embedding layers. Moreover, embedding layer computations exhibit low arithmetic intensity (operations per byte), rendering them memory-bound. This limitation necessitates a shift from conventional von Neumann architectures to processing-in-memory (PIM) architectures. To address these challenges, this paper proposes (1) XOR-based Masking Compression (XMC), a lossless compression algorithm specialized for sLLM embedding layers, and (2) PriME, which integrates XMC with PIM architecture to alleviate memory bottlenecks in embedding layers. XMC enhances zero-bit representation in 16-bit FP data using ADD and XOR masking, achieving an average compression ratio of$1.49 \times$while being implementable with a 3-cycle decompression delay. PriME enables parallel processing of compressed data within PIM, accelerating embedding computations by an average of$4.0 \times$and up to$5.26 \times$compared to GPUs, while simultaneously reducing energy consumption by over 30 %, leading to an average energy efficiency improvement of$6.29 \times$. Designed for broad applicability, PriME is compatible with various sLLMs and holds scalability for extension to multimodal small vision language models, demonstrating its versatility for efficient AI acceleration. Junghyeok Lee, Jihoon Jang 0001, Hyun Kim 0001 |
ICCD | 2 |
| 2025 | XNC: XOR and NOT-Based Lossless Compression for Optimizing Unquantized Embedding Layers in Large Language ModelsabstractAlthough 4-bit quantized small LLMs have been proposed recently, many studies have retained FP16 precision for embedding layers, as they constitute a relatively small proportion of the overall model in existing LLMs and suffer from severe accuracy degradation when quantized. However, in quantized small LLMs, the embedding layer accounts for a substantial proportion of the total model parameters, necessitating its compression. Since embedding layers are sensitive to approximation, lossless compression is more desirable than lossy compression methods such as quantization. While existing lossless compression methods efficiently compress patterns such as zeros, narrow values, or frequently occurring values, embedding layers typically lack these patterns, making effective compression more challenging. In this paper, we propose XOR and NOT-based lossless compression (XNC), which applies XOR operations between adjacent 16-bit blocks and then performs a NOT operation on the result, effectively truncating the upper and lower bits to compress the embedding layer to 9-bit without any loss. The proposed method leverages XOR and NOT operations, enabling easy hardware implementation, with only four cycles required for compression and three cycles for decompression, ensuring efficient data compression without performance degradation. As a result, the proposed compression technique achieves an average compression ratio of 1.34× for the embedding layers of small LLMs without any loss, effectively reducing the model size of 4-bit quantized LLMs by an average of 9.91%. The code is available at https://github.com/IDSL-SeoulTech/XNC. Junghyeok Lee, Jihoon Jang 0001, Hyun Kim 0001 |
ISCAS | 2 |
| 2025 | PIM-BEACON: A Benchmarking and Emulation Framework Supporting Adaptive CONfigurations in DRAM-Based Processing-in-Memory SystemsabstractThe growing demand for data storage and memory bandwidth in large-scale deep neural networks has exacerbated the data movement bottleneck in traditional von Neumann architectures. Processing-in-memory (PIM) technology addresses this challenge by integrating computation within memory, reducing data transfer overhead. Prior research has predominantly relied on in-house PIM simulators for evaluation. However, these simulators often exhibit limited versatility and slow evaluation runtimes, constraining their effectiveness for comprehensive design space exploration. We present PIM-BEACON, a highspeed trace-based PIM emulation platform that ensures high reliability, fidelity, and versatility to address these limitations. PIM-BEACON adopts a modular design by configuring the PIM controller and emulator regions separately and employs SystemVerilog for FPGA-based high-fidelity implementation. To improve evaluation runtime, we introduce an efficient BRAM management mechanism that maximizes FPGA resource utilization. Supporting a wide range of DRAM-based PIM architectures, PIM-BEACON achieves up to 133.61 × faster runtime with only a 1.57 % average performance cycle error rate. Inseong Hwang, Jihoon Jang 0001, Hyun Kim 0001 |
ISPASS | 2 |
| 2025 | HyDRASim: A Versatile and Cycle-Accurate Simulator for Hybrid DRAM PIM-CPU SystemsabstractRecent advancements in deep neural networks (DNNs) have exacerbated the data movement bottleneck inherent in traditional von Neumann architectures. Processing-in-memory (PIM) has emerged as a promising paradigm by integrating computation within memory to alleviate this challenge. However, the lack of versatile and cycle-accurate simulation frameworks significantly limits the evaluation and optimization of diverse PIM designs. Existing in-house PIM simulators tend to be narrowly tailored to specific architectures, lacking generality and extensibility across hybrid computing environments. In this paper, we introduce HyDRASim, a cycle-accurate and extensible PIM-CPU simulator designed to support a wide range of PIM architectural features. HyDRASim integrates ZSim for detailed CPU modeling with DRAMSim3 for accurate memory modeling to enable flexible simulation across CPU-only, PIM-only, and PIM-CPU hybrid configurations. Validation experiments using GEMV workloads of varying sizes demonstrate that HyDRASim achieves cycle-level fidelity, with average latency and power errors of just $5.92 \%$ and $5.61 \%$, respectively, compared to established in-house baselines. Furthermore, system-level performance evaluations with diverse DNN workloads reveal that transformer-based models achieve $2.21 \times$ greater acceleration compared to convolutional neural networks under hybrid configurations modeled with HyDRASim. These results establish HyDRASim as a reliable and powerful tool for accurately modeling, evaluating, and optimizing emerging PIM-CPU hybrid architectures, providing critical insights into the future of memory-centric system design. We open-source HyDRASim at https://github.com/IDSL-SeoulTech/HyDRASim Jihoon Jang 0001, Inseong Hwang, Hyun Kim 0001 |
MASCOTS | 1 |
| 2024 | EDeN: Enabling Low-Power CNN Inference on Edge Devices Using Prefetcher-assisted NVM SystemsabstractThe accuracy of convolutional neural networks (CNNs) has significantly improved over the years. Meanwhile, due to the high portability and usefulness of edge devices, the demand for artificial intelligence (AI) based applications on edge computing devices has been soaring recently. Accordingly, CNN inference has become one of the mainstream AI applications on edge devices. However, the continually increasing leakage power of edge devices drags down the wide deployment of CNN inference applications, as the technology node scales down. Jihoon Jang 0001, Hyokeun Lee, Hyun Kim 0001 |
ISLPED | 1 |