EDBT 2026 Demo / reviewers in the wild / expert
Po-Kai Hsu
dblp:199/5514
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-7518-9472ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Proxima: Near-Storage Acceleration for Graph-Based Approximate Nearest Neighbor Search in 3D NANDabstractApproximate nearest neighbor search (ANNS) plays an indispensable role in a wide variety of applications, including recommendation systems, information retrieval, and semantic search. Among the cutting-edge ANNS algorithms, graph-based approaches provide superior accuracy and scalability on massive datasets. However, the best-performing graph-based ANNS solutions incur tens of hundreds of memory footprints as well as costly distance computation, thus hindering their efficient deployment at scale. The 3D NAND flash is emerging as a promising device for data-intensive applications due to its high density and nonvolatility. In this work, we present the near-storage processing (NSP)-based ANNS solution Proxima to accelerate graph-based ANNS with algorithm-hardware co-design in 3D NAND flash. Proxima significantly reduces the complexity of graph search by leveraging the distance approximation and early termination. On top of the algorithmic enhancement, we implement the Proxima search algorithm in 3D NAND flash using the heterogeneous integration technique. To maximize 3D NAND’s bandwidth utilization, we present a customized dataflow and optimized data allocation scheme. Our evaluation results show that, compared to graph ANNS on CPU and GPU, Proxima achieves a magnitude improvement in throughput or energy efficiency. Proxima yields 7× to 13× speedup over existing ASIC designs. Furthermore, Proxima achieves a good balance between accuracy, efficiency, and storage density compared to previous NSP-based accelerators. Po-Kai Hsu, Jaeyoung Kang 0001, Minxuan Zhou, Sumukh Pinge, Shimeng Yu, Tajana Rosing |
IEEE Trans. Computers | 3 |
| 2026 | SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive ThresholdingabstractLarge language models (LLMs), composed of Transformer decoders, have demonstrated unparalleled proficiency in understanding and generating human language. However, efficient LLM inference on resource-constraint embedded devices remains a challenge because of the sheer model size and memory-intensive operations that arise from feedforward network (FFN) and multi-head attention (MHA) layers. Existing accelerations offload LLM inference to heterogeneous computing systems comprising expensive memory and processing units. However, recent studies show that most hardware resources are not used because LLM exhibits significant sparsity during inference. The sparsity of LLMs provides a good opportunity to perform memory-efficient inference. In this work, we propose SLIM, an algorithm and hardware co-design optimized for sparse LLM serving on the edge. SLIM exploits LLM’s sparsity by only fetching activated neurons to significantly reduce data movement. To this end, the efficient inference algorithm based on adaptive thresholding is proposed to support runtime configurable sparsity at the cost of negligible accuracy loss. Then, we present the SLIM heterogeneous hardware architecture that combines the best of both near-storage processing (NSP) and processing-in-memory (PIM). SLIM stores FFN weights in high-density 3D NAND and computes FFN layers in NSP units, alleviating high memory requirements caused by FFN weights. The memory-intensive MHA with low arithmetic density is processed in the PIM module. By leveraging the inherent sparsity observed in LLM operations and integrating NSP with PIM techniques within SSDs, SLIM significantly reduces memory footprint, data movement, and energy consumption. Meanwhile, we present the software support for integrating design into existing SSD system. Our comprehensive analysis and system-level optimization demonstrate the effectiveness of our sparsity-tailored accelerator, offering 13-18× throughput improvements over SSD-GPU system and 9-10× better energy efficiency over DRAM-GPU system while maintaining low latency. Haein Choi, Po-Kai Hsu, Shimeng Yu, Tajana Rosing |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | FeNOMS: Enhancing Open Modification Spectral Library Search with In-Storage Processing on Ferroelectric NAND (FeNAND) FlashabstractThe rapid expansion of mass spectrometry (MS) data, now exceeding hundreds of terabytes, poses significant challenges for efficient, large-scale library search — a critical component for drug discovery. Traditional processors struggle to handle this data volume efficiently, making in-storage computing (ISP) a promising alternative. This work introduces an ISP architecture leveraging a 3D Ferroelectric NAND (FeNAND) structure, providing significantly higher density, faster speeds, and lower voltage requirements compared to traditional NAND flash. Despite its superior density, the NAND structure has not been widely utilized in ISP applications due to limited throughput associated with row-by-row reads from serially connected cells. To overcome these limitations, we integrate hyperdimensional computing (HDC), a brain-inspired paradigm that enables highly parallel processing with simple operations and strong error tolerance. By combining HDC with the proposed dual-bound approximate matching (D-BAM) distance metric, tailored to the FeNAND structure, we parallelize vector computations to enable efficient MS spectral library search, achieving 43× speedup and 21× higher energy efficiency over state-of-the-art 3D NAND methods, while maintaining comparable accuracy. Sumukh Pinge, Ashkan Moradifirouzabadi, Keming Fan, Prasanna Venkatesan Ravindran, Tanvir H. Pantha, Po-Kai Hsu, Zihan Xia 0002, Flavio Ponzina, Winston Chern, Taeyoung Song, Priyankka Gundlapudi Ravikumar, Mengkun Tian, Lance Fernandes, Hari Jayasankar, Chinsung Park, Amrit Garlapati, Kijoon Kim, Jongho Woo, Suhwan Lim, Wanki Kim, Daewon Ha, Duygu Kuzum, Shimeng Yu, Tajana Rosing, Mingu Kang |
ICCAD | 6 |
| 2025 | Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE ServingabstractAs Large Language Models (LLMs) continue to evolve, Mixture of Experts (MoE) architecture has emerged as a prevailing design for achieving state-of-the-art performance across a wide range of tasks.MoE models use sparse gating to activate only a handful of expert sub-networks per input, achieving billion-parameter capacity with inference costs akin to much smaller models.However, such models often pose challenges for hardware deployment due to the massive data volume introduced by the MoE layers.To address the challenges of serving MoE models, we propose Stratum, a system-hardware co-design approach that combines the novel memory technology Monolithic 3D-Stackable DRAM (Mono3D DRAM), near-memory processing (NMP), and GPU acceleration.The logic and Mono3D DRAM dies are connected through hybrid bonding, whereas the Mono3D DRAM stack and GPU are interconnected via silicon interposer.Mono3D DRAM offers higher internal bandwidth than HBM thanks to the dense vertical interconnect pitch enabled by its monolithic structure, which supports implementations of higher-performance near-memory processing.Furthermore, we tackle the latency differences introduced by aggressive vertical scaling of Mono3D DRAM along the 𝑧-dimension by constructing internal memory tiers and assigning data across layers based on * Equal contribution Yue Pan 0009, Zihan Xia 0002, Po-Kai Hsu, Lanxiang Hu, Hyungyo Kim, Janak Sharda, Minxuan Zhou, Nam Sung Kim, Shimeng Yu, Tajana Rosing, Mingu Kang |
MICRO | 3 |
| 2025 | Enhancing Virtualization Security Through System Call-Based Anomaly Detection in ContainersabstractIn the current era of micro-services, containerized applications face unprecedented security challenges due to shared kernels and limited isolation. This research proposes a container security framework based on monitoring system call sequences to detect anomalies in micro-service containers. We introduce a custom dataset named XXXX, which capture container captures system call sequences behavior in micro-services containers and simulated attacks. The framework includes real-time system call monitors, parsers, dashboards, and an unsupervised anomaly detection model using unsupervised learning with autoencoders to enhance the detection capability of unknown vulnerabilities. It leverages containerization benefits-simplicity, scalability, and automation. Our evaluation emphasizes false alarm rate and average detection time. Results show that the attack detection performance of most containers meets expectations, though the detection time of one subset had slightly longer detection time due to the intrinsic complexity of vulnerabilities. This work offers valuable insights for improving container security in microservice systems. Kuan-Chieh Wang, Po-Kai Hsu, Jhen-Jie Hsieh, Po-Shen Chen, Tze-Rong Jian, Kun-Hsiang Huang, Min-Te Sun, Chun-Ying Huang |
TENCON | 3 |
| 2024 | HyperGen: compact and efficient genome sketching using hyperdimensional vectorsabstractMOTIVATION: Genomic distance estimation is a critical workload since exact computation for whole-genome similarity metrics such as Average Nucleotide Identity (ANI) incurs prohibitive runtime overhead. Genome sketching is a fast and memory-efficient solution to estimate ANI similarity by distilling representative k-mers from the original sequences. In this work, we present HyperGen that improves accuracy, runtime performance, and memory efficiency for large-scale ANI estimation. Unlike existing genome sketching algorithms that convert large genome files into discrete k-mer hashes, HyperGen leverages the emerging hyperdimensional computing (HDC) to encode genomes into quasi-orthogonal vectors (Hypervector, HV) in high-dimensional space. HV is compact and can preserve more information, allowing for accurate ANI estimation while reducing required sketch sizes. In particular, the HV sketch representation in HyperGen allows efficient ANI estimation using vector multiplication, which naturally benefits from highly optimized general matrix multiply (GEMM) routines. As a result, HyperGen enables the efficient sketching and ANI estimation for massive genome collections. RESULTS: We evaluate HyperGen's sketching and database search performance using several genome datasets at various scales. HyperGen is able to achieve comparable or superior ANI estimation error and linearity compared to other sketch-based counterparts. The measurement results show that HyperGen is one of the fastest tools for both genome sketching and database search. Meanwhile, HyperGen produces memory-efficient sketch files while ensuring high ANI estimation accuracy. AVAILABILITY AND IMPLEMENTATION: A Rust implementation of HyperGen is freely available under the MIT license as an open-source software project at https://github.com/wh-xu/Hyper-Gen. The scripts to reproduce the experimental results can be accessed at https://github.com/wh-xu/experiment-hyper-gen. Po-Kai Hsu, Niema Moshiri, Shimeng Yu, Tajana Rosing |
Bioinform. | 2 |
| 2024 | A Heterogeneous Platform for 3D NAND-Based In-Memory Hyperdimensional Computing Engine for Genome Sequencing ApplicationsabstractHyperdimensional (HD) computing is a promising paradigm for large-scale genome sequencing. In prior work, we proposed a 3D NAND-based HD computing engine as an energy-efficient solution for sequencing several gigabytes or terabytes of genomic data. In this work, we introduce an improved HD computing engine for genome sequencing that leverages heterogeneous 3D integration techniques. We employ Cu-Cu hybrid bonding and CMOS under array (CuA) technologies to integrate the digital logic tier with the 3D NAND-based associative memory tier. We benchmark the performance of the proposed hardware design using a dataset of 10,575 microorganism genomes. The results indicate the robustness of the classification accuracy despite device non-idealities. Compared to conventional methods, the proposed design reduces the system-level energy consumption by$1000\times $, provided the data is not offloaded from the solid-state drive to the computing units. Po-Kai Hsu, Vaidehi Garg, Anni Lu, Shimeng Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Design of Computing-in-Memory (CIM) with Vertical Split-Gate Flash Memory for Deep Neural Network (DNN) Inference AcceleratorabstractComputing-In-Memory (CIM) using Flash memory is a potential solution to support a heavy-weight DNN inference accelerator for edge computing applications. Flash memory provides the best high-density and low-cost non-volatile memory solution to store the weights, while CIM functions of Flash memory can compute AI neural network calculations inside the memory chip. Our analysis indicates that Flash CIM can save data movements by ~85% as compared with the conventional Von-Neumann architecture. In this work, we propose a detail device and design co-optimizations to realize Flash CIM, using a novel vertical split-gate Flash device. Our device supports low-voltage (<; 1V) read at WL's and BL's, tight and tunable cell current (Icell) ranging from 150nA to 1.5uA, extremely large Icell ON/OFF ratio ~ 7 orders, small RTN noise and negligible read disturb to provide a high-performance and highly-reliable CIM solution. Hang-Ting Lue, Han-Wen Hu, Tzu-Hsuan Hsu, Po-Kai Hsu, Keh-Chung Wang, Chih-Yuan Lu |
ISCAS | 4 |
| 2017 | The VLSI Architecture of a Highly Efficient Deblocking Filter for HEVC SystemsabstractThis paper presents the VLSI architecture and hardware implementation of a highly efficient deblocking filter (DBF) for High Efficiency Video Coding systems. In order to reduce the number of data accesses and thus to enhance the timing efficiency, novel data structures and memory access schemes for image pixels are proposed. Furthermore, a novel edge-fetching order is presented to strike a balance between the processing throughput and complexity. Based on the proposed structure and access pattern, a six-stage pipelined two-line DBF engine with low-latency data access sequence is designed, aiming to achieve high processing throughput while at the same time maintaining low complexity. The detailed storage structure and data access scheme are illustrated and VLSI architecture for the DBF engine is depicted in this paper. In addition, the proposed DBF is implemented using TSMC 90-nm standard cell library. The experimental results based on postlayout estimations show that the proposed design can achieve 60 frames/s for a frame resolution of 4096 × 2048 pixels (ultra high definition resolution) assuming an operating frequency of 100 MHz. Moreover, this design occupies an area complexity of 466.5 kGE with a power consumption of 26.26 mW. In comparison with prior designs targeting similar system specification and throughput, the proposed design results in a significantly reduced area complexity. Po-Kai Hsu, Chung-An Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |