Keming Fan

dblp:377/2536 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-6659-9971ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
abstract
All-pairs shortest paths (APSP) remains a major bottleneck for large-scale graph analytics, as data movement with cubic complexity overwhelms the bandwidth of conventional memory hierarchies. We propose RAPID-Graph, a processing-in-memory (PIM) system co-designed across algorithm, architecture, and device levels to address this challenge. At the algorithm level, we introduce a recursion-aware partitioner that enables an exact APSP computation by decomposing graphs into vertex tiles to reduce data dependency, such that both Floyd-Warshall and Min-Plus kernels execute fully in-place within digital PIM arrays. At the architecture and device levels, we design a 2.5D PIM stack integrating two phase-change memory compute dies, a logic die, and high-bandwidth scratchpad memory within a unified advanced package. An external non-volatile storage stack stores large APSP results persistently. The design achieves both tile-level and unit-level parallel processing to sustain high throughput. On the 2.45M-node OGBN-Products dataset, RAPID-Graph is 5.8× faster and 1 186× more energy efficient than state-of-the-art GPU clusters, while exceeding prior PIM accelerators by 8.3× in speed and 104× in efficiency. It further delivers up to 42.8× speedup and 392× energy savings over an NVIDIA H100 GPU.
Keming Fan, Runyang Tian, John Hsu, Minxuan Zhou, Tajana Rosing
DATE3
2026 FHEIns: Fully Homomorphic Encryption Acceleration for Large Data Applications with In-Storage Processing
abstract
Recently, the significance of data privacy protection has been growing rapidly. Homomorphic encryption (HE) enables computation directly on ciphertexts, making it attractive for privacy-sensitive databases in cloud datacenters. Although FHE enables privacy-preserving compute, ciphertext expansion and long-latency primitives drive up memory footprint and delay, worsening compute and memory pressure for database search. In practice, encrypted databases span hundreds of gigabytes to terabytes, making the storage I/O the dominant bottleneck. However, most prior FHE accelerators optimize on-chip computation and the main memory traffic while assuming working sets fit in HBM. Therefore, in this work, we present FHEIns, an in-storage processing architecture that executes FHE kernels close to data inside the NAND flash-based solid-state drives (SSDs) to exploit the internal bandwidth of the SSD. FHEIns achieves up to 24.7× and 2.67× speedup compared to the state-of-the-art FHE ASIC accelerators on trending FHE-based database benchmarks.
Xuan Wang 0040, Keming Fan, Augusto Vega, Minxuan Zhou, Tajana Rosing
DATE3
2026 NOVA-PIM: Noise-Aware Hyperdimensional Processing in Memory with Optimized Vector Allocation and Minimal ADCs
abstract
Hyperdimensional computing (HDC) is an emerging brain-inspired paradigm that enables highly efficient and robust inference and learning. Analog processing in memory (PIM) has become a promising solution to accelerate HDC by processing lengthy hypervectors (HVs) directly in memory, thereby reducing costly data movement and leveraging massive parallelism. Despite its efficiency, analog PIM suffers from non-idealities that reduce reliability and accuracy. Although the similarity search stage in HDC is inherently error-tolerant given the high dimensionality of HVs, the encoding stage, which transforms raw input data into HVs, remains sensitive to analog noise. Moreover, encoding accounts for a dominant portion of energy consumption, creating a long-standing bottleneck that limits the overall efficiency of analog PIM-based HDC systems. To overcome this challenge, we propose a noise-aware partitioning scheme that improves HDC inference accuracy by processing a critical subset of HV dimensions digitally, while offloading most of the non-critical dimensions to analog PIM. To further synergize the PIM operations across the two consecutive stages, we eliminate the analog-to-digital converters (ADCs) overhead for encoding by employing pulse width modulation (PWM), allowing direct interfacing with the subsequent similarity search stage. The proposed system achieves a 2.6 × reduction in area, 1.5 × –10.3 × lower energy consumption, and 4.2 × –6.5 × speedup compared with state-of-the-art, while maintaining inference accuracy.
Keming Fan, Chang Eun Song, Xuan Wang 0040, Tajana Rosing, Mingu Kang
ACM Great Lakes Symposium on VLSI1
2026 HyperMetric: Efficient Hyperdimensional Computing With Metric Learning for Robust Edge Intelligence
abstract
Hyperdimensional computing (HDC) is emerging as an efficient and robust computing paradigm that has strong resilience to various types of errors. The error robustness nature of HDC makes it a good match for error-prone memory systems. However, the mechanisms behind HDCs robustness are not fully understood. In this work, we propose HyperMetric, a framework to train highly robust and hardware-friendly HDC models. We found that HDC’s error resilience is driven by Hamming distance margin between hypervectors. Based on this, we propose HyperMetric training that is based on metric learning in order to optimize for high robustness. The experiments show that HyperMetric trained HDC models deliver up to 17W larger Hamming distance margin and up to 14.3 We accelerate HyperMetric trained models using ReRAM. As compared to state-of-the-art HDC algorithms OnlineHD and HyDREA, HyperMetric ReRAM accelerator is > 20% more accurate for computing-in-memory (CIM) errors and > 10% more accurate for bit errors even in the face of variations. Furthermore, HyperMetric hardware is 35% more accurate in comparison with existing tinyHD and GENERIC accelerators in the face of 3× ReRAM resistance variance, and 20% more accurate with BER of up to 20% due to voltage scaling while keeping a good balance between area, power, and processing laten
Sean Fuhrman, Keming Fan, Sumukh Pinge, Wei-Chen Chen, Tajana Rosing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 Clo-HDnn: Continual On-Device Learning Accelerator with Hyperdimensional Computing via Progressive Search
abstract
Clo-HDnn is an on-device learning (ODL) accelerator designed for emerging continual learning (CL) tasks. Clo-HDnn integrates hyperdimensional computing (HDC) along with low-cost Kronecker HD Encoder and weight clustering feature extraction (WCFE) to optimize accuracy and efficiency. Clo-HDnn adopts gradient-free CL to efficiently update and store the learned knowledge in the form of class hypervectors. Its dual-mode operation enables bypassing costly feature ex- traction for simpler datasets, while progressive search reduces complexity by up to $61 \%$ by encoding and comparing only partial query hypervectors. Achieving 4.66 TFLOPS/W (FE) and 3.78 TOPS/W (classifier), Clo-HDnn delivers $7.77 \times$ and $4.85 \times$ higher energy efficiency compared to SOTA ODL accelerators.
Chang Eun Song, Keming Fan, Soumil Jain, Gopabandhu Hota, Haichao Yang, Leo Liu, Meng-Fan Chang, Carlos H. Diaz, Gert Cauwenberghs, Tajana Rosing, Mingu Kang
HCS3
2025 FeNOMS: Enhancing Open Modification Spectral Library Search with In-Storage Processing on Ferroelectric NAND (FeNAND) Flash
abstract
The rapid expansion of mass spectrometry (MS) data, now exceeding hundreds of terabytes, poses significant challenges for efficient, large-scale library search — a critical component for drug discovery. Traditional processors struggle to handle this data volume efficiently, making in-storage computing (ISP) a promising alternative. This work introduces an ISP architecture leveraging a 3D Ferroelectric NAND (FeNAND) structure, providing significantly higher density, faster speeds, and lower voltage requirements compared to traditional NAND flash. Despite its superior density, the NAND structure has not been widely utilized in ISP applications due to limited throughput associated with row-by-row reads from serially connected cells. To overcome these limitations, we integrate hyperdimensional computing (HDC), a brain-inspired paradigm that enables highly parallel processing with simple operations and strong error tolerance. By combining HDC with the proposed dual-bound approximate matching (D-BAM) distance metric, tailored to the FeNAND structure, we parallelize vector computations to enable efficient MS spectral library search, achieving 43× speedup and 21× higher energy efficiency over state-of-the-art 3D NAND methods, while maintaining comparable accuracy.
Sumukh Pinge, Ashkan Moradifirouzabadi, Keming Fan, Prasanna Venkatesan Ravindran, Tanvir H. Pantha, Po-Kai Hsu, Zihan Xia 0002, Flavio Ponzina, Winston Chern, Taeyoung Song, Priyankka Gundlapudi Ravikumar, Mengkun Tian, Lance Fernandes, Hari Jayasankar, Chinsung Park, Amrit Garlapati, Kijoon Kim, Jongho Woo, Suhwan Lim, Wanki Kim, Daewon Ha, Duygu Kuzum, Shimeng Yu, Tajana Rosing, Mingu Kang
ICCAD3
2024 Efficient Open Modification Spectral Library Searching in High-Dimensional Space with Multi-Level-Cell Memory
abstract
Open Modification Search (OMS) is a promising algorithm for mass spectrometry analysis that enables the discovery of modified peptides. However, OMS encounters challenges as it exponentially extends the search scope. Existing OMS accelerators either have limited parallelism or struggle to scale effectively with growing data volumes. In this work, we introduce an OMS accelerator utilizing multi-level-cell (MLC) RRAM memory to enhance storage capacity by 3x. Through in-memory computing, we achieve up to 77x faster data processing with two to three orders of magnitude better energy efficiency. Testing was done on a fabricated MLC RRAM chip. We leverage hyperdimensional computing to tolerate up to 10% memory errors while delivering massive parallelism in hardware.
Keming Fan, Wei-Chen Chen, Sumukh Pinge, H.-S. Philip Wong, Tajana Rosing
DAC1