VLDB 2026 Research / reviewers in the wild / expert
Jaeyoung Kang 0001
dblp:123/5488-1
· DBLP profile ↗
21ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-1048-1285ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 7 first-author · 20 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Proxima: Near-Storage Acceleration for Graph-Based Approximate Nearest Neighbor Search in 3D NANDabstractApproximate nearest neighbor search (ANNS) plays an indispensable role in a wide variety of applications, including recommendation systems, information retrieval, and semantic search. Among the cutting-edge ANNS algorithms, graph-based approaches provide superior accuracy and scalability on massive datasets. However, the best-performing graph-based ANNS solutions incur tens of hundreds of memory footprints as well as costly distance computation, thus hindering their efficient deployment at scale. The 3D NAND flash is emerging as a promising device for data-intensive applications due to its high density and nonvolatility. In this work, we present the near-storage processing (NSP)-based ANNS solution Proxima to accelerate graph-based ANNS with algorithm-hardware co-design in 3D NAND flash. Proxima significantly reduces the complexity of graph search by leveraging the distance approximation and early termination. On top of the algorithmic enhancement, we implement the Proxima search algorithm in 3D NAND flash using the heterogeneous integration technique. To maximize 3D NAND’s bandwidth utilization, we present a customized dataflow and optimized data allocation scheme. Our evaluation results show that, compared to graph ANNS on CPU and GPU, Proxima achieves a magnitude improvement in throughput or energy efficiency. Proxima yields 7× to 13× speedup over existing ASIC designs. Furthermore, Proxima achieves a good balance between accuracy, efficiency, and storage density compared to previous NSP-based accelerators. Po-Kai Hsu, Jaeyoung Kang 0001, Minxuan Zhou, Sumukh Pinge, Shimeng Yu, Tajana Rosing |
IEEE Trans. Computers | 4 |
| 2025 | RelHDx: Hyperdimensional Computing for Learning on Graphs With FeFET AccelerationabstractGraph neural networks (GNNs) are a powerful machine learning (ML) method to analyze graph data. The training of GNN has compute and memory-intensive phases along with irregular data movements, which makes in-memory acceleration challenging. We present a hyperdimensional computing (HDC)-based graph ML framework called RelHDx that aggregates node features and graph structure, along with representing node and edge information in high-dimensional space. RelHDx enables single-pass training and inference with simple arithmetic operations, resulting in the efficient design of graph-based ML tasks: node classification and link prediction. We accelerate RelHDx using scalable processing in-memory (PIM) architecture based on emerging ferroelectric FET (FeFET) technology. Our accelerator uses a data allocation optimization and operation scheduler to address the irregularity of the graph and maximize the performance. Evaluation results show that RelHDx offers comparable accuracy to popular GNN-based algorithms while achieving up to$63.8\boldsymbol{\times}$faster speed on GPU. Our FeFET-based accelerator, RelHDx-PIM, is$32\boldsymbol{\times}$faster for node classification, while for link prediction it is$65.4\boldsymbol{\times}$faster than when running on GPU. Furthermore, RelHDx-PIM improves energy efficiency by four orders of magnitude over GPU. Compared to the state-of-the-art in-memory processing-based GNN accelerator, PIM-GCN[1], RelHDx-PIM is$10\boldsymbol{\times}$faster and$986\boldsymbol{\times}$more energy-efficient on average. Jaeyoung Kang 0001, Minxuan Zhou, Tajana Rosing |
IEEE Trans. Computers | 1 |
| 2024 | HygHD: Hyperdimensional Hypergraph LearningabstractHypergraphs can model real-world data that has higher-order relationships. Graph neural network (GNN)-based solutions emerged as a hypergraph learning solution, but they face non-uniform memory accesses and accompany memory-intensive and compute-intensive operations, making the acceleration with near-data processing challenging. We propose a hyperdimensional computing (HDC)-based hypergraph learning framework called HygHD, which consists of highly parallelizable and lightweight HDC operations. HygHD accelerates both the training and inference on ferroelectric field-effect transistor (FeFET)-based processing-in-memory (PIM) hardware. Furthermore, we devise a hardware-friendly block-level concatenation and fine-grained block-level scheduler for high efficiency. Our evaluation results show that HygHD offers comparable accuracy to existing GNN-based solutions. Also, HygHD on GPU is up to 443× (7.67×) faster and 142× (2.78×) more energy efficient in training (inference) than the fastest GNN-based approach [1] on GPU. The HygHD accelerator further accelerates the HygHD algorithm, providing an average speedup of 40.0× (3.41×) on training (inference) compared to the HygHD GPU implementation. Jaeyoung Kang 0001, Youhak Lee, Minxuan Zhou, Tajana Rosing |
DATE | 1 |
| 2024 | SpecHD: Hyperdimensional Computing Framework for FPGA-Based Mass Spectrometry ClusteringabstractMass spectrometry-based proteomics is a key enabler for personalized healthcare, providing a deep dive into the complex protein compositions of biological systems. This technology has vast applications in biotechnology and biomedicine but faces significant computational bottlenecks. Current methodologies often require multiple hours or even days to process extensive datasets, particularly in the domain of spectral clustering. To tackle these inefficiencies, we introduce SpecHD, a hyperdimensional computing (HDC) framework supplemented by an FPGA-accelerated architecture with integrated near-storage preprocessing. Utilizing streamlined binary operations in an HDC environment, SpecHD capitalizes on the low-latency and parallel capabilities of FPGAs. This approach markedly improves clustering speed and efficiency, serving as a catalyst for real-time, high-throughput data analysis in future healthcare applications. Our evaluations demonstrate that SpecHD not only maintains but often surpasses existing clustering quality metrics while drastically cutting computational time. Specifically, it can cluster a large-scale human proteome dataset-comprising 25 million MS/MS spectra and 131 GB of MS data-in just 5 minutes. With energy efficiency exceeding 31x and a speedup factor that spans a range of 6x to 54x over existing state-of-the-art solutions, SpecHD emerges as a promising solution for the rapid analysis of mass spectrometry data with great implications for personalized healthcare. Sumukh Pinge, Jaeyoung Kang 0001, Niema Moshiri, Wout Bittremieux, Tajana Rosing |
DATE | 3 |
| 2024 | AttBind: Memory-Efficient Acceleration for Long-Range Attention Using Vector-Derived Symbolic BindingabstractTransformer models have achieved a number of breakthrough results in a variety of complex tasks. Transformer's promising performance originates from multi-head attention (MHA), which can model long-range sequence data dependency. Better performance has been demonstrated to be obtained by increasing the sequence length$N$. However, scaling up the sequence length is extremely challenging for memory-constrained hardware because the naive Transformer requires quadratic$O(N^{2})$complexity. In this work, we address this challenge by leveraging the binding operation in vector symbolic architecture (VSA). We propose the memory-efficient MHA algorithm to simplify the MHA computation at the cost of linear complexity. Then, we present the ASIC hardware architecture with optimized timing and dataflow to accelerate the proposed algorithm. We extensively evaluate our design across various long-range attention tasks. Our experiments show that the accuracy is competitive to state-of-the-art MHA optimization approaches with lower memory consumption and inference latency. The proposed algorithm achieves 7.8× speedup and 4.5× reduction in data movement over the naive Transformer on ASIC. Meanwhile, our design supports 8 to 16 × sequence lengths compared to existing hardware accelerators. Jaeyoung Kang 0001, Tajana Rosing |
DATE | 2 |
| 2024 | DRAM-Based Acceleration of Open Modification Search in Hyperdimensional SpaceabstractMass spectrometry, commonly used for protein identification, generates a massive number of spectra that need to be matched against a large database. In reality, most of them remain unidentified or mismatched due to unexpected post-translational modifications. Open modification search (OMS) has been proposed as a strategy to improve the identification rate by considering changes in spectra, but it expands the search space exponentially. In this work, we propose HyperOMS, an algorithm-hardware co-design for boosted OMS, to cope with the enlarged database and expanded search space. HyperOMS encodes spectral data into binary vectors and performs the efficient OMS in high-dimensional space. We accelerate the HyperOMS algorithm using a DRAM-based PIM accelerator, which combines processing-using-memory and near-memory processing technologies. In order to maximize the parallelization and efficiency of the accelerator, we optimize the data allocation and devise an approximation strategy for similarity computation. Experimental results show that the HyperOMS accelerator yields up to 3.8× speedup and 119W higher energy efficiency compared to running HyperOMS on GPU, and up to 99× speedup and 1984× higher energy efficiency over the state-of-the-art OMS tool, ANN-SoLo 1, while providing comparable search quality to competing tools. Jaeyoung Kang 0001, Wout Bittremieux, Niema Moshiri, Tajana Rosing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | FSL-HD: Accelerating Few-Shot Learning on ReRAM using Hyperdimensional ComputingabstractFew-shot learning (FSL) is a promising meta-learning paradigm that trains classification models on the fly with a few training samples. However, existing FSL classifiers are either computationally expensive, or are not accurate enough. In this work, we propose an efficient in-memory FSL classifier, FSL-HD, based on hyperdimensional computing (HDC) that achieves state-of-the-art FSL accuracy and efficiency. We devise an HDC-based FSL framework with efficient HDC encoding and search to reduce high complexity caused by the large dimensionality. Also, we design a scalable in-memory architecture to accelerate FSL-HD on ReRAM with distributed dataflow and organization that maximizes the data parallelism and hardware utilization. The evaluation shows that FSL-HD achieves 4.2% higher accuracy compared to other FSL classifiers. FSL-HD achieves$100-1000\times$better energy efficiency and$9-66\times$speedup over the CPU and GPU baselines. Moreover, FSL-HD is more accurate, scalable and$2.5\times$faster than the state-of-the-art ReRAM-based FSL design, SAPIENS, while requiring 85% less area. Jaeyoung Kang 0001, Tajana Rosing |
DATE | 2 |
| 2023 | Accelerating open modification spectral library searching on tensor core in high-dimensional spaceabstractMOTIVATION: Driven by technological advances, the throughput and cost of mass spectrometry (MS) proteomics experiments have improved by orders of magnitude in recent decades. Spectral library searching is a common approach to annotating experimental mass spectra by matching them against large libraries of reference spectra corresponding to known peptides. An important disadvantage, however, is that only peptides included in the spectral library can be found, whereas novel peptides, such as those with unexpected post-translational modifications (PTMs), will remain unknown. Open modification searching (OMS) is an increasingly popular approach to annotate modified peptides based on partial matches against their unmodified counterparts. Unfortunately, this leads to very large search spaces and excessive runtimes, which is especially problematic considering the continuously increasing sizes of MS proteomics datasets. RESULTS: We propose an OMS algorithm, called HOMS-TC, that fully exploits parallelism in the entire pipeline of spectral library searching. We designed a new highly parallel encoding method based on the principle of hyperdimensional computing to encode mass spectral data to hypervectors while minimizing information loss. This process can be easily parallelized since each dimension is calculated independently. HOMS-TC processes two stages of existing cascade search in parallel and selects the most similar spectra while considering PTMs. We accelerate HOMS-TC on NVIDIA's tensor core units, which is emerging and readily available in the recent graphics processing unit (GPU). Our evaluation shows that HOMS-TC is 31× faster on average than alternative search engines and provides comparable accuracy to competing search tools. AVAILABILITY AND IMPLEMENTATION: HOMS-TC is freely available under the Apache 2.0 license as an open-source software project at https://github.com/tycheyoung/homs-tc. Jaeyoung Kang 0001, Wout Bittremieux, Niema Moshiri, Tajana Rosing |
Bioinform. | 1 |
| 2022 | Massively Parallel Open Modification Spectral Library Searching with Hyperdimensional ComputingabstractMass spectrometry, for protein identification, generates a massive number of spectra that need to be matched against a large database. In reality, most spectra remain mismatched due to unexpected post-translational modifications. Open modification search (OMS) improves the identification rate by considering every possible change in spectra, but it expands the search space exponentially. We propose HyperOMS, which redesigns OMS based on hyperdimensional computing to cope with such challenges. HyperOMS encodes floating-point spectral data with high-dimensional binary vectors, enabling the massive parallelism in OMS. Experimental results show that HyperOMS on GPU is up to 17× faster and 6.4× more energy efficient than the state-of-the-art GPU-based OMS tool [2] while providing comparable search quality. Jaeyoung Kang 0001, Wout Bittremieux, Tajana Rosing |
PACT | 1 |
| 2022 | XCelHD: An Efficient GPU-Powered Hyperdimensional Computing with Parallelized TrainingabstractHyperdimensional Computing (HDC) is an emerging lightweight machine learning method alternative to deep learning. One of its key strengths is the ability to accelerate it in hardware, as it offers massive parallelisms. Prior work primarily focused on FPGA and ASIC, which do not provide the seamless flexibility required for HDC applications. Few studies that attempted GPU designs are inefficient, partly due to the complexity of accelerating HDC on GPUs because of the bit-level operations of HDC. Besides, HDC training exhibited low hardware utilization due to sequential operations. In this paper, we present XCelHD, a high-performance GPU-powered framework for HDC. XCelHD uses a novel training method to maximize the training speed of the HDC model while fully utilizing hardware. We propose memory optimization strategies specialized for GPU-based HDC, minimizing the access time to different memory subsystems and redundant operations. We show that the proposed training method reduces the required number of training epochs by four-fold to achieve comparable accuracy. Our evaluation results on NVIDIA Jetson TX2 show that XCelHD is up to$35\times$faster than the state-of-the-art TensorFlow-based HDC implementation. Jaeyoung Kang 0001, Behnam Khaleghi, Yeseong Kim, Tajana Rosing |
ASP-DAC | 1 |
| 2022 | FHDnn: communication efficient and robust federated learning for AIoT networksabstractThe advent of IoT and advances in edge computing inspired federated learning, a distributed algorithm to enable on device learning. Transmission costs, unreliable networks and limited compute power all of which are typical characteristics of IoT networks pose a severe bottleneck for federated learning. In this work we propose FHDnn, a synergetic federated learning framework that combines the salient aspects of CNNs and Hyperdimensional Computing. FHDnn performs hyperdimensional learning on features extracted from a self-supervised contrastive learning framework to accelerate training, lower communication costs, and increase robustness to network errors by avoiding the transmission of the CNN and training only the hyperdimensional component. Compared to CNNs, we show through experiments that FHDnn reduces communication costs by 66X, local client compute and energy consumption by 1.5 - 6X, while being highly robust to network errors with minimal loss in accuracy. Rishikanth Chandrasekaran, Kazim Ergun, Dhanush Nanjunda, Jaeyoung Kang 0001, Tajana Rosing |
DAC | 5 |
| 2022 | GENERIC: highly efficient learning engine on edge using hyperdimensional computingabstractHyperdimensional Computing (HDC) mimics the brain's basic principles in performing cognitive tasks by encoding the data to high-dimensional vectors and employing non-complex learning techniques. Conventional processing platforms such as CPUs and GPUs are incapable of taking full advantage of the highly-parallel bit-level operations of HDC. On the other hand, existing HDC encoding techniques do not cover a broad range of applications to make a custom design plausible. In this paper, we first propose a novel encoding that achieves high accuracy for diverse applications. Thereafter, we leverage the proposed encoding and design a highly efficient and flexible ASIC accelerator, dubbed GENERIC, suited for the edge domain. GENERIC supports both classification (train and inference) and clustering for unsupervised learning on edge. Our design is flexible in the input size (hence it can run various applications) and hypervectors dimensionality, allowing it to trade off the accuracy and energy/performance on-demand. We augment GENERIC with application-opportunistic power-gating and voltage over-scaling (thanks to the notable error resiliency of HDC) for further energy reduction. GENERIC encoding improves the prediction accuracy over previous HDC and ML techniques by 3.5% and 6.5%, respectively. At 14 nm technology node, GENERIC occupies an area of 0.30 mm2, and consumes 0.09 mW static and 1.97 mW active power. Compared to the previous inference-only accelerator, GENERIC reduces the energy consumption by 4.1×. Behnam Khaleghi, Jaeyoung Kang 0001, Hanyang Xu 0002, Justin Morris, Tajana Rosing |
DAC | 2 |
| 2022 | PatterNet: explore and exploit filter patterns for efficient deep neural networksabstractWeight clustering is an effective technique for compressing deep neural networks (DNNs) memory by using a limited number of unique weights and low-bit weight indexes to store clustering information. In this paper, we propose PatterNet, which enforces shared clustering topologies on filters. Cluster sharing leads to a greater extent of memory reduction by reusing the index information. PatterNet effectively factorizes input activations and post-processes the unique weights, which saves multiplications by several orders of magnitude. Furthermore, PatterNet reduces the add operations by harnessing the fact that filters sharing a clustering pattern have the same factorized terms. We introduce techniques for determining and assigning clustering patterns and training a network to fulfill the target patterns. We also propose and implement an efficient accelerator that builds upon the patterned filters. Experimental results show that PatterNet shrinks the memory and operation count up to 80.2% and 73.1%, respectively, with similar accuracy to the baseline models. PatterNet accelerator improves the energy efficiency by 107x over Nvidia 1080 1080 GTX and 2.2x over state of the art. Behnam Khaleghi, Uday Mallappa, Duygu Yaldiz, Haichao Yang, Monil Shah, Jaeyoung Kang 0001, Tajana Rosing |
DAC | 6 |
| 2022 | A near-storage framework for boosted data preprocessing of mass spectrum clusteringabstractMass spectrometry (MS) has been a key to proteomics and metabolomics due to its unique ability to identify and analyze protein structures. Modern MS equipment generates massive amount of tandem mass spectra with high redundancy, making spectral analysis the major bottleneck in design of new medicines. Mass spectrum clustering is one promising solution as it greatly reduces data redundancy and boosts protein identification. However, state-of-the-art MS tools take many hours to run spectrum clustering. Spectra loading and preprocessing consumes average 82% execution time and energy during clustering. We propose a near-storage framework, MSAS, to speed up spectrum preprocessing. Instead of loading data into host memory and CPU, MSAS processes spectra near storage, thus reducing the expensive cost of data movement. We present two types of accelerators that leverage internal bandwidth at two storage levels: SSD and channel. The accelerators are optimized to match the data rate at each storage level with negligible overhead. Our results demonstrate that the channel-level design yields the best performance improvement for preprocessing - it is up to 187X and 1.8X faster than the CPU and the state-of-the-art in-storage computing solution, INSIDER, respectively. After integrating channel-level MSAS into existing MS clustering tools, we measure system level improvements in speed of 3.5X to 9.8X with 2.8X to 11.9X better energy efficiency. Jaeyoung Kang 0001, Tajana Rosing |
DAC | 2 |
| 2022 | TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for TransformerabstractTransformer-based models are state-of-the-art for many machine learning (ML) tasks. Executing Transformer usually requires a long execution time due to the large memory footprint and the low data reuse rate, stressing the memory system while under-utilizing the computing resources. Memory-based processing technologies, including processing in-memory (PIM) and near-memory computing (NMC), are promising to accelerate Transformer since they provide high memory bandwidth utilization and extensive computation parallelism. However, the previous memory-based ML accelerators mainly target at optimizing dataflow and hardware for compute-intensive ML models (e.g., CNNs), which do not fit the memory-intensive characteristics of Transformer. In this work, we propose TransPIM, a memory-based acceleration for Transformer using software and hardware co-design. In the software-level, TransPIM adopts a token-based dataflow to avoid the expensive inter-layer data movements introduced by previous layer-based dataflow. In the hardware-level, TransPIM introduces lightweight modifications in the conventional high bandwidth memory (HBM) architecture to support PIM-NMC hybrid processing and efficient data communication for accelerating Transformer-based models. Our experiments show that TransPIM is 3.7× to 9.1× faster than existing memory-based acceleration. As compared to conventional accelerators, TransPIM is 22.1× to 114.9× faster than GPUs and provides 2.0× more throughput than existing ASIC-based accelerators. Minxuan Zhou, Jaeyoung Kang 0001, Tajana Rosing |
HPCA | 3 |
| 2022 | RelHD: A Graph-based Learning on FeFET with Hyperdimensional ComputingabstractAdvances in graph neural network (GNN)-based algorithms enable machine learning on relational data. GNNs are computationally demanding since they rely upon backpropagation over the graph data that has sparse and irregular characteristics. In this paper, we propose a lightweight graph-based machine learning framework based on hyperdimensional computing (HDC) called RelHD. It maps the features of each node into a high-dimensional space and embeds relationships between nodes. Using lightweight HDC operations, RelHD enables both training and inference on graph data without backpropagation. Furthermore, we design a scalable processing in-memory (PIM) architecture based on the emerging FeFET technology to accelerate the proposed algorithm. Our strategy optimizes data allocation and operation scheduling that maximizes the accelerator performance by addressing the sparseness and irregularity of the graph. Experimental results show that RelHD offers comparable accuracy to the popular GNN-based algorithms while being up to 32× faster on GPU. Also, our FeFET-based accelerator achieves 33× of speedup and 59287× energy efficiency improvement on average over the GPU. It is 10× faster and 986× more energy efficient on average compared to the state-of-the-art in-memory processing-based GNN accelerator. Jaeyoung Kang 0001, Minxuan Zhou, Abhinav Bhansali, Anthony Thomas, Tajana Rosing |
ICCD | 1 |
| 2022 | COSMO: Computing with Stochastic Numbers in MemoryabstractStochastic computing (SC) reduces the complexity of computation by representing numbers with long streams of independent bits. However, increasing performance in SC comes with either an increase in area or a loss in accuracy. Processing in memory (PIM) computes data in-place while having high memory density and supporting bit-parallel operations with low energy consumption. In this article, we propose COSMO, an architecture for co mputing with s tochastic numbers in me mo ry, which enables SC in memory. The proposed architecture is general and can be used for a wide range of applications. It is a highly dense and parallel architecture that supports most SC encodings and operations in memory. It maximizes the performance and energy efficiency of SC by introducing several innovations: (i) in-memory parallel stochastic number generation, (ii) efficient implication-based logic in memory, (iii) novel memory bit line segmenting, (iv) a new memory-compatible SC addition operation, and (v) enabling flexible block allocation. To show the generality and efficiency of our stochastic architecture, we implement image processing, deep neural networks (DNNs), and hyperdimensional (HD) computing on the proposed hardware. Our evaluations show that running DNN inference on COSMO is 141× faster and 80× more energy efficient as compared to GPU. Saransh Gupta, Mohsen Imani, Joonseop Sim, Andrew Huang 0001, Jaeyoung Kang 0001, Yeseong Kim, Tajana Rosing |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2022 | OpenHD: A GPU-Powered Framework for Hyperdimensional ComputingabstractHyperdimensional computing (HDC) has emerged as an alternative lightweight learning solution to deep neural networks. A key characteristic of HDC is the great extent of parallelism that can facilitate hardware acceleration. However, previous hardware implementations of HDC seldom focus on GPU designs, which were also inefficient partly due to the complexity of accelerating HDC on GPUs. In this paper, we present OpenHD, a flexible and high-performance GPU-powered framework for automating the mapping of general HDC applications including classification and clustering to GPUs. OpenHD takes advantage of memory optimization strategies specialized for HDC, minimizing the access time to different memory subsystems, and removing redundant operations. We also propose a novel training method to enable data parallelism in HDC training. Our evaluation result shows that the proposed training rapidly achieves the target accuracy, reducing the required training epochs by 4×. With OpenHD, users can deploy GPU-accelerated HDC applications without domain expert knowledge. Compared to the state-of-the-art GPU-powered HDC implementation, our evaluation on NVIDIA Jetson TX2 shows that OpenHD is up to 10.5× and 314× faster for HDC-based classification and clustering, respectively. Compared with non-HDC classification and clustering on GPUs, OpenHD-based HDC is 11.7× and 53× faster at comparable accuracy. OpenHD is available at:https://github.com/UCSD-SEELab/openhd. Jaeyoung Kang 0001, Behnam Khaleghi, Tajana Rosing, Yeseong Kim |
IEEE Trans. Computers | 1 |
| 2022 | Store-n-Learn: Classification and Clustering with Hyperdimensional Computing across Flash HierarchyabstractProcessing large amounts of data, especially in learning algorithms, poses a challenge for current embedded computing systems. Hyperdimensional (HD) computing (HDC) is a brain-inspired computing paradigm that works with high-dimensional vectors called hypervectors . HDC replaces several complex learning computations with bitwise and simpler arithmetic operations at the expense of an increased amount of data due to mapping the data into high-dimensional space. These hypervectors, more often than not, cannot be stored in memory, resulting in long data transfers from storage. In this article, we propose Store-n-Learn, an in-storage computing solution that performs HDC classification and clustering by implementing encoding, training, retraining, and inference across the flash hierarchy. To hide the latency of training and enable efficient computation, we introduce the concept of batching in HDC. We also present on-chip acceleration for HDC encoding in flash planes. This enables us to exploit the high parallelism provided by the flash hierarchy and encode multiple data points in parallel in both batched and non-batched fashion. Store-n-Learn also implements a single top-level FPGA accelerator with novel implementations for HDC classification training, retraining, inference, and clustering on the encoded data. Our evaluation over 10 popular datasets shows that Store-n-Learn is on average 222× (543×) faster than CPU and 10.6× (7.3×) faster than the state-of-the-art in-storage computing solution, INSIDER for HDC classification (clustering). Saransh Gupta, Behnam Khaleghi, Sahand Salamat, Justin Morris, Ranganathan Ramkumar, Jeffrey Yu, Aniket Tiwari, Jaeyoung Kang 0001, Mohsen Imani, Baris Aksanli, Tajana Rosing |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2021 | HyperRec: Efficient Recommender Systems with Hyperdimensional ComputingabstractRecommender systems are important tools for many commercial applications such as online shopping websites. There are several issues that make the recommendation task very challenging in practice. The first is that an efficient and compact representation is needed to represent users, items and relations. The second issue is that the online markets are changing dynamically, it is thus important that the recommendation algorithm is suitable for fast updates and hardware acceleration. In this paper, we propose a new hardware-friendly recommendation algorithm based on Hyperdimensional Computing, called HyperRec. Unlike existing solutions which leverages floating-point numbers for the data representation, in HyperRec, users and items are modeled with binary vectors in a high dimension. The binary representation enables to perform the reasoning process of the proposed algorithm only using Boolean operations, which is efficient on various computing platforms and suitable for hardware acceleration. In this work, we show how to utilize GPU and FPGA to accelerate the proposed HyperRec. When compared with the state-of-the-art methods for rating prediction, the CPU-based HyperRec implementation is 13.75x faster and consumes 87% less memory, while decreasing the mean squared error (MSE) for the prediction by as much as 31.84%. Our FPGA implementation is on average 67.0x faster and has 6.9x higher energy efficient as compared to CPU. Our GPU implementation further achieves on average 3.1x speedup as compared to FPGA, while providing only 1.2x lower energy efficiency. Yunhui Guo, Mohsen Imani, Jaeyoung Kang 0001, Sahand Salamat, Justin Morris, Baris Aksanli, Yeseong Kim, Tajana Rosing |
ASP-DAC | 3 |
| 2021 | FPGA Acceleration of Protein Back-Translation and AlignmentabstractIdentifying genome functionality changes our understanding of humans and helps us in disease diagnosis; as well as drug, bio-material, and genetic engineering of plants and animals. Comparing the structure of the protein sequences, when only sequence information is available, against a database with known functionality helps us to identify and recognize the functionality of the unknown sequence. The process of predicting the possible RNA sequence that a specific protein has originated from is called back-translation. Aligning the back-translated RNA sequence against the database locates the most similar sequences, which is used to predict the functionality of the unknown protein sequence. Providing massive parallelism, FPGAs can accelerate bioinformatics applications substantially. In this paper, we propose, FabP11FabP is also the name of a family of proteins, “Fatty-Acid-Binding Proteins”., an optimized FPGA-based accelerator for aligning a back-translated protein sequence against a database of DNA/RNA sequences. FabP is deeply optimized to fully utilize the FPGA resources and the DRAM memory bandwidth to maximize the performance. FabP on a mid-range FPGA provides 8.1 % and 23.3× (24.8× and 266.8 ×) speedup and higher energy efficiency as compared to the GPU-based implementation on a high-end NVIDIA GPU (state-of-the-art CPU implementation), respectively. Sahand Salamat, Jaeyoung Kang 0001, Yeseong Kim, Mohsen Imani, Niema Moshiri, Tajana Rosing |
DATE | 2 |