Yeseong Kim

dblp:142/9828 · DBLP profile ↗
← Back
68ranked-venue papers
9as first author
46since 2021 · last 2026
0000-0001-5947-9632ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 64 · 8 first-author · 44 since 2021Software engineering, systems software and programming languages · 21 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MeshHD: Near-Linear Encoding for Hyperdimensional Computing via Multi-Scale Bases and Kronecker Factorization
abstract
Hyperdimensional (HD) computing is attractive for low-power platforms, but common encoders flatten inputs and treat neighboring features as independent, discarding spatial structure and inflating the cost of a dense F×D apply. We present MESHHD, a spatially aware, relative and multi-scale base that maps 2D coordinates with random Fourier features to approximate a distance kernel; nearby locations receive similar hypervectors regardless of absolute position. We further introduce a compact Kronecker-structured apply that realizes the bundled base with three small GEMMs, reducing arithmetic and weight movement from O(FD) toward a near-linear form while preserving encoder semantics. Our experimental results show that MESHHD consistently improves accuracy over the state-of-the-art nonlinear HD encoders, especially at smaller D, and reduces per-batch encoding time by ~ 3×, with up to 10× savings in encoder MACs/state at D=10,000.
Woongjae Han, Jiseung Kim 0005, Hyukjun Kwon, Hojeong Kim, Selim An, Shinhyoung Jang, Yeseong Kim
DATE7
2026 Enhanced CXL Pooled Memory System for Scalable AI via Embedding Access Prediction
abstract
The embedding operation, pivotal in modern AI applications such as recommendation systems and natural language processing, transforms high-dimensional sparse data into dense vector representations. However, embedding tables are memory-intensive and pose significant challenges in DRAM-based architectures due to their substantial size. This paper introduces Sage, a scalable architecture for embedding operations in CXL-based pooled memory systems. Sage employs advanced caching and prefetching strategies, leveraging an online clustering algorithm to predict embedding table access patterns, and selectively uses Near-Data Processing (NDP) to mitigate the latency associated with CXL memory access. Our comprehensive evaluation demonstrates that Sage significantly enhances throughput and efficiency, providing a cost-effective solution for large-scale AI models. Our experimental results demonstrate that Sage enhances throughput by 2.84 × as compared to conventional memory management systems.
Hoyeon Lee, Minho Ha, Byungil Koh, Jungmin Choi, Yeseong Kim
DATE7
2026 Million-Scale Text-to-Video Retrieval with Hyperdimensional Computing
abstract
Scalable video retrieval is increasingly challenging as datasets reach tens of millions of videos. Current text-to-video retrieval (T2VR) methods either compress videos into single dense vectors, losing segment-level detail, or expand them into multi-frame representations, incurring prohibitive storage and search costs. We propose a binary hyperdimensional representation that encodes each video into a compact 3,072-dimension hypervector, preserving semantic fidelity while reducing memory via bit-packing. To leverage the properties of hypervectors for sublinear search, we introduce Hypervector Retrieval (HVR), a frequency-aware inverted index that prioritizes rare informative positions and refines candidates using GPU-accelerated Hamming search. Experiments show that our approach matches or exceeds dense baselines for T2VR and surpasses state-of-the-art partially relevant video retrieval (PRVR) by over 5% Recall@ 10 on ActivityNet. At scale, HVR processes over 2,000 queries per second on 10M videos, maintains recall within 1% of exact search, and achieves 5.3× greater storage capacity than CLIP4Clip and over 2,116× over MS-SL.
Hyunsei Lee, Jaewoo Gwak, Shinhyoung Jang, Yeseong Kim
EuroSys5
2025 Bit-Level Semantics: Scalable RAG Retrieval with Neurosymbolic Hyperdimensional Computing
abstract
Retrieval-Augmented Generation (RAG) systems typically rely on dense floating-point embeddings to retrieve relevant documents, but this approach incurs significant memory and compute costs at scale. We propose a Hyperdimensional Computing (HDC) framework that projects transformer token embeddings into high-dimensional binary hypervectors, which are aggregated into compact document representations. To support sublinear search, we introduce HD-NSW, a graph-based index inspired by navigable small-world networks. HD-NSW clusters similar hypervectors into bundled centroids and connects them with sparse Hammingdistance edges, enabling efficient, beam-guided traversal entirely in the binary domain. Across 15 BEIR benchmarks and synthetic Gaussian mixture corpora, HD-NSW achieves over 99% of dense retrieval quality, reduces memory usage by $8 \times$, and supports over 860 queries per second at 10 million documents while maintaining over 80% throughput at 40 million documents. At five million documents, HD-NSW achieves $7.68 \times$ higher throughput compared to state of the art approximate nearest neighbor methods. Beyond this point, competing baselines encounter memory exhaustion, while HD-NSW continues scaling and maintains high throughput at larger corpus sizes.
Hyunsei Lee, Shinhyoung Jang, Jaewoo Gwak, Yeseong Kim
PACT5
2025 PersonalizedHD: Hyperdimensional Online Learning with Scalable Personalization and Memory-Efficient Replay
Shinhyoung Jang, Hyunsei Lee, Ilhong Suh, Yeseong Kim
IEEE Big Data6
2025 Late Breaking Results: A Diffusion-Based Framework for Configurable and Realistic Multi-Storage Trace Generation
abstract
We propose DiTTO, a novel diffusion-based framework for generating realistic, precisely configurable, and diverse multi-device storage traces. Leveraging advanced diffusion techniques, DiTTO enables the synthesis of high-fidelity continuous traces that capture temporal dynamics and inter-device dependencies with user-defined configurations. Our experimental results demonstrate that DiTTO can generate traces with high fidelity and diversity while aligning closely with guided configurations with only 8% errors.
Jinhyung Koo, Yeseong Kim
DAC6
2025 Late Breaking Results: Hyperdimensional Regression with Fine-Grained and Scalable Confidence-Based Learning
abstract
We propose an advanced hyperdimensional computing (HDC) framework for regression tasks, addressing the limitations of existing methods through three key innovations: fine-grained feature encoding, confidence-based inference, and dimension-split boosting for scalable training. By preserving inter-feature relationships and enabling efficient computation on high-dimensional spaces, the framework achieves superior accuracy and efficiency across diverse benchmarks. Our evaluation demon-strates that HB R F achieves significant improvements in prediction quality and computational efficiency as compared to the state-of-the-art HDC- based regression by 31% and 54.8 %, respectively.
Jiseung Kim 0005, Hyunsei Lee, Tajana Rosing, Mohsen Imani, Yeseong Kim
DATE5
2025 Exploiting Boosting in Hyperdimensional Computing for Enhanced Reliability in Healthcare
abstract
Hyperdimensional computing (HDC) enables efficient data encoding and processing in high-dimensional spaces, benefiting machine learning and data analysis. However, under-utilization of these spaces can lead to overfitting and reduced model reliability, especially in data-limited systems-a critical issue in sectors like healthcare that demand robustness and consistent performance. We introduce BoostHD, an approach that applies boosting algorithms to partition the hyperdimensional space into subspaces, creating an ensemble of weak learners. By integrating boosting with HDC, BoostHD enhances performance and reliability beyond existing HDC methods. Our analysis highlights the importance of efficient utilization of hyperdimensional spaces for improved model performance. Experiments on healthcare datasets show that BoostHD outperforms state-of-the-art methods. On the WESAD dataset, it achieved an accuracy of 98.37% ± 0.32%, surpassing Random Forest, XGBoost, and On-lineHD. BoostHD also demonstrated superior inference efficiency and stability, maintaining high accuracy under data imbalance and noise. In person-specific evaluations, it achieved an average accuracy of 96.19%, outperforming other models. By addressing the limitations of both boosting and HDC, BoostHD expands the applicability of HDC in critical domains where reliability and precision are paramount.
Sungheon Jeong 0001, Hamza Errahmouni Barkam, Sanggeon Yun, Yeseong Kim, Shaahin Angizi, Mohsen Imani
DATE4
2025 Late Breaking Results: Dynamically Scalable Pruning for Transformer-Based Large Language Models
abstract
We propose Matryoshka, a novel framework for transformer model pruning, enabling dynamic runtime controls while maintaining competitive accuracy to modern large language models (LLMs). Matryoshka incrementally constructs submodels with varying complexities, allowing runtime adaptation without maintaining separate models. Our evaluations on LLaMA-7B demonstrate that Matryoshka achieves up to 34% speedup and outperforms the quality of state-of-the-art pruning methods, providing a flexible solution for deploying LLMs.
Shinhyoung Jang, Ilhong Suh, Hoon Sung Chwa, Yeseong Kim
DATE7
2025 Hyperdimensional Computing-Based Federated Learning in Mobile Robots Through Synthetic Oversampling
abstract
Traditional federated learning frameworks, often reliant on deep neural networks, face challenges related to computational demands and privacy risks. In this paper, we present a novel Hyperdimensional (HD) Computing-based federated learning framework designed for resource-constrained mobile robots. Unlike other HD-based learning, our approach introduces dynamic encoding, which improves both model accuracy and privacy by continuously updating hypervector representations. To further address the issue of imbalanced data, especially prevalent in robotics tasks, we propose a hypervector oversampling technique, enhancing model robustness. Extensive evaluations on LiDAR-equipped mobile robots demonstrate that our oversampling method outperforms state-of-the-art HD computing frameworks, achieving up to a 22.9% increase in accuracy while maintaining computational efficiency.
Hyunsei Lee, Woongjae Han, Hojeong Kim, Hyukjun Kwon, Shinhyoung Jang, Ilhong Suh, Yeseong Kim
ICRA7
2025 FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF Rendering
abstract
Neural Radiance Fields (NeRF), an AI-driven approach for 3D view reconstruction, has demonstrated impressive performance, sparking active research across fields.As a result, a range of advanced NeRF models has emerged, leading on-device applications to increasingly adopt NeRF for highly realistic scene reconstructions.With the advent of diverse NeRF models, NeRF-based applications leverage a variety of NeRF frameworks, creating the need for hardware capable of efficiently supporting these models.However, GPUs fail to meet the performance, power, and area (PPA) cost demanded by these on-device applications, or are specialized for specific NeRF algorithms, resulting in lower efficiency when applied to other NeRF models.To address this limitation, in this work, we introduce FlexNeRFer, an energy-efficient versatile NeRF accelerator.The key components enabling the enhancement of FlexNeRFer include: i) a flexible network-on-chip (NoC) supporting multi-dataflow and sparsity on precision-scalable MAC array, and ii) efficient data storage using an optimal sparsity format based on the sparsity ratio and precision modes.To evaluate the effectiveness of FlexNeRFer, we performed a layout implementation using 28nm CMOS technology.Our evaluation shows that FlexNeRFer achieves 8.2∼243.3×speedup and 24.1∼520.3×improvement in energy efficiency over a GPU (i.e., NVIDIA RTX 2080 Ti), while demonstrating 4.2∼86.9×speedup and 2.3∼47.5×improvement in energy efficiency compared to a state-of-the-art NeRF accelerator (i.e., NeuRex).
Seock-Hwan Noh, Banseok Shin, Jeik Choi, Seungpyo Lee, Yeseong Kim
ISCA6
2025 Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
abstract
In this work, we introduce an area- and energy-efficient multiply-accumulate (MAC) unit, named Jack Unit, that is a jack-of-all-trades, supporting various data formats such as integer (INT), floating point (FP), and microscaling data format (MX). It provides bit-level flexibility and enhances hardware efficiency by i) replacing the carry-save multiplier (CSM) in the FP multiplier with a precision-scalable CSM, ii) performing the adjustment of significands based on the exponent differences within the CSM, and iii) utilizing 2D sub-word parallelism. To assess effectiveness, we implemented the layout of the Jack unit and three baseline MAC units. Additionally, we designed an AI accelerator equipped with our Jack units to compare with a state-of-the-art AI accelerator supporting various data formats. The proposed MAC unit achieves an area reduction of 14.53∼50.25% and a power reduction of 4.76 ∼ 45.65% compared to the baseline MAC units. On five AI benchmarks, the accelerator de-signed with our Jack units improves energy efficiency by 1.32 ∼ 5.41× over the baseline across various data formats.
Seock-Hwan Noh, Sungju Kim, Daehoon Kim 0001, Jaeha Kung 0001, Yeseong Kim
ISLPED6
2025 DeepPM: Predicting Performance and Energy Consumption of Program Binaries Using Transformers
abstract
Accurate estimation of performance and energy consumption is critical for optimizing application efficiency on diverse hardware platforms. Traditional methods often rely on profiling and measurements, requiring at least one execution, making them time-consuming and resource-intensive. This article introduces the Deep Power Meter (DeepPM) framework, leveraging deep learning, specifically the Transformer architecture, to predict performance and energy consumption of basic blocks directly from compiled binaries, eliminating the need for explicit measurement processes. The DeepPM model effectively learns the performance and energy consumption of basic blocks, enabling accurate predictions for each. Furthermore, the framework enhances applicability across different ISAs and microarchitectures, addressing limitations of state-of-the-art ML-based techniques restricted to specific processor architectures. Experimental results using the SPEC CPU 2017 benchmark suite show that DeepPM achieves significantly lower prediction errors compared to state-of-the-art ML-based techniques, with a 24% improvement in performance and an 18% improvement in energy consumption for x86 basic blocks, and similar gains for ARM processors. Fine-tuning with minimal data from the Phoronix Test Suite further validates DeepPM’s robustness, achieving an error of approximately 13.7%, close to the fully trained model’s 13.3% error. These findings demonstrate DeepPM’s ability to enhance the accuracy and efficiency of performance and energy consumption predictions, making it a valuable tool for optimizing computing systems across diverse hardware environments.
Jun S. Shim, Hyeonji Chang, Yeseong Kim, Jihong Kim 0001
ACM Trans. Design Autom. Electr. Syst.3
2024 NDPipe: Exploiting Near-data Processing for Scalable Inference and Continuous Training in Photo Storage
abstract
This paper proposes a novel photo storage system called NDPipe, which accelerates the performance of training and inference for image data by leveraging near-data processing in photo storage servers. NDPipe distributes storage servers with inexpensive commodity GPUs in a data center and uses their collective intelligence to perform inference and training near image data. By efficiently partitioning deep neural network (DNN) models and exploiting the data parallelism of many storage servers, NDPipe can achieve high training throughput with low synchronization costs. NDPipe optimizes the near-data processing engine to maximally utilize system components in each storage server. Our results show that, given the same energy budget, NDPipe exhibits 1.39× higher inference throughput and 2.64× faster training speed than typical photo storage systems.
Jungwoo Kim 0004, Seonggyun Oh, Jaeha Kung 0001, Yeseong Kim, Sungjin Lee 0001
ASPLOS (3)4
2024 Towards Forward-Only Learning for Hyperdimensional Computing
abstract
Hyperdimensional (HD) Computing is a lightweight representation system that symbolizes data as high-dimensioned vectors. HD computing has been growing in popularity in recent years as an alternative to deep neural networks mainly due to its simple and efficient operations. In HD-based learning frameworks, the encoding of the high dimensional representations are widely cited to be the most contributing procedure to accuracy and efficiency. However, throughout HD computing's history, the encoder has largely remained static. In this work, we explore methods for a dynamic encoder that yields better representations as training progresses. Our proposed method, SEP, achieves accuracies comparable to state-of-the-art HD-based methods proposed in the literature; more notably, our solutions outperform existing work at lower dimensions while maintaining a relatively small dimension of$D=3,000$, which equates to an average of$3.32\times$faster inference.
Hyunsei Lee, Hyukjun Kwon, Jiseung Kim 0005, Mohsen Imani, Yeseong Kim
DATE6
2024 Efficient Forward-Only Training for Brain-Inspired Hyperdimensional Computing
abstract
Hyperdimensional (HD) computing is an emerging paradigm inspired by human cognition, utilizing high-dimensional vectors to represent and learn information in a lightweight manner based on its simple and efficient operations. In HD-based learning frameworks, the encoding of the high dimensional representations is the most contributing procedure to accuracy and efficiency. However, throughout HD computing's history, the encoder has largely remained static, which leads to sub-optimal hypervector representations and excessive dimensionality requirements. In this paper, we propose novel forward-only training methods for HD encoders, Stochastic Error Projection (SEP) and Input Modulated Projection (IMP), which dynamically adjust the encoding process during training. Our methods achieve accuracies comparable to state-of-the-art HD-based techniques, with SEP and IMP outperforming existing methods by 5.49% on average at a reduced dimensionality of D = 3,000. This reduction in dimensionality results in a 3.32x faster inference.
Hyunsei Lee, Jiseung Kim 0005, Hyukjun Kwon, Mohsen Imani, Ilhong Suh, Yeseong Kim
ICCD7
2024 Brain-Inspired Hyperdimensional Computing in the Wild: Lightweight Symbolic Learning for Sensorimotor Controls of Wheeled Robots
abstract
Efficiency and performance are significant challenges in applying Machine Learning (ML) to robotics, especially in energy-constrained real-world scenarios. In this context, Hyperdimensional Computing offers an energy-efficient alternative but has been underexplored in robotics. We introduce ReactHD, an HDC-based framework tailored for perception-action-based learning for sensorimotor controls of robot tasks. ReactHD employs hypervectors to encode sensory inputs and learn the suitable high-dimensional pattern for robot actions. It also integrates two HD-based lightweight symbolic learning techniques: HDC-based supervised learning by demonstration (HDC-IL) and HD-Reinforcement Learning (HDC-RL) to enable precise, reactive robot behaviors in complex environments. Our empirical evaluations show that ReactHD achieves robust and accurate learning outcomes comparable to state-of-the-art deep learning while substantially improving the performance and energy consumption efficiency by 14.2× and 15.3×. To the best of our knowledge, ReactHD is the first HDC-based framework deployed in real-world settings.
Hyukjun Kwon, Kangwon Kim, Hyunsei Lee, Jiseung Kim 0005, Jinhyung Kim, Yongnyeon Kim, Yang Ni 0001, Mohsen Imani, Ilhong Suh, Yeseong Kim
ICRA12
2024 Advancing Hyperdimensional Computing Based on Trainable Encoding and Adaptive Training for Efficient and Accurate Learning
abstract
Hyperdimensional computing (HDC) is a computing paradigm inspired by the mechanisms of human memory, characterizing data through high-dimensional vector representations, known as hypervectors. Recent advancements in HDC have explored its potential as a learning model, leveraging its straightforward arithmetic and high efficiency. The traditional HDC frameworks are hampered by two primary static elements: randomly generated encoders and fixed learning rates. These static components significantly limit model adaptability and accuracy. The static, randomly generated encoders, while ensuring high-dimensional representation, fail to adapt to evolving data relationships, thereby constraining the model’s ability to accurately capture and learn from complex patterns. Similarly, the fixed nature of the learning rate does not account for the varying needs of the training process over time, hindering efficient convergence and optimal performance. This article introducesTrainableHD, a novel HDC framework that enables dynamic training of the randomly generated encoder depending on the feedback of the learning data, thereby addressing the static nature of conventional HDC encoders.TrainableHDalso enhances the training performance by incorporating adaptive optimizer algorithms in learning the hypervectors. We further refineTrainableHDwith effective quantization to enhance efficiency, allowing the execution of the inference phase in low-precision accelerators. Our evaluations demonstrate thatTrainableHDsignificantly improves HDC accuracy by up to 27.99% (averaging 7.02%) without additional computational costs during inference, achieving a performance level comparable to state-of-the-art deep learning models. Furthermore,TrainableHDis optimized for execution speed and energy efficiency. Compared to deep learning on a low-power GPU platform like NVIDIA Jetson Xavier,TrainableHDis 56.4 times faster and 73 times more energy efficient. This efficiency is further augmented through the use of Encoder Interval Training (EIT) and adaptive optimizer algorithms, enhancing the training process without compromising the model’s accuracy.
Jiseung Kim 0005, Hyunsei Lee, Mohsen Imani, Yeseong Kim
ACM Trans. Design Autom. Electr. Syst.4
2023 Comprehensive Integration of Hyperdimensional Computing with Deep Learning towards Neuro-Symbolic AI
abstract
HD computing is a symbolic representation system which performs various learning tasks in a highly-parallelizable and binary-centric way by drawing inspiration from concepts in human long-term memory. However, the current HD computing is ineffective in extracting high-level feature information for image data. In this paper, we present a neuro-symbolic approach called NSHD, which integrates CNNs and Hyperdimensional (HD) learning techniques to provide efficient learning with state-of-the-art quality. We devise the HD training procedure, which fully integrates knowledge from the deep learning model through a distillation process with optimized computation costs due to the integration. Our experimental results show that NSHD provides high energy efficiency as compared to CNN, e.g., up to 64% with comparable accuracy, and can outperform the learning quality when more computing resources are allowed. We also show the symbolic nature of the NSHD can make the learning humnan-interpretable by exploiting the property of HD computing.
Hyunsei Lee, Jiseung Kim 0005, Hanning Chen, Ariela Zeira, Narayan Srinivasa, Mohsen Imani, Yeseong Kim
DAC7
2023 Sidekick: Near Data Processing for Clustering Enhanced by Automatic Memory Disaggregation
abstract
Near Data Processing (NDP) is a promising solution for data mining/analysis techniques, which extract useful information from big data. In this paper, we propose a novel NDP-enabled memory disaggregation system called Sidekick, based on a type-2 CXL device and enhanced by an automated allocation technique for clustering algorithms. The key enabler of our migration technique is to understand clustering workflows in a unit of the program context, which is the function call stack for functions, threads, and memory allocations to drive the automated decision. The proposed technique relates the migrated computation tasks with a series of function calls and performs GA-based optimization to identify the optimal allocation scenario for a target clustering algorithm. In Scikit-learn, a popular machine learning library, we use the genetic algorithm to find the optimal memory allocation policy and the operation offloading policy using the program context. The results show that the proposed technique increases the clustering performance as compared to the case, which only uses disaggregated memory without NDP cores, by up to 92% in terms of execution time, while reducing the majority of remote CXL memory accesses.
Minho Ha, Byungil Koh, Kyoung Park, Yeseong Kim
DAC6
2023 Efficient Hyperdimensional Learning with Trainable, Quantizable, and Holistic Data Representation
abstract
Hyperdimensional computing (HDC) is a computing paradigm that draws inspiration from human memory models. It represents data in the form of high-dimensional vectors. Recently, many works in literature have tried to use HDC as a learning model due to its simple arithmetic and high efficiency. However, learning frameworks in HDC use encoders that are randomly generated and static, resulting in many parameters and low accuracy. In this paper, we propose TrainableHD, a framework for HDC that utilizes a dynamic encoder with effective quantization for higher efficiency. Our model considers errors gained from the HD model and dynamically updates the encoder during training. Our evaluations show that TrainableHD improves the accuracy of the HDC by up to 22.26% (on average 3.62%) without any extra computation costs, achieving a comparable level to state-of-the-art deep learning. Also, the proposed solution is 56.4 x faster and 73 x more energy efficient as compared to the deep learning on NVIDIA Jetson Xavier, a low-power GPU platform.
Jiseung Kim 0005, Hyunsei Lee, Mohsen Imani, Yeseong Kim
DATE4
2023 Efficient Off-Policy Reinforcement Learning via Brain-Inspired Computing
abstract
Reinforcement Learning (RL) has opened up new opportunities to enhance existing smart systems that generally include a complex decision-making process. However, modern RL algorithms, e.g., Deep Q-Networks (DQN), are based on deep neural networks, resulting in high computational costs. In this paper, we propose QHD, an off-policy value-based Hyperdimensional Reinforcement Learning, that mimics brain properties toward robust and real-time learning. QHD relies on a lightweight brain-inspired model to learn an optimal policy in an unknown environment. On both desktop and power-limited embedded platforms, QHD achieves significantly better overall efficiency than DQN while providing higher or comparable rewards. QHD is also suitable for highly-efficient reinforcement learning with great potential for online and real-time learning. Our solution supports a small experience replay batch size that provides 12.3 times speedup compared to DQN while ensuring minimal quality loss. Our evaluation shows QHD capability for real-time learning, providing 34.6 times speedup and significantly better quality of learning than DQN.
Yang Ni 0001, Danny Abraham, Mariam Issa, Yeseong Kim, Pietro Mercati, Mohsen Imani
ACM Great Lakes Symposium on VLSI4
2023 Hierarchical, Distributed and Brain-Inspired Learning for Internet of Things Systems
abstract
In this paper, we propose EdgeHD, a hierarchy-aware learning solution that performs online training and inference in a highly distributed, cost-effective way. We use brain-inspired hyperdimensional (HD) computing as the key enabler. HD computing performs the computation tasks on a high-dimensional space to emulate functionalities of the human memory, such as inter-data relationship reasoning and information aggregation. EdgeHD exploits HD computing to effectively learn the classification models on individual devices and combine the models through the hierarchical IoT nodes without high communication costs. We also propose a hardware design that accelerates EdgeHD on low-power FPGA platforms. We evaluated EdgeHD for a wide range of real-world classification applications. The evaluation shows that EdgeHD provides highly efficient computation with reduced communication. For example, EdgeHD achieves on average$3.4\times$and$11.7\times (1.9\times$and$7.8\times$) speedup and energy efficiency improvement during the training (inference) as compared to the centralized learning approach. It reduces the communication costs by 85% for the training and 78% for the inference.
Mohsen Imani, Yeseong Kim, Behnam Khaleghi, Justin Morris, Haleh Alimohamadi, Farhad Imani, Hugo Latapie
ICDCS2
2023 Algorithm-Hardware Co-Design for Efficient Brain-Inspired Hyperdimensional Learning on Edge (Extended Abstract)
abstract
In this paper, we propose an efficient framework to accelerate a lightweight brain-inspired learning solution, hyperdimensional computing (HDC), on existing edge systems. Through algorithm-hardware co-design, we optimize the HDC models to run them on the low-power host CPU and machine learning accelerators like Edge TPU. By treating the lightweight HDC learning model as a hyper-wide neural network, we exploit the capabilities of the accelerator and machine learning platform, while reducing training runtime costs by using bootstrap aggregating. Our experimental results conducted on mobile CPU and the Edge TPU demonstrate that our framework achieves 4.5 times faster training and 4.2 times faster inference than the baseline platform. Furthermore, compared to the embedded ARM CPU, Raspberry Pi, with similar power consumption, our framework achieves 19.4 times faster training and 8.9 times faster inference.
Yang Ni 0001, Yeseong Kim, Tajana Rosing, Mohsen Imani
IJCAI2
2023 Sparsity Controllable Hyperdimensional Computing for Genome Sequence Matching Acceleration
abstract
In this paper, we propose a Hyper-Dimensional genome analysis platform. Instead of working with original sequences, our method maps the genome sequences into high-dimensional space and performs sequence matching with simple and parallel similarity searches. At the algorithm level, we revisit the sequence searching with brain-like memorization that Hyper-Dimensional computing natively supports. Instead of working on the original data, we map all data points into high-dimensional space, enabling the main sequence searching operations to process in a hardware-friendly way. We accordingly design a density-aware FPGA implementation. Our solution searches the similarity of an encoded query and large-scale genome library through different chunks. We exploit the holographic representation of patterns to stop search operations on libraries with a lower chance of a match. This translates our computation from dense to highly sparse just after a few chuck-based searches. Our evaluation shows that our accelerator can provide 46× speedup and 188× energy efficiency improvement compared to a state-of-the-art GPU implementation. Results show that our accelerator achieves up to 3440.6 GCUPS using a single Xilinx Alveo U280 board.
Hanning Chen, Yeseong Kim, Elaheh Sadredini, Saransh Gupta, Hugo Latapie, Mohsen Imani
VLSI-SoC2
2022 XCelHD: An Efficient GPU-Powered Hyperdimensional Computing with Parallelized Training
abstract
Hyperdimensional Computing (HDC) is an emerging lightweight machine learning method alternative to deep learning. One of its key strengths is the ability to accelerate it in hardware, as it offers massive parallelisms. Prior work primarily focused on FPGA and ASIC, which do not provide the seamless flexibility required for HDC applications. Few studies that attempted GPU designs are inefficient, partly due to the complexity of accelerating HDC on GPUs because of the bit-level operations of HDC. Besides, HDC training exhibited low hardware utilization due to sequential operations. In this paper, we present XCelHD, a high-performance GPU-powered framework for HDC. XCelHD uses a novel training method to maximize the training speed of the HDC model while fully utilizing hardware. We propose memory optimization strategies specialized for GPU-based HDC, minimizing the access time to different memory subsystems and redundant operations. We show that the proposed training method reduces the required number of training epochs by four-fold to achieve comparable accuracy. Our evaluation results on NVIDIA Jetson TX2 show that XCelHD is up to$35\times$faster than the state-of-the-art TensorFlow-based HDC implementation.
Jaeyoung Kang 0001, Behnam Khaleghi, Yeseong Kim, Tajana Rosing
ASP-DAC3
2022 Neural computation for robust and holographic face detection
abstract
Face detection is an essential component of many tasks in computer vision with several applications. However, existing deep learning solutions are significantly slow and inefficient to enable face detection on embedded platforms. In this paper, we propose HDFace, a novel framework for highly efficient and robust face detection. HDFace exploits HyperDimensional Computing (HDC) as a neurally-inspired computational paradigm that mimics important brain functionalities towards high-efficiency and noise-tolerant computation. We first develop a novel technique that enables HDC to perform stochastic arithmetic computations over binary hypervectors. Next, we expand these arithmetic for efficient and robust processing of feature extraction algorithms in hyperspace. Finally, we develop an adaptive hyperdimensional classification algorithm for effective and robust face detection. We evaluate the effectiveness of HDFace on large-scale emotion detection and face detection applications. Our results indicate that HDFace provides, on average, 6.1X (4.6X) speedup and 3.0X (12.1X) energy efficiency as compared to neural networks running on CPU (FPGA), respectively.
Mohsen Imani, Ali Zakeri, Hanning Chen, Prathyush Poduval, Hyunsei Lee, Yeseong Kim, Elaheh Sadredini, Farhad Imani
DAC7
2022 QuiltNet: efficient deep learning inference on multi-chip accelerators using model partitioning
abstract
We have seen many successful deployments of deep learning accelerator designs on different platforms and technologies, e.g., FPGA, ASIC, and Processing In-Memory platforms. However, the size of the deep learning models keeps increasing, making computations a burden on the accelerators. A naive approach to resolve this issue is to design larger accelerators; however, it is not scalable due to high resource requirements, e.g., power consumption and off-chip memory sizes. A promising solution is to utilize multiple accelerators and use them as needed, similar to conventional multiprocessing. For example, for smaller networks, we may use a single accelerator, while we may use multiple accelerators with proper network partitioning for larger networks. However, partitioning DNN models into multiple parts leads to large communication overheads due to inter-layer communications. In this paper, we propose a scalable solution to accelerate DNN models on multiple devices by devising a new model partitioning technique. Our technique transforms a DNN model into layer-wise partitioned models using an autoencoder. Since the autoencoder encodes a tensor output into a smaller dimension, we can split the neural network model into multiple pieces while significantly reducing the communication overhead to pipeline them. Our evaluation results conducted on state-of-the-art deep learning models show that the proposed technique significantly improves performance and energy efficiency. Our solution increases performance and energy efficiency by up to 30.5% and 28.4% with minimal accuracy loss as compared to running the same model on pipelined multi-block accelerators without the autoencoder.
Hyukjun Kwon, Seowoo Kim, Minho Ha, Eui-Cheol Lim, Mohsen Imani, Yeseong Kim
DAC8
2022 Adaptive neural recovery for highly robust brain-like representation
abstract
Today's machine learning platforms have major robustness issues dealing with insecure and unreliable memory systems. In conventional data representation, bit flips due to noise or attack can cause value explosion, which leads to incorrect learning prediction. In this paper, we propose RobustHD, a robust and noise-tolerant learning system based on HyperDimensional Computing (HDC), mimicking important brain functionalities. Unlike traditional binary representation, RobustHD exploits a redundant and holographic representation, ensuring all bits have the same impact on the computation. RobustHD also proposes a runtime framework that adaptively identifies and regenerates the faulty dimensions in an unsupervised way. Our solution not only provides security against possible bit-flip attacks but also provides a learning solution with high robustness to noises in the memory. We performed a cross-stacked evaluation from a conventional platform to emerging processing in-memory architecture. Our evaluation shows that under 10% random bit flip attack, RobustHD provides a maximum of 0.53% quality loss, while deep learning solutions are losing over 26.2% accuracy.
Prathyush Poduval, Yang Ni 0001, Yeseong Kim, Kai Ni 0004, Raghavan Kumar, Rosario Cammarota, Mohsen Imani
DAC3
2022 Algorithm-Hardware Co-Design for Efficient Brain-Inspired Hyperdimensional Learning on Edge
abstract
Machine learning methods have been widely utilized to provide high quality for many cognitive tasks. Running sophisticated learning tasks requires high computational costs to process a large amount of learning data. Brain-inspired Hyperdimensional Computing (HDC) is introduced as an alternative solution for lightweight learning on edge devices. However, HDC models still rely on accelerators to ensure realtime and efficient learning. These hardware designs are not commercially available and need a relatively long period to synthesize and fabricate after deriving the new applications. In this paper, we propose an efficient framework for accelerating the HDC at the edge by fully utilizing the available computing power. We optimize the HDC through algorithm-hardware co-design of the host CPU and existing low-power machine learning accelerators, such as Edge TPU. We interpret the lightweight HDC learning model as a hyper-wide neural network to take advantage of the accelerator and machine learning platform. We further improve the runtime cost of training by employing a bootstrap aggregating algorithm called bagging while maintaining the learning quality. We evaluate the performance of the proposed framework with several applications. Joint experiments on mobile CPU and the Edge TPU show that our framework achieves 4.5 × faster training and 4.2 × faster inference compared to the baseline platform. In addition, our framework achieves 19.4 × faster training and 8.9 × faster inference as compared to embedded ARM CPU, Raspberry Pi, that consumes similar power consumption.
Yang Ni 0001, Yeseong Kim, Tajana Rosing, Mohsen Imani
DATE2
2022 Online Performance and Power Prediction for Edge TPU via Comprehensive Characterization
abstract
In this paper, we characterize and model the performance and power consumption of Edge TPU, which efficiently accelerates deep learning (DL) inference in a low-power environment. Systolic array, as a high throughput computation architecture, its usage in the edge excites our interest in its performance and power pattern. We perform an extensive study for various neural network settings and sizes using more than 10,000 DL models. Through comprehensive exploration, we profile which factors highly influence the inference time and power to run DL Models. We show our key remarks for the relation between the performance/power and DL model complexity to enable hardware-aware optimization and design decisions. For example, our measurement shows that energy/performance is not linearly-proportional to the number of MAC operations. In fact, as the computation and DL model size increase, the performance follows a stepped pattern. Hence, the accurate estimate should consider other features of DL models such as on-chip/off-chip memory usages. Based on the characterization, we propose a modeling framework, called PETET, which perform online predictions for the performance and power of Edge TPU. The proposed method automatically identifies the relationship of the performance, power, and memory usages to the DL model settings based on machine learning techniques.
Yang Ni 0001, Yeseong Kim, Tajana Rosing, Mohsen Imani
DATE2
2022 DeepPM: Transformer-based Power and Performance Prediction for Energy-Aware Software
abstract
Many system-level management and optimization techniques need accurate estimates of power consumption and performance. Earlier research has proposed many high-level/source-level estimation modeling works, particularly for basic blocks. However, most of them still need to execute the target software at least once on a fine-grained simulator or real hardware to extract required features. This paper proposes a performance/power prediction framework, called Deep Power Meter (DeepPM), which estimates them accurately only using the compiled binary. Inspired by the deep learning techniques in natural language processing, we convert the program instructions in the form of vectors and predict the average power and performance of basic blocks based on a transformer model. In addition, unlike existing works based on a Long Short-Term Memory (LSTM) model structure, which only works for basic blocks with a small number of instructions, DeepPM provides highly accurate results for long basic blocks, which takes the majority of the execution time for actual application runs. In our evaluation conducted with SPEC2006 benchmark suite, we show that DeepPM can provide accurate prediction for performance and power consumption with 10.2% and 12.3% error, respectively. DeepPM also outperforms the LSTM-based model by up to 67.2% and 34.9% error for performance and power, respectively.
Jun S. Shim, Bogyeong Han, Yeseong Kim, Jihong Kim 0001
DATE3
2022 DeepSketch: A New Machine Learning-Based Reference Search Technique for Post-Deduplication Delta Compression
Jisung Park 0001, Jeonggyun Kim, Yeseong Kim, Sungjin Lee 0001, Onur Mutlu
FAST3
2022 BioHD: an efficient genome sequence search platform using HyperDimensional memorization
abstract
In this paper, we propose BioHD, a novel genomic sequence searching platform based on Hyper-Dimensional Computing (HDC) for hardware-friendly computation. BioHD transforms inherent sequential processes of genome matching to highly-parallelizable computation tasks. We exploit HDC memorization to encode and represent the genome sequences using high-dimensional vectors. Then, it combines the genome sequences to generate an HDC reference library. During the sequence searching, BioHD performs exact or approximate similarity check of an encoded query with the HDC reference library. Our framework simplifies the required sequence matching operations while introducing a statistical model to control the alignment quality. To get actual advantage from BioHD inherent robustness and parallelism, we design a processing in-memory (PIM) architecture with massive parallelism and compatible with the existing crossbar memory. Our PIM architecture supports all essential BioHD operations natively in memory with minimal modification on the array. We evaluate BioHD accuracy and efficiency on a wide range of genomics data, including COVID-19 databases. Our results indicate that PIM provides 102.8× and 116.1× (9.3× and 13.2×) speedup and energy efficiency compared to the state-of-the-art pattern matching algorithm running on GeForce RTX 3060 Ti GPU (state-of-the-art PIM accelerator).
Zhuowen Zou, Hanning Chen, Prathyush Poduval, Yeseong Kim, Mahdi Imani, Elaheh Sadredini, Rosario Cammarota, Mohsen Imani
ISCA4
2022 COSMO: Computing with Stochastic Numbers in Memory
abstract
Stochastic computing (SC) reduces the complexity of computation by representing numbers with long streams of independent bits. However, increasing performance in SC comes with either an increase in area or a loss in accuracy. Processing in memory (PIM) computes data in-place while having high memory density and supporting bit-parallel operations with low energy consumption. In this article, we propose COSMO, an architecture for co mputing with s tochastic numbers in me mo ry, which enables SC in memory. The proposed architecture is general and can be used for a wide range of applications. It is a highly dense and parallel architecture that supports most SC encodings and operations in memory. It maximizes the performance and energy efficiency of SC by introducing several innovations: (i) in-memory parallel stochastic number generation, (ii) efficient implication-based logic in memory, (iii) novel memory bit line segmenting, (iv) a new memory-compatible SC addition operation, and (v) enabling flexible block allocation. To show the generality and efficiency of our stochastic architecture, we implement image processing, deep neural networks (DNNs), and hyperdimensional (HD) computing on the proposed hardware. Our evaluations show that running DNN inference on COSMO is 141× faster and 80× more energy efficient as compared to GPU.
Saransh Gupta, Mohsen Imani, Joonseop Sim, Andrew Huang 0001, Jaeyoung Kang 0001, Yeseong Kim, Tajana Rosing
ACM J. Emerg. Technol. Comput. Syst.7
2022 OpenHD: A GPU-Powered Framework for Hyperdimensional Computing
abstract
Hyperdimensional computing (HDC) has emerged as an alternative lightweight learning solution to deep neural networks. A key characteristic of HDC is the great extent of parallelism that can facilitate hardware acceleration. However, previous hardware implementations of HDC seldom focus on GPU designs, which were also inefficient partly due to the complexity of accelerating HDC on GPUs. In this paper, we present OpenHD, a flexible and high-performance GPU-powered framework for automating the mapping of general HDC applications including classification and clustering to GPUs. OpenHD takes advantage of memory optimization strategies specialized for HDC, minimizing the access time to different memory subsystems, and removing redundant operations. We also propose a novel training method to enable data parallelism in HDC training. Our evaluation result shows that the proposed training rapidly achieves the target accuracy, reducing the required training epochs by 4×. With OpenHD, users can deploy GPU-accelerated HDC applications without domain expert knowledge. Compared to the state-of-the-art GPU-powered HDC implementation, our evaluation on NVIDIA Jetson TX2 shows that OpenHD is up to 10.5× and 314× faster for HDC-based classification and clustering, respectively. Compared with non-HDC classification and clustering on GPUs, OpenHD-based HDC is 11.7× and 53× faster at comparable accuracy. OpenHD is available at:https://github.com/UCSD-SEELab/openhd.
Jaeyoung Kang 0001, Behnam Khaleghi, Tajana Rosing, Yeseong Kim
IEEE Trans. Computers4
2021 HyperRec: Efficient Recommender Systems with Hyperdimensional Computing
abstract
Recommender systems are important tools for many commercial applications such as online shopping websites. There are several issues that make the recommendation task very challenging in practice. The first is that an efficient and compact representation is needed to represent users, items and relations. The second issue is that the online markets are changing dynamically, it is thus important that the recommendation algorithm is suitable for fast updates and hardware acceleration. In this paper, we propose a new hardware-friendly recommendation algorithm based on Hyperdimensional Computing, called HyperRec. Unlike existing solutions which leverages floating-point numbers for the data representation, in HyperRec, users and items are modeled with binary vectors in a high dimension. The binary representation enables to perform the reasoning process of the proposed algorithm only using Boolean operations, which is efficient on various computing platforms and suitable for hardware acceleration. In this work, we show how to utilize GPU and FPGA to accelerate the proposed HyperRec. When compared with the state-of-the-art methods for rating prediction, the CPU-based HyperRec implementation is 13.75x faster and consumes 87% less memory, while decreasing the mean squared error (MSE) for the prediction by as much as 31.84%. Our FPGA implementation is on average 67.0x faster and has 6.9x higher energy efficient as compared to CPU. Our GPU implementation further achieves on average 3.1x speedup as compared to FPGA, while providing only 1.2x lower energy efficiency.
Yunhui Guo, Mohsen Imani, Jaeyoung Kang 0001, Sahand Salamat, Justin Morris, Baris Aksanli, Yeseong Kim, Tajana Rosing
ASP-DAC7
2021 DP-Sim: A Full-stack Simulation Infrastructure for Digital Processing In-Memory Architectures
abstract
Digital processing in-memory (DPIM) is a promising technology that significantly reduces data movements while providing high parallelism. In this work, we design and implement the first full-stack DPIM simulation infrastructure, DP-Sim, which evaluates a comprehensive range of DPIM-specific design space concerning both software and hardware. DP-Sim provides a C++ library to enable DPIM acceleration in general programs while supporting several aspects of software-level exploration by a convenient interface. The DP-Sim software front-end generates specialized instructions that can be processed by a hardware simulator based on a new DPIM-enabled architecture model which is 10.3% faster than conventional memory simulation models. We use DP-Sim to explore the DPIM-specific design space of acceleration for various emerging applications. Our experiments show that bank-level control is 11.3x faster than conventional channel-level control because of higher computing parallelism. Furthermore, cost-aware memory allocation can provide at least 2.2x speedup vs. heuristic methods, showing the importance of data layout in DPIM acceleration.
Minxuan Zhou, Mohsen Imani, Yeseong Kim, Saransh Gupta, Tajana Rosing
ASP-DAC3
2021 CascadeHD: Efficient Many-Class Learning Framework Using Hyperdimensional Computing
abstract
The brain-inspired hyperdimensional computing (HDC) gains attention as a light-weight and extremely parallelizable learning solution alternative to deep neural networks. Prior research shows the effectiveness of HDC-based learning on less powerful systems such as edge computing devices. However, the many-class classification problem is beyond the focus of mainstream HDC research; the existing HDC would not provide sufficient quality and efficiency due to its coarse-grained training. In this paper, we propose an efficient many-class learning framework, called CascadeHD, which identifies latent high-dimensional patterns of many classes holistically while learning a hierarchical inference structure using a novel meta-learning algorithm for high efficiency. Our evaluation conducted on the NVIDIA Jetson device family shows that CascadeHD improves the accuracy for many-class classification by up to 18% while achieving 32% speedup compared to the existing HDC.
Yeseong Kim, Jiseung Kim 0005, Mohsen Imani
DAC1
2021 A Framework for Efficient and Binary Clustering in High-Dimensional Space
abstract
Today's applications generate a large amount of data where the majority of the data are not associated with any labels. Clustering methods are the most commonly used algorithms for data analysis, especially in healthcare. However, running clustering algorithms on embedded devices is significantly slow as the computation involves a large amount of complex pairwise similarity measurements. In this paper, we proposed FebHD, an adaptive framework for efficient and fully binary clustering in high-dimensional space. Instead of using complex similarity metrics, e.g., Euclidean distance, FebHD introduces a nonlinear encoder to map data points into sparse high-dimensional space. FebHD encoder simplifies the similarity search, the most costly and frequent clustering operation, to Hamming distance, which can be accelerated in today's hardware. FebHD performs clustering by assigning each data point to a set of initialized centers. It then updates the centers adaptively based on: (i) data points assigned to each cluster, and (ii) the confidence of the model on the clustering prediction. This adaptive update enables FebHD to provide a high quality of clustering with very few learning iterations. We also propose an end-to-end hardware accelerator that parallelizes the entire FebHD computation by exploiting FPGA bit-level granularity. Our evaluation shows that FebHD provides comparable accuracy to state-of-the-art clustering algorithms, while providing 6.2× and 9.1× (4.7× and 5.8×) faster and higher energy efficiency when running on the same FPGA (GPU) platform.
Alejandro Hernández-Cano, Yeseong Kim, Mohsen Imani
DATE2
2021 FPGA Acceleration of Protein Back-Translation and Alignment
abstract
Identifying genome functionality changes our understanding of humans and helps us in disease diagnosis; as well as drug, bio-material, and genetic engineering of plants and animals. Comparing the structure of the protein sequences, when only sequence information is available, against a database with known functionality helps us to identify and recognize the functionality of the unknown sequence. The process of predicting the possible RNA sequence that a specific protein has originated from is called back-translation. Aligning the back-translated RNA sequence against the database locates the most similar sequences, which is used to predict the functionality of the unknown protein sequence. Providing massive parallelism, FPGAs can accelerate bioinformatics applications substantially. In this paper, we propose, FabP11FabP is also the name of a family of proteins, “Fatty-Acid-Binding Proteins”., an optimized FPGA-based accelerator for aligning a back-translated protein sequence against a database of DNA/RNA sequences. FabP is deeply optimized to fully utilize the FPGA resources and the DRAM memory bandwidth to maximize the performance. FabP on a mid-range FPGA provides 8.1 % and 23.3× (24.8× and 266.8 ×) speedup and higher energy efficiency as compared to the GPU-based implementation on a high-end NVIDIA GPU (state-of-the-art CPU implementation), respectively.
Sahand Salamat, Jaeyoung Kang 0001, Yeseong Kim, Mohsen Imani, Niema Moshiri, Tajana Rosing
DATE3
2021 ManiHD: Efficient Hyper-Dimensional Learning Using Manifold Trainable Encoder
abstract
Hyper-Dimensional (HD) computing emulates the human short memory functionality by computing with hyper-vectors as an alternative to computing with numbers. The main goal of HD computing is to map data points into sparse high-dimensional space where the learning task can perform in a linear and hardware-friendly way. The existing HD computing algorithms are using static and non-trainable encoder; thus, they require very high-dimensionality to provide acceptable accuracy. However, this high dimensionality results in high computational cost, especially over the realistic learning problems. In this paper, we proposed ManiHD that supports adaptive and trainable encoder for efficient learning in high-dimensional space. ManiHD explicitly considers non-linear interactions between the features during the encoding. This enables ManiHD to provide maximum learning accuracy using much lower dimensionality. ManiHD not only enhances the learning accuracy but also significantly improves the learning efficiency during both training and inference phases. ManiHD also enables online learning by sampling data points and capturing the essential features in an unsupervised manner. We also propose a quantization method that trades accuracy and efficiency for optimal configuration. Our evaluation of a wide range of classification tasks shows that ManiHD provides 4.8% higher accuracy than the state-of-the-art HD algorithms. In addition, ManiHD provides, on average, 12.3× (3.2×) faster and 19.3× (6.3×) more energy-efficient training (inference) as compared to the state-of-the-art learning algorithms.
Zhuowen Zou, Yeseong Kim, M. Hassan Najafi, Mohsen Imani
DATE2
2021 Revisiting HyperDimensional Learning for FPGA and Low-Power Architectures
abstract
Today's applications are using machine learning algorithms to analyze the data collected from a swarm of devices on the Internet of Things (IoT). However, most existing learning algorithms are overcomplex to enable real-time learning on IoT devices with limited resources and computing power. Recently, Hyperdimensional computing (HDC) is introduced as an alternative computing paradigm for enabling efficient and robust learning. HDC emulates the cognitive task by representing the values as patterns of neural activity in high-dimensional space. HDC first encodes all data points to high-dimensional vectors. It then efficiently performs the learning task using a well-defined set of operations. Existing HDC solutions have two main issues that hinder their deployments on low-power embedded devices: (i) the encoding module is costly, dominating 80% of the entire training performance, (ii) the HDC model size and the computation cost grow significantly with the number of classes in online inference.In this paper, we proposed a novel architecture, LookHD, which enables real-time HDC learning on low-power edge devices. LookHD exploits computation reuse to memorize the encoding module and simplify its computation with single memory access. LookHD also address the inference scalability by exploiting HDC governing mathematics that compresses the HDC trained model into a single hypervector. We present how the proposed architecture can be implemented on the existing low power architectures: ARM processor and FPGA design. We evaluate the efficiency of the proposed approach on a wide range of practical classification problems such as activity recognition, face recognition, and speech recognition. Our evaluations show that LookHD can achieve, on average, $ 28.3\times$ faster and $ 97.4\times$ more energy-efficient training as compared to the state-of-the-art HDC implemented on the FPGA. Similarly, in the inference, LookHD is $ 2.2\times$ faster, $ 4.1\times$ more energy-efficient, and has $ 6.3\times$ smaller model size than the same state-of-the-art algorithms.
Mohsen Imani, Zhuowen Zou, Samuel Bosch, Sanjay Anantha Rao, Sahand Salamat, Venkatesh Kumar, Yeseong Kim, Tajana Rosing
HPCA7
2021 Massively Parallel Big Data Classification on a Programmable Processing In-Memory Architecture
abstract
With the emergence of Internet of Things, massive data created in the world pose huge technical challenges for efficient processing. Processing in-memory (PIM) technology has been widely investigated to overcome expensive data movements between processors and memory blocks. However, existing PIM designs incur large area overhead to enable computing capability via additional near-data processing cores and analog/mixed signal circuits. In this paper, we propose a new massively-parallel processing in-memory (PIM) architecture, called CHOIR, based on emerging nonvolatile memory technology for big data classification. Unlike existing PIM designs which demand large analog/mixed signal circuits, we support the parallel PIM instructions for conditional and arithmetic operations in an area-efficient way. As a result, the classification solution performs both training and testing on the PIM architecture by fully utilizing the massive parallelism. Our design significantly improves the performance and energy efficiency of the classification tasks by 123× and 52× respectively as compared to the state-of-the-art tree boosting library running on GPU.
Yeseong Kim, Mohsen Imani, Saransh Gupta, Minxuan Zhou, Tajana Rosing
ICCAD1
2021 Efficient Brain-Inspired Hyperdimensional Learning with Spatiotemporal Structured Data
abstract
Brain-inspired hyperdimensional (HD) computing is a new computing paradigm based on theoretical neuroscience to enable efficient learning. In HD computing, the original data are encoded to points in a high-dimensional space to perform learning with lightweight algebra. In this paper, we propose STEMHD that elicits key features from spatiotemporal data along with a hardware design that empowers computation reuse. Our evaluation shows that STEMHD successfully interprets structural data at a low cost achieving higher accuracy than the state-of-the-art methods. Our evaluation shows that STEMHD improves performance and energy efficiency during the model training by 16.3% and 19.7%, respectively, with a negligible accuracy loss of less than 0.25%. For the model inference, we observe the inference speedup of 1.96× on average.
Jiseung Kim 0005, Hyunsei Lee, Mohsen Imani, Yeseong Kim
MASCOTS4
2021 Scalable edge-based hyperdimensional learning system with brain-like neural adaptation
abstract
In the Internet of Things (IoT) domain, many applications are running machine learning algorithms to assimilate the data collected in the swarm of devices. Sending all data to the powerful computing environment, e.g., cloud, poses significant efficiency and scalability issues. A promising way is to distribute the learning tasks onto the IoT hierarchy, often referred to edge computing; however, the existing sophisticated algorithms such as deep learning are often overcomplex to run on less-powerful and unreliable embedded IoT devices. Hyperdimensional Computing (HDC) is a brain-inspired learning approach for efficient and robust learning on today's embedded devices. Encoding, or transforming the input data into high-dimensional representation, is the key first step of HDC before performing a learning task. All existing HDC approaches use a static encoder; thus, they still require very high dimensionality, resulting in significant efficiency loss for the edge devices with limited resources. In this paper, we have developed NeuralHD, a new HDC approach with a dynamic encoder for adaptive learning. Inspired by human neural regeneration study in neuroscience, NeuralHD identifies insignificant dimensions and regenerates those dimensions to enhance the learning capability and robustness. We also present a scalable learning framework to distribute NeuralHD computation over edge devices in IoT systems. Our solution enables edge devices capable of real-time learning from both labeled and unlabeled data. Our evaluation on a wide range of practical classification tasks shows that NeuralHD provides 5.7X and 6.1X (12.3X and 14.1X) faster and more energy-efficient training compared to the HD-based algorithms (DNNs) running on the same platform. NeuralHD also provides 4.2X and 11.6X higher robustness to noise in the unreliable network and hardware of IoT environments as compared to DNNs.
Zhuowen Zou, Yeseong Kim, Farhad Imani, Haleh Alimohamadi, Rosario Cammarota, Mohsen Imani
SC2
2020 GenieHD: Efficient DNA Pattern Matching Accelerator Using Hyperdimensional Computing
abstract
DNA pattern matching is widely applied in many bioinformatics applications. The increasing volume of the DNA data exacerbates the runtime and power consumption to discover DNA patterns. In this paper, we propose a hardware-software co-design, called GenieHD, which efficiently parallelizes the DNA pattern matching task. We exploit brain-inspired hyperdimensional (HD) computing which mimics pattern-based computations in human memory. We transform inherent sequential processes of the DNA pattern matching to highly-parallelizable computation tasks using HD computing. The proposed technique first encodes the whole genome sequence and target DNA pattern to high-dimensional vectors. Once encoded, a light-weight operation on the high-dimensional vectors can identify if the target pattern exists in the whole sequence. We also design an accelerator architecture which effectively parallelizes the HD-based DNA pattern matching while significantly reducing the number of memory accesses. The architecture can be implemented on various parallel computing platforms to meet target system requirements, e.g., FPGA for low-power devices and ASIC for high-performance systems. We evaluate GenieHD on practical large-size DNA datasets such as human and Escherichia Coli genomes. Our evaluation shows that GenieHD significantly accelerates the DNA matching procedure, e.g., 44.4× speedup and 54.1× higher energy efficiency as compared to a state-of-the-art FPGA-based design.
Yeseong Kim, Mohsen Imani, Niema Moshiri, Tajana Rosing
DATE1
2020 Deep Learning Acceleration with Neuron-to-Memory Transformation
abstract
Deep neural networks (DNN) have demonstrated effectiveness for various applications such as image processing, video segmentation, and speech recognition. Running state-of-theart DNNs on current systems mostly relies on either generalpurpose processors, ASIC designs, or FPGA accelerators, all of which suffer from data movements due to the limited on-chip memory and data transfer bandwidth. In this work, we propose a novel framework, called RAPIDNN, which performs neuron-to-memory transformation in order to accelerate DNNs in a highly parallel architecture. RAPIDNN reinterprets a DNN model and maps it into a specialized accelerator, which is designed using non-volatile memory blocks that model four fundamental DNN operations, i.e., multiplication, addition, activation functions, and pooling. The framework extracts representative operands of a DNN model, e.g., weights and input values, using clustering methods to optimize the model for in-memory processing. Then, it maps the extracted operands and their pre-computed results into the accelerator memory blocks. At runtime, the accelerator identifies computation results based on efficient in-memory search capability which also provides tunability of approximation to improve computation efficiency further. Our evaluation shows that RAPIDNN achieves 68.4×, 49.5× energy efficiency improvement and 48.1×, 10.9× speedup as compared to ISAAC and PipeLayer, the state-of-the-art DNN accelerators, while ensuring less than 0.5% quality loss.
Mohsen Imani, Mohammad Samragh Razlighi, Yeseong Kim, Saransh Gupta, Farinaz Koushanfar, Tajana Rosing
HPCA3
2020 SHEARer: highly-efficient hyperdimensional computing by software-hardware enabled multifold approximation
abstract
Hyperdimensional computing (HD) is an emerging paradigm for machine learning based on the evidence that the brain computes on high-dimensional, distributed, representations of data. The main operation of HD is encoding, which transfers the input data to hyperspace by mapping each input feature to a hypervector, followed by a bundling procedure that adds up the hypervectors to realize the encoding hypervector. The operations of HD are simple and highly parallelizable, but the large number of operations hampers the efficiency of HD in embedded domain. In this paper, we propose SHEARer, an algorithmhardware co-optimization to improve the performance and energy consumption of HD computing. We gain insight from a prudent scheme of approximating the hypervectors that, thanks to error resiliency of HD, has minimal impact on accuracy while provides high prospect for hardware optimization. Unlike previous works that generate the encoding hypervectors in full precision and then and then perform ex-post quantization, we compute the encoding hypervectors in an approximate manner that saves resources yet affords high accuracy. We also propose a novel FPGA architecture that achieves striking performance through massive parallelism with low power consumption. Moreover, we develop a software framework that enables training HD models by emulating the proposed approximate encodings. The FPGA implementation of SHEARer achieves an average throughput boost of 104,904× (15.7×) and energy savings of up to 56,044× (301×) compared to state-of-the-art encoding methods implemented on Raspberry Pi 3 (GeForce GTX 1080 Ti) using practical machine learning datasets.
Behnam Khaleghi, Sahand Salamat, Anthony Thomas, Fatemeh Asgarinejad, Yeseong Kim, Tajana Rosing
ISLPED5
2020 DUAL: Acceleration of Clustering Algorithms using Digital-based Processing In-Memory
abstract
Today's applications generate a large amount of data that need to be processed by learning algorithms. In practice, the majority of the data are not associated with any labels. Unsupervised learning, i.e., clustering methods, are the most commonly used algorithms for data analysis. However, running clustering algorithms on traditional cores results in high energy consumption and slow processing speed due to a large amount of data movement between memory and processing units. In this paper, we propose DUAL, a Digital-based Unsupervised learning AcceLeration, which supports a wide range of popular algorithms on conventional crossbar memory. Instead of working with the original data, DUAL maps all data points into high-dimensional space, replacing complex clustering operations with memory-friendly operations. We accordingly design a PIM-based architecture that supports all essential operations in a highly parallel and scalable way. DUAL supports a wide range of essential operations and enables in-place computations, allowing data points to remain in memory. We have evaluated DUAL on several popular clustering algorithms for a wide range of large-scale datasets. Our evaluation shows that DUAL provides a comparable quality to existing clustering algorithms while using a binary representation and a simplified distance metric. DUAL also provides 58.8× speedup and 251.2× energy efficiency improvement as compared to the state-of-the-art solution running on GPU.
Mohsen Imani, Saikishan Pampana, Saransh Gupta, Minxuan Zhou, Yeseong Kim, Tajana Rosing
MICRO5
2019 A Framework for Collaborative Learning in Secure High-Dimensional Space
abstract
As the amount of data generated by the Internet of the Things (IoT) devices keeps increasing, many applications need to offload computation to the cloud. However, it often entails risks due to security and privacy issues. Encryption and decryption methods add to an already significant computational burden. In this paper, we propose a novel framework, called SecureHD, which provides a secure learning solution based on the idea of high-dimensional (HD) computing. We encode original data into secure, high-dimensional vectors. The training is performed with the encoded vectors. Thus, applications can send their data to the cloud with no security concerns, while the cloud can perform the offloaded tasks without additional decryption steps. In particular, we propose a novel HD-based classification algorithm which is suitable to handle a large amount of data that the cloud typically processes. In addition, we also show how SecureHD can recover the encoded data in a lossless manner. In our evaluation, we show that the proposed SecureHD framework can perform the encoding and decoding tasks 145.6× and 6.8× faster than a state-of-the-art encryption/decryption library running on the contemporary CPU. In addition, our learning method achieves high accuracy of 95% on average for diverse practical classification tasks including cloud-scale datasets.
Mohsen Imani, Yeseong Kim, M. Sadegh Riazi, John Messerly, Patric Liu, Farinaz Koushanfar, Tajana Rosing
CLOUD2
2019 GRAM: graph processing in a ReRAM-based computational memory
abstract
The performance of graph processing for real-world graphs is limited by inefficient memory behaviours in traditional systems because of random memory access patterns. Offloading computations to the memory is a promising strategy to overcome such challenges. In this paper, we exploit the resistive memory (ReRAM) based processing-in-memory (PIM) technology to accelerate graph applications. The proposed solution, GRAM, can efficiently executes vertex-centric model, which is widely used in large-scale parallel graph processing programs, in the computational memory. The hardware-software co-design used in GRAM maximizes the computation parallelism while minimizing the number of data movements. Based on our experiments with three important graph kernels on seven real-world graphs, GRAM provides 122.5X and 11.1x speedup compared with an in-memory graph system and optimized multithreading algorithms running on a multi-core CPU. Compared to a GPU-based graph acceleration library and a recently proposed PIM accelerator, GRAM improves the performance by 7.1X and 3.8X respectively.
Minxuan Zhou, Mohsen Imani, Saransh Gupta, Yeseong Kim, Tajana Rosing
ASP-DAC4
2019 HDCluster: An Accurate Clustering Using Brain-Inspired High-Dimensional Computing
abstract
Internet of things has increased the rate of data generation. Clustering is one of the most important tasks in this domain to find the latent correlation between data. However, performing today's clustering tasks is often inefficient due to the data movement cost between cores and memory. We propose HDCluster, a brain-inspired unsupervised learning algorithm which clusters input data in a high-dimensional space by fully mapping and processing in memory. Instead of clustering input data in either fixed-point or floating-point representation, HDCluster maps data to vectors with dimension in thousands, called hypervectors, to cluster them. Our evaluation shows that HDCluster provides better clustering quality for the tasks that involve a large amount of data while providing a potential for accelerating in a memory-centric architecture.
Mohsen Imani, Yeseong Kim, Thomas Worley, Saransh Gupta, Tajana Rosing
DATE2
2019 Application Performance Prediction and Optimization Under Cache Allocation Technology
abstract
Many applications running on high-performance computing systems share limited resources such as the last-level cache, often resulting in lower performance. Intel recently introduced a new control mechanism, called cache allocation technology (CAT), which controls the cache size used by each application. To intelligently utilize this technology for automated management, it is essential to accurately identify application performance behavior for different cache allocation scenarios. In this work, we show a novel approach which automatically builds a prediction model for application performance changes with CAT. We profile the workload characteristics based on Intel Top-down Microarchitecture Analysis Method (TMAM), and train the model using machine learning. The model predicts instructions per cycle (IPC) across available cache sizes allocated for the applications. We also design a dynamic cache management technique which utilizes the prediction model and intelligently partitions the cache resource to improve application throughput. We implemented and evaluated the proposed framework in Intel PMU profiling tool running on Xeon Platinum 8186 Skylake processor. In our evaluation, we show that the proposed model accurately predicts the IPC changes of applications with 4.7% error on average for different cache allocation scenarios. Our predictive online cache managements achieves improvements on application performance of up to 25% as compared to a prediction-agnostic policy.
Yeseong Kim, Ankit More, Emily Shriver, Tajana Rosing
DATE1
2019 DigitalPIM: Digital-based Processing In-Memory for Big Data Acceleration
abstract
In this work, we design, DigitalPIM, a Digital-based Processing In-Memory platform capable of accelerating fundamental big data algorithms in real time with orders of magnitude more energy efficient operation. Unlike the existing near-data processing approach such as HMC 2.0, which utilizes additional low-power processing cores next to memory blocks, the proposed platform implements the entire algorithm directly in memory blocks without using extra processing units. In our platform, each memory block supports the essential operations including: bitwise operation, addition/multiplication, and search operation internally in memory without reading any values out of the block. This significantly mitigates the processing costs of the new architecture, while providing high scalability and parallelism for performing the extensive computations. We exploit these essential operations to accelerate popular big data applications entirely in memory such as machine learning algorithms, query processing, and graph processing. Our evaluations show that for all tested applications, the performance can be accelerated significantly by eliminating the memory access bottleneck
Mohsen Imani, Saransh Gupta, Yeseong Kim, Minxuan Zhou, Tajana Rosing
ACM Great Lakes Symposium on VLSI3
2019 UPIM: Unipolar Switching Logic for High Density Processing-in-Memory Applications
abstract
Internet of Things (IoT) has built a network with billions of connected devices which generate massive volumes of data. Processing large data on existing systems requires significant costs for data movements between processors and memory due to limited cache capacity and memory bandwidth. Processing-In-Memory (PIM) is a promising solution to address the issue. Prior techniques that enable the computation in non-volatile memory (NVM) are designed on a bipolar switching mode, which suffers from a high sneak current in a crossbar array (CBA) structure. In this paper, we propose a unipolar-switching logic for high-density PIM applications, called UPIM. Our design exploits a unipolar-switching mode of memristor devices which can be operated in 1D1R structure hence suppresses the sneak current that exists in prior PIM technologies. Moreover, UPIM takes advantages of a 3D vertical crossbar array (CBA) structure to increase memory utilization per unit area for high-density applications. Our evaluation on a wide range of applications shows that the UPIM achieves up to 31.3× energy saving and 113.8× energy-delay product (EDP) improvement as compared to a recent GPGPU architecture. As compared to the state-of-the-art PIM design based on the bipolar switching mode, our design achieves 3.1× lower energy consumption.
Joonseop Sim, Saransh Gupta, Mohsen Imani, Yeseong Kim, Tajana Rosing
ACM Great Lakes Symposium on VLSI4
2019 FloatPIM: in-memory acceleration of deep neural network training with high precision
abstract
Processing In-Memory (PIM) has shown a great potential to accelerate inference tasks of Convolutional Neural Network (CNN). However, existing PIM architectures do not support high precision computation, e.g., in floating point precision, which is essential for training accurate CNN models. In addition, most of the existing PIM approaches require analog/mixed-signal circuits, which do not scale, exploiting insufficiently reliable multi-bit Non-Volatile Memory (NVM). In this paper, we propose FloatPIM, a fully-digital scalable PIM architecture that accelerates CNN in both training and testing phases. FloatPIM natively supports floating-point representation, thus enabling accurate CNN training. FloatPIM also enables fast communication between neighboring memory blocks to reduce internal data movement of the PIM architecture. We evaluate the efficiency of FloatPIM on ImageNet dataset using popular large-scale neural networks. Our evaluation shows that FloatPIM supporting floating point precision can achieve up to 5.1% higher classification accuracy as compared to existing PIM architectures with limited fixed-point precision. FloatPIM training is on average 303.2× and 48.6× (4.3× and 15.8×) faster and more energy efficient as compared to GTX 1080 GPU (PipeLayer [1] PIM accelerator). For testing, FloatPIM also provides 324.8× and 297.9× (6.3× and 21.6×) speedup and energy efficiency as compared to GPU (ISAAC [2] PIM accelerator) respectively.
Mohsen Imani, Saransh Gupta, Yeseong Kim, Tajana Rosing
ISCA3
2019 ROAD: Routability Analysis and Diagnosis Framework Based on SAT Techniques
abstract
Routability diagnosis has increasingly become the bottleneck in detailed routing for sub-10nm technology due to the limited tracks, high density, and complex design rules. The conventional ways to examine the routability of detailed routing are ILP- and SAT-based techniques. However, once we identify the routability, the diagnosis remains an open problem for physical designers. In this paper, we propose a novel framework, called ROAD, which diagnoses explicit reasons for routing failures. The proposed ROAD framework utilizes a diagnosis-friendly SAT formulation to represent design's layout and diagnoses the routability with SAT solving techniques. Based on the diagnosis, ROAD provides human-interpretable explanations for conflicted routing conditions. To show the practical value of our framework, we also generate comprehensive test-sets that enable exhaustive exploration of layouts based on Rent's rule. We demonstrate that ROAD successfully examines conflict causes for diverse pin layouts. Throughout extensive diagnosis, we also present several key findings for design failure. ROAD performs routability diagnosis within 2 minutes on average for 90 grids testsets, while diagnosing the exact causes of routing failures in terms of congestion and conditional design rules.
Dongwon Park, Ilgweon Kang, Yeseong Kim, Sicun Gao, Bill Lin 0001, Chung-Kuan Cheng
ISPD3
2017 MPIM: Multi-purpose in-memory processing using configurable resistive memory
abstract
Running Internet of Things applications on general purpose processors results in a large energy and performance overhead, due to the high cost of data movement. Processing in-memory is a promising solution to reduce the data movement cost by processing the data locally inside the memory. In this paper, we design a Multi-Purpose In-Memory Processing (MPIM) system, which can be used as main memory and for processing. MPIM consists of multiple crossbar memories with the capability of efficient in-memory computations. Instead of transferring the large dataset to the processors, MPIM provides two important in-memory processing capabilities: i) data searching for the nearest neighbor ii) bitwise operations including OR, AND and XOR with small analog sense amplifiers. The experimental results show that the MPIM can achieve up to 5.5× energy savings and 19× speedup for the search operations as compared to AMD GPU-based implementation. For bitwise vector processing, we present 11000× energy improvements with 62× speedup over the SIMD-based computation, while outperforming other state-of-the-art in-memory processing techniques.
Mohsen Imani, Yeseong Kim, Tajana Rosing
ASP-DAC2
2017 Efficient neural network acceleration on GPGPU using content addressable memory
abstract
Recently, neural networks have been demonstrated to be effective models for image processing, video segmentation, speech recognition, computer vision and gaming. However, high energy computation and low performance are the primary bottlenecks of running the neural networks. In this paper, we propose an energy/performance-efficient network acceleration technique on General Purpose GPU (GPGPU) architecture which utilizes specialized resistive nearest content addressable memory blocks, called NNCAM, by exploiting computation locality of the learning algorithms. NNCAM stores highly frequent patterns corresponding to neural network operations and searches for the most similar patterns to reuse the computation results. To improve NNCAM computation efficiency and accuracy, we proposed layer-based associative update and selective approximation techniques. The layer-based update improves data locality of NNCAM blocks by filling NNCAM values based on the frequent computation patterns of each neural network layer. To guarantee the appropriate level of computation accuracy while providing maximum energy saving, our design adaptively allocates the neural network operations to either NNCAM or GPGPU floating point units (FPUs). The selective approximation relaxes computation on neural network layers by considering the impact on accuracy. In evaluation, we integrate NNCAM blocks with the modern AMD southern Island GPU architecture. Our experimental evaluation shows that the enhanced GPGPU can result in 68% energy savings and 40% speedup running on four popular convolutional neural networks (CNN), ensuring acceptable < 2% quality loss.
Mohsen Imani, Daniel Peroni, Yeseong Kim, Abbas Rahimi, Tajana Rosing
DATE3
2017 ORCHARD: Visual object recognition accelerator based on approximate in-memory processing
abstract
In recent years, machine learning for visual object recognition has been applied to various domains, e.g., autonomous vehicle, heath diagnose, and home automation. However, the recognition procedures still consume a lot of processing energy and incur a high cost of data movement for memory accesses. In this paper, we propose a novel hardware accelerator design, called ORCHARD, which processes the object recognition tasks inside memory. The proposed design accelerates both the image feature extraction and boosting-based learning algorithm, which are key subtasks of the state-of-the-art image recognition approaches. We optimize the recognition procedures by leveraging approximate computing and emerging non-volatile memory (NVM) technology. The NVM-based in-memory processing allows the proposed design to mitigate the CMOS-based computation overhead, highly improving the system efficiency. In our evaluation conducted on circuit- and device-level simulations, we show that ORCHARD successfully performs practical image recognition tasks, including text, face, pedestrian, and vehicle recognition with 0.3% of accuracy loss made by computation approximation. In addition, our design significantly improves the performance and energy efficiency by up to 376x and 1896x, respectively, compared to the existing processor-based implementation.
Yeseong Kim, Mohsen Imani, Tajana Rosing
ICCAD1
2017 P4: Phase-based power/performance prediction of heterogeneous systems via neural networks
abstract
The emergence of Internet of Things increases the complexity and the heterogeneity of computing platforms. Migrating workload between various platforms is one way to improve both energy efficiency and performance. Effective migration decisions require accurate estimates of its costs and benefits. To date, these estimates were done by either instrumenting the source code/binaries, thus causing high overhead, or by using power estimates from hardware performance counters, which work well for individual machines, but until now have not been accurate for predicting across different architectures. In this paper, we propose P4, a new Phase-based Power and Performance Prediction framework which identifies cross-platform application power and performance at runtime for heterogeneous computing systems. P4analyzes and detects machine-independent application phases by characterizing computing platforms offline with a set of benchmarks, and then builds neural network-based models to automatically identify and generalize the complex cross-platform relationships for each benchmark phase. It then leverages these models along with performance counter measurements collected at runtime to estimate performance and power consumption if it were running on a completely different computing platform, including a different CPU architecture, without ever having to run it on there. We evaluate the proposed framework on four commercial heterogeneous platforms, ranging from X86 servers to mobile ARM-based architecture, with 129 industry-standard benchmarks. Our experimental results show that P4can predict the power and performance changes with only 6.8% and 5.6% error, respectively, even for completely different architectures from the ones applications ran on.
Yeseong Kim, Pietro Mercati, Ankit More, Emily Shriver, Tajana Rosing
ICCAD1
2017 Enabling efficient system design using vertical nanowire transistor current mode logic
abstract
Vertical Nanowire-FET (VNFET) is a promising candidate to succeed in industry mainstream due to its superior suppression of short-channel-effects and area efficiency. However, to design logic gates, CMOS is not an appropriate solution due to the process incompatibility with VNFET, which creates a technical challenge for mass production. In this work, we propose a novel VNFET-based logic design, called VnanoCML (Vertical Nanowire Transistor-based Current Mode Logic), which addresses the process issue while significantly improving power and performance of diverse logic designs. Unlike the CMOS-based logic, our design exploits current mode logic to overcome the fabrication issue. Furthermore, we reduce drain-to-source resistance of VnanoCML, which results in higher performance improvement without compromising the subthreshold swing. In order to show the impact of the proposed VnanoCML, we present key logic designs which are SRAM, full adder and multiplier, and also evaluate the application-level effectiveness of digital designs for image processing and mathematical computation. Our proposed design improves the fundamental circuit characteristics including output swing, delay time and power consumption compared to conventional planar MOSFET (PFET)-based circuits. Consequentially our architecture-level results show that VnanoCML can enhance the performance and power by 16.4× and 1.15×, respectively. Furthermore, we show that VnanoCML improves the energy-delay product by 38.5× on average compared to PFET-based designs.
Joonseop Sim, Mohsen Imani, Yeseong Kim, Tajana Rosing
VLSI-SoC3
2016 ACAM: Approximate Computing Based on Adaptive Associative Memory with Online Learning
abstract
The Internet of Things (IoT) dramatically increases the amount of data to be processed for many applications including multimedia. Unlike traditional computing environment, the workload of IoT significantly varies overtime. Thus, an efficient runtime profiling is required to extract highly frequent computations and pre-store them for memory-based computing. In this paper, we propose an approximate computing technique using a low-cost adaptive associative memory, named ACAM, which utilizes runtime learning and profiling. To recognize the temporal locality of data in real-world applications, our design exploits a reinforcement learning algorithm with a least recently use (LRU) strategy to select images to be profiled; the profiler is implemented using an approximate concurrent state machine. The profiling results are then stored into ACAM for computation reuse. Since the selected images represent the observed input dataset, we can avoid redundant computations thanks to high hit rates displayed in the associative memory. We evaluate ACAM on the recent AMD Southern Island GPU architecture, and the experimental results shows that the proposed design achieves by 34.7% energy saving for image processing applications with an acceptable quality of service (i.e., PSNR>30dB).
Mohsen Imani, Yeseong Kim, Abbas Rahimi, Tajana Rosing
ISLPED2
2016 A Personalized Network Activity-Aware Approach to Reducing Radio Energy Consumption of Smartphones
abstract
The radio energy consumption takes a large portion of the total energy consumption in smartphones. However, a significant portion of radio energy is wasted in a special waiting interval, known as the tail time after a transmission is completed while waiting for a subsequent transmission. In order to reduce the wasted energy in the tail time, the fast dormancy feature allows a quick release of a radio connection in the tail time. For supporting the fast dormancy efficiently, it is important to accurately predict whether a subsequent transmission will occur in the tail time. In this paper, we show that there are strong personal characteristics on how user interacts with a radio network within the tail time. Based on these observations, we propose a novel personalized network activity-aware predictive dormancy technique, called Personalized Diapause (pD). By automatically identifying user-specific tail-time transmission characteristics for various network activities, our proposed technique takes advantages of personalized high-level network usage patterns in deciding when to release radio connections. Our experimental results using real network usage logs from 25 users show that pD can reduce the amount of the wasted tail time energy by 51 percent on average, thus saving the total radio energy consumption by 23 percent with less than 10 percent reconnection increase.
Yeseong Kim, Boyeong Jeon, Jihong Kim 0001
IEEE Trans. Mob. Comput.1
2015 CAUSE: Critical Application Usage-Aware Memory System using Non-volatile Memory for Mobile Devices
abstract
Mobile devices are severely limited in memory, which affects critical user-experience metrics such as application service time. Emerging non-volatile memory (NVM) technologies such as STT-RAM and PCM are ideal candidates to provide higher memory capacity with negligible energy overhead. However, existing memory management systems overlook mobile users application usage which provides crucial cues for improving user experience. In this paper, we propose CAUSE, a novel memory system based on DRAM-NVM hybrid memory architecture. CAUSE takes explicit account of the application usage patterns to distinguish data criticality and identify suitable swap candidates. We also devise NVM hardware design optimized for the access characteristics of the swapped pages. We evaluate CAUSE on a real Android smartphone and NVSim simulator using user application usage logs. Our experimental results show that the proposed technique achieves 32% faster launch time for mobile applications while reducing energy cost by 90% and 44% on average over non-optimized STT-RAM and PCM, respectively.
Yeseong Kim, Mohsen Imani, Shruti Patil, Tajana Rosing
ICCAD1
2015 Smartphone Analysis and Optimization based on User Activity Recognition
abstract
Behavior of smartphone systems is highly influenced by user interactions, such as `zooming' and `scrolling', which determine the execution phases within applications and lead to different power and performance demands. Current power and thermal management algorithms are agnostic to these behaviors. We propose a novel user activity recognition framework that enables user activity-aware system decisions. The proposed framework carefully monitors system events initiated by user interactions and identifies the current user activity based on an online activity model. We implemented the proposed framework in Android platform, and tested it on Qualcomm MDP 8660 smartphone. To show the practical value of our recognition strategy, we design effective power and thermal management policies that adapt system settings to user activity changes. Our experimental results using 10 real mobile applications show that the proposed proactive management technique can reduce the CPU energy by up to 28% while meeting a given thermal constraint.
Yeseong Kim, Francesco Paterna, Sameer Tilak, Tajana Rosing
ICCAD1
2014 Personalized optimization for android smartphones
abstract
As a highly personalized computing device, smartphones present a unique new opportunity for system optimization. For example, it is widely observed that a smartphone user exhibits very regular application usage patterns (although different users are quite different in their usage patterns). User-specific high-level app usage information, when properly managed, can provide valuable hints for optimizing various system design requirements. In this article, we describe the design and implementation of a personalized optimization framework for the Android platform that takes advantage of user's application usage patterns in optimizing the performance of the Android platform. Our optimization framework consists of two main components, the application usage modeling module and the usage model-based optimization module. We have developed two novel application usage models that correctly capture typical smartphone user's application usage patterns. Based on the application usage models, we have implemented an app-launching experience optimization technique which tries to minimize user-perceived delays, extra energy consumption, and state loss when a user launches apps. Our experimental results on the Nexus S Android reference phones show that our proposed optimization technique can avoid unnecessary application restarts by up to 78.4% over the default LRU-based policy of the Android platform.
Wook Song, Yeseong Kim, Hakbong Kim, Jehun Lim, Jihong Kim 0001
ACM Trans. Embed. Comput. Syst.2