EDBT 2026 Demo / reviewers in the wild / expert
Minxuan Zhou
dblp:176/6533
· DBLP profile ↗
40ranked-venue papers
12as first author
33since 2021 · last 2026
0000-0002-5523-7270ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 12 first-author · 30 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on GraphsabstractAll-pairs shortest paths (APSP) remains a major bottleneck for large-scale graph analytics, as data movement with cubic complexity overwhelms the bandwidth of conventional memory hierarchies. We propose RAPID-Graph, a processing-in-memory (PIM) system co-designed across algorithm, architecture, and device levels to address this challenge. At the algorithm level, we introduce a recursion-aware partitioner that enables an exact APSP computation by decomposing graphs into vertex tiles to reduce data dependency, such that both Floyd-Warshall and Min-Plus kernels execute fully in-place within digital PIM arrays. At the architecture and device levels, we design a 2.5D PIM stack integrating two phase-change memory compute dies, a logic die, and high-bandwidth scratchpad memory within a unified advanced package. An external non-volatile storage stack stores large APSP results persistently. The design achieves both tile-level and unit-level parallel processing to sustain high throughput. On the 2.45M-node OGBN-Products dataset, RAPID-Graph is 5.8× faster and 1 186× more energy efficient than state-of-the-art GPU clusters, while exceeding prior PIM accelerators by 8.3× in speed and 104× in efficiency. It further delivers up to 42.8× speedup and 392× energy savings over an NVIDIA H100 GPU. Keming Fan, Runyang Tian, John Hsu, Minxuan Zhou, Tajana Rosing |
DATE | 7 |
| 2026 | FHEIns: Fully Homomorphic Encryption Acceleration for Large Data Applications with In-Storage ProcessingabstractRecently, the significance of data privacy protection has been growing rapidly. Homomorphic encryption (HE) enables computation directly on ciphertexts, making it attractive for privacy-sensitive databases in cloud datacenters. Although FHE enables privacy-preserving compute, ciphertext expansion and long-latency primitives drive up memory footprint and delay, worsening compute and memory pressure for database search. In practice, encrypted databases span hundreds of gigabytes to terabytes, making the storage I/O the dominant bottleneck. However, most prior FHE accelerators optimize on-chip computation and the main memory traffic while assuming working sets fit in HBM. Therefore, in this work, we present FHEIns, an in-storage processing architecture that executes FHE kernels close to data inside the NAND flash-based solid-state drives (SSDs) to exploit the internal bandwidth of the SSD. FHEIns achieves up to 24.7× and 2.67× speedup compared to the state-of-the-art FHE ASIC accelerators on trending FHE-based database benchmarks. Xuan Wang 0040, Keming Fan, Augusto Vega, Minxuan Zhou, Tajana Rosing |
DATE | 5 |
| 2026 | PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAMabstractAll-pairs shortest paths is a fundamental algorithm used for routing, logistics, and network analysis, but the cubic time complexity and heavy data movement of the canonical Floyd-Warshall algorithm severely limits its scalability on conventional CPUs or GPUs. In this paper, we propose PIM-FW, a novel co-designed hardware architecture and dataflow leveraging processing in and near memory to accelerate the blocked FW algorithm on an HBM3 stack. To enable fine-grained parallelism, we propose a massively parallel array of specialized bit-serial bank and channel PEs designed to accelerate core min-plus operations. Our dataflow complements this hardware, employing an interleaved mapping policy for superior load balancing and a hybrid memory computing model for efficient computation and reduction. This in-bank computing approach allows all distance updates to be performed and stored locally, a key contribution which eliminates the data-movement bottleneck inherent in GPU-based approaches. We implement a full hardware-software co-design using a cycle-accurate simulator for an 8-channel, 4-Hi HBM3 stack on real road-network traces. Experimental results show that, for an 8, 192 × 8, 192 graph, PIM-FW achieves an 18.7 × speedup and consumes 3200 × lower memory-stack/accelerator-side energy under our modeled PIM stack assumptions compared to a state-of-the-art GPU-only Floyd-Warshall. Tsung-Han Lu, Minxuan Zhou, John Hsu, Tajana Rosing |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | SoK: Can Fully Homomorphic Encryption Support General AI Computation? A Functional and Cost AnalysisabstractArtificial intelligence (AI) increasingly powers sensitive applications in domains such as healthcare and finance, relying on both extit{linear operations} (e.g., matrix multiplications in large language models) and extit{non-linear operations} (e.g., sorting in retrieval-augmented generation). Fully homomorphic encryption (FHE) has emerged as a promising tool for privacy-preserving computation, but it remains unclear whether existing methods can support the full spectrum of AI workloads that combine these operations. In this SoK, we ask: extit{Can FHE support general AI computation?} We provide both a functional analysis and a cost analysis. First, we categorize ten distinct FHE approaches and evaluate their ability to support general computation. We then identify three promising candidates and benchmark workloads that mix linear and non-linear operations across different bit lengths and SIMD parallelization settings. Finally, we evaluate five real-world, privacy-sensitive AI applications that instantiate these workloads. Our results quantify the costs of achieving general computation in FHE and offer practical guidance on selecting FHE methods that best fit specific AI application requirements. Our codes are available at https://github.com/UCF-ML-Research/FHE-AI-Generality. Wei Zhang 0076, Mengxin Zheng, Minxuan Zhou, Yushun Dong, Dongjie Wang 0001, Jiafeng Xie, David Mohaisen, Hongyi Wu, Qian Lou |
Proc. Priv. Enhancing Technol. | 6 |
| 2026 | Proxima: Near-Storage Acceleration for Graph-Based Approximate Nearest Neighbor Search in 3D NANDabstractApproximate nearest neighbor search (ANNS) plays an indispensable role in a wide variety of applications, including recommendation systems, information retrieval, and semantic search. Among the cutting-edge ANNS algorithms, graph-based approaches provide superior accuracy and scalability on massive datasets. However, the best-performing graph-based ANNS solutions incur tens of hundreds of memory footprints as well as costly distance computation, thus hindering their efficient deployment at scale. The 3D NAND flash is emerging as a promising device for data-intensive applications due to its high density and nonvolatility. In this work, we present the near-storage processing (NSP)-based ANNS solution Proxima to accelerate graph-based ANNS with algorithm-hardware co-design in 3D NAND flash. Proxima significantly reduces the complexity of graph search by leveraging the distance approximation and early termination. On top of the algorithmic enhancement, we implement the Proxima search algorithm in 3D NAND flash using the heterogeneous integration technique. To maximize 3D NAND’s bandwidth utilization, we present a customized dataflow and optimized data allocation scheme. Our evaluation results show that, compared to graph ANNS on CPU and GPU, Proxima achieves a magnitude improvement in throughput or energy efficiency. Proxima yields 7× to 13× speedup over existing ASIC designs. Furthermore, Proxima achieves a good balance between accuracy, efficiency, and storage density compared to previous NSP-based accelerators. Po-Kai Hsu, Jaeyoung Kang 0001, Minxuan Zhou, Sumukh Pinge, Shimeng Yu, Tajana Rosing |
IEEE Trans. Computers | 5 |
| 2025 | Rhychee-FL: Robust and Efficient Hyperdimensional Federated Learning with Homomorphic Encryption
Yujin Nam, Abhishek Moitra, Yeshwanth Venkatesha, Xiaofan Yu 0001, Gabrielle De Micheli, Xuan Wang 0040, Minxuan Zhou, Augusto Vega, Priyadarshini Panda, Tajana Rosing |
DATE | 7 |
| 2025 | Efficient Privacy-Preserving Recommendation on Sparse Data using Fully Homomorphic EncryptionabstractIn today’s data-driven world, recommendation systems personalize user experiences across industries but rely on sensitive data, raising privacy concerns. Fully homomorphic encryption (FHE) can secure these systems, but a significant challenge in applying FHE to recommendation systems is efficiently handling the inherently large and sparse user-item rating matrices. FHE operations are computationally intensive, and naively processing various sparse matrices in recommendation systems would be prohibitively expensive. Additionally, the communication overhead between parties remains a critical concern in encrypted domains. We propose a novel approach combining Compressed Sparse Row (CSR) representation with FHE-based matrix factorization that efficiently handles matrix sparsity in the encrypted domain while minimizing communication costs. Our experimental results demonstrate high recommendation accuracy with encrypted data while achieving the lowest communication costs, effectively preserving user privacy. Moontaha Nishat Chowdhury, André Bauer 0001, Minxuan Zhou |
eScience | 3 |
| 2025 | PATHE: A Privacy-Preserving Database Pattern Search Platform with Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) enables secure computation on encrypted data without decryption, allowing a great opportunity for privacy-preserving computation. Many companies maintain extensive, high-quality databases to deliver services, making preserving data privacy during the database pattern searches crucial. With FHE, the server can take encrypted queries from clients and search through the reference database on the server without decryption, thus guaranteeing data security for all parties. While FHE provides a promising solution to data privacy, it has severe drawbacks of explosive memory requirements and excessive latency, which amplify the computational and memory inefficiencies for database search applications.To address these, we propose PATHE that exploits FHE and hyperdimensional computing (HDC), which provides high parallelism, excellent robustness to errors, for high-performance privacy-preserving database search. On the software side, we propose an FHE-friendly PATHE algorithm that leverages efficient FHE-HDC search and a scheme-switching-based argmax to support database search and maintain comparable accuracy to the state-of-the-art. On the hardware side, PATHE proposes an efficient and scalable FHE accelerator system using Compute Express Link (CXL) for large-scale FHE database search, along with a novel, storage-aware dataflow designed to optimize memory and storage transfers for large database workloads. We evaluate PATHE on the large-scale encrypted database of protein mass spectra, PATHE achieves 2.1× speedup and 1.7× better energy efficiency compared to the baseline system. Xuan Wang 0040, Minxuan Zhou, Gabrielle De Micheli, Yujin Nam, Sumukh Pinge, Augusto Vega, Tajana Rosing |
ICCAD | 2 |
| 2025 | HPVM-HDC: A Heterogeneous Programming System for Accelerating Hyperdimensional ComputingabstractHyperdimensional Computing (HDC), a technique inspired by cognitive models of computation, has been proposed as an efficient and robust alternative basis for machine learning.HDC programs are often manually written in low-level and target specific languages targeting CPUs, GPUs, and FPGAs-these codes cannot be easily retargeted onto HDC-specific accelerators.No previous programming system enables productive development of HDC programs and generates efficient code for several hardware targets.We propose a heterogeneous programming system for HDC: a novel programming language, HDC++, for writing applications using a unified programming model, including HDC-specific primitives to improve programmability, and a heterogeneous compiler, HPVM-HDC, that provides an intermediate representation for compiling HDC programs to many hardware targets.We implement two tuning optimizations, automatic binarization and reduction perforation, that exploit the error resilient nature of HDC.Our evaluation shows that HPVM-HDC generates performance-competitive code for CPUs and GPUs, achieving a geomean speed-up of 1.17x over optimized baseline CUDA implementations with a geomean * Equally contributing authors. Russel Arbore, Xavier Routh, Abdul Rafae Noor, Akash Kothari, Haichao Yang, Sumukh Pinge, Minxuan Zhou, Tajana Rosing, Vikram S. Adve |
ISCA | 8 |
| 2025 | OptiPIM: Optimizing Processing-in-Memory Acceleration Using Integer Linear ProgrammingabstractProcessing-in-memory (PIM) accelerators provide superior performance and energy efficiency to conventional architectures by minimizing off-chip data movement and exploiting extensive internal memory bandwidth for computation.However, efficient PIM acceleration requires careful software-hardware mapping that transforms application algorithms into PIM operations and data layout.Unfortunately, existing PIM accelerators adopt manually tuned heuristics or exhaustive search to determine the mappings on PIM accelerators, leading to under-optimized performance and/or long optimization time.In this work, we propose OptiPIM, a novel optimization framework based on Integer Linear Programming (ILP) to efficiently generate the optimal mapping for data-intensive applications on PIM accelerators.The proposed framework adopts a PIM-friendly mapping representation with accurate cost modeling and a concise description of the entire design space, allowing us to formulate an efficient and effective ILP problem and optimize the mapping on PIM architectures.We implement OptiPIM in the opensource MLIR framework, enabling OptiPIM to generate optimized mappings for PyTorch workloads on PIM accelerators.We evaluate widely used machine learning workloads on two state-of-the-art PIM accelerators.Our experiments show that OptiPIM can generate optimal mappings within 4 minutes.Mappings generated by OptiPIM are at least 1.9× faster than those generated by heuristics. Minxuan Zhou, Yue Pan 0002, Chien-Yi Yang, Lana Josipovic, Tajana Rosing |
ISCA | 2 |
| 2025 | Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE ServingabstractAs Large Language Models (LLMs) continue to evolve, Mixture of Experts (MoE) architecture has emerged as a prevailing design for achieving state-of-the-art performance across a wide range of tasks.MoE models use sparse gating to activate only a handful of expert sub-networks per input, achieving billion-parameter capacity with inference costs akin to much smaller models.However, such models often pose challenges for hardware deployment due to the massive data volume introduced by the MoE layers.To address the challenges of serving MoE models, we propose Stratum, a system-hardware co-design approach that combines the novel memory technology Monolithic 3D-Stackable DRAM (Mono3D DRAM), near-memory processing (NMP), and GPU acceleration.The logic and Mono3D DRAM dies are connected through hybrid bonding, whereas the Mono3D DRAM stack and GPU are interconnected via silicon interposer.Mono3D DRAM offers higher internal bandwidth than HBM thanks to the dense vertical interconnect pitch enabled by its monolithic structure, which supports implementations of higher-performance near-memory processing.Furthermore, we tackle the latency differences introduced by aggressive vertical scaling of Mono3D DRAM along the 𝑧-dimension by constructing internal memory tiers and assigning data across layers based on * Equal contribution Yue Pan 0009, Zihan Xia 0002, Po-Kai Hsu, Lanxiang Hu, Hyungyo Kim, Janak Sharda, Minxuan Zhou, Nam Sung Kim, Shimeng Yu, Tajana Rosing, Mingu Kang |
MICRO | 7 |
| 2025 | RelHDx: Hyperdimensional Computing for Learning on Graphs With FeFET AccelerationabstractGraph neural networks (GNNs) are a powerful machine learning (ML) method to analyze graph data. The training of GNN has compute and memory-intensive phases along with irregular data movements, which makes in-memory acceleration challenging. We present a hyperdimensional computing (HDC)-based graph ML framework called RelHDx that aggregates node features and graph structure, along with representing node and edge information in high-dimensional space. RelHDx enables single-pass training and inference with simple arithmetic operations, resulting in the efficient design of graph-based ML tasks: node classification and link prediction. We accelerate RelHDx using scalable processing in-memory (PIM) architecture based on emerging ferroelectric FET (FeFET) technology. Our accelerator uses a data allocation optimization and operation scheduler to address the irregularity of the graph and maximize the performance. Evaluation results show that RelHDx offers comparable accuracy to popular GNN-based algorithms while achieving up to$63.8\boldsymbol{\times}$faster speed on GPU. Our FeFET-based accelerator, RelHDx-PIM, is$32\boldsymbol{\times}$faster for node classification, while for link prediction it is$65.4\boldsymbol{\times}$faster than when running on GPU. Furthermore, RelHDx-PIM improves energy efficiency by four orders of magnitude over GPU. Compared to the state-of-the-art in-memory processing-based GNN accelerator, PIM-GCN[1], RelHDx-PIM is$10\boldsymbol{\times}$faster and$986\boldsymbol{\times}$more energy-efficient on average. Jaeyoung Kang 0001, Minxuan Zhou, Tajana Rosing |
IEEE Trans. Computers | 2 |
| 2025 | Fast-OverlaPIM: A Fast Overlap-Driven Mapping Framework for Processing In-Memory Neural Network AccelerationabstractProcessing in-memory (PIM) is promising to accelerate neural networks (NNs) because it minimizes data movement and provides large computational parallelism. Similar to machine learning accelerators, application mapping, which determines the operation scheduling and data layout, plays a critical role in the NN acceleration on PIM. The mapping optimization of the previous NN accelerators focused on optimizing the latency of sequential execution. However, PIM accelerators feature a distinct design space of application mapping from conventional NN accelerators, due to the spatial execution of NN layers across different memory locations. This enables opportunities for overlapping execution of consecutive NN layers to improve the latency, where the succeeding layer can start execution before the preceding layer fully completes the computation. In this article, we propose Fast-OverlaPIM framework that incorporates computational overlapping optimization into the deep neural network mapping exploration process on PIM architectures. Fast-OverlaPIM includes analytical algorithms for fast and accurate overlap analysis. Furthermore, it proposes a novel mapping search strategy and a transformation mechanism to enable efficient design space exploration on the overlap-based mapping for the whole network. Our framework demonstrates a significant improvement in runtime performance from$3.4\times $to$323.1\times $compared to the previous state-of-the-art overlap-based framework. Our experiments show that Fast-OverlaPIM can efficiently produce mappings that are$4.6\times $to$18.1\times $faster than the state-of-the-art mapping optimization framework under the same architecture constraints. Xuan Wang 0040, Minxuan Zhou, Tajana Rosing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | PRIMATE: Processing in Memory Acceleration for Dynamic Token-pruning TransformersabstractAttention-based models such as Transformers represent the state of the art for various machine learning (ML) tasks. Their superior performance is often overshadowed by the substantial memory requirements and low data reuse opportunities. Processing in Memory (PIM) is a promising solution to accelerate Transformer models due to its massive parallelism, low data movement costs, and high memory bandwidth utilization. Existing PIM accelerators lack the support for algorithmic optimizations like dynamic token pruning that can significantly improve the efficiency of Transformers. We identify two challenges to enabling dynamic token pruning on PIM-based architectures: the lack of an in-memory top-k token selection mechanism and the memory underutilization problem from pruning. To address these challenges, we propose PRIMATE, a software-hardware co-design PIM framework based on High Bandwidth Memory (HBM). We initiate minor hardware modifications to conventional HBM to enable Transformer model computation and top-k selection. For software, we introduce a pipelined mapping scheme and an optimization framework for maximum throughput and efficiency. PRIMATE achieves $30.6\times$ improvement in throughput, $29.5\times$ improvement in space efficiency, and $4.3\times$ better energy efficiency compared to the current state-of-the-art PIM accelerator for Transformers. Minxuan Zhou, Chonghan Lee, Rishika Kushwah, Narayanan Vijaykrishnan, Tajana Rosing |
ASPDAC | 2 |
| 2024 | Efficient Host Intrusion Detection using Hyperdimensional ComputingabstractModern host-based intrusion detection systems (HIDS) rely on querying provenance graphs—graph representations of activity history on a system—to detect and respond to security threats present on a system. However, as the complexity and number of applications running on a system increase, the size of provenance graphs also increase, and thus the latency to query them. State-of-the-art designs deliver query latencies that are impractical for modern threat detection. In this paper, we introduce a hyper-dimensional computing (HDC) approach to querying provenance graphs for HIDS. By encoding provenance graphs and attack patterns/signatures into hyper-dimensional vectors, we can implement a query engine using simple vector operations. Our approach is hardware accelerator compatible, providing further speedups under resource-constrained environments. Our evaluation on a real-world dataset shows that our approach achieves > 90% detection accuracy and up to 4, 242× speedups over the state-of-the-art. This shows that HDC-based approaches can effectively deal with scaling issues in modern HIDS. Yujin Nam, Quinn Burke 0002, Minxuan Zhou, Patrick D. McDaniel, Tajana Rosing |
IEEE Big Data | 4 |
| 2024 | RL-PTQ: RL-based Mixed Precision Quantization for Hybrid Vision TransformersabstractExisting quantization approaches incur significant accuracy loss when compressing hybrid convolution and transformer models with low bit-width. This paper presents RL-PTQ, a novel post-training quantization (PTQ) framework utilizing reinforcement learning (RL). Our focus is on determining the most effective bit-width and observer for quantization configurations tailored for mixed precision by grouping layers and addressing the challenges of quantization of hybrid transformers. We achieved the highest quantized accuracy for MobileViTs compared to the previous PTQ methods [5--7]. Furthermore, our quantized model on Processing In Memory (PIM) architecture exhibited an energy efficiency enhancement of 10.1× and 22.6× compared to the baseline model, on the state-of-the-art PIM accelerator [15] and GPU, respectively. Eunji Kwon, Minxuan Zhou, Tajana Rosing, Seokhyeong Kang |
DAC | 2 |
| 2024 | HygHD: Hyperdimensional Hypergraph LearningabstractHypergraphs can model real-world data that has higher-order relationships. Graph neural network (GNN)-based solutions emerged as a hypergraph learning solution, but they face non-uniform memory accesses and accompany memory-intensive and compute-intensive operations, making the acceleration with near-data processing challenging. We propose a hyperdimensional computing (HDC)-based hypergraph learning framework called HygHD, which consists of highly parallelizable and lightweight HDC operations. HygHD accelerates both the training and inference on ferroelectric field-effect transistor (FeFET)-based processing-in-memory (PIM) hardware. Furthermore, we devise a hardware-friendly block-level concatenation and fine-grained block-level scheduler for high efficiency. Our evaluation results show that HygHD offers comparable accuracy to existing GNN-based solutions. Also, HygHD on GPU is up to 443× (7.67×) faster and 142× (2.78×) more energy efficient in training (inference) than the fastest GNN-based approach [1] on GPU. The HygHD accelerator further accelerates the HygHD algorithm, providing an average speedup of 40.0× (3.41×) on training (inference) compared to the HygHD GPU implementation. Jaeyoung Kang 0001, Youhak Lee, Minxuan Zhou, Tajana Rosing |
DATE | 3 |
| 2024 | Multi-Objective Software-Hardware Co-Optimization for HD-PIM via Noise-Aware Bayesian OptimizationabstractIn hardware accelerator design, software-hardware co-optimization requires intricate trade-offs and tight integration between software algorithms and hardware design to optimize performance, power efficiency, and area (PPA) while ensuring high accuracy. Furthermore, the inherent non-ideality in some emerging hardware technologies poses extra challenges to the co-optimization problem. This paper proposes a novel software-hardware co-optimization framework for hyperdimensional (HD) computing accelerators with emerging ReRAM-based processing in-memory (PIM) technologies, which have shown superior performance and energy efficiency over conventional machine learning accelerators. We first comprehensively characterize the non-trivial trade-offs between design parameters in HD-PIM and PPA and accuracy metrics in HD-PIM. Then, we develop a multi-objective noise-aware Bayesian optimization algorithm to find the Pareto set (optimal trade-offs between metrics) of the HD-PIM design. Our methodology uniquely addresses the stochastic nature of ReRAM by integrating error characteristics into the optimization process, thereby enhancing the quality of the generated designs. Experimental results show that our configurations achieve up to 4.28% accuracy improvement, 35.38% power reduction, 49x timing improvement, and 10% area reduction over a non-optimized design. Chien-Yi Yang, Minxuan Zhou, Flavio Ponzina, Suraj Sathya Prakash, Raid Ayoub, Pietro Mercati, Mahesh Subedar, Tajana Rosing |
ICCAD | 2 |
| 2024 | UFC: A Unified Accelerator for Fully Homomorphic EncryptionabstractFully homomorphic encryption (FHE) is crucial for post-quantum privacy-preserving computing. Researchers have proposed various FHE schemes that excel at different encrypted computations, such as single-instruction multiple-data (SIMD) arithmetic or arbitrary single-data functions. Hybrid-scheme FHE, which exploits appropriate schemes for specific tasks, is essential for real-world applications requiring optimal performance and accuracy. However, existing FHE accelerators only adopt scheme-specific custom designs, leading to inefficiency or lack of capability to support applications in hybrid FHE settings. In this work, we propose a Unified FHE aCcelerator (UFC) that provides better performance and cost-efficiency than prior scheme-specific accelerators on hybrid FHE applications. Our design process involves a comprehensive analysis of processing flows to abstract the primitives covering all operations in hybrid FHE applications. The UFC architecture primarily comprises hardware function units for these primitives, diverging from the deeply pipelined units in previous designs. This approach enables high hardware utilization across different FHE schemes. Further-more, we propose several algorithm-hardware co-optimizations to minimize the hardware cost of supporting various data shuffling patterns in FHE. This enables high-throughput implementation of function units that provide good cost efficiency. We also propose several compiler-level optimizations to achieve high hardware utilization of the unified architecture for computing FHE data in various algorithmic parameter settings. We evaluate the performance of UFC on different FHE programs, including scheme-specific and hybrid-scheme workloads. Our experiments show that UFC provides up to 6.0 × speedup and 1.6 × delay-energy-area efficiency improvement over state-of-the-art FHE accelerators. Minxuan Zhou, Yujin Nam, Xuan Wang 0040, Youhak Lee, Chris Wilkerson, Raghavan Kumar, Sachin Taneja, Sanu Mathew, Rosario Cammarota, Tajana Rosing |
MICRO | 1 |
| 2024 | Abakus: Accelerating k-mer Counting with Storage TechnologyabstractThis work seeks to leverage Processing-with-storage-technology (PWST) to accelerate a key bioinformatics kernel called k -mer counting, which involves processing large files of sequence data on the disk to build a histogram of fixed-size genome sequence substrings and thereby entails prohibitively high I/O overhead. In particular, this work proposes a set of accelerator designs called Abakus that offer varying degrees of tradeoffs in terms of performance, efficiency, and hardware implementation complexity. The key to these designs is a set of domain-specific hardware extensions to accelerate the key operations for k -mer counting at various levels of the SSD hierarchy, with the goal of enhancing the limited computing capabilities of conventional SSDs, while exploiting the parallelism of the multi-channel, multi-way SSDs. Our evaluation suggests that Abakus can achieve 8.42×, 6.91×, and 2.32× speedup over the CPU-, GPU-, and near-data processing solutions. Lingxi Wu, Minxuan Zhou, Ashish Venkat, Tajana Rosing, Kevin Skadron |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | Lightning Talk: Private and Secure Edge AI with Hyperdimensional ComputingabstractAs a lightweight and robust brain-inspired computing paradigm, Hyperdimensional Computing (HDC) serves as a promising solution for the next-generation edge AI. However, the basic form of HDC is vulnerable to privacy leaks and cyber attacks. In this paper, we breifly review and discuss the recent contributions to privacy and security of HDC. We first summarize existing HDC designs to protect against privacy leaks, such as differential privacy. Next, we review the data encryption techniques for collaborative learning using HDC based on Multi-Party Computation and Homomorphic Encryption. Finally, we discuss the HDC-based designs for combating cyber attacks in a malicious environment. More research on private and secure HDC-based methods are needed for future large-scale edge deployment. Xiaofan Yu 0001, Minxuan Zhou, Fatemeh Asgarinejad, Onat Güngör, Baris Aksanli, Tajana Rosing |
DAC | 2 |
| 2023 | OverlaPIM: Overlap Optimization for Processing In-Memory Neural Network AccelerationabstractProcessing in-memory (PIM) can accelerate neural networks (NNs) for its extensive parallelism and data movement minimization. The performance of NN acceleration on PIM heavily depends on software-to-hardware mapping, which indicates the order and distribution of operations across the hardware resources. Previous works optimize the mapping problem by exploring the design space of per-layer and cross-layer data layout, achieving speedup over manually designed mappings. However, previous works do not consider computation overlapping across consecutive layers. By overlapping computation, we can process a layer before its preceding layer fully completes, decreasing the execution latency of the whole network. The mapping optimization without overlap analysis can result in sub-optimal performance. In this work, we propose OverlaPIM, a new framework that integrates the overlap analysis with the DNN mapping optimization on PIM architectures. OverlaPIM adopts several techniques to enable efficient overlap analysis and optimization for the whole network mapping on PIM architectures. We test OverlaPIM on popular DNN networks and compare the results to non-overlap optimization. Our experiments show that OverlaPIM can efficiently produce mappings that are 2.10 x to 4.11 x faster than the state-of-the-art mapping optimization framework. Minxuan Zhou, Xuan Wang 0040, Tajana Rosing |
DATE | 1 |
| 2023 | Efficient Machine Learning on Encrypted Data Using Hyperdimensional ComputingabstractFully Homomorphic Encryption (FHE) enables arbitrary computations on encrypted data without decryption, thus protecting data in cloud computing scenarios. However, FHE adoption has been slow due to the significant computation and memory overhead it introduces. This becomes particularly challenging for end-to-end processes, including training and inference, for conventional neural networks on FHE-encrypted data. Additionally, machine learning tasks require a high throughput system due to data-level parallelism. However, existing FHE accelerators only utilize a single SoC, disregarding the importance of scalability. In this work, we address these challenges through two key innovations. First, at an algorithmic level, we combine hyperdimensional Computing (HDC) with FHE. The machine learning formulation based on HDC, a brain-inspired model, provides lightweight operations that are inherently well-suited for FHE computation. Consequently, FHE-HD has significantly lower complexity while maintaining comparable accuracy to the state-of-the-art. Second, we propose an efficient and scalable FHE system for FHE-based machine learning. The proposed system adopts a novel interconnect network between multiple FHE accelerators, along with an automated scheduling and data allocation framework to optimize throughput and hardware utilization. We evaluate the value of the proposed FHE-HD system on the MNIST dataset and demonstrate that the expected training time is 4.7 times faster compared to state-of-the-art MLP training. Furthermore, our system framework exhibits up to 38.2 times speedup and 13.8 times energy efficiency improvement over the baseline scalable FHE systems that use the conventional data-parallel processing flow. Yujin Nam, Minxuan Zhou, Saransh Gupta, Gabrielle De Micheli, Rosario Cammarota, Chris Wilkerson, Daniele Micciancio, Tajana Rosing |
ISLPED | 2 |
| 2022 | PIMProf: An Automated Program Profiler for Processing-in-Memory Offloading DecisionsabstractProcessing-in-memory (PIM) architectures reduce the data movement overhead by bringing computation closer to the memory. However, a key challenge is to decide which code regions of a program should be offloaded to PIM for the best performance. The goal of this work is to help programmers leverage PIM architectures by automatically profiling legacy workloads to find PIM-friendly code regions for offloading. We propose PIMProf11The source code of PIMProf can be found at https://github.com/Systems-ShiftLab/PIMProf, an automated profiling and offloading tool to determine PIM offloading regions for CPU-PIM hybrid architectures. PIMProf efficiently models the comprehensive cost related to PIM offloading and makes the offloading decision by an effective and computational-tractable algorithm. We demonstrate the effectiveness of PIMProf by evaluating the GAP graph benchmark suite and the PARSEC benchmark suite under different PIM and CPU configurations. Our evaluation shows that, compared to the CPU baseline and a PIM-only configuration, the offloading decisions by PIMProf provides$5.33\times$and$1.39\times$speedup in the GAP graph workloads, respectively;$2.22\times$and$1.74\times$speedup in the PARSEC benchmarks, respectively. Yizhou Wei, Minxuan Zhou, Sihang Liu 0001, Korakit Seemakhupt, Tajana Rosing, Samira Manabi Khan |
DATE | 2 |
| 2022 | TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for TransformerabstractTransformer-based models are state-of-the-art for many machine learning (ML) tasks. Executing Transformer usually requires a long execution time due to the large memory footprint and the low data reuse rate, stressing the memory system while under-utilizing the computing resources. Memory-based processing technologies, including processing in-memory (PIM) and near-memory computing (NMC), are promising to accelerate Transformer since they provide high memory bandwidth utilization and extensive computation parallelism. However, the previous memory-based ML accelerators mainly target at optimizing dataflow and hardware for compute-intensive ML models (e.g., CNNs), which do not fit the memory-intensive characteristics of Transformer. In this work, we propose TransPIM, a memory-based acceleration for Transformer using software and hardware co-design. In the software-level, TransPIM adopts a token-based dataflow to avoid the expensive inter-layer data movements introduced by previous layer-based dataflow. In the hardware-level, TransPIM introduces lightweight modifications in the conventional high bandwidth memory (HBM) architecture to support PIM-NMC hybrid processing and efficient data communication for accelerating Transformer-based models. Our experiments show that TransPIM is 3.7× to 9.1× faster than existing memory-based acceleration. As compared to conventional accelerators, TransPIM is 22.1× to 114.9× faster than GPUs and provides 2.0× more throughput than existing ASIC-based accelerators. Minxuan Zhou, Jaeyoung Kang 0001, Tajana Rosing |
HPCA | 1 |
| 2022 | RelHD: A Graph-based Learning on FeFET with Hyperdimensional ComputingabstractAdvances in graph neural network (GNN)-based algorithms enable machine learning on relational data. GNNs are computationally demanding since they rely upon backpropagation over the graph data that has sparse and irregular characteristics. In this paper, we propose a lightweight graph-based machine learning framework based on hyperdimensional computing (HDC) called RelHD. It maps the features of each node into a high-dimensional space and embeds relationships between nodes. Using lightweight HDC operations, RelHD enables both training and inference on graph data without backpropagation. Furthermore, we design a scalable processing in-memory (PIM) architecture based on the emerging FeFET technology to accelerate the proposed algorithm. Our strategy optimizes data allocation and operation scheduling that maximizes the accelerator performance by addressing the sparseness and irregularity of the graph. Experimental results show that RelHD offers comparable accuracy to the popular GNN-based algorithms while being up to 32× faster on GPU. Also, our FeFET-based accelerator achieves 33× of speedup and 59287× energy efficiency improvement on average over the GPU. It is 10× faster and 986× more energy efficient on average compared to the state-of-the-art in-memory processing-based GNN accelerator. Jaeyoung Kang 0001, Minxuan Zhou, Abhinav Bhansali, Anthony Thomas, Tajana Rosing |
ICCD | 2 |
| 2021 | PIM-DL: Boosting DNN Inference on Digital Processing In-Memory Architectures via Data Layout OptimizationsabstractDigital processing in-memory (DPIM) provides very low overhead, highly parallel computation in conventional memory, which significantly accelerates data-intensive workloads like deep neural networks (DNNs). DPIM-based DNN accelerators require that data be properly laid out to make the best use of the available in-memory operations. However, existing DPIM accelerators tend to optimize for a particular DNN dataflow, neglecting the large design space of data layout. This work systematically investigates the data layout for DPIM DNN acceleration. We propose a mapping framework to represent the whole design space of DPIM data layout for general DNN models. Our investigation shows that an exhaustive exploration on the whole design space for mapping a DNN application to DPIM architecture is not computationally tractable. Therefore, we propose a compiler-level optimization, PIM-DL, that finds highly efficient data layouts for DPIM DNN acceleration using a two-level dynamic programming algorithm and a heuristic-based search. Our experiments show that DNN DPIM solutions created by our PIM-DL provide 3.7× and 4.3× better performance and energy efficiency as compared to the state of the art under the same hardware constraints. Minxuan Zhou, Guoyang Chen, Mohsen Imani, Saransh Gupta, Weifeng Zhang 0003, Tajana Rosing |
PACT | 1 |
| 2021 | Ultra Efficient Acceleration for De Novo Genome Assembly via Near-Memory ComputingabstractDe novo assembly of genomes for which there is no reference, is essential for novel species discovery and metagenomics. In this work, we accelerate two key performance bottlenecks of DBG-based assembly, graph construction and graph traversal, with a near-data processing (NDP) architecture based on 3D-stacking. The proposed framework distributes key operations across NDP cores to exploit a high degree of parallelism and high memory bandwidth. We propose several optimizations based on domain-specific properties to improve the performance of our design. We integrate the proposed techniques into an existing DBG assembly tool, and our simulation-based evaluation shows that the proposed NDP implementation can improve the performance of graph construction by 33× and traversal by 16× compared to the state-of-the-art. Minxuan Zhou, Lingxi Wu, Muzhou Li, Niema Moshiri, Kevin Skadron, Tajana Rosing |
PACT | 1 |
| 2021 | DP-Sim: A Full-stack Simulation Infrastructure for Digital Processing In-Memory ArchitecturesabstractDigital processing in-memory (DPIM) is a promising technology that significantly reduces data movements while providing high parallelism. In this work, we design and implement the first full-stack DPIM simulation infrastructure, DP-Sim, which evaluates a comprehensive range of DPIM-specific design space concerning both software and hardware. DP-Sim provides a C++ library to enable DPIM acceleration in general programs while supporting several aspects of software-level exploration by a convenient interface. The DP-Sim software front-end generates specialized instructions that can be processed by a hardware simulator based on a new DPIM-enabled architecture model which is 10.3% faster than conventional memory simulation models. We use DP-Sim to explore the DPIM-specific design space of acceleration for various emerging applications. Our experiments show that bank-level control is 11.3x faster than conventional channel-level control because of higher computing parallelism. Furthermore, cost-aware memory allocation can provide at least 2.2x speedup vs. heuristic methods, showing the importance of data layout in DPIM acceleration. Minxuan Zhou, Mohsen Imani, Yeseong Kim, Saransh Gupta, Tajana Rosing |
ASP-DAC | 1 |
| 2021 | MAT: Processing In-Memory Acceleration for Long-Sequence AttentionabstractAttention-based machine learning is used to model long-term dependencies in sequential data. Processing these models on long sequences can be prohibitively costly because of the large memory consumption. In this work, we propose MAT, a processing in-memory (PIM) framework, to accelerate long-sequence attention models. MAT adopts a memory-efficient processing flow for attention models to process sub-sequences in a pipeline with much smaller memory footprint. MAT utilizes a reuse-driven data layout and an optimal sample scheduling to optimize the performance of PIM attention. We evaluate the efficiency of MAT on two emerging long-sequence tasks including natural language processing and medical image processing. Our experiments show that MAT is $2.7 \times$ faster and $3.4 \times$ more energy efficient than the state-of-the-art PIM acceleration. As compared to TPU and GPU, MAT is $5.1 \times$ and $16.4 \times$ faster while consuming $27.5 \times$ and $41.0 \times$ less energy. Minxuan Zhou, Yunhui Guo, Bin Li 0064, Kevin W. Eliceiri, Tajana Rosing |
DAC | 1 |
| 2021 | HyGraph: Accelerating Graph Processing with Hybrid Memory-centric ComputingabstractGraph applications are challenging to run efficiently on conventional systems because of their large and irregular data. Several works have exploited near-data processing (NDP) based on emerging 3D-stacked memory to accelerate graph processing applications by offloading computations to massively parallel cores in the memory chip. Even though NDP can efficiently support parallel operations in a memory scalable way, it still requires data movement between memory and near-memory cores. Such data movement introduces large overhead because of the random data pattern in graph workloads. Furthermore, the parallelism provided by NDP systems is still insufficient for graph applications because of the limited number of processing cores. In this work, we tackle these challenges by integrating processing in-memory (PIM) technology in the NDP-based accelerator. We propose HyGraph, a software-hardware co-design for graph acceleration that exploits hybrid memory-centric computing technologies, including NDP and PIM. The design of HyGraph includes an optimization algorithm for hybrid memory layout, a run-time system combining both NDP and PIM processing flows, and customized hardware for efficiently enabling PIM functionality in NDP systems. Our experimental results show that HyGraph is up to 1.9× faster and 2.4× more energy-efficient than state-of-the-art memory-centric graph accelerators on several widely used graph algorithms with various real-world graphs. Minxuan Zhou, Muzhou Li, Mohsen Imani, Tajana Rosing |
DATE | 1 |
| 2021 | Massively Parallel Big Data Classification on a Programmable Processing In-Memory ArchitectureabstractWith the emergence of Internet of Things, massive data created in the world pose huge technical challenges for efficient processing. Processing in-memory (PIM) technology has been widely investigated to overcome expensive data movements between processors and memory blocks. However, existing PIM designs incur large area overhead to enable computing capability via additional near-data processing cores and analog/mixed signal circuits. In this paper, we propose a new massively-parallel processing in-memory (PIM) architecture, called CHOIR, based on emerging nonvolatile memory technology for big data classification. Unlike existing PIM designs which demand large analog/mixed signal circuits, we support the parallel PIM instructions for conditional and arithmetic operations in an area-efficient way. As a result, the classification solution performs both training and testing on the PIM architecture by fully utilizing the massive parallelism. Our design significantly improves the performance and energy efficiency of the classification tasks by 123× and 52× respectively as compared to the state-of-the-art tree boosting library running on GPU. Yeseong Kim, Mohsen Imani, Saransh Gupta, Minxuan Zhou, Tajana Rosing |
ICCAD | 4 |
| 2021 | FPRA: A Fine-grained Parallel RRAM ArchitectureabstractEmerging resistive memory (RRAM) based crossbar array is a promising technology to accelerate neural network applications. RRAM-based CNN accelerators support a high-degree of intra-layer and inter-layer parallelism. The intra-layer parallelism duplicates kernels for each network layer while the inter-layer parallelism allows execution of each layer when a portion of input data is available. However, previously proposed RRAM-based accelerators do not leverage data sharing between duplicate kernels leading to significant idleness of crossbar arrays during inference. This shared data creates data dependencies that stall the processing of the next layer in the pipeline. To address these issues, we propose Fine-grained Parallel RRAM Architecture (FPRA), a novel architectural design, to improve parallelism for pipeline-enabled RRAM-based accelerators. FPRA addresses the data sharing issue with kernel batching and data sharing aware memory. Kernel batching rearranges the layout of the kernels and minimizes the data dependencies created by the input shared data. The data sharing aware memory uniformly buffers the input and output data for each layer, efficiently dispatching data to duplicate kernels while reducing the amount of data transferred between layers. We evaluate FPRA on eight popular image recognition CNN models with various configurations in a cycle-accurate simulator. We find that FPRA manages to achieve 2.0 $\times$ average latency speedup, and 2.1 $\times$ average throughput increase, as compared to the state-of-the-art RRAM-based accelerators. Xiao Liu 0033, Minxuan Zhou, Rachata Ausavarungnirun, Sean Eilert, Ameen Akel, Tajana Rosing, Narayanan Vijaykrishnan, Jishen Zhao |
ISLPED | 2 |
| 2020 | DUAL: Acceleration of Clustering Algorithms using Digital-based Processing In-MemoryabstractToday's applications generate a large amount of data that need to be processed by learning algorithms. In practice, the majority of the data are not associated with any labels. Unsupervised learning, i.e., clustering methods, are the most commonly used algorithms for data analysis. However, running clustering algorithms on traditional cores results in high energy consumption and slow processing speed due to a large amount of data movement between memory and processing units. In this paper, we propose DUAL, a Digital-based Unsupervised learning AcceLeration, which supports a wide range of popular algorithms on conventional crossbar memory. Instead of working with the original data, DUAL maps all data points into high-dimensional space, replacing complex clustering operations with memory-friendly operations. We accordingly design a PIM-based architecture that supports all essential operations in a highly parallel and scalable way. DUAL supports a wide range of essential operations and enables in-place computations, allowing data points to remain in memory. We have evaluated DUAL on several popular clustering algorithms for a wide range of large-scale datasets. Our evaluation shows that DUAL provides a comparable quality to existing clustering algorithms while using a binary representation and a simplified distance metric. DUAL also provides 58.8× speedup and 251.2× energy efficiency improvement as compared to the state-of-the-art solution running on GPU. Mohsen Imani, Saikishan Pampana, Saransh Gupta, Minxuan Zhou, Yeseong Kim, Tajana Rosing |
MICRO | 4 |
| 2020 | Temperature-Aware DRAM Cache Management - Relaxing Thermal Constraints in 3-D SystemsabstractHigh bandwidth 3-D-stacked dynamic random access memory (DRAM) has been proposed to address the memory wall in modern systems, especially when it is used as a large last-level cache (LLC). However, stacking DRAM directly on top of the processor significantly impedes the efficiency of cooling, potentially causing thermal issues both in the processor and DRAM. Dynamic thermal management (DTM) based on DRAM temperature can be heavily intrusive because the normal working temperature for DRAM is lower than the processor temperature limit. This paper shows that in many cases it is better to disable hot portions of the cache rather than apply DTM and slow down the processor. Three temperature-aware cache management mechanisms are proposed to decrease the performance impact of DTM on 3-D systems. Our experiments show these techniques can improve the performance of DRAM-targeted DTM by 26.1% on average which make 3-D systems more practical for the future high-performance computing. Minxuan Zhou, Andreas Prodromou, Rui Wang 0014, Hailong Yang 0002, Depei Qian 0001, Dean M. Tullsen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | GRAM: graph processing in a ReRAM-based computational memoryabstractThe performance of graph processing for real-world graphs is limited by inefficient memory behaviours in traditional systems because of random memory access patterns. Offloading computations to the memory is a promising strategy to overcome such challenges. In this paper, we exploit the resistive memory (ReRAM) based processing-in-memory (PIM) technology to accelerate graph applications. The proposed solution, GRAM, can efficiently executes vertex-centric model, which is widely used in large-scale parallel graph processing programs, in the computational memory. The hardware-software co-design used in GRAM maximizes the computation parallelism while minimizing the number of data movements. Based on our experiments with three important graph kernels on seven real-world graphs, GRAM provides 122.5X and 11.1x speedup compared with an in-memory graph system and optimized multithreading algorithms running on a multi-core CPU. Compared to a GPU-based graph acceleration library and a recently proposed PIM accelerator, GRAM improves the performance by 7.1X and 3.8X respectively. Minxuan Zhou, Mohsen Imani, Saransh Gupta, Yeseong Kim, Tajana Rosing |
ASP-DAC | 1 |
| 2019 | Thermal-Aware Design and Management for Search-based In-Memory AccelerationabstractRecently, Processing-In-Memory (PIM) techniques exploiting resistive RAM (ReRAM) have been used to accelerate various big data applications. ReRAM-based in-memory search is a powerful operation which efficiently finds required data in a large data set. However, such operations result in a large amount of current which may create serious thermal issues, especially in state-of-the-art 3D stacking chips. Therefore, designing PIM accelerators based on in-memory searches requires a careful consideration of temperature. In this work, we propose static and dynamic techniques to optimize the thermal behavior of PIM architectures running intensive in-memory search operations. Our experiments show the proposed design significantly reduces the peak chip temperature and dynamic management overhead. We test our proposed design in two important categories of applications which benefit from the search-based PIM acceleration - hyper-dimensional computing and database query. Validated experiments show that the proposed method can reduce the steady-state temperature by at least 15.3 °C which extends the lifetime of the ReRAM device by 57.2% on average. Furthermore, the proposed fine-grained dynamic thermal management provides 17.6% performance improvement over state-of-the-art methods. Minxuan Zhou, Mohsen Imani, Saransh Gupta, Tajana Rosing |
DAC | 1 |
| 2019 | DigitalPIM: Digital-based Processing In-Memory for Big Data AccelerationabstractIn this work, we design, DigitalPIM, a Digital-based Processing In-Memory platform capable of accelerating fundamental big data algorithms in real time with orders of magnitude more energy efficient operation. Unlike the existing near-data processing approach such as HMC 2.0, which utilizes additional low-power processing cores next to memory blocks, the proposed platform implements the entire algorithm directly in memory blocks without using extra processing units. In our platform, each memory block supports the essential operations including: bitwise operation, addition/multiplication, and search operation internally in memory without reading any values out of the block. This significantly mitigates the processing costs of the new architecture, while providing high scalability and parallelism for performing the extensive computations. We exploit these essential operations to accelerate popular big data applications entirely in memory such as machine learning algorithms, query processing, and graph processing. Our evaluations show that for all tested applications, the performance can be accelerated significantly by eliminating the memory access bottleneck Mohsen Imani, Saransh Gupta, Yeseong Kim, Minxuan Zhou, Tajana Rosing |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | GAS: A Heterogeneous Memory Architecture for Graph ProcessingabstractGraph processing has become important for various applications in today's big data era. However, most graph processing applications suffer from large memory overhead due to random memory accesses. Such random memory access pattern provides little temporal and spatial locality which cannot be accelerated by the conventional hierarchical memory system. In this work, we propose GAS, a heterogeneous memory architecture, to accelerate graph applications implemented in message-based vertex program model, which is widely used in various graph processing systems. GAS utilizes the specialized content-addressable memory (CAM) to store random data, and determine exact access patterns by a series of associative search. Thus, GAS not only removes the inefficiency of random accesses but also reduces the memory access latency by accurate prefetching. We test the efficiency of GAS with three important graph processing kernels on five well-known graphs. Our experimental results show that GAS can significantly reduce cache miss rate and improve the bandwidth utilization as compared to a conventional system with a state-of-the-art graph-specific prefetching mechanism. These enhancements result in 34% and 27% reduction in energy consumption and execution time, respectively. Minxuan Zhou, Mohsen Imani, Saransh Gupta, Tajana Rosing |
ISLPED | 1 |
| 2016 | HV2M: A novel approach to boost inter-VM network performance for Xen-based HVMs
Yuebin Bai, Yongwang Zhao, Duo Lu, Yuanfeng Peng, Minxuan Zhou |
J. Syst. Softw. | 7 |