VLDB 2026 Research / reviewers in the wild / expert
Yizhuo Wang 0001
dblp:31/1736-1 · also Yi-Zhuo Wang 0001
· DBLP profile ↗
34ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0002-1288-331XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 9 first-author · 12 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HRPF: A parallel programming framework for recursive algorithms on heterogeneous CPU-GPU systems
Yizhuo Wang 0001, Senhao Shao, Jianhua Gao 0001, Weixing Ji, Hongbo Xing |
Parallel Comput. | 1 |
| 2025 | Adaptive point cloud compression based on precision-aware floating-point encoding
Yanpeng Han, Yizhuo Wang 0001, Fawang Liu, Jianhua Gao 0001, Weixing Ji |
CCF Trans. High Perform. Comput. | 2 |
| 2025 | RaNAS: Resource-Aware Neural Architecture Search for Edge ComputingabstractNeural architecture search (NAS) for edge devices is often time-consuming because of long-latency deploying and testing on edge devices. The ability to accurately predict the computation cost and memory requirement for convolutional neural networks (CNNs) in advance holds substantial value. Existing work primarily relies on analytical models, which can result in high prediction errors. This article proposes a resource-aware NAS (RaNAS) model based on various features. Additionally, a new graph neural network is introduced to predict inference latency and maximum memory requirements for CNNs on edge devices. Experimental results show that, within the error bound of ±1%, RaNAS achieves an accuracy improvement of approximately 8% for inference latency prediction and about 25% for maximum memory occupancy prediction over the state-of-the-art approaches. Jianhua Gao 0001, Zeming Liu, Yizhuo Wang 0001, Weixing Ji |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | PTPS: Precision-Aware Task Partitioning and Scheduling for SpMV on CPU-FPGA Heterogeneous PlatformsabstractThe CPU-FPGA heterogeneous computing architecture is extensively employed in the embedded domain due to its low cost and power efficiency, with numerous sparse matrix-vector multiplication (SpMV) acceleration efforts already targeting this architecture. However, existing work rarely includes collaborative SpMV computations between CPU and FPGA, which limits the exploration of hybrid architectures that could potentially offer enhanced performance and flexibility. This article introduces an FPGA architecture design that supports multiprecision SpMV computations, including FP16, FP32, and FP64. Building on this, PTPS, a precision-aware SpMV task partitioning and dynamic scheduling algorithm tailored for the CPU-FPGA heterogeneous architecture, is proposed. The core idea of PTPS is lossless partitioning of sparse matrices across multiple precisions, prioritizing low-precision SpMV computations on the FPGA and high-precision computations on the CPU. PTPS not only leverages the strengths of CPU and FPGA for collaborative SpMV computations but also reduces data transmission overhead between them, thereby improving the overall computational efficiency. Experimental evaluation demonstrates that the proposed approach offers an average speedup of$1.57\times $over the CPU-only approach and$2.58\times $over the FPGA-only approach. Jianhua Gao 0001, Xingze Huang, Yizhuo Wang 0001, Weixing Ji |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | X's Day: Personality-Driven Virtual Human Behavior GenerationabstractDeveloping convincing and realistic virtual human behavior is essential for enhancing user experiences in virtual reality (VR) and augmented reality (AR) settings. This paper introduces a novel task focused on generating long-term behaviors for virtual agents, guided by specific personality traits and contextual elements within 3D environments. We present a comprehensive framework capable of autonomously producing daily activities autoregressively. By modeling the intricate connections between personality characteristics and observable activities, we establish a hierarchical structure of Needs, Task, and Activity levels. Integrating a Behavior Planner and a World State module allows for the dynamic sampling of behaviors using large language models (LLMs), ensuring that generated activities remain relevant and responsive to environmental changes. Extensive experiments validate the effectiveness and adaptability of our approach across diverse scenarios. This research makes a significant contribution to the field by establishing a new paradigm for personalized and context-aware interactions with virtual humans, ultimately enhancing user engagement in immersive applications. Our project website is at: https://behavior.agent-x.cn/. Wei Liang 0008, Yizhuo Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | A Multi-View Deep Learning Method for Predicting Blood-Brain Barrier Permeability of PeptidesabstractThe blood-brain barrier (BBB) plays a crucial role in protecting brain health by acting as a barrier between the brain and blood vessels. This barrier also presents challenges for delivering peptide drugs to brain targets. There is a pressing need for computational methods to accurately predict the permeability of peptides across the BBB. However, existing approaches face challenges due to limited real experimentally data and incomplete molecular information within peptide sequences. In this paper, we introduce MultiB3Pred, a multi-view deep learning method designed to address these challenges. Our method makes three key contributions. Firstly, we employ a effective amino acid replacement strategy for data augmentation. Secondly, We utilize sequence embeddings from a biologically pretrained model ProtT5 [1], further refined by a Transformer to capture dependencies on our specific dataset, leading to better sequence representations for the sequence predictor. Lastly, we derive SMILES from the sequences and train a novel SMILES learner. Precisely, the physicochemical properties of the molecules with the graph representation captured by the graph neural network from the molecular graphs are integrated through multilayer perceptron. The predicted probabilities from two sub-predictors are averaged to obtain the final result. Experiments demonstrate that MultiB3Pred achieves state-of-the-art accuracy and Matthews correlation coefficient of 94.4% and 89.9% respectively, showcasing its excellent performance in predicting blood-brain barrier penetration. At the same time, the stability of the model is confirmed by the good results of the 5-fold crossover experiment. Yizhuo Wang 0001, Chunfeng Li, Weixing Ji |
BIBM | 2 |
| 2024 | Load Balancing Optimizations for Distributed GMRES Algorithm
Shuaizhe Guo, Jianhua Gao 0001, Weixing Ji, Yizhuo Wang 0001 |
ICA3PP (6) | 5 |
| 2024 | JLeaks: A Featured Resource Leak Repository Collected From Hundreds of Open-Source Java ProjectsabstractHigh-quality defect repositories are vital in defect detection, localization, and repair. However, existing repositories collected from open-source projects are either small-scale or inadequately labeled and packed. This paper systematically summarizes the programming APIs of system resources (i.e., file, socket, and thread) in Java. Additionally, this paper demonstrates the exceptions that may cause resource leaks in the chained and nested streaming operations. A semi-automatic toolchain is built to improve the efficiency of defect extraction, including automatic building for large legacy Java projects. Accordingly, 1,094 resource leaks were collected from 321 open-source projects on GitHub. This repository, named JLeaks, was built by round-by-round filtering and cross-validation, involving the review of approximately 3,185 commits from hundreds of projects. JLeaks is currently the largest resource leak repository, and each defect in JLeaks is well-labeled and packed, including causes, locations, patches, source files, and compiled bytecode files for 254 defects. We have conducted a detailed analysis of JLeaks for defect distribution, root causes, and fix approaches. We compare JLeaks with two well-known resource leak repositories, and the results show that JLeaks is more informative and complete, with high availability, uniqueness, and consistency. Additionally, we show the usability of JLeaks in two application scenarios. Future studies can leverage our repository to encourage better design and implementation of defect-related algorithms and tools. Weixing Ji, Wuhuang Yao, Yizhuo Wang 0001, Hui Liu 0003, Haiyang Peng |
ICSE | 5 |
| 2024 | pSpMv: precision-based sparse matrix partition and SpMV optimizationabstractAbstract The new generation of computing devices tends to support multiple floating-point formats and different computing precision. Besides single and double precision, half precision is embraced and widely supported by new computing devices. Low-precision representations have compact memory size and lightweight computing strength, and they also bring opportunities to the optimization of BLAS routines. This paper proposes a new sparse matrix partition approach based on IEEE 754 standard floating-point format. An input sparse matrix in double precision is partitioned and transformed into several sub-matrices in different precision without loss of accuracy. Most non-zero elements can be stored in half or single precision, if the most significant bits of exponent and the least significant bits of mantissa are zeros in double-precision representation. Based on this mixed-precision representation of sparse matrix, we also present a new SpMV algorithm pSpMV for GPU devices. pSpMV not only reduces the memory access overhead, but also reduces the computing strength of floating-point numbers. Experimental results on two GPU devices show that pSpMV achieves a geometric mean speedup of 1.39x on Tesla V100 and 1.45x on Tesla P100 over double-precision SpMV for 2,554 sparse matrices. Yizhuo Wang 0001, Jianhua Gao 0001, Weixing Ji |
CCF Trans. High Perform. Comput. | 2 |
| 2024 | Revisiting thread configuration of SpMV kernels on GPU: A machine learning based approach
Jianhua Gao 0001, Weixing Ji, Yizhuo Wang 0001, Feng Shi 0009 |
J. Parallel Distributed Comput. | 4 |
| 2024 | Optimization of Large-Scale Sparse Matrix-Vector Multiplication on Multi-GPU SystemsabstractSparse matrix-vector multiplication (SpMV) is one of the important kernels of many iterative algorithms for solving sparse linear systems. The limited storage and computational resources of individual GPUs restrict both the scale and speed of SpMV computing in problem-solving. As real-world engineering problems continue to increase in complexity, the imperative for collaborative execution of iterative solving algorithms across multiple GPUs is increasingly apparent. Although the multi-GPU-based SpMV takes less kernel execution time, it also introduces additional data transmission overhead, which diminishes the performance gains derived from parallelization across multi-GPUs. Based on the non-zero elements distribution characteristics of sparse matrices and the tradeoff between redundant computations and data transfer overhead, this article introduces a series of SpMV optimization techniques tailored for multi-GPU environments and effectively enhances the execution efficiency of iterative algorithms on multiple GPUs. First, we propose a two-level non-zero elements-based matrix partitioning method to increase the overlap of kernel execution and data transmission. Then, considering the irregular non-zero elements distribution in sparse matrices, a long-row-aware matrix partitioning method is proposed to hide more data transmissions. Finally, an optimization using redundant and inexpensive short-row execution to exchange costly data transmission is proposed. Our experimental evaluation demonstrates that, compared with the SpMV on a single GPU, the proposed method achieves an average speedup of 2.00× and 1.85× on platforms equipped with two RTX 3090 and two Tesla V100-SXM2, respectively. The average speedup of 2.65× is achieved on a platform equipped with four Tesla V100-SXM2. Jianhua Gao 0001, Weixing Ji, Yizhuo Wang 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Optimization of Sparse Matrix Computation for Algebraic Multigrid on GPUsabstractAMG is one of the most efficient and widely used methods for solving sparse linear systems. The computational process of AMG mainly consists of a series of iterative calculations of generalized sparse matrix-matrix multiplication (SpGEMM) and sparse matrix-vector multiplication (SpMV). Optimizing these sparse matrix calculations is crucial for accelerating solving linear systems. In this paper, we first focus on optimizing the SpGEMM algorithm in AmgX, a popular AMG library for GPUs. We propose a new algorithm called SpGEMM-upper, which achieves an average speedup of 2.02× on Tesla V100 and 1.96× on RTX 3090 against the original algorithm. Next, through experimental investigation, we conclude that no single SpGEMM library or algorithm performs optimally for most sparse matrices, and the same holds true for SpMV. Therefore, we build machine learning-based models to predict the optimal SpGEMM and SpMV used in the AMG calculation process. Finally, we integrate the prediction models, SpGEMM-upper, and other selected algorithms into a framework for adaptive sparse matrix computation in AMG. Our experimental results prove that the framework achieves promising performance improvements on the test set. Yizhuo Wang 0001, Fangli Chang, Bingxin Wei, Jianhua Gao 0001, Weixing Ji |
ACM Trans. Archit. Code Optim. | 1 |
| 2023 | An Automatic Deployment Method for Hybrid Cloud Simulation Platform
Xilai Yao, Yizhuo Wang 0001, Weixing Ji, Qiurui Chen |
APPT | 2 |
| 2022 | TaiChi: A Hybrid Compression Format for Binary Sparse Matrix-Vector Multiplication on GPUabstractBinary Sparse Matrix-Vector Multiplication (SpMV) is a heavy computational kernel in weblink analysis, integer factorization, compressed sensing, spectral graph theory, and other domains. Testing several popular GPU-based SpMV implementations on 400 sparse matrices, we observed that data transfer to GPU memory accounts for a large part of the total computation time. The transfer of constant value “1”s can be easily eliminated for binary sparse matrices. However, compressing index arrays has always been a great challenge. This article proposes a new compression format TaiChi to further reduce index data copies and improve the performance of SpMV, especially for diagonally dominant binary sparse matrices. Input matrices are first partitioned into relatively dense and ultra-sparse areas. Then the dense areas are encoded inversely by marking “0”s, while the ultra-sparse area is encoded by marking “1”s. We also designed a new SpMV algorithm only using addition and subtraction for binary matrices based on our partition and encoding format. Evaluation results on real-world binary sparse matrices show that our hybrid encoding for binary matrix significantly reduces the data transfer and speeds up the kernel execution. It achieves the highest transfer and kernel execution speedups of 5.63x and 3.84x on GTX 1080 Ti, 3.39x and 3.91x on Tesla V100. Jianhua Gao 0001, Weixing Ji, Zhaonian Tan, Yizhuo Wang 0001, Feng Shi 0009 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | AMF-CSR: Adaptive Multi-Row Folding of CSR for SpMV on GPUabstractSpMV is a cost-dominant operation used in many iterative methods for solving large-scale sparse linear systems. However, irregular memory access of SpMV to the multiplied vector leads to low data locality and then harms the performance. This paper presents an adaptive multi-row folding of CSR (AMF-CSR) format for SpMV calculation on GPU. This new storage format supports the folding of the variable number of rows in order to achieve better load balancing in computation. AMF-CSR not only increases the density of non-zero elements in a folded row, thereby improving the access locality of the multiplied vector, but also merges an approximately equal number of nonzero elements in a folded row, hence achieving load balancing. The performance evaluation using 28 sparse matrices shows that the proposed SpMV algorithm based on AMF-CSR achieves the highest speedup of 4.11x and 3.62x on GTX 1080 Ti and Tesla V100 respectively against a fixed multi-row folding-based SpMV algorithm. Evaluation results using 450 regular sparse matrices and 450 irregular sparse matrices also show that AMF-CSR is superior to other SpMV implementations. Jianhua Gao 0001, Weixing Ji, Senhao Shao, Yizhuo Wang 0001, Feng Shi 0009 |
ICPADS | 5 |
| 2021 | Towards Optimal Fast Matrix Multiplication on CPU-GPU Platforms
Senhao Shao, Yizhuo Wang 0001, Weixing Ji, Jianhua Gao 0001 |
PDCAT | 2 |
| 2020 | MMSparse: 2D partitioning of sparse matrix based on mathematical morphology
Zhaonian Tan, Weixing Ji, Jianhua Gao 0001, Yueyan Zhao, Akrem Benatia, Yizhuo Wang 0001, Feng Shi 0009 |
Future Gener. Comput. Syst. | 6 |
| 2020 | Cube-based incremental outlier detection for streaming computing
Jianhua Gao 0001, Weixing Ji, Anmin Li, Yizhuo Wang 0001, Zongyu Zhang |
Inf. Sci. | 5 |
| 2018 | Exploiting Task-Based Parallelism for Parallel Discrete Event SimulationabstractToday large-scale simulation applications are becoming common in research and industry. A significant fraction of them run on multi-core clusters. Current parallel simulation kernels use multi-process and multi-thread to exploit inter-node parallelism and intra-node parallelism on multi-core clusters. We exploit task-base parallelism in parallel discrete event simulation (PDES) kernels, which is more fine-grained than thread-level and process-level parallelism. In our system, every simulation event is wrapped to a task. Work-stealing task scheduling scheme is applied to achieve dynamic load balancing among the multi-cores, and a graph partitioning approach is applied in partitioning simulation entities among the cluster nodes. Experimental results show that our PDES kernel outperforms existing PDES kernels by fully exploiting task parallelism. Yizhuo Wang 0001, Zhiwei Gao 0001, Weixing Ji, Duzheng Qing |
PDP | 1 |
| 2018 | BestSF: A Sparse Meta-Format for Optimizing SpMV on GPUabstractThe Sparse Matrix-Vector Multiplication (SpMV) kernel dominates the computing cost in numerous scientific applications. Many implementations based on different sparse formats were proposed to improve this kernel on the recent GPU architectures. However, it has been widely observed that there is no “best-for-all” sparse format for the SpMV kernel on GPU. Indeed, serious performance degradation of an order of magnitude can be observed without a careful selection of the sparse format to use. To address this problem, we propose in this article BestSF (Best Sparse Format), a new learning-based sparse meta-format that automatically selects the most appropriate sparse format for a given input matrix. To do so, BestSF relies on a cost-sensitive classification system trained using Weighted Support Vector Machines (WSVMs) to predict the best sparse format for each input sparse matrix. Our experimental results on two different NVIDIA GPU architectures using a large number of real-world sparse matrices show that BestSF achieved a noticeable overall performance improvement over using a single sparse format. While BestSF is trained to select the best sparse format in terms of performance (GFLOPS), our further experimental investigations revealed that using BestSF also led, in most of the test cases, to the best energy efficiency (MFLOPS/W). To prove its practical effectiveness, we also evaluate the performance and energy efficiency improvement achieved when using BestSF as a building block in a GPU-based Preconditioned Conjugate Gradient (PCG) iterative solver. Akrem Benatia, Weixing Ji, Yizhuo Wang 0001, Feng Shi 0009 |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | Machine Learning Approach for the Predicting Performance of SpMV on GPUabstractSparse Matrix-Vector Multiplication (SpMV) kernel dominates the computing cost in numerous scientific applications. Many implementations based on different sparse formats were proposed recently for optimizing this kernel on the GPU side. Since the performance of the SpMV varies significantly according to the sparsity characteristics of the input matrix and the hardware features, developing an accurate performance model for this kernel is a challenging task. The traditional approach of building such models by analytical modeling is difficult in practice and requires a thorough understanding of the interaction between the GPU hardware and the sparse code. In this paper, we propose to use a machine learning approach to predict the performance of the SpMV kernel using several sparse formats (COO, CSR, ELL, and HYB) on GPU. We used two popular machine learning algorithms, Support Vector Regression (SVR) and Multilayer Perceptron neural network (MLP). Our experimental results on two different GPUs (Fermi GTX 512 and Maxwell GTX 980 Ti) show that the SVR models deliver the best accuracy with average prediction error ranging between 7% and 14%. Akrem Benatia, Weixing Ji, Yizhuo Wang 0001, Feng Shi 0009 |
ICPADS | 3 |
| 2016 | Sparse Matrix Format Selection with Multiclass SVM for SpMV on GPUabstractSparse Matrix-Vector Multiplication (SpMV) kernel dominates the computing cost in numerous scientific applications. Many implementations based on different sparse formats were proposed recently for this kernel on the GPU side. Since the performance of these sparse formats varies significantly according to the sparsity characteristics of the input matrix and the hardware specifications, no one of them can be considered as the best one to use for every sparse matrix. In this paper, we address the problem of selecting the best representation for a given sparse matrix on GPU by using a machine learning approach. First, we present some interesting and easy to compute features for characterizing the sparse matrices on GPU. Second, we use a multiclass Support Vector Machine (SVM) classifier to select the best format for each input matrix. We consider in this paper four popular formats (COO, CSR, ELL, and HYB), but our work can be extended to support more sparse representations. Experimental results on two different GPUs (Fermi GTX 580 and Maxwell GTX 980 Ti) show that we achieved more than 98% of the performance possible with a perfect selection. Akrem Benatia, Weixing Ji, Yizhuo Wang 0001, Feng Shi 0009 |
ICPP | 3 |
| 2015 | Task Parallel Implementation of Matrix Multiplication on Multi-socket Multi-core Architectures
Yizhuo Wang 0001, Weixing Ji, Xu Chen 0016, Sensen Hu |
ICA3PP (3) | 1 |
| 2015 | Memory-Aware NoC Application Mapping Based on Adaptive Genetic Algorithm
Yizhuo Wang 0001, Zhibiao Zhang, Lifu Huang, Weixing Ji |
ICA3PP (1) | 1 |
| 2015 | Refactoring for Separation of Concurrent Concerns
Yang Zhang 0037, Dongwen Zhang, Weixing Ji, Yizhuo Wang 0001 |
ICA3PP (3) | 4 |
| 2014 | A Compilation and Run-Time Framework for Maximizing Performance of Self-scheduling Algorithms
Yizhuo Wang 0001, Laleh Aghababaie Beni, Alexandru Nicolau, Alexander V. Veidenbaum, Rosario Cammarota |
NPC | 1 |
| 2014 | An adaptive and hierarchical task scheduling scheme for multi-core clusters
Yizhuo Wang 0001, Yang Zhang 0037, Xiaojun Wang 0005, Xu Chen 0016, Weixing Ji, Feng Shi 0009 |
Parallel Comput. | 1 |
| 2013 | A work-stealing scheduling framework supporting fault toleranceabstractFault tolerance and load balancing are critical points for executing long-running parallel applications on multicore clusters. This paper addresses both fault tolerance and load balancing on multicore clusters by presenting a novel work-stealing task scheduling framework which supports hardware fault tolerance. In this framework, both transient and permanent faults are detected and recovered at task granularity. We incorporate task-based fault detection and recovery mechanisms into a hierarchical work-stealing scheme to establish the framework. This framework provides low-overhead fault-tolerance and optimal load balancing by fully exploiting task parallelism. Yizhuo Wang 0001, Weixing Ji, Feng Shi 0009, Qi Zuo |
DATE | 1 |
| 2012 | A fault tolerant self-scheduling scheme for parallel loops on shared memory systemsabstractAs the number of cores per chip increases, significant speedup for many applications could be achieved by exploiting loop level parallelism (LLP). Meanwhile, ever scaling device size makes multicore/multiprocessor systems suffer from increased reliability problems. Scheduling scheme plays a key role to exploit LLP. In existing dynamic loop scheduling schemes, self-scheduling is the most commonly used scheme1. This paper presents FTSS, a fault tolerant self-scheduling scheme which aims to execute parallel loops efficiently in the presence of hardware faults on shared memory systems. Our technique transforms a loop to ensure the correctness of the re-execution of loop iterations by buffering variables with anti-dependences, which make it possible to design a fault tolerant loop scheduling scheme without checkpointing. FTSS combines work-stealing with self-scheduling, and uses a bidirectional execution model when work is stolen from a faulty core. Experimental results show that FTSS achieve better load balancing than existing self-scheduling schemes. Compared with checkpoint/restart implementations that save a checkpoint before executing each chunk of iterations and restart the whole chunk running on a faulty core, FTSS exhibits better runtime performance. In addition, FTSS greatly outperforms existing self-scheduling schemes in terms of performance and stability in heavy loaded runtime environment. Yizhuo Wang 0001, Alexandru Nicolau, Rosario Cammarota, Alexander V. Veidenbaum |
HiPC | 1 |
| 2012 | Exploring Object-Level Parallelism on Chip Multi-processors
Weixing Ji, Yizhuo Wang 0001, Junqing Zhao |
ICA3PP (2) | 2 |
| 2012 | Communication Locality Analysis of Triplet-Based Hierarchical Interconnection Network in Chip Multiprocessor
Shahnawaz Talpur, Feng Shi 0009, Yizhuo Wang 0001 |
NPC | 3 |
| 2012 | Knowledge-Based Adaptive Self-Scheduling
Yizhuo Wang 0001, Weixing Ji, Feng Shi 0009, Qi Zuo, Ning Deng 0002 |
NPC | 1 |
| 2012 | A Hierarchical Work-Stealing Framework for Multi-core ClustersabstractWork-stealing has been widely used in task-based parallel programming for dynamic load balancing. The overhead of work-stealing on distributed memory systems is much higher than that on shared memory systems. To minimize the overhead of work-stealing on a multi-core cluster, we propose a hierarchical work-stealing framework, in which work-stealing is performed inside a node before across the node boundary. Two key techniques used in our framework to reduce the inter-node steals are: a) adaptive initial partitioning for different task parallel patterns; b) centralized control for inter-node work-stealing, which improves the efficiency of victim selection and termination detection. We compare our technique to the classical work-stealing scheme and a state-of-the-art work-stealing scheme [1] for multi-core clusters. Our technique outperforms them by 19% and 8% respectively. Yizhuo Wang 0001, Weixing Ji, Qi Zuo, Feng Shi 0009 |
PDCAT | 1 |
| 2009 | A Novel Adaptive Scratchpad Memory Management StrategyabstractScratchpad Memory (SPM) is a fast and small software-managed SRAM. Its current extensive uses in embedded processors are motivated by the advantages of power saving, small area and low access time compared with cache. However, existing SPM management methods depend heavily on profiling and compilers. The dependence on compiler also makes embedded applications hard to transplant. This paper presents a novel strategy to manage the scratchpad memory without compiler support. Based on the memory reference locality theory, a hardware random sampling module is adopted to dynamically identify the frequently accessed addresses at runtime. The consequential data movement and address redirection are handled by software operation with the assistance of memory management unit (MMU). We evaluate our method on 10 typical embedded applications and compare the results to a cache reference system. Experimental results show that, on average, our scheme can achieve 33:5% reduction in energy consumption with only slight (<1%) decrease in throughput versus the reference system. Ning Deng 0002, Weixing Ji, Feng Shi 0009, Yizhuo Wang 0001 |
RTCSA | 5 |