EDBT 2026 Demo / reviewers in the wild / expert
Xi Li 0003
dblp:46/2311-3
· DBLP profile ↗
102ranked-venue papers
3as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 70 · 3 first-author · 14 since 2021Software engineering, systems software and programming languages · 14 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Computer networks · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Late Breaking Results: A Fast Nearest Neighbor Search Acceleration for 3D Point CloudabstractThis paper presents FastNN, a novel accelerator architecture for efficient K-Nearest Neighbors (KNN) search in point clouds. FastNN leverages a locality-sensitive E2LSH partitioning method and a precomparator module to significantly reduce the candidate search space and minimize the number of Euclidean distance calculations. Compared to octree-based partitioning methods, our approach reduces candidate points by 58.57% to 86.17% and achieves a $10.04 \times$ acceleration in processing throughput relative to the BitNN comparator subsystem. The proposed design effectively enhances search throughput, resource utilization, and precision, highlighting its potential for accelerating KNN search on FPGA platforms. Jinao Li, Qianyu Cheng, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
DAC | 7 |
| 2025 | TSI: A Time-Semantic Instruction Set for Deterministic Data-Flow Execution in Real-Time Embedded SystemsabstractReal-Time Embedded Systems (RTES) are widely used in safety-critical devices, where deterministic data flow is essential to system verification and reliable execution. It requires that each consumer task instance reads data from the deterministic producer task instance. In software based on general-purpose computing instruction sets, communication-related instruction execution order couples data flow among tasks, necessitating a deterministic execution order of these instructions to preserve data-flow determinism. However, enforcing this order complicates software, and suffers from priority inversion and variable execution overheads, which significantly increases task worst-case response times (WCRT) and response time variability. This paper identifies the cause of above issues as the semantics of general-purpose instruction sets, under which dataflow determinism relies on the deterministic execution order of communication-related instructions. To address this, we make the following contributions. First, we propose Time-Semantic Instruction set (TSI), which supports memory access using both addresses and timestamps. TSI enables data-flow determinism without strict instruction ordering. Second, we design a TSI-enabled implementation compatible with conventional memory systems. Third, we provide two TSI-based deterministic data-flow programming paradigms, along with correctness proofs. Finally, we evaluate TSI hardware cost and implement a cycle-accurate simulator based on a TSI-extended RISC-V. Experiments demonstrate that, under reasonable memory overhead, our approach reduces programming complexity and achieves up to$21.6 \times$reduction in WCRT and up to$89.6 \times$reduction in response time variability compared to existing methods. Yinkang Gao, Yixuan Zhu, Lei Gong 0003, Wenqi Lou, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
RTSS | 8 |
| 2025 | Work-in-Progress: A Timing-Anomaly Free Dynamic Scheduling on Heterogeneous SystemsabstractHeterogeneous systems commonly adopt dynamic scheduling algorithms to improve resource utilization and enhance scheduling flexibility. However, it may introduce timing anomalies, wherein locally reduced tasks' actual execution times can lead to an increase in the overall system execution time. This phenomenon significantly complicates the analysis of WorstCase Response Time (WCRT), rendering conventional analysis either overly pessimistic or unsafe, and often necessitating exhaustive state-space exploration to ensure correctness. To address this challenge, this paper presents the first timing-anomalyfree dynamic scheduling algorithm for heterogeneous systems, referred to as Deterministic Dynamic Execution. The core idea is to apply deterministic execution constraints, which partially restrict the resource allocation and execution order of tasks at runtime. It achieves a safe and tight WCRT through a single offline simulation execution. In this paper, we provide preliminary experimental validation of the timing-anomaly-free property of our algorithm and outline the basic idea of a formal proof. Yixuan Zhu, Yinkang Gao, Binze Jiang, Xiaohang Gong, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
RTSS | 7 |
| 2025 | Optimizing utilization in logical execution time system with preserved externally-observable timed I/O semantics
Caixu Zhao, Yinkang Gao, Yixuan Zhu, Lei Gong 0003, Wenqi Lou, Xi Li 0003 |
J. Syst. Archit. | 8 |
| 2025 | Hardware Accelerated Vision Transformer via Heterogeneous Architecture Design and Adaptive Dataflow MappingabstractVision transformer (ViT) models have demonstrated remarkable advantages in visual tasks. However, the ViT model contains various types of operators, and its sophisticated model structure imposes substantial computational complexity and storage burden. Existing hardware solutions still fail to fully unleash the ViT acceleration potential due to the mismatch between operators and hardware architectures, suffering from inefficient dataflow mapping. This work proposes HDViT, a full-fledged heterogeneous hardware accelerator on FPGA, to enhance the ViT acceleration by comprehensively analyzing and addressing the challenges of heterogeneous architecture design. Specifically, HDViT first develops a heterogeneous architecture design that is composed of multiple processing engines (PEs) to accelerate various operators in the ViT model. Then, HDViT devises a hybrid-oriented dataflow mapping strategy to reduce data transmission granularity and alleviate storage resource pressure. Lastly, to achieve the latency balancing among multiple PEs, we formulate the HDViT architecture and implement an automated exploration process to identify optimized parallelism parameters that satisfy computation and storage demands while enhancing the heterogeneous architectural performance. Experimental results indicate that HDViT achieves significant performance speedups of 2.16$\times$and 3.51$\times$compared to previous heterogeneous and unified accelerators, respectively. HDViT also achieves a maximum of 98.46% hardware utilization. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Dong Dai 0001, Yang Yang 0080, Xianglan Chen, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Computers | 8 |
| 2025 | Advancing Neuromorphic Architecture Toward Emerging Spiking Neural Network on FPGAabstractSpiking neural networks (SNNs) replace the multiply-and-accumulate operations in traditional artificial neural networks (ANNs) with lightweight mask-and-accumulate operations, achieving greater performance. Existing SNN architectures are primarily designed based on fully-connected or convolutional SNN topologies and still struggle with low task accuracy, limiting their practical applications. Recently, transformer SNN (TSNN) models have shown promise in matching the accuracy of nonspiking ANNs and demonstrated potential application prospects. However, their diverse computation pattern and sophisticated network structure with high computation and memory footprints impede their efficient deployment. Thus, in this work, we move our attention to heterogeneous architecture design and propose SpikeTA, the first neuromorphic hardware accelerator explicitly designed for the TSNN model on FPGA. First, SpikeTA enables parameterizable hardware engines (HEs) designed for the network layers in TSNN, enhancing compatibility between HEs and network layers. Second, SpikeTA optimizes arithmetic operations between binary spikes and synaptic weights by presenting a DSP-efficient addition tree. By analyzing the inherent data characteristics, SpikeTA further introduces a depth-aware buffer management strategy to provide sufficient access ports. Third, SpikeTA employs a streaming dataflow mapping to optimize data transmission granularity and leverages a split-engine dataflow mapping to facilitate pipelined latency balancing. Experimental results demonstrate that SpikeTA achieves significant performance speedups of$140.73\times $–$1023.53\times $and$2.97\times $–$7.29\times $over architectures running on the AMD EPYC 7542 CPU and NVIDIA A100 GPU, respectively. SpikeTA also outperforms state-of-the-art SNN and Transformer accelerators by$2.79\times $and$2.66\times $in architecture performance while achieving a peak performance of 28.99 TOPs. Yingxue Gao, Yang Yang 0080, Lei Gong 0003, Xianglan Chen, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | Magnifier: A Chiplet Feature-Aware Test Case Generation Method for Deep Learning AcceleratorsabstractThe development of deep learning has led to increasing demands for computation and memory, making multi-chiplet accelerators a powerful solution. Multi-chiplet accelerators require more precise consideration of hardware configurations and mapping schemes in terms of computation, memory, and communication patterns compared to monolithic designs, in order to avoid underutilization of performance. However, there is currently a lack of performance testing methods specifically tailored for multi-chiplet accelerators. Existing testing methods primarily focus on correctness testing and do not address potential performance issues from a hardware perspective. To address these issues, this paper proposes Magnifier: a test case generation method for performance testing of multi-chiplet accelerators. Firstly, we analyze typical multi-chiplet accelerator prototype from the perspectives of computation, memory, and communication patterns, and summarize a chiplet feature-aware operator task set. Next, we define the test evaluation metric IPPstd and use a candidate operator set to construct a sampling space for model-level test cases. Finally, we build a GAN to learn the distribution of high-diversity test cases, enabling the rapid generation of high-quality test cases. We validate the proposed method on both simulated and real multi-chiplet accelerators. Experiments show that Magnifier can improve the metric of test cases by up to 3.42 times and significantly reduce generation time, providing valuable insights for optimizing the hardware and software of multi-chiplet accelerators. Boyu Li 0006, Zongwei Zhu, Weihong Liu, Qianyue Cao, Changlong Li 0006, Cheng Ji 0002, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | Fine-Grained Shared Cache Interference Analysis Using Basic Block's Execution TimeabstractShared last-level caches in multi-core architectures may lead to mutual interference in memory accesses between cores, resulting in additional access latency. To ensure the accuracy of programs' worst-case execution time (WCET) analysis, the inter-core interference of shared caches must be analyzed. This involves getting the time when memory access may occur between cores (i.e., inter-core context). But due to the uncertainty of inputs, describing inter-core context becomes extremely challenging. Existing works often consider the entire execution time of program as the time when all memory accesses may occur to enhance scalability, which ignores the timing information within programs and may lead to an overestimation of interference. This paper proposes a fine-grained shared cache interference analysis method based on the execution time of basic blocks. It employs a path-based strategy to estimate the execution time of basic blocks and use it as the lifecycles of memory accesses within the blocks, which can effectively eliminate impossible interferences. Experiments show that compared to the all interference method, we can reduce WCET by 65% in the best case and by 14% on average, with only a 130% increase in average analysis time. Yixuan Zhu, Wenqi Lou, Yinkang Gao, Binze Jiang, Xiaohang Gong, Xi Li 0003 |
ICCD | 6 |
| 2024 | Enhancing Graph Random Walk Acceleration via Efficient Dataflow and Hybrid Memory ArchitectureabstractGraph random walk sampling is becoming increasingly important with the widespread popularity of graph applications. It aims to capture the desirable graph properties by launching multiple walkers to collect feature paths. However, previous research suffers long sampling latency and severe memory access bottlenecks due to intrinsic data dependency and skewed vertex distribution. Thus, in this paper, we propose FastRW, a dedicated accelerator to boost graph random walk operation on FPGAs. Specifically, FastRW first integrates multiple parallel processing engines to achieve data-level parallelism, where each processing engine also leverages dataflow scheduling to resolve data dependency and hide long sampling latency. Secondly, FastRW leverages a combination of multiple storage resources to implement a hybrid memory architecture adapted to skewed vertex distribution. By integrating the above optimizations, FastRW develops a performance model to take advantage of the balance between computation parallelism and bandwidth demand. We evaluate FastRW with two classic sampling algorithms on a wide range of real-world graph datasets. The experimental results show that FastRW achieves a speedup of 37.52$\boldsymbol{\times}$on average over the system running on two 8-core Intel CPUs. FastRW also achieves an average of 28.04$\boldsymbol{\times}$speedup over the architecture implemented on V100 GPU. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Yiqing Hu, Zhongming Liu, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Computers | 8 |
| 2023 | Work-in-Progress: NAPMAE: Generalized Data-Efficient Neural Architecture Predictor with Masked Autoencoder
Qiaochu Liang, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou, Xi Li 0003 |
CODES+ISSS | 5 |
| 2023 | FastRW: A Dataflow-Efficient and Memory-Aware Accelerator for Graph Random Walk on FPGAsabstractGraph random walk (GRW) sampling is becoming increasingly important with the widespread popularity of graph applications. It involves some walkers that wander through the graph to capture the desirable properties and reduce the size of the original graph. However, previous research suffers long sampling latency and severe memory access bottlenecks due to intrinsic data dependency and irregular vertex distribution. This paper proposes FastRW, a dedicated accelerator to release GRW acceleration on FPGAs. FastRW first schedules walkers' execution to address data dependency and mask long sampling latency. Then, FastRW leverages pipeline specialization and bit-level optimization to customize a processing engine with five modules and achieve a pipelining dataflow. Finally, to alleviate the differential accesses caused by irregular vertex distribution, FastRW implements a hybrid memory architecture to provide parallel access ports according to the vertex's degree. We evaluate FastRW with two classic GRW algorithms on a wide range of real-world graph datasets. The experimental results show that FastRW achieves a speedup of 14.13× on average over the system running on two 8-core Intel CPUs. FastRW also achieves 3.28×∼198.24× energy efficiency over the architecture implemented on V100 GPU. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
DATE | 5 |
| 2023 | NeuralMAE: Data-Efficient Neural Architecture Predictor with Masked Autoencoder
Qiaochu Liang, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou, Xi Li 0003 |
PRCV (8) | 5 |
| 2023 | Enabling Fast and Memory-Efficient Acceleration for Pattern Matching Workloads: The Lightweight Automata Processing EngineabstractGrowing pattern matching applications are employing finite automata as their basic processing model. These applications match tens to thousands of patterns on a large amount of data, which brings a great challenge to conventional processors. Therefore hardware-based solutions have emerged frequently and achieved high throuphput automata processing. However, existing methods are generally difficult to achieve both processing speed and storage efficiency, and are often too heavy to be integrated into a small chip and have to rely on off-chip DRAMs or other high capacity memories even on some simple data sets, leading to the potential area and power consumption issues. In this paper, we focus on building a more lightweight automata processing engine, hoping to store the whole automata model into on-chip memory and run effectively and independently. We propose LAP, a lightweight automata processing engine. Powered with a novel automata model (A-DFA) and efficient packing algorithms, extremely high storage efficiency compared with traditional DFA is achieved in LAP. Meanwhile, we identify the key parallelization factors in the A-DFA model and then propose a specialized microarchitecture with novel instructions to further accelerate the state transition process. As a result, LAP can obtain more effective trade-off between processing speed and storage efficiency. Evaluation results show that LAP achieves extremely high storage efficiency on simple data sets, exceeding IBM's RegX by 8×, and achieves significant improvements in processing speed ranging from 1.32× to 1.91× compared with previous lightweight hardware implementations. Moreover, LAP has good scalability in hardware architecture. It is easy to build an acceleration system with higher throughput by increasing the number of cores. We prototype a 16-core system into Xilinx ZC702 FPGA and a 64-core system into Xilinx ZCU102 FPGA respectively. The prototype system on ZC702 on average achieves 3.5 GB/s throughput on simple data sets, and the prototype system on ZCU102 can obtain higher throughput and compute density values on part of large datasets in ANMLZoo compared with modern in-memory NFA-based solutions. Lei Gong 0003, Chao Wang 0003, Haojun Xia, Xianglan Chen, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Computers | 5 |
| 2023 | Algorithm/Hardware Co-Optimization for Sparsity-Aware SpMM Acceleration of GNNsabstractIn recent years, graph neural networks (GNNs) have achieved impressive performance in various application fields by extracting information from graph-structured data. It contains extensive feature aggregation operations and has become a performance bottleneck, which can be abstracted as a specialized sparse-dense matrix multiplication (SpMM) operation. Previous works have leveraged the inner product or outer product to accelerate the feature aggregation process. However, inefficient execution leads to extremely unbalanced workloads and extensive intermediate data, hampering the performance of previous processors. So in this article, we demonstrate an algorithm/hardware co-optimization chance to enhance SpMM acceleration for GNNs. First, the algorithm part develops a dataflow-efficient SpMM algorithm that integrates three optimization methods to mitigate computation and memory access inefficiencies. Specifically, 1) the proposed equal-value partition method achieves fine-grained data partition and enables load balancing during data movement; 2) after observing the vertex aggregation phenomenon, a vertex-clustering optimization method is presented to enable significant data locality; and 3) the adaptive dataflow based on Gustavson’s algorithm is further implemented to enable the efficient distribution of sparse elements and improves computing resource utilization. Then, the hardware part features the proposed SpMM algorithm and customizes SDMA, a flexible and efficient accelerator to boost SpMM acceleration, which follows the adaptive dataflow to eliminate sparsity and explore the regular parallelism dimension. Finally, we prototype SDMA on the Xilinx Alveo U280 FPGA accelerator card. The results demonstrate that SDMA achieves$5.68\times $–$14.68\times $energy efficiency over the previous GPU implementations deployed on the Nvidia GTX 1080Ti and$1.32\times $higher throughput over the state-of-the-art FPGA prototype. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | IRGA: An Intelligent Implicit Real-time Gait Authentication System in Heterogeneous Complex ScenariosabstractGait authentication as a technique that can continuously provide identity recognition on mobile devices for security has been investigated by academics in the community for decades. However, most of the existing work achieves insufficient generalization to complex real-world environments due to the complexity of the noisy real-world gait data. To address this limitation, we propose an intelligent Implicit Real-time Gait Authentication (IRGA) system based on Deep Neural Networks (DNNs) for enhancing the adaptability of gait authentication in practice. In the proposed system, the gait data (whether with complex interference signals) will first be processed sequentially by the imperceptible collection module and data preprocessing module for improving data quality. In order to illustrate and verify the suitability of our proposal, we provide analysis of the impact of individual gait changes on data feature distribution. Finally, a fusion neural network composed of a Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) is designed to perform feature extraction and user authentication. We evaluate the proposed IRGA system in heterogeneous complex scenarios and present start-of-the-art comparisons on three datasets. Extensive experiments demonstrate that the IRGA system achieves improved performance simultaneously in several different metrics. Li Yang 0005, Xi Li 0003, Zhuoru Ma, Lu Li 0008, Naixue Xiong, Jianfeng Ma 0001 |
ACM Trans. Internet Techn. | 2 |
| 2022 | Work-in-Progress: Scheduler for Collaborated FPGA-GPU-CPU Based on Intermediate LanguageabstractFPGA-GPU-CPU collaboration compromise high performance and low cost in modern computing systems. However, the large mapping space between modules and heterogeneous processors brings complexity to the scheduling algorithm. This paper proposes a uniform-pipeline-based real-time oriented scheduling algorithm and a servant execution-flow model (SEFM) optimized for this scheduler. SEFM at runtime generates the target code from the intermediate language (IL) and scheduler-controlled parameters. The algorithms such as contrast stretching, etc., are accelerated by 1.4-2.7×, 1.9-3.8×, 2.7-10.5× respectively on CPU, GPU, and FPGA over OpenCV baseline. A case study of 3D waveform oscilloscope using scheduling solution on collaborated processors achieves 1.5× resource utilization than the pure FPGA. Chao Wang 0003, Xuehai Zhou, Xi Li 0003 |
CODES+ISSS | 4 |
| 2021 | UH-JLS: A Parallel Ultra-High Throughput JPEG-LS Encoding Architecture for Lossless Image CompressionabstractThe lossless image compression technique has a great application value in distortion-sensitive applications. JPEG-LS, as a mature lossless compression standard, is widely adopted for its excellent compression ratio. Many hardware JPEG-LS compressors are proposed on FPGAs and ASICs to achieve high energy efficiency and low cost. However, JPEG-LS has a contextual Read-After-Write (RAW) issue, making previous hardware either insufficiently explore its parallelism potential or induce other defects while parallelizing, such as compression ratio dropping and compatibility problems. In this paper, we propose a hardware/software co-design method for high-performance JPEG-LS compressor design. At the software level, we propose a pixel grouping scheduling scheme and the Pseudo-LS method to tap the parallelism aiming at the RAW issue. At the hardware level, we discuss the high-performance design methods of these software-level schemes and propose a design space exploration method to constrain the resource usage introduced by parallelization. To our knowledge, our architecture, UH-JLS, is the first pixel-level parallelization streaming image compressor based on the standard JPEG-LS. The experiments show that in the lossless manner and the Pseudo-LS manner, UH-JLS respectively achieves 5.6x and 7.1x speedup than the previous state-of-the-art FPGA-based JPEG-LS compressor. Xuan Wang 0020, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
ICCD | 4 |
| 2021 | GenSeq+: A Scalable High-Performance Accelerator for Genome SequencingabstractGenome sequencing is one of the most challenging problems in computational biology and bioinformatics. As a traditional algorithm, the string match meets a challenge with the development of the massive volume of data because of gene sequencing. Surveys show that there will be a huge amount of short read segments during the process of gene sequencing and the need for a highly efficient is urgent. As a classic fast and exact single pattern matching algorithm, Knuth-Morris-Pratt (KMP) algorithm has been demonstrated in network security and computational biology. However, with the increasing amount of data in the modern society, it becomes increasingly important and essential to provide a High-performance implementation of KMP algorithm. In this article, we implement a scalable KMP accelerator based on FPGA, named GeneKMP. The accelerator is composed of different computing units to achieve a pipelined organization for higher throughput with satisfying scalability. A novel programming model is provided to alleviate the burden of the high-level programmers. We provide a greedy-based partitioning algorithm for the software/hardware design paradigms. Experimental results on the state-of-the-art Xilinx FPGA hardware prototype show that our accelerator can achieve up to a promising speedup with insignificant hardware cost and power consumption. Chao Wang 0003, Lei Gong 0003, Shiming Lei, Haijie Fang, Xi Li 0003, Aili Wang 0003, Xuehai Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | Improving HW/SW Adaptability for Accelerating CNNs on FPGAs Through A Dynamic/Static Co-Reconfiguration ApproachabstractWith the continuous evolution of Convolutional Neural Networks (CNNs) and the improvement of the computing capability of FPGAs, the deployment of CNN accelerator based on FPGA has become more and more popular in various computing scenarios. The key element of implementing these accelerators is to take full advantage of underlying hardware characteristics to adapt to the computational features of the software-level CNN model. To achieve this goal, however, previous designs mainly focus on the static hardware reconfiguration pattern, which is not flexible enough and can hardly make the accelerator architecture and the CNN features fully fit, resulting in inefficient computations and data communications. By leveraging the dynamic partial reconfiguration technology equipped in the modern FPGA devices, in this article, we propose a new accelerator architecture for implementing CNNs on FPGAs in which static and dynamic reconfigurabilities of the hardware are cooperatively utilized to maximize the acceleration efficiency. Based on this architecture, we further present a systematic design and optimization methodology for implementing the specific CNN model in the particular computing scenario, in which a static design space exploration method and a reinforcement learning-based decision method are proposed to obtain the optimal static hardware configuration and run-time reconfiguration strategy respectively. We evaluate our proposal by implementing three widely used CNN models, AlexNet, VGG16C, and ResNet34, on the Xilinx ZCU102 FPGA platform. Experimental results show that our implementations on average can achieve 683 GOPS under 16-bit fixed data type and 1.37 TOPS under 8-bit fixed data type for three targeted CNN models, and improve the computational density from 1.1× to 1.91× compared with previous implementations on the same type of FPGA platform. Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | SOLAR: Services-Oriented Deep Learning Architectures-Deep Learning as a ServiceabstractDeep learning has been an emerging field of machine learning during past decades. However, the diversity and large scale data size have posed significant challenge to construct a flexible and high performance implementations of deep learning neural networks. In order to improve the performance as well to maintain the scalability, in this paper we present SOLAR, a services-oriented deep learning architecture using various accelerators like GPU and FPGA. SOLAR provides a uniform programming model to users so that the hardware implementation and the scheduling is invisible to the programmers. At runtime, the services can be executed either on the software processors or the hardware accelerators. To leverage the trade-offs between the metrics among performance, power, energy, and efficiency, we present a multitarget design space exploration. Experimental results on the real state-of-the-art FPGA board demonstrate that the SOLAR is able to provide a ubiquitous framework for diverse applications without increasing the burden of the programmers. Moreover, the speedup of the GPU and FPGA hardware accelerator in SOLAR can achieve significant speedup comparing to the conventional Intel i5 processors with great scalability. Chao Wang 0003, Lei Gong 0003, Xi Li 0003, Aili Wang 0003, Patrick C. K. Hung, Xuehai Zhou |
IEEE Trans. Serv. Comput. | 3 |
| 2020 | WooKong: A Ubiquitous Accelerator for Recommendation Algorithms With Custom Instruction Sets on FPGAabstractRecommendation algorithms, such as Neighborhood-based Collaborative- Filtering (CF), have been widely applied in various emerging machine learning applications. However, under the circumstance of the explosive big data, it poses significant challenges to CF recommendation algorithms as it is becoming quite time and energy-consuming. It has to be optimized and accelerated by powerful engines to process on large data scale. To solve these problems, in this article, we propose WooKong, a ubiquitous accelerator architecture for the collaborative-filtering recommendation on FPGA. It is able to accommodate three types of CF recommendation algorithms, including User-based CF, Item-based CF, and SlopeOne recommendations algorithms, with five different similarity analysis metrics including Jaccard, Cosine, CosineIR, euclidean, and Pearson. To maintain flexibility for these different CF algorithms and metrics, we adopt custom instruction sets to manipulate the learning and prediction accelerators. We implement a hardware prototype on a real Xilinx Zynq FPGA development board. Experimental results show that the proposed learning and prediction accelerators can achieve 8.0X speedup and 1.7X speedup compared with an Intel i7 processor respectively. The accelerator has the energy benefits of up to 137.4X compared with an NVIDIA Tesla K40C GPU, with the affordable hardware cost. Chao Wang 0003, Lei Gong 0003, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Computers | 4 |
| 2020 | A Ubiquitous Machine Learning Accelerator With Automatic Parallelization on FPGAabstractMachine learning has been widely applied in various emerging data-intensive applications, and has to be optimized and accelerated by powerful engines to process very large scale data. Recently, the instruction set based accelerators on Field Progarmmable Gate Arrays (FPGAs) have been a promising topic for machine learning applications. The customized instructions can be further scheduled to achieve higher instruction-level parallelism. In this article, we design a ubiquitous accelerator with out-of-order automatic parallelization for large-scale data-intensive applications. The accelerator accommodates four representative applications, including clustering algorithms, deep neural networks, genome sequencing, and collaborative filtering. In order to improve the coarse-grained instruction-level parallelism, the accelerator employs an out-of-order scheduling method to enable parallel dataflow computation. We use Colored Petri Net (CPN) tools to analyze the dependences in the applications, and build a hardware prototype on the real FPGA platform. For cluster applications, the accelerator can support four different algorithms, including K-Means, SLINK, PAM, and DBSCAN. For collaborative filtering applications, it accommodates Tanimoto, euclidean, Cosine, and Pearson Correlation as Similarity metrics. For deep learning applications, we implement hardware accelerators for both training process and inference process. Finally, for genome sequencing, we design a hardware accelerator for the BWA-SW algorithm. Experimental results show that the accelerator architecture can reach up to 25X speedup against Intel processors with affordable hardware cost, insignificant power consumption, and high flexibility. Chao Wang 0003, Lei Gong 0003, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | DCW: A Reactive and Predictable Programming Framework for LET-Based Distributed Real-Time SystemsabstractReal-time systems continuously interact with the physical environment and often have to satisfy stringent timing constraints imposed by their interactions. Those systems involve two main properties: reactivity and predictability. Reactivity allows the system to continuously react to a non-deterministic external environment, while predictability guarantees the deterministic execution of safety-critical parts of applications. However, with the increase in software complexity, traditional approaches to develop real-time systems make temporal behaviors difficult to infer, especially when the system is required to address non-deterministic aperiodic events from the physical environment. In this article, we propose a reactive and predictable programming framework, Distributed Clockwerk (DCW), for distributed real-time systems. DCW introduces the Servant, which is a non-preemptible execution entity, to implement periodic tasks based on the Logical Execution Time (LET) model. Furthermore, a joint schedule policy, based on the slack stealing algorithm, is proposed to efficiently address aperiodic events with no violated hard-time constraints. To further support predictable communication among distributed nodes, DCW implements the Time-Triggered Controller Area Network (TTCAN) to avoid collisions while accessing the shared communication medium. Moreover, a programming framework implements to provide a set of programming APIs for defining timing and functional behaviors of concurrent tasks. An example is further implemented to illustrate the DCW design flow. The evaluation results demonstrate that our proposal can improve both periodic and aperiodic reactivity compared with existing work, and the implemented DCW can also ensure the system predictability by achieving extremely low overheads. Xi Li 0003, Caixu Zhao, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2018 | Delayed Wake-Up Mechanism Under Suspend Mode of Smartphone
Bo Chen 0010, Xi Li 0003, Xuehai Zhou, Zongwei Zhu |
CollaborateCom | 2 |
| 2018 | RTMUSRT: a real-time testbed for empirically comparing real-time multicore schedulers: work-in-progressabstractIn this paper, we present a real-time testbed RTMUSRTto eliminate unpredictable behaviors and improve dependability when empirically evaluating multicore real-time scheduling algorithms. Experimental results are obtained by measuring kernel overheads of RTMUSRT, and demonstrate that RTMUSRThas fewer disturbances in task executions than other existing work. Xi Li 0003, Kaiqi Zhou, Caixu Zhao, Chao Wang 0003, Xuehai Zhou |
EMSOFT | 3 |
| 2018 | Domino: An Asynchronous and Energy-efficient Accelerator for Graph Processing: (Abstract Only)abstractLarge-scale graphs processing, which draws attentions of researchers, applies in a large range of domains, such as social networks, web graphs, and transport networks. However, processing large-scale graphs on general processors suffers from difficulties including computation and memory inefficiency. Therefore, the research of hardware accelerator for graph processing has become a hot issue recently. Meanwhile, as a power-efficiency and reconfigurable resource, FPGA is a potential solution to design and employ graph processing algorithms. In this paper, we propose Domino, an asynchronous and energy-efficient hardware accelerator for graph processing. Domino adopts the asynchronous model to process graphs, which is efficient for most of the graph algorithms, such as Breadth-First Search, Depth-First Search, and Single Source Shortest Path. Domino also proposes a specific data structure based on row vector, named Batch Row Vector, to present graphs. Our work adopts the naive update mechanism and bisect update mechanism to perform asynchronous control. Ultimately, we implement Domino on an advanced Xilinx Virtex-7 board, and experimental results demonstrate that Domino has significant performance and energy improvement, especially for graphs with a large diameter(e.g., roadNet-CA and USA-Road). Case studies in Domino achieve 1.47x-7.84x and 0.47x-2.52x average speedup for small-diameter graphs(e.g., com-youtube, WikiTalk, and soc-LiveJournal), over GraphChi on the Intel Core2 and Core i7 processors, respectively. Besides, compared to Intel Core i7 processors, Domino also performs significant energy-efficiency that is 2.03x-10.08x for three small-diameter graphs and 27.98x-134.50x for roadNet-CA which is a graph with relatively large diameter. Chongchong Xu, Chao Wang 0003, Yiwei Zhang 0001, Lei Gong 0003, Xi Li 0003, Xuehai Zhou |
FPGA | 5 |
| 2018 | MuDBN: An Energy-Efficient and High-Performance Multi-FPGA Accelerator for Deep Belief NetworksabstractWith the increasing size of neural networks, state-of-the-art deep neural networks (DNNs) have hundreds of millions of parameters. Due to multiple fully-connected layers, DNNs are compute-intensive and memory-intensive, making them hard to deploy on embedded devices with limited power budgets and hardware resources. Therefore, this paper presents a deep belief network accelerator based on multi-FPGA. Two different schemes, the division between layers (DBL) and the division inside layers (DIL), are adopted to map the DBN to the multi-FPGA system. Experimental results demonstrate that the accelerator can achieve 4.24x (DBL) -6.20x (DIL) speedup comparing to the Intel Core i7 CPU and save 119x (DBL) -90x (DIL) power consumption comparing to the Tesla K40C GPU. Yuming Cheng, Chao Wang 0003, Xianglan Chen, Xuehai Zhou, Xi Li 0003 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2018 | Domino: Graph Processing Services on Energy-Efficient Hardware AcceleratorabstractLarge-scale graphs processing, which draws attentions of researchers, applies in a large range of domains. However, large-scale graphs processing on traditional platforms suffers from difficulties including computation and memory inefficiency. To enhance the computation-efficiency and energy-efficiency, in this paper, we exploit graph processing services on the energy-efficient hardware accelerator, called Domino. Domino adopts the asynchronous model to process graphs, which is efficient for many graph algorithms, such as Breadth-First Search, Depth-First Search, and Single Source Shortest Path. Domino also proposes a specific data structure based on row vectors to present graphs, named Batch Row Vector. Besides, our work employs naive update mechanism and bisect update mechanism to perform the asynchronous control. Ultimately, we implement Domino on an advanced Xilinx Virtex-7 board, and experimental results demonstrated that Domino has a significant performance and energy improvement. Case studies in Domino achieve 1.47x-7.84x and 0.47x-2.52x average speedup for small-diameter graphs(e.g., com-youtube, WikiTalk, and soc-LiveJournal), over GraphChi on the Intel Core2 and Core i7 processors, respectively. Besides, compared to Intel Core i7 processors, Domino also performs a significant energy-efficiency that is 2.03x-10.08x for three small-diameter graphs and 27.98x-134.50x for roadNet-CA which is a graph with relatively large diameter. Chongchong Xu, Chao Wang 0003, Lei Gong 0003, Lihui Jin, Xi Li 0003, Xuehai Zhou |
ICWS | 5 |
| 2018 | Model checking of MARTE/CCSL time behaviors using timed I/O automata
Bo Chen 0010, Xi Li 0003, Xuehai Zhou |
J. Syst. Archit. | 2 |
| 2018 | Refactoring Network Functions Modules to Reduce Latencies and Improve Fault Tolerance in NFVabstractNetwork functions virtualization (NFV) allows service providers to deliver new services to their customers more quickly by adopting software-centric network functions implementation over commercial, off-the-shelf hardwares. This NFV-based software-centric approach cannot use dedicated mechanisms implemented over custom built boxes to reduce latencies and tolerate faults. We present a case study of IP multimedia subsystem (IMS), which is the most complex NFV instance, requires extremely low end-to-end latency (40 msec), and demands system availability as high as five nines. Through an empirical study, we discover that highly modular IMS network functions implementation over virtualized platform: 1) incurs latencies and 2) does not tolerate faults. NFV-based IMS modules incur high latencies by creating a feedback loop among each other while executing delay sensitive data-plane traffic. These IMS modules are also susceptible to failure, causing the control-plane to terminate the application session while keeping the data-plane to forward data packets. To address these issues, we propose to refactor network function modules. We reduce latencies by pipelining the IMS modules, and recover failed modules by reconfiguring their neighboring modules. We build our system prototype of open source IMS over OpenStack platform. Our results show that our scheme reduces latencies and failure recovery time up to 12× and 10×, respectively, when compared with the state-of-the-art virtualized IMS implementation. Muhammad Taqi Raza, Songwu Lu, Mario Gerla, Xi Li 0003 |
IEEE J. Sel. Areas Commun. | 4 |
| 2018 | MALOC: A Fully Pipelined FPGA Accelerator for Convolutional Neural Networks With All Layers Mapped on ChipabstractRecently, field-programmable gate arrays (FPGAs) have been widely used in the implementations of hardware accelerator for convolutional neural networks (CNNs). However, most of these existing accelerators are designed in the same idea as their ASIC counterparts, in which all operations from different layers are mapped to the same hardware units and working in a multiplexed way. This manner does not take full advantage of reconfigurability and customizability of FPGAs, resulting in a certain degree of computational efficiency degradation. In this paper, we propose a new architecture for FPGA-based CNN accelerator that maps all the layers to their own on-chip units and working concurrently as a pipeline. A comprehensive mapping and optimizing methodology based on establishing roofline model oriented optimization model is proposed, which can achieve maximum resource utilization as well as optimal computational efficiency. Besides, to ease the programming burden, we propose a design framework which can provide a one-stop function for developers to generate the accelerator with our optimizing methodology. We evaluate our proposal by implementing different modern CNN models on Xilinx Zynq-7020 and Virtex-7 690t FPGA platforms. Experimental results show that our implementations can achieve a peak performance of 910.2 GOPS on Virtex-7 690t, and 36.36 GOP/s/W energy efficiency on Zynq-7020, which are superior to the previous approaches. Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Huaping Chen 0001, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Device-Customized Multi-Carrier Network Access on Commodity SmartphonesabstractAccessing multiple carrier networks (T-Mobile, Sprint, AT&T, and so on) offers a promising paradigm for smartphones to boost its mobile network quality. However, the current practice does not achieve the full potential of this approach because it has not utilized fine-grained, cellular-specific domain knowledge. Our experiments and code analysis discover three implementation-independent issues: 1) it may not trigger the anticipated switch when the serving carrier network is poor; 2) the switch takes a much longer time than needed; and 3) the device fails to choose the high-quality network (e.g., selecting 3G rather than 4G). To address them, we propose iCellular, which exploits low-level cellular information at the device to improve multi-carrier access. iCellular is proactive and adaptive in its multi-carrier selection by leveraging existing end-device mechanisms and standards-complaint procedures. It performs adaptive monitoring to ensure responsive selection and minimal service disruption and enhances carrier selection with online learning and runtime decision fault prevention. It is readily deployable on smartphones without infrastructure/hardware modifications. We implement iCellular on commodity phones and harness the efforts of Project Fi to assess multi-carrier access over two U.S. carriers: T-Mobile and Sprint. Our evaluation shows that, iCellular boosts the devices' throughput with up to 3.74× throughput improvement, 6.9× suspension reduction, and 1.9× latency decrement over the state of the art, with moderate CPU, and memory and energy overheads. Yuanjie Li, Chunyi Peng 0001, Haotian Deng 0001, Zengwen Yuan, Guan-Hua Tu, Songwu Lu, Xi Li 0003 |
IEEE/ACM Trans. Netw. | 8 |
| 2017 | Clockwerk: A Predictable and Efficient Extension of Logical Execution Time ModelabstractReal-time systems focus on achieving fast response time, which, however, is usually accompanied by sacrificing the predictability. The Logical Execution Time (LET) model achieves I/O-predictability by fixing the timing of reading inputs and writing outputs. However, LET lacks the mechanism to react to aperiodic requests which have frequently been used in soft real-time applications. To improve aperiodic responsiveness as well as obtain predictability, in this paper, we present a component-based LET extension-Clockwerk. Clockwerk introduces the concept of Servant as basic components to make both communication and computation parts of periodic tasks predictable. By combining the LET model with the aperiodic server, Clockwerk achieves the improvement of aperiodic responsiveness. Consequently, our proposal can reduce the complexity of schedulability analysis on periodic tasks to polynomial time, and preliminary simulation results show that it also achieves an aperiodic responsiveness speedup of 4.1X-6.5X compared with existing work. Xi Li 0003, Kaiqi Zhou, Haizhao Luo, Chao Wang 0003, Xianglan Chen, Xuehai Zhou |
APSEC | 2 |
| 2017 | A Power-Efficient Accelerator for Convolutional Neural NetworksabstractConvolutional neural networks(CNNs) have been widely applied in various applications. However, the computation-intensive convolutional layers and memory-intensive fully connected layers have brought many challenges to the implementation of CNN on embedded platforms. To overcome this problem, this work proposes a power-efficient accelerator for CNNs, and different methods are applied to optimize the convolutional layers and fully connected layers. For the convolutional layer, the accelerator first rearranges the input features into matrix on-the-fly when storing them to the on-chip buffers. Thus the computation of convolutional layer can be completed through matrix multiplication. For the fully connected layer, the batch-based method is used to reduce the required memory bandwidth, which also can be completed through matrix multiplication. Then a two-layer pipelined computation method for matrix multiplication is proposed to increase the throughput. As a case study, we implement a widely used CNN model, LeNet-5, on an embedded device. It can achieve a peak performance of 34.48 GOP/s and the power efficiency with the value of 19.45 GOP/s/W under 100MHz clock frequency which outperforms previous approaches. Chao Wang 0003, Lei Gong 0003, Chongchong Xu, Yiwei Zhang 0001, Yuntao Lu, Xi Li 0003, Xuehai Zhou |
CLUSTER | 7 |
| 2017 | OmniGraph: A Scalable Hardware Accelerator for Graph ProcessingabstractLarge-scale graphs processing attracts more and more attentions, and it has been widely applied in many application domains. FPGA is a promising platform to implement graph processing algorithms with high power-efficiency and parallelism. In this paper, we propose OmniGraph, a scalable hardware accelerator for graph processing. OmniGraph can process graphs with different sizes adaptively and is adaptable to various graph algorithms. OmniGraph improves the preprocessing methodology based on Interval-Shard and consists of three computation engines, vertices on-chip && edges on-chip engine, vertices on-chip && edges off-chip engine, and vertices off-chip && edges off-chip engine. Experimental results on the state-of-the-art Xilinx Virtex-7 board demonstrate that case studies in OmniGraph achieve 1.03x-8.13x average speedup comparing to GraphChi on Intel core2 processors. Chongchong Xu, Chao Wang 0003, Lei Gong 0003, Yuntao Lu, Yiwei Zhang 0001, Xi Li 0003, Xuehai Zhou |
CLUSTER | 7 |
| 2017 | A Power-Efficient Accelerator Based on FPGAs for LSTM NetworkabstractToday, artificial neural networks (ANNs) are widely used in a variety of applications, including speech recognition, face detection, disease diagnosis, etc. And as the emerging field of ANNs, Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) which contains complex computational logic. To achieve high accuracy, researchers always build large-scale LSTM networks which are time-consuming and power-consuming. In this paper, we present a hardware accelerator for the LSTM neural network layer based on FPGA Zedboard and use pipeline methods to parallelize the forward computing process. We also implement a sparse LSTM hidden layer, which consumes fewer storage resources than the dense network. Our accelerator is power-efficient and has a higher speed than ARM Cortex-A9 processor. Yiwei Zhang 0001, Chao Wang 0003, Lei Gong 0003, Yuntao Lu, Chongchong Xu, Xi Li 0003, Xuehai Zhou |
CLUSTER | 7 |
| 2017 | GenServ: Genome Sequencing Services on Scalable Energy Efficient AcceleratorsabstractAs a traditional algorithm, the string match meets a challenge with the development of the massive volume of data because of gene sequencing. Surveys show that there will be a huge amount of short read segments during the process of gene sequencing and the need for a highly efficient is urgent. The BWA is an effective algorithm to deal with the short read mapping. Compared with other short read mapping algorithms, the BWA algorithm has a smaller size, and this does not influence its effect. However, there is still not a system is used to accelerate the BWA algorithm especially. Thus we decide to build a system to expedite the algorithm and make it satisfied with the application of gene sequencing. In this paper, we present genome sequencing services on scalable energy-efficient accelerators. Especially, we first introduce the BWA algorithm and claim the reason for the choice of the algorithm. Then, we implement an accelerator based on FPGA to improve the performance of the algorithm. Compared to the other major platforms in accelerating the algorithm, we discuss the advantages of the FPGA platform and the limit of the other platform. Last, we build our hardware platform with a Xilinx ZYNQ FPGA development board, and the result shows that our accelerator can achieve a promising speedup and resource utilization and make it balanced between power and cost. Chao Wang 0003, Haijie Fang, Shiming Lei, Lei Gong 0003, Aili Wang 0003, Xi Li 0003, Xuehai Zhou |
ICWS | 6 |
| 2017 | xFilter: A Temporal Locality Accelerator for Intrusion Detection System ServicesabstractThe Intrusion Detection Systems (IDS) is becoming important and quite timing/space consuming due to the increasing volume of explosive data flood. During the past decades, there have been plenty of studies proposing software mechanisms to exploit the temporal locality in the IDS systems. However, it requires considerable memory blocks to store the redundancy table, therefore, the performance as well as the memory utilization is still worth pursuing. To tackle the above weakness, in this paper, we present xFilter, which explores the temporal locality to capture the redundancy, and propose a novel architecture to store and operate the redundancy table on FPGA. To demonstrate the performance of the xFilter structure, we designed a high efficient accelerator for Aho-Corasick (AC) algorithm used in Snort to detect the attack strings. To show the performance of xFilter, we implement a hardware prototype using Xilinx Zynq FPGA platform. Experimental results show that the xFilter accelerator can achieve 5.1x speedup against software implementation with insignificant hardware cost. Furthermore, the proposed hardware redundancy table mechanism can achieve 1.6x speedup against the traditional hardware accelerator. Chao Wang 0003, Jinhong Zhou, Lei Gong 0003, Xi Li 0003, Aili Wang 0003, Xuehai Zhou |
ICWS | 4 |
| 2017 | Evaluation and Trade-offs of Graph Processing for Cloud ServicesabstractLarge-scale data is often represented as graphs in the field of modern cloud computing. Graph processing attracts more and more attentions when utilizing the cloud computing service. With the increasing attentions to process massive graphs (e.g., social networks, web graphs, transport networks, and bioinformatics), many state-of-the-art open source graph computing systems on a single node have been proposed, including GraphChi, X-Stream, and GridGraph. GraphChi adopts a vertex-centric model while the latter two adopt an edge-centric model. However, there is a lack of evaluations and analyses to the performance of these systems, which makes it difficult for users to choose the best system for their applications. In this paper, to make the graph processing provide excellent cloud services to users, we propose an evaluation framework, conduct a series of extensive experiments to evaluate the performance and analyze the bottlenecks of these systems on graphs with different characteristics and different kinds of algorithms. The metrics we adopt in this paper are principles to design graph computing systems on a single node, such as RunTime, CPU Utilization, and Data Locality. The results demonstrate the trade-offs among different graph frameworks and X-Stream is more suitable to process transport networks on WCC and BFS, compared to GridGraph. Besides, we present several discussions on GridGraph. The results of our work are concluded as a reference for users, researchers, and developers. Chongchong Xu, Jinhong Zhou, Yuntao Lu, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
ICWS | 7 |
| 2017 | Work-in-Progress: TTI: A Timing ISA for LET Model in Safety-Critical SystemsabstractSafety-critical systems have suffered a complexity growth as the number of services continuously increases in these systems. The Logical Execution Time (LET) model is applied to tackle this issue due to its simple strategies and deterministic timed behaviors. However, existing implementations of LET usually rely on periodic timer interrupts of operating systems, yielding limited time precision and enormous jitter in kernel's executions. In this paper, we propose a time-triggered instruction set - TTI to augment ISA with timing properties. TTI implements as a processor architecture extension using co-processor2 interfaces in standard MIPS32. The extension mainly comprises a task management module, and a timed I/O-behaviors management module. Preliminary results show that our approach can significantly reduce overheads of LET kernel and jitters of the LET-based tasks compared to traditional implementations, which can achieve cycle-level precise timed behaviors as a result. Xi Li 0003, Haizhao Luo, Chao Wang 0003, Xianglan Chen, Xuehai Zhou |
RTSS | 2 |
| 2017 | Hot spots profiling and dataflow analysis in custom dataflow computing SoftProcessors
Chao Wang 0003, Xi Li 0003, Huizhen Zhang, Aili Wang 0003, Xuehai Zhou |
J. Syst. Softw. | 2 |
| 2017 | DLAU: A Scalable Deep Learning Accelerator Unit on FPGAabstractAs the emerging field of machine learning, deep learning shows excellent ability in solving complex learning problems. However, the size of the networks becomes increasingly large scale due to the demands of the practical applications, which poses significant challenge to construct a high performance implementations of deep learning neural networks. In order to improve the performance as well as to maintain the low power cost, in this paper we design deep learning accelerator unit (DLAU), which is a scalable accelerator architecture for large-scale deep learning networks using field-programmable gate array (FPGA) as the hardware prototype. The DLAU accelerator employs three pipelined processing units to improve the throughput and utilizes tile techniques to explore locality for deep learning applications. Experimental results on the state-of-the-art Xilinx FPGA board demonstrate that the DLAU accelerator is able to achieve up to 36.1× speedup comparing to the Intel Core2 processors, with the power consumption at 234 mW. Chao Wang 0003, Lei Gong 0003, Xi Li 0003, Yuan Xie 0001, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | SuperMIC: Analyzing Large Biological Datasets in Bioinformatics with Maximal Information CoefficientabstractThe maximal information coefficient (MIC) has been proposed to discover relationships and associations between pairs of variables. It poses significant challenges for bioinformatics scientists to accelerate the MIC calculation, especially in genome sequencing and biological annotations. In this paper, we explore a parallel approach which uses MapReduce framework to improve the computing efficiency and throughput of the MIC computation. The acceleration system includes biological data storage on HDFS, preprocessing algorithms, distributed memory cache mechanism, and the partition of MapReduce jobs. Based on the acceleration approach, we extend the traditional two-variable algorithm to multiple variables algorithm. The experimental results show that our parallel solution provides a linear speedup comparing with original algorithm without affecting the correctness and sensitivity. Chao Wang 0003, Dong Dai 0001, Xi Li 0003, Aili Wang 0003, Xuehai Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Service-Oriented Architecture on FPGA-Based MPSoCabstractThe integration of software services-oriented architecture (SOA) and hardware multiprocessor system-on-chip (MPSoC) has been pursued for several years. However, designing and implementing a service-oriented system for diverse applications on a single chip has posed significant challenges due to the heterogeneous architectures, programming interfaces, and software tool chains. To solve the problem, this paper proposes SoSoC, a service-oriented system-on-chip framework that integrates both embedded processors and software defined hardware accelerators s as computing services on a single chip. Modeling and realizing the SOA design principles, SoSoC provides well-defined programming interfaces for programmers to utilize diverse computing resources efficiently. Furthermore, SoSoC can provide task level parallelization and significant speedup to MPSoC chip design paradigms by providing out-of-order execution scheme with hardware accelerators. To evaluate the performance of SoSoC, we implemented a hardware prototype on Xilinx Virtex5 FPGA board with EEMBC benchmarks. Experimental results demonstrate that the service componentization over original version is less than 3 percent, while the speedup for typical software Benchmarks is up to 372x. To show the portability of SoSoC, we implement the convolutional neural network as a case study on both Xilinx Zynq and Altera DE5 FPGA boards. Results show the SoSoC outperforms state-of-the-art literature with great flexibility. Chao Wang 0003, Xi Li 0003, Yunji Chen, Youhui Zhang, Oliver Diessel, Xuehai Zhou |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | A Classroom Scheduling Service for Smart ClassesabstractDuring past decades, the classroom scheduling problem has posed significant challenges to educational programmers and teaching secretaries. In order to alleviate the burden of the programmers, this paper presents SmartClass, which allows the programmers to solve this problem using web services. By introducing service-oriented architecture (SOA), SmartClass is able to provide classroom scheduling services with back-stage design space exploration and greedy algorithms. Furthermore, the SmartClass architecture can be dynamically coupled to different scheduling algorithms (e.g. Greedy, DSE, etc.) to fit in specific demands. A typical case study demonstrates that SmartClass provides a new efficient paradigm to the traditional classroom scheduling problem, which could achieve high flexibility by software services reuse and ease the burden of educational programmers. Evaluation results on efficiency, overheads and scheduling performance demonstrate the SmartClass has lower scheduling overheads with higher efficiency. Chao Wang 0003, Xi Li 0003, Aili Wang 0003, Xuehai Zhou |
IEEE Trans. Serv. Comput. | 2 |
| 2016 | Display power reduction for mobile closed-source gamesabstractWith the rapid development of mobile games, power consumption and battery life of mobile platforms become a crucial problem, therefore, saving power for display which consumes a lot of power without decreasing the user experience becomes necessary. Content-centric techniques are most commonly used to save display power. However, they cannot be applied to closed-source games as they need to modify the game images. In this paper, we explore a backlight power model and propose a novel method to make a trade-off between the game image quality and display power. Different from content-centric methods, based on the proposed trade-off model, we propose a backlight dimming algorithm which maintains the user game experience using game-state information and saves display power for mobile games without modifying any game image. We implement the proposed trade-off model and backlight dimming policy in a contemporary mobile platform where the evaluation results show that, maintaining a specific game image quality level, our policy can save system power up to 10.43% compared with the static policy without decreasing games' performances. Zhinan Cheng, Xi Li 0003, Jiachen Song, Beilei Sun, Xuehai Zhou, Chao Wang 0003 |
ASAP | 2 |
| 2016 | FCM: Towards Fine-Grained GPU Power Management for Closed Source Mobile GamesabstractContemporary mobile platforms employ embedded graphic processing units (GPUs) for graphics-intensive games, and dynamic voltage and frequency scaling (DVFS) policies are used to save energy without sacrificing quality. However, current GPU DVFS policies result in unnecessary power waste due to defective workload estimations of embedded GPUs during game play. In this paper, we propose the Frame-Complexity Model (FCM), a fine-grained estimation of the GPU workload in a game frame, to quantify the GPU workload with the real runtime demand for GPU computing resources of a game frame. In FCM, three constituents of a game frame (i.e., structure, textures and computation) are quantified without modification of mobile games. Preliminary experiments show that, compared with the default policy, the FCM-directed GPU DVFS policy can reduce more power consumption of games (11.3% to 25.8%) with good Quality of Service (QoS). Jiachen Song, Xi Li 0003, Beilei Sun, Zhinan Cheng, Chao Wang 0003, Xuehai Zhou |
ACM Great Lakes Symposium on VLSI | 2 |
| 2016 | PIE: A Pipeline Energy-Efficient Accelerator for Inference Process in Deep Neural NetworksabstractIt has been a new research hot topic to speed up the inference process of deep neural networks (DNNs) by hardware accelerators based on field programmable gate arrays (FPGAs). Because of the layer-wise structure and data dependency between layers, previous studies commonly focus on the inherent parallelism of a single layer to reduce the computation time but neglect the parallelism between layers. In this paper, we propose a pipeline energy-efficient accelerator named PIE to accelerate the DNN inference computation by pipelining two adjacent layers. Through realizing two adjacent layers in different calculation orders, the data dependency between layers can be weakened. As soon as a layer produces an output, the next layer reads the output as an input and starts the parallel computation immediately in another calculation method. In such a way, computations between adjacent layers are pipelined. We conduct our experiments on a Zedboard development kit using Xilinx Zynq-7000 FPGA, compared with Intel Core i7 4.0GHz CPU and NVIDIA K40C GPU. Experimental results indicate that PIE is 4.82x faster than CPU and can reduce the energy consumptions of CPU and GPU by 355.35x and 12.02x respectively. Besides, compared with the none-pipelined method that layers are processed in serial, PIE improves the performance by nearly 50%. Xuda Zhou, Xuehai Zhou, Xi Li 0003, Chao Wang 0003 |
ICPADS | 5 |
| 2016 | SOLAR: Services-Oriented Learning ArchitecturesabstractDeep learning has been an emerging field of machine learning during past decades. However, the diversity and large scale data sizes have posed significant challenge to construct a flexible and high efficient implementations of deep learning neural networks. In order to improve the performance as well to maintain the scalability, in this paper we present SOLAR, a services-oriented deep learning architecture using various accelerators like GPU and FPGA based approaches. SOLAR provides a uniform programming model to users so that the hardware implementation and the scheduling is invisible to the programmers. At runtime, the services can be executed either on the software processors or the hardware accelerators. Experimental results on the real state-of-the-art FPGA board demonstrate that the SOLAR is able to provide a ubiquitous framework for diverse applications without increasing the burden of the programmers. Moreover, the speedup of the GPU and FPGA hardware accelerator in SOLAR can achieve significant speedup comparing to the conventional Intel i5 processors with great scalability. Chao Wang 0003, Xi Li 0003, Aili Wang 0003, Patrick C. K. Hung, Xuehai Zhou |
ICWS | 2 |
| 2016 | FairPlay: Services Migration with Lock-Free Mechanisms for Load Balancing in Cloud ArchitecturesabstractDuring past few years, how to achieve load balance using efficient software on cloud architecture is posing significant challenges to the research community. Due to the access conflictions among the shared hardware resources like distributed file systems and database transactions, creditable measures like mutex based locks, semaphore schemes, and global run queues have been widely applied. However, growing with the data scale and integration of multiprocessors in the cloud computing environment, each processor has to obtain the global lock of the system run queue, which brings inevitable burden for the runtime support. In this paper, we propose a novel lock free structure, named FairPlay, which is able to support service migration through duplex buffers between processors. Based on the buffer based structure, a load balance scheduling algorithm is presented to handle the service allocation asynchronously. Experimental results on the modified Linux operating system kernel demonstrate that the lock free mechanism could efficiently reduce the overheads on the locks with great scalability and affordable overheads. Chao Wang 0003, Jinhong Zhou, Xi Li 0003, Aili Wang 0003, Xuehai Zhou |
ICWS | 3 |
| 2016 | Behavior-Aware Integrated CPU-GPU Power Management for Mobile GamesabstractSince game applications have spilled over on the modern mobile platforms equipped with Multiprocessor Systemon-Chips and highlighted the power consumption and battery life problem of these platforms, reducing the game power for mobile devices becomes meaningful. The design of independent CPU-GPU power managements in contemporary platforms results in power consumption waste due to the failure of consideration of CPU-GPU interaction and game workload behaviors. Through analyzing the Application-Operating System (APP-OS) interaction and CPU-GPU interaction, we extract the system-call information and OpenGL API information to characterize the game workload in a low-complexity way. In this paper, based on identifying the game workload behavior and performance bottleneck, we propose a behavior-aware integrated CPU-GPU power management approach for mobile games. We also implement our power saving policy in the real platform, where the evaluation results show that our behavior-aware policy can significantly reduce power and improve game performance. Our policy provides 18% and 5% higher power-efficiency on average compared with the current policy used in our platform and the state-of-the-art policy respectively. Zhinan Cheng, Xi Li 0003, Beilei Sun, Jiachen Song, Chao Wang 0003, Xuehai Zhou |
MASCOTS | 2 |
| 2016 | Brief Announcement: MIC++: Accelerating Maximal Information Coefficient Calculation with GPUs and FPGAsabstractTo discover relationships and associations between pairs of variables in large data sets have become one of the most significant challenges for bioinformatics scientists. To tackle this problem, maximal information coefficient (MIC) is widely applied as a measure of the linear or non-linear association between two variables. To improve the performance of MIC calculation, in this work we present MIC++, a parallel approach based on the heterogeneous accelerators including Graphic Processing Unit (GPU) and Field Programmable Gate Array (FPGA) engines, focusing on both coarse-grained and fine-grained parallelism. As the evaluation of MIC++, we have demonstrated the performance on the state-of-the-art GPU accelerators and the FPGA-based accelerators. Preliminary estimated results show that the proposed parallel implementation can significantly achieve more than 6X-14X speedup using GPU, and 4X-13X using FPGA-based accelerators. Chao Wang 0003, Xi Li 0003, Aili Wang 0003, Xuehai Zhou |
SPAA | 2 |
| 2016 | Definitions of predictability for Cyber Physical Systems
Beilei Sun, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Xianglan Chen |
J. Syst. Archit. | 2 |
| 2016 | Hardware Implementation on FPGA for Task-Level Parallel Dataflow Execution EngineabstractHeterogeneous multicore platform has been widely used in various areas to achieve both power efficiency and high performance. However, it poses significant challenges to researchers to uncover more coarse-grained task level parallelization. In order to support automatic task parallel execution, this paper proposes a FPGA implementation of a hardware out-of-order scheduler on heterogeneous multicore platform. The scheduler is capable of exploring potential inter-task dependency, leading to a significant acceleration of dependence-aware applications. With the help of renaming scheme, the task dependencies are detected automatically during execution, and then task-level Write-After-Write (WAW) and Write-After-Read (WAR) dependencies can be eliminated dynamically. We extended the instruction level renaming techniques to perform task-level out-of-order execution, and implemented a prototype on a state-of-art Xilinx Virtex-5 FPGA device. Given the reconfigurable characteristic of FPGA, our scheduler supports changing accelerators at runtime to improve the flexibility. Experimental results demonstrate that our scheduler is efficient at both performance and resources usage. Chao Wang 0003, Junneng Zhang, Xi Li 0003, Aili Wang 0003, Xuehai Zhou |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Evaluation and Tradeoffs for Out-of-Order Execution on Reconfigurable Heterogeneous MPSoCabstractOut-of-order (OoO) execution schemes show incredible promise for task-level parallelism in multiprocessor system-on-chip (MPSoC) designs. However, the main challenge of the OoO execution lies in the analysis of the intertask dependences. In this paper, we address this challenge by applying the instruction-level scoreboarding algorithm at the task level. Furthermore, we introduce both software-based static and dynamic implementations on top of a heterogeneous MPSoC prototyped on a field-programmable gate array fabric. Our experimental results show that our approach can achieve up to 94.75% and 97.68% of the theoretical speedup. Finally, we present an in-depth analysis of the tradeoff between our dynamic and static approaches. Qi Guo 0001, Xi Li 0003, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Automatic frame rate-based DVFS of gameabstractThe rapid development of mobile games highlights the power consumption problem in the mobile platform. Most of the power saving techniques use the prediction-based dynamic voltage frequency scaling (DVFS) scheme. However, the prediction could be inaccurate resulting from the frequent interactions of user when playing games. We have observed that frame rate is near-linear to CPU frequency, but there is a bottleneck, frame rate will not increase as CPU frequency increases when CPU frequency reaches this threshold. Moreover, previous research has shown that utilizing the information of game state can reduce the influence of game interactive characterization to DVFS policy. We explore a method to automatically detect the game state. We propose the Automatic Frame Rate-Based DVFS policy, which can learn the threshold of frame rate online and utilize the information of game state and frame rate to scale the frequency without prediction. Our evaluation result shows that, compared with the prediction-based Android default Interactive DVFS policy, our policy saves more power in all the testing games. Up to 15.2% more power can be saved by Automatic Frame Rate-Based DVFS policy. Zhinan Cheng, Xi Li 0003, Beilei Sun, Ce Gao, Jiachen Song |
ASAP | 2 |
| 2015 | A Deep Learning Prediction Process Accelerator Based FPGAabstractRecently, machine learning is widely used in applications and cloud services. And as the emerging field of machine learning, deep learning shows excellent ability in solving complex learning problems. To give users better experience, high performance implementations of deep learning applications seem very important. As a common means to accelerate algorithms, FPGA has high performance, low power consumption, small size and other characteristics. So we use FPGA to design a deep learning accelerator, the accelerator focuses on the implementation of the prediction process, data access optimization and pipeline structure. Compared with Core 2 CPU 2.3GHz, our accelerator can achieve promising result. Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
CCGRID | 4 |
| 2015 | An FPGA-Based Accelerator for Neighborhood-Based Collaborative Filtering Recommendation AlgorithmsabstractNeighborhood-based Collaborative Filtering (CF) is a kind of techniques in the field of recommendation algorithms and has been widely used in lots of personalized recommender systems. In the big data era, the increasing data amounts make these CF recommendation algorithms become time-consuming and energy-wasted. At present, Cloud computing and Graphic Processing Unit (GPU) are the two major platforms to accelerate CF algorithms. However, both platforms exist some remarkable shortcomings such as efficiency and power. To solve these problems, in our work, we investigate three neighborhood-based CF algorithms and design a general and flexible accelerator for them based on Field Programmable Gate Array (FPGA). This accelerator cooperates with host CPU and could accelerates primary time-consuming parts that these algorithms share. Experimental results show that our accelerator could significantly improve the acceleration efficiency with the affordable hardware cost and less energy consumption. Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
CLUSTER | 4 |
| 2015 | SODA: software defined FPGA based accelerators for big data
Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
DATE | 2 |
| 2015 | RapidPath: Accelerating Constrained Shortest Path Finding in Graphs on FPGA (Abstract Only)abstractEmerging applications, such as Software Defined Network (SDN), Social Media, and Location Based System (LBS), are typical big graph based applications. Due to the explosive network flood, it is essential to speedup the computation process in the big graph application, such as Constrained Shortest Path Finding (CSPF) algorithm is one of the most challenging part. Meanwhile, FPGA has been an effective and efficient platform in novel big data architectures and systems, due to its computing power and low power consumption. It enables the researchers to deploy massive accelerators within one single chip. In this paper, we present RapidPath, an acceleration method for CSPF algorithm in software defined networks, which decomposes a large and complex system of programs into small single-purpose source code libraries that perform specialized tasks in parallel. Only the CSPF step is implemented in hardware and the rest steps run on the processor. We have built a prototyping system on Zynq with CSPF case studies. The ARM processor uses a shared memory with the FPGA based accelerator using DMA based channels. Control signals are transferred via AXI bus interfaces. Experimental results depict that RapidPath is able to achieve up to 43.75X speedup at 128 nodes, comparing to the software execution (without cache) on Xilinx Zynq board. Furthermore, hardware cost and overheads reveal that the RapidPath architecture can achieve high speedup with insignificant cost. Chao Wang 0003, Xi Li 0003, Qi Guo 0001, Xuehai Zhou |
FPGA | 2 |
| 2015 | SAKMA: Specialized FPGA-Based Accelerator Architecture for Data-Intensive K-Means Algorithms
Fahui Jia, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
ICA3PP (2) | 3 |
| 2015 | CRAIS: A Crossbar-Based Interconnection Scheme on FPGA for Big Data
Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
J. Comput. Sci. Technol. | 2 |
| 2015 | A case study of parallel JPEG encoding on an FPGA
Chao Wang 0003, Xi Li 0003, Peng Chen 0004, Xuehai Zhou |
J. Parallel Distributed Comput. | 2 |
| 2015 | Architecture Support for Task Out-of-Order Execution in MPSoCsabstractMulti-processor system on chip (MPSoC) has been widely applied in embedded systems in the past decades. However, it has posed great challenges to efficiently design and implement a rapid prototype for diverse applications due to heterogeneous instruction set architectures (ISA), programming interfaces and software tool chains. In order to solve the problem, this paper proposes a novel high level architecture support for automatic out-of-order (OoO) task execution on FPGA based heterogeneous MPSoCs. The architecture support is composed of a hierarchical middleware with an automatic task level OoO parallel execution engine. Incorporated with a hierarchical OoO layer model, the middleware is able to identify the parallel regions and generate the sources codes automatically. Besides, a runtime middleware Task-Scoreboarding analyzes the inter-task data dependencies and automatically schedules and dispatches the tasks with parameter renaming techniques. The middleware has been verified by the prototype built on FPGA platform. Examples and a JPEG case study demonstrate that our model can largely ease the burden of programmers as well as uncover the task level parallelism. Chao Wang 0003, Xi Li 0003, Junneng Zhang, Peng Chen 0004, Yunji Chen, Xuehai Zhou, Ray C. C. Cheung |
IEEE Trans. Computers | 2 |
| 2015 | Heterogeneous Cloud Framework for Big Data Genome SequencingabstractThe next generation genome sequencing problem with short (long) reads is an emerging field in numerous scientific and big data research domains. However, data sizes and ease of access for scientific researchers are growing and most current methodologies rely on one acceleration approach and so cannot meet the requirements imposed by explosive data scales and complexities. In this paper, we propose a novel FPGA-based acceleration solution with MapReduce framework on multiple hardware accelerators. The combination of hardware acceleration and MapReduce execution flow could greatly accelerate the task of aligning short length reads to a known reference genome. To evaluate the performance and other metrics, we conducted a theoretical speedup analysis on a MapReduce programming platform, which demonstrates that our proposed architecture have efficient potential to improve the speedup for large scale genome sequencing applications. Also, as a practical study, we have built a hardware prototype on the real Xilinx FPGA chip. Significant metrics on speedup, sensitivity, mapping quality, error rate, and hardware cost are evaluated, respectively. Experimental results demonstrate that the proposed platform could efficiently accelerate the next generation sequencing problem with satisfactory accuracy and acceptable hardware cost. Chao Wang 0003, Xi Li 0003, Peng Chen 0004, Aili Wang 0003, Xuehai Zhou, Hong Yu 0011 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | FreeRider: Non-Local Adaptive Network-on-Chip Routing with Packet-Carried Propagation of Congestion InformationabstractNon-local adaptive routing techniques, which utilize statuses of both local and distant links to make routing decisions, have recently been shown to be effective solutions for promoting the performance of Network-on-Chip (NoC). The essence of non-local adaptive routing was an additional network dedicated to propagate congestion information of distant links on the NoC. While the dedicated Congestion Propagation Network (CPN) helps routers to make promising routing decisions, it incurs additional wiring and power costs and becomes an unnecessary decoration when the load of NoC is light. Moreover, the CPN has to be extended if one would utilize more sophisticated congestion information to enhance the performance of NoC, bringing in even larger wiring and power costs. This paper proposes an innovative non-local adaptive routing technique called FreeRider, which does not use a dedicated CPN but instead leverages free bits in head flits of existing packets to carry and propagate rich congestion information without introducing additional wires or flits. In order to balance the network load, FreeRider adopts a novel three-stage strategy of output link selection, which adequately utilizes the propagated information to make routing decisions. Experimental results on both synthetic traffic patterns and application traces show that FreeRider achieves better throughput, shorter latency, and smaller power consumption than a state-of-the-art adaptive routing technique with dedicated CPN. Shaoli Liu, Tianshi Chen 0002, Ling Li 0001, Xi Li 0003, Mingzhe Zhang 0005, Chao Wang 0003, Haibo Meng, Xuehai Zhou, Yunji Chen |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2014 | Co-processing with dynamic reconfiguration on heterogeneous MPSoC: practices and design tradeoffs (abstract only)abstractReconfiguration technique has been considered as one of the most promising electronic design automation (EDA) technologies in MPSoC design paradigms. However, due to the unavoidable latency in the reconfiguration procedure, it still poses a significant challenge to efficiently analyze the trade-offs for the software/hardware execution, static reconfiguration and dynamic reconfiguration. In this paper we first present a heterogeneous MPSoC middleware to support state-of-the-art dynamic partial reconfigurable technologies. Furthermore, we evaluate the reconfiguration latency and analyze the trade-off for the dynamic partial reconfiguration technologies. Chao Wang 0003, Xi Li 0003, Xuehai Zhou, Yunji Chen, Koen Bertels |
FPGA | 2 |
| 2014 | Big data genome sequencing on Zynq based clusters (abstract only)abstractNext-generation sequencing (NGS) problems have attracted many attentions of researchers in biological and medical computing domains. The current state-of-the-art NGS computing machines are dramatically lowering the cost and increasing the throughput of DNA sequencing. In this paper, we propose a practical study that uses Xilinx Zynq board to summarize acceleration engines using FPGA accelerators and ARM processors for the state-of-the-art short read mapping approaches. The heterogeneous processors and accelerators are coupled with each other using a general Hadoop distributed processing framework. First the reads are collected by the central server, and then distributed to multiple accelerators on the Zynq for hardware acceleration. Therefore, the combination of hardware acceleration and Map-Reduce execution flow could greatly accelerate the task of aligning short length reads to a known reference genome. Our approach is based on preprocessing the reference genomes and iterative jobs for aligning the continuous incoming reads. The hardware acceleration is based on the creditable read-mapping algorithm RMAP software approach. Furthermore, the speedup analysis on a Hadoop cluster, which concludes 8 development boards, is evaluated. Experimental results demonstrate that our proposed architecture and methods has the speedup of more than 112X, and is scalable with the number of accelerators. Finally, the Zynq based cluster has efficient potential to accelerate even general large scale big data applications. Chao Wang 0003, Xi Li 0003, Xuehai Zhou, Yunji Chen, Ray C. C. Cheung |
FPGA | 2 |
| 2014 | Temperature-Aware Scheduling Based on Dynamic Time-Slice Scaling
Gangyong Jia, Youwei Yuan, Jian Wan 0001, Congfeng Jiang, Xi Li 0003, Dong Dai 0001 |
ICA3PP (1) | 5 |
| 2014 | Behavior Gaps and Relations between Operating System and Applications on Accessing DRAMabstractDetailed analyses of the behaviors of operating system and applications are significant for taking full advantage of the precious hardware resources and improving performance. This paper focus on their DRAM access behaviors based on access proportion and row-buffer miss ratio (RBM). The access proportions of Kernel and User vary greatly in different stages throughout the lifetime of a process. Most of the row-buffer misses are caused by the one having higher access proportion. By analyzing the RBM series through ARMA model, we found that User's DRAM accesses only have short-term influences on its behavior, while the Kernel's influences are relatively deeper. The ARMA model for the RBM series is able to predict the future RBMs, which are profound basis to schedule the DRAM access commands. The results of Gaussian Fitting show that Kernel and User are tightly correlated on accessing DRAM, especially in the steady stage and the end stage of a process's life cycle. Based on this close relation, it is possible to estimate the DRAM access behaviors of the other one according to the one whose behaviors have been known. System-calls that obviously affect the access proportions and RBMs are also revealed in this paper. Beilei Sun, Xi Li 0003, Zongwei Zhu, Xuehai Zhou |
ICECCS | 2 |
| 2014 | A Thread Behavior-Based Memory Management Framework on Multi-core SmartphoneabstractMemory management systems have significantly affected the overall performance of modern multi-core smartphone systems. Android, as one of the most popular smartphone operating systems, adopts a global buddy system with the FCFS (first come, first served) principle for memory allocation, and releases requests to manage external fragmentations and maintain the memory allocation efficiency. However, extensive experimental study on thread behaviors indicates that memory external fragmentation is no longer the crucial bottleneck in most Android applications. Specifically, a thread usually allocates or releases memory in bursts, resulting in serious memory locks and inefficient memory allocation. Furthermore, the pattern of such bursting behaviors varies throughout the life cycle of a thread. The conventional FCFS policy of Android buddy system fails to adapt to such variations and thus suffers from performance degradation. In this paper, we propose a novel memory management framework, called Memory Management Based on Thread Behaviors (MMBTB), for multi-core smartphone systems. It adapts to various thread behaviors through targeted optimizations to provide efficient memory allocation. The efficiency and effectiveness of this new memory management scheme on multicore architecture is proved by a theoretical emulation model. Our experimental studies on the real Android system show that MMBTB can improve the efficiency of memory allocation by 12%-20%, confirming the theoretical analysis results. Zongwei Zhu, Xi Li 0003, Hengchang Liu, Cheng Ji 0002, Xuehai Zhou, Beilei Sun |
ICECCS | 2 |
| 2014 | Kernel-User Space Separation in DRAM MemoryabstractPerformance of software is increasingly restricted by the Memory Wall instead of CPU. Many studies focus on alleviating the DRAM latency by improving the row-buffer hit rate. But most of them treat the Kernel and User equally. Data used by Operating System and User applications spread in different rows of the same bank, leading to the contentions for the row-buffer when they access the bank successively. We find that contentions between Kernel and User make up of a great proportion of all the row-buffer misses. To alleviate the contentions between Kernel and User, we divide the united DRAM memory space into Kernel-Space and User-Space. A new page-allocation-system, the K/U-Aware page-allocation-system, is proposed to manage Kernel-Space and User-Space in DRAM memory in different address mapping schemes of DRAM memory controller. In the new system, pages are allocated from different spaces according to applicants (Kernel or User). Sizes of the two spaces increase and decrease dynamically as required. For benchmarks in PARSEC suites, the proposed system reduces the contentions of Kernel and User effectively, producing significant improvements of row-buffer hit rate. The execution time is reduced by 9.45% (max. 20.45%) and 6.51% (max. 18.05%) respectively in two typical address mapping schemes. Xi Li 0003, Beilei Sun, Zongwei Zhu, Chao Wang 0003, Xuehai Zhou |
ISPA | 1 |
| 2014 | Multi-objective aware design flow for coarse-grained systems on chipabstractThis paper presents a software and hardware co-design flow for the coarse-grained systems on chip. It enables a multi-target design space exploration (MT-DSE) algorithm with multiple objectives such as chip area utilization, energy consumption, core efficiency, interconnection structure, application workload and speedup aware. With the help of the MT-DSE tool, the proposed design flow can supply a valuable assistance for architecture designers to develop a well trade-off multi-processor system. In contrast, most of state-of-the-art design space exploration tools rely on varieties of simulations or implementations that are quite time-consuming. Benefit from no such dependences, the MT-DSE method could turn out the optimized alternative very fast with multiple factors balanced. Besides, the tool is employed at a very early stage in the component based systems design and only needs a little profiling information which can greatly reduce the development term of the design. As an illustration, the JPEG compression algorithm is chosen to demonstrate how the tool exploits a given application and guides to build the most desired architecture. Peng Chen 0004, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
RTCSA | 3 |
| 2014 | PUMA: Pseudo unified memory architecture for single-ISA heterogeneous multi-core systemsabstractSingle-ISA heterogeneous multi-core processors have advantages over cost-equivalent homogeneous ones, which integrate cores having the same instruction set architecture (ISA) but offer different performance and power characteristics. When these cores share the off-chip main memory, requests from different cores will interfere with each other, leading to low system performance and unfairness even starvation. Unfortunately, state-of-the-art memory scheduling and thread scheduling algorithms are ineffective at solving these problems. This paper proposes a fundamentally new memory architecture of pseudo unified memory (PUMA), which partitions the memory into regions according cores' different performance, each core mostly requests only one memory region seldom exceeding, reducing interfere among cores while retaining bank level parallelism for improving performance and fairness. We evaluate the design trade-offs involved in our PUMA and compare it against three state-of-the-art memory management methods. Our experimental results show that PUMA improves both system performance and fairness among cores while reducing memory power. Gangyong Jia, Liang Shi 0001, Jian Wan 0001, Youwei Yuan, Xi Li 0003, Dong Dai 0001 |
RTCSA | 5 |
| 2014 | Memory power optimization on different memory address mapping schemasabstractSince memory accounts for a large and increasing fraction of the energy consumed by computers, memory manufacturers have developed memory devices with different power/work modes. For taking full advantage of these modes, more and more creditable hardware or software power mode control algorithms have been proposed. In this paper, by analyzing the effects of power mode control polices on different memory address mapping schemas (schema is used to translate a given physical address to a specific memory cell in DRAM system), we find that most previous power mode control policies are sensitive to mapping schemas. Therefore, in order to manage these power modes on different mapping schemas more effectively, we divide them into two categories: high-bit multi-access cross memory (HMCM) and low-bit multi-access cross memory (LMCM), and then take a targeted optimization. For the former schema, a rank-sensitive buddy system (RS-Buddy) was proposed to cluster pages together to prolong memory modules' low power time. For the latter, we introduce a comprehensive solution named as MSPA. It adopts a memory address segmentation module (MASM) to split memory into many regions configured as different mapping schemas. And with the help of an OS power-aware memory allocator (PAMA), MSPA can dynamically allocate one application's memory from its preferred region to balance power and performance. By performing extensive experiments on practical platform for HMCM while on simulator for LMCM, the results of HMCM show that RS-Buddy can optimize the power efficiency from 2% to 22%. Furthermore, the simulation results of LMCM demonstrate that MSPA can further improve the power efficiency from 3% to 17% when combined with other previous state-of-the-art studies. Zongwei Zhu, Xi Li 0003, Chao Wang 0003, Xuehai Zhou |
RTCSA | 2 |
| 2014 | Colored Petri Net model with automatic parallelization on real-time multicore architectures
Chao Wang 0003, Xiaojing Feng, Xi Li 0003, Xuehai Zhou, Peng Chen 0004 |
J. Syst. Archit. | 3 |
| 2014 | Accelerating the Next Generation Long Read Mapping with the FPGA-Based SystemabstractTo compare the newly determined sequences against the subject sequences stored in the databases is a critical job in the bioinformatics. Fortunately, recent survey reports that the state-of-the-art aligners are already fast enough to handle the ultra amount of short sequence reads in the reasonable time. However, for aligning the long sequence reads (>400 bp) generated by the next generation sequencing (NGS) technology, it is still quite inefficient with present aligners. Furthermore, the challenge becomes more and more serious as the lengths and the amounts of the sequence reads are both keeping increasing with the improvement of the sequencing technology. Thus, it is extremely urgent for the researchers to enhance the performance of the long read alignment. In this paper, we propose a novel FPGA-based system to improve the efficiency of the long read mapping. Compared to the state-of-the-art long read aligner BWA-SW, our accelerating platform could achieve a high performance with almost the same sensitivity. Experiments demonstrate that, for reads with lengths ranging from 512 up to 4,096 base pairs, the described system obtains a 10x -48x speedup for the bottleneck of the software. As to the whole mapping procedure, the FPGA-based platform could achieve a 1.8x -3:3x speedup versus the BWA-SW aligner, reducing the alignment cycles from weeks to days. Peng Chen 0004, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2013 | Acceleration of the long read mapping on a PC-FPGA architecture (abstract only)abstractThe genome sequence alignment, whereby ultra scale of sequence reads should be compared to an enormous long reference, has been one central challenge to the biologists for a long period. For recent years, new sequencing technology makes it possible to generate longer reads (sequences of genome fragments) which seem more valuable for the life science research. It has been foreseen that long genome reads (length longer than 200 base pairs) will dominate the field in the near future. Unfortunately, most of the state-of-art aligners nowadays are optimized and only applicable for the short read mapping while present long read aligners are still not satisfying at the aspect of speed. In this paper, we propose a novel PC-FPGA hybrid system to improve the performance of the long read mapping. The BWA-SW algorithm is chosen as the alignment approach and by accelerating the bottleneck of the algorithm, our solution could archive a significant improvement in term of speed. Experiments demonstrate that the described system is as accurate as the BWA-SW aligner and about 1.41-2.73 times faster than it for reads with lengths ranging from 500bp to 2000bp. Peng Chen 0004, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
FPGA | 3 |
| 2013 | Custom instruction generation and mapping for reconfigurable instruction set processors (abstract only)abstractReconfigurable instruction set processors (RISP) is an emerging research field for state-of-the-art adaptive systems. However, it still poses significant challenges to generate and map the custom instructions to the original codes. This paper proposes a generation and mapping scheme to extend custom instructions for adaptive RISP. First a target function blocks (basic blocks) are generated from a dynamic profiler. Then the selected hot spot will be considered as a custom instruction and implemented in reconfigurable hardware logic units. With respect to the instruction selection, an instruction generator is utilized to provide a mapping mechanism from hot blocks to hardware implementations, using data flow analysis, instruction clustering, subgraph enumerating and subgraph merging techniques. Finally the original executable files are recompiled and regenerated by a customized GCC compiler. To demonstrate the effectiveness and performance of the framework, a prototype instruction generator has been implemented to verify the correctness and efficiency of the mapping mechanism. Chao Wang 0003, Xi Li 0003, Huizhen Zhang, Jinsong Ji, Xuehai Zhou |
FPGA | 2 |
| 2013 | Genome sequencing using mapreduce on FPGA with multiple hardware accelerators (abstract only)abstractThe genome sequencing problem with short reads is an emerging field with seemingly limitless possibilities for advances in numerous scientific research and application domains. It has been the hot topic during the past few years. Growing with the data population and the ease to access for personal users, how to shorten the response interval for short read mapping at a large scale computing domain is extremely important. In this paper we propose a novel FPGA-based acceleration solution with Map-Reduce framework on multiple hardware acceleration engines. The combination of hardware accelerators and Map-Reduce execution flow could greatly expedite the task of aligning short length reads to a known reference genome. Our approach is based on preprocessing the reference genomes and iterative jobs for aligning the continuous incoming reads. The read-mapping algorithm is modeled after the creditable RMAP software approach. Furthermore, theoretical speedup analysis on a MapReduce programming platform is presented, which demonstrates that our proposed architecture has efficient potential to reduce the average waiting time for large scale short reads applications. Chao Wang 0003, Xi Li 0003, Xuehai Zhou, Jim Martin 0001, Ray C. C. Cheung |
FPGA | 2 |
| 2013 | Hardware acceleration for the banded Smith-Waterman algorithm with the cycled systolic arrayabstractThe Smith-Waterman is one of the most popular algorithms in the molecular sequence alignment. It is often used to find the best local alignment between two strings by calculating the similarity score of the pair of strings. The algorithm is of great potential to be parallelized and has been employed by a lot of FPGA-based solutions, mostly with the systolic array manner. However, the architecture designers always find the number of the process elements (PE) in their implementation quite limited by the resources available on the FPGA devices. They either make decomposition or fold the implementation of their applications when facing a large requirement for the process elements number. In this paper, we put forward a novel FPGA-based architecture which could address the problem with a bounded number of PEs to realize any lengths of systolic array. It is mainly based on the idea of the banded Smith-Waterman but with a key distinguish that it reuses the PEs which are beyond the boundary. Analysis shows that the approach is as fast as the normal systolic fabric and obtains quite considerable resource reduction. Peng Chen 0004, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
FPT | 3 |
| 2013 | Coordinate Task and Memory Management for Improving Power Efficiency
Gangyong Jia, Xi Li 0003, Jian Wan 0001, Chao Wang 0003, Dong Dai 0001, Congfeng Jiang |
ICA3PP (1) | 2 |
| 2013 | Power-aware buddy system and task group schedulerabstractMemory is responsible for a large and increasing fraction of the energy consumed by computers. To address this challenge, memory manufacturers have developed memory devices with different power states. In order to more effectively manage the power states in the operating system, in this paper, we propose a rank-sensitive buddy system (RS-Buddy) which clusters pages together to prolong the idle time of memory ranks without breaking defragmentation characteristics. For the purpose of decreasing unnecessary frequent mode transitions, we introduce a power-aware task group scheduler (PATGS) that groups the threads which access the same rank together to schedule while sustaining system fairness. Finally, we integrate state-of-the-art mode control policies with our RS-Buddy and PATGS, with experimental results demonstrating that our algorithms can improve the power efficiency from 25.31% to 27.35% compared with state-of-the-art studies. Xi Li 0003, Zongwei Zhu, Gangyong Jia, Xuehai Zhou |
ISCAS | 1 |
| 2013 | FPGA implementation of a scheduler supporting parallel dataflow executionabstractHeterogeneous multicore platform has been widely used in various areas to achieve both power efficiency and high performance. This paper proposes a FPGA implementation of a hardware scheduler supporting parallel dataflow execution on heterogeneous multicore platform. The scheduler has the capability to explore potential parallelism, leading to a high acceleration of dependence-aware applications. Given the reconfigurable characteristic of FPGA platform, our scheduler supports changing accelerators during runtime to increase the flexibility of the platform. We implement and optimize the scheduler on a state-of-art Xilinx Virtex-5 FPGA board, experimental results show that our scheduler is efficient at both performance and resources usage. Junneng Zhang, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
ISCAS | 3 |
| 2013 | SOBA: A Services-Oriented Browser Architecture with Distributed URL-Filtering Mechanisms for TeenagersabstractIn order to protect the teenagers in online cyberspace, this paper presents SOBA, novel services-oriented internet browser architecture with distributed filtering mechanisms. SOBA is the first literature that introduces SOA concepts into the web browser design paradigm. It contains a server cluster, which employs URL filter functions to verify the URL access, while whitelists and validation results are wrapped as services. Mean-while, a customized SOBA client is in charge of data conversion and web site navigation while verification workloads are de-ployed on the server side. Administrators can manage URL data-bases through back stage websites. A prototyping software browser of SOBA demonstrates that SOA concepts can greatly improve the safety for teenage users with high flexibility and modularity. Aili Wang 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
SERVICES | 3 |
| 2013 | A Semantics-based Translation Method for Automated Verification of SystemC TLM Designs
Yanyan Gao 0001, Xi Li 0003 |
J. Electron. Test. | 2 |
| 2013 | Location-aware private service discovery in pervasive computing environment
Chen Yu 0003, Dezhong Yao 0002, Xi Li 0003, Yan Zhang 0002, Laurence T. Yang, Naixue Xiong, Hai Jin 0001 |
Inf. Sci. | 3 |
| 2013 | MP-Tomasulo: A Dependency-Aware Automatic Parallel Execution Engine for Sequential ProgramsabstractThis article presents MP-Tomasulo, a dependency-aware automatic parallel task execution engine for sequential programs. Applying the instruction-level Tomasulo algorithm to MPSoC environments, MP-Tomasulo detects and eliminates Write-After-Write (WAW) and Write-After-Read (WAR) inter-task dependencies in the dataflow execution, therefore to operate out-of-order task execution on heterogeneous units. We implemented the prototype system within a single FPGA. Experimental results on EEMBC applications demonstrate that MP-Tomasulo can execute the tasks out-of-order to achieve as high as 93.6% to 97.6% of ideal peak speedup. A comparative study against a state-of-the-art dataflow execution scheme is illustrated with a classic JPEG application. The promising results show MP-Tomasulo enables programmers to uncover more task-level parallelism on heterogeneous systems, as well as to ease the burden of programmers. Chao Wang 0003, Xi Li 0003, Junneng Zhang, Xuehai Zhou, Xiaoning Nie |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | Cloud Based Short Read Mapping ServiceabstractBioinformatics is an emerging field with seemingly limitless possibilities for advances in numerous scientific research and applications domains. In this paper, we summaries the explosive cutting-edge acceleration engines for the emerging short read mapping problems. What's more, we propose a novel Cloud based web service solution to the short read mapping problem in DNA sequencing, which greatly accelerates the task of aligning continuous incoming short length reads to uncertain known reference genomes. This approach is based on the pre-process of the reference genomes and iterative MapReduce jobs for aligning the continuous incoming reads. The MapReduce-based read-mapping algorithm is modeled after RMAP. Preliminary experimental results on incorporated MapReduce programming framework demonstrate that our proposed architecture and methods efficiently reduces the waiting time for large scale short reads applications. This architecture would be much important and efficient in future commercial personal gnome sequencing service. Dong Dai 0001, Xi Li 0003, Chao Wang 0003, Xuehai Zhou |
CLUSTER | 2 |
| 2012 | Cache Promotion Policy Using Re-reference Interval PredictionabstractThe last-level cache (LLC) mitigates the long latencies of memory access in today's chip multi-core processor (CMP). The promotion policy in the LLC largely affects cache efficiency, while an inappropriate promotion policy may lead useless blocks to remain in the cache longer than necessary, in turn result into inefficiency. Currently state-of-the-art promotion policies are unaware of the re-reference interval of cache accesses. Applications that exhibit a long re-reference interval perform poorly with these promotion policies. In this paper, we propose a promotion policy that uses re-reference interval prediction (RRIP) information. Such technique requires minor hardware modification over the least-recently-used (LRU) replacement policy. Our evaluation shows that RRIP improves IPCsumby 2.58%, Weighted Speedup by 3.54% and IPCnorm_hmeanby 6.2% on average over single-step promotion policy. Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu |
CLUSTER | 2 |
| 2012 | Memory Affinity: Balancing Performance, Power, Thermal and Fairness for Multi-core SystemsabstractMain memory is expected to grow significantly in both speed and capacity for it is a major shared resource among cores in a multi-core system, which will lead to increasing power consumption. Therefore, it is critical to address the power issue without seriously decreasing performance in the memory subsystem. In this paper, we firstly propose memory affinity which retains the active and low power memory ranks as long as possible to avoid frequently switching between active and low power status, and then present a memory affinity aware scheduling (MAS) to balance performance, power, thermal and fairness for multi-core systems. Experimental results demonstrate our memory affinity aware scheduling algorithms well adapt to system loading to maximize power saving and avoid memory hotspot at the same time while sustaining the system bandwidth demand and preserving fairness among threads. Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu |
CLUSTER | 2 |
| 2012 | Phase Detection for Loop-Based Programs on Multicore ArchitecturesabstractPhase detection and behavior analysis have been major concerned to improve the performance as well as the system throughputs. However, for the distributed acceleration engines, the execution among different phases is much more difficult to be analyzed, especially for the loop based programs. With respect to the tasks in different iterations, how to efficiently detect the phases belonging to the same loop iteration or even across iterations is posing significant challenge. In this paper we propose a phase detection method for loop-based programs on multiprocessor system-on-chip (MPSoC). A cross compiling tool based on state-of-the-art ARM RVDS is employed to locate the hot spot function of the program. Based on the hot spots, we target the function optimization on a hadoop cluster for performance evaluation. The preliminary experimental results demonstrate that our proposed techniques can extract the hot block function with high accuracy and modest overheads. The method can be applied to guide the optimization and adaptive mapping scheme on MPSoC architectures. Chao Wang 0003, Xi Li 0003, Dong Dai 0001, Gangyong Jia, Xuehai Zhou |
CLUSTER | 2 |
| 2012 | CaaS: Core as a service realizing hardware sercices on reconfigurable MPSoCSabstractService-oriented architecture (SOA) has been proved as an efficient way for high level programming paradigms. This paper realizes services into reconfigurable MPSoC to organize CaaS: a core as a service framework, which implements hardware services on state-of-the-art reconfigurable multi-processor system-on-chip (MPSoC) platform for high level parallelization. The integration of SOA concepts can provide structural programming models to ease the burden of high level programming. For demonstration, a prototype with JPEG application has been built on an FPGA, regarding embedded processors and IP cores as computing servants. The experimental results demonstrate the CaaS can achieve high flexibility with dynamic reconfigurable techniques. Chao Wang 0003, Xi Li 0003, Junneng Zhang, Peng Chen 0004, Xuehai Zhou |
FPL | 2 |
| 2012 | Parallel dataflow execution for sequential programs on reconfigurable hybrid MPSoCsabstractReconfigurable hybrid multi-processor systems-on-chips (MPSoCs) are very powerful computing platforms. However, it has been quite challenging to schedule and map tasks to different function units of the MPSoCs, especially for tasks with inter-task dependencies. This paper introduces a parallel dataflow execution support, called ReArc, for the FPGA based reconfigurable hybrid MPSoCs. It constructs a hierarchical model for the high level programming with a parallel execution flow and dynamic reconfigurations. A prototype has been built on a Xilinx FPGA with a state-of-the-art software-hardware co-design paradigm. Experimental results demonstrate that ReArc could significantly facilitate researchers to construct a high-level, application oriented FPGA implementation with acceptable hardware utilizations and reconfiguration overheads. Chao Wang 0003, Xi Li 0003, Xuehai Zhou, Yajun Ha |
FPT | 2 |
| 2012 | A task-level OoO framework for heterogeneous systemsabstractThis paper proposes a framework targeting the problem of task-level out-of-order (OoO) execution for heterogeneous systems. The framework consists of three layers: 1) Programming model; 2) OoO task scheduler; 3) Processing Elements. In order to uncover task-level parallelism automatically, renaming scheme is applied from instruction-level parallelism (ILP) to task-level parallelism (TLP). With the help of renaming scheme, inter-task data dependencies can be detected automatically during execution, and then task-level WAW and WAR dependencies can be eliminated dynamically. We applied Tomasulo algorithm from ILP to perform task-level OoO execution, and implemented a prototype on a state-of-art reconfigurable FPGA platform. Experimental results show that the framework is efficient for heterogeneous systems. Junneng Zhang, Chao Wang 0003, Xi Li 0003, Peng Chen 0004, Xiaojing Feng, Xuehai Zhou |
FPT | 3 |
| 2012 | Share memory aware scheduler: balancing performance and fairnessabstractOptimizing system performance through scheduling has received a lot of attention. However, none of the existing approaches can balance the system performance improvement and the fair share of CPU time among threads. We present in this paper a share memory aware scheduler (SMAS). The key idea is to adopt thread group scheduling which partitions threads based on memory address space to reduce switching overhead and to give each thread a fair chance to occupy CPU time. There are three main contributions: 1) SMAS does well in balancing system performance and fairness among all threads; 2) to our knowledge, this is the first attempt to use share memory aware scheduler for system performance improvement; 3) we implement SMAS both in testbed and simulator for evaluation. The testbed results on a 2-core processor show that our proposed scheduler can improve performance of different performance parameters with neglected overhead in fairness, which reduced 0.128% in cache miss rate, 2.62% in run time, 13.15% in DTBL misses, 31.68% in ITLB misses and 46.15% in ITLB flushes maximum. Furthermore, our extensive simulation results for 4 and 8 cores demonstrate that SMAS is highly scalable. Xi Li 0003, Gangyong Jia, Zongwei Zhu, Xuehai Zhou |
ACM Great Lakes Symposium on VLSI | 1 |
| 2012 | A Dependency Aware Task Partitioning and Scheduling Algorithm for Hardware-Software Codesign on MPSoCs
Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Fangling Zeng |
ICA3PP (1) | 2 |
| 2012 | Behavior Aware Data Locality for CachesabstractOptimizing cache performance through improving data locality has been receiving a lot of attention. However, none of the existing approaches can combine each task's behavior to optimize data locality for caches. We present a behavior aware data locality (BADL) to optimize cache performance in this paper. The key idea is to add each task's behavior when allocating memory, which can take advantage of each task's different locality to optimize cache performance. There are five main contributions: 1. to our best knowledge, this is the first attempt to improve cache performance through combining task behavior, 2. BADL detailed analyzes low performance derived from internal of the cache line, which is more fine-grained than the current state-of-the-art fine-grained in hardware angle, 3. BADL optimizes the cache performance through improving internal of cache line efficiency, 4. we implement BADL both in single-threaded application and multi-threaded applications scenarios, 5. BADL can be combined to most of the cache optimizing researches. The experiment results show our proposed BADL can improve 18.6% performance on average in single-threaded application situation and improve 20.8% performance on average in multi-threaded application situation. Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu |
ICPADS | 2 |
| 2012 | Frequency Affinity: Analyzing and Maximizing Power Efficiency in Multi-core SystemsabstractPerformance optimization and energy efficiency are the major challenges in multi-core system design. Of the state-of-the-art approaches, cache affinity aware scheduling and techniques based on dynamic voltage frequency scaling (DVFS) are widely applied to improve performance and save energy consumptions respectively. In modern operating systems, schedulers exploit high cache affinity by allocating a process on a recently used processor whenever possible. When a process runs on a high-affinity processor it will find most of its states already in the cache and will thus achieve more efficiency. However, most state-of-the-art DVFS techniques do not concentrate on the cost analysis for DVFS mechanism. In this paper, we firstly propose frequency affinity which retains the voltage frequency as long as possible to avoid frequently switching, and then present a frequency affinity aware scheduling (FAS) to maximize power efficiency for multi-core systems. Experimental results demonstrate our frequency affinity aware scheduling algorithms are much more power efficient than single-ISA heterogeneous multi-core processors. Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu |
MASCOTS | 2 |
| 2012 | Analyzing Parallelization and Program Performance in Heterogeneous MPSoCsabstractIn this paper we extend and analyze Amdahl's law to general heterogeneous MPSoC era, to find out how the speedup is affected by the parameters, including amount and speedup for microprocessors and accelerators, as well as the task partition characteristics. We also analyze the theoretical results about how the extended Amdahl's Law is applied to leverage load balancing of a heterogeneous MPSoC without the abstract limitation of base core equivalents (BCEs). A prototype on FPGA is constructed with Microblaze processors and JPEG hardware accelerators. The experimental results demonstrate that our extended model reinforces state-of-the-art performance evaluation methods for hybrid MPSoC architectures and also provide creditable new insights on the heterogeneous research communities, in particular for scalable FPGA based reconfigurable MPSoCs. Chao Wang 0003, Xi Li 0003, Junneng Zhang, Gangyong Jia, Peng Chen 0004, Xuehai Zhou |
MASCOTS | 2 |
| 2012 | A star network approach in heterogeneous multiprocessors system on chip
Chao Wang 0003, Xi Li 0003, Junneng Zhang, Xuehai Zhou, Aili Wang 0003 |
J. Supercomput. | 2 |
| 2006 | A Fast Instruction Set Evaluation Method for ASIP Designs
Angela Yun Zhu, Xi Li 0003, Laurence T. Yang, Jun Yang 0002 |
EUC | 2 |