VLDB 2026 Research / reviewers in the wild / expert
Mei Wen
dblp:91/6049
· DBLP profile ↗
84ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0002-5875-3297ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 65 · 5 first-author · 29 since 2021Software engineering, systems software and programming languages · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Artificial intelligence and machine learning · 4 · 2 since 2021Computer networks · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DyGen: A Constant-Time Kernel Generator for Dynamic-Shape Neural NetworksabstractIn recent years, dynamic-shape neural networks have been widely adopted in intelligent applications, such as Mixture-of-Experts based large language models and computer vision tasks. However, in dynamic scenarios, operator shapes are determined at runtime. This leads to prohibitively expensive compilation times for existing static compilers, as they must search across a vast optimization space to identify the best configuration. To address the need for efficient optimization of dynamic-shape neural networks, we present DyGen (Dynamic-shape Kernel Generator)—a lightweight, two-stage compiler plug-in on GPU platforms. In the offline stage, DyGen employs deliberately crafted pruning rules to construct a compact candidate configuration set for the target hardware, then select the configuration of the high-performance kernel to train a configuration generation model. During the online stage, dynamic operator information is directly fed into the generator, which can quickly produce efficient kernel configurations without the need for costly search. Compared to state-of-the-art tensor compilers, DyGen improves inference performance by an average of 36%, while significantly reducing generation overhead from 9 seconds to 0.3 seconds. Yuhan Kang, Dong Chen 0015, Yang Shi 0008, Jianchao Yang, Zeyu Xue, Mei Wen |
DATE | 8 |
| 2026 | Harmonia: A Unified Hierarchical Scheduling Framework for Sparse Matrix Multiplication
Jingkui Yang, Fangxin Liu, Ning Yang 0012, Chenyang Guan, Zongwu Wang, Mei Wen, Li Jiang 0002, Haibing Guan |
ISCA | 8 |
| 2026 | Precision boundary modeling for area-efficient Block Floating Point accumulation
Xin Ju 0005, Yasong Cao, Zhongdi Luo, Jianchao Yang, Jingkui Yang, Dong Chen 0015, Mei Wen |
J. Syst. Archit. | 11 |
| 2026 | VersaAccel: A Versatile Configurable Accelerator for Diverse Sparse-Dense Matrix OperatorsabstractMatrix operators are fundamental to various applications, particularly in deep learning. While early models relied on dense operations, techniques like pruning have introduced sparsity, leading to a mix of dense and sparse operator types. Most existing accelerators are specialized for specific operators and perform poorly in mixed scenarios, while those supporting multiple operators often lack flexibility and suffer from suboptimal performance. To overcome these limitations, we propose VersaAccel, a configurable accelerator for sparse and dense matrix operators. It supports four distinct configurations, each optimized for a set of operators. A key feature of our design is its adaptive configuration selection mechanism, driven by a lightweight cost model that explicitly evaluates the performance-energy trade-off between available options. This allows VersaAccel to dynamically choose the most efficient configuration—opting for higher performance when the gain outweighs the energy cost, or prioritizing energy efficiency when appropriate. Experimental results demonstrate that VersaAccel achieves an average performance/area improvement of 3.10× across multiple operators (MV, MM, SPMV, SpMM, SpMSpM) and 4.06× on full model evaluations (ResNet18, VGG16, LLaMA2-7B, BERT-Base), compared to mainstream accelerators. Minjin Tang, Mei Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | C-CIM: A Multi-Mode Convolution-Capable SRAM-CIMabstractSRAM is widely used in computing-in-memory (CIM) neural network accelerators because of its relatively mature technology and good compatibility with complementary metal oxide semiconductor logic process. Digital SRAM-CIM is favored by researchers because of its stability and accuracy. However, the current digital SRAM-CIM macro only supports the weight-stationary dataflows, which means the repeated movement of graph data. Some special deep neural network layers, such as depth-wise, make the utilization of computing resources inside CIM low. To overcome these problems, we propose C-CIM, which can switch between input-stationary and weight-stationary dataflows and support matrix multiplication as well as convolution operations with multiple mainstream convolution kernel sizes (1×1, 3×3, 5×5 and 7×7). The C-CIM achieves an average performance of 27.31TOPS/W@8b at a frequency of 1GHz. Experimental results show that our proposed SRAM-CIM successfully outperforms baseline in terms of performance optimization, achieving up to 7.6× performance speedup and up to 86.84% reduction in activation relocation. Renyu Yang, Xin Ju 0005, Mei Wen, Jinjin Deng, Junzhong Shen, Tianyu Wang 0003, Zhaoyan Shen, Zili Shao |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | WinAcc: Window-based Acceleration of Neural Networks Using Block Floating PointabstractDeep Neural Networks (DNNs) impose significant computational demands, necessitating optimizations for computational and energy efficiencies. Per-vector scaling, which applies a scaling factor to blocks of elements using narrow integer types, effectively reduces storage and computational overhead. However, the frequent occurrence of floating-point accumulations between vectors limits further improvements in energy efficiency. State-of-the-art accelerators address this challenge by grouping and summing vector products based on their exponent differences, thereby reducing the overhead associated with intra-group shifting and accumulation. Nevertheless, this approach increases the complexity of register usage and grouping logic, leading to limited energy benefits and hardware efficiency. In this context, we introduce WinAcc, a novel algorithm and architecture co-designed solution that utilizes a low-cost accumu-lator to handle the majority of data in DNNs, offering low area overhead and high energy efficiency gains. Our key insight is that the data of DNNs follows a Laplace-like distribution, which enables the use of a customized data format with a narrow dynamic range to encode most of the data. This allows for the design of a low-cost accumulator with narrow shifters and adders, significantly reducing reliance on floating-point accumulator and consequently improving energy efficiency. Compared with state-of-the-art architecture Bucket, WinAcc achieves 33.95% energy reduction across seven representative DNNs and reduces area by 9.5% while maintaining superior model performance. Xin Ju 0005, Mei Wen, Yasong Cao, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008 |
DATE | 3 |
| 2025 | SparSynergy: Unlocking Flexible and Efficient DNN Acceleration Through Multi-Level SparsityabstractTo more effectively address the computational and memory requirements of deep neural networks (DNNs), leveraging multi-level sparsity-including value-level and bit-level sparsity-has emerged as a pivotal strategy. While substantial research has been dedicated to exploring value-level and bit-level sparsity individually, the combination of both has largely been overlooked until now. In this paper, we propose SparSynergy, which-to the best of our knowledge-is the first accelerator that synergistically integrates multi-level sparsity into a unified framework, maximizing computational efficiency and minimizing memory usage. However, jointly considering multi-level sparsity is non-trivial, as it presents several challenges: (1) increased hardware overhead due to the complexity of incorporating multiple sparsity levels, (2) bandwidth-intensive data transmission during multiplexing, and (3) decreased throughput and scalability caused by bottlenecks in bit-serial computation. Our proposed SparSynergy addresses these challenges by introducing a unified sparsity format and a cooptimized hardware design. Experimental results demonstrate that SparSynergy achieves a 5.38 x geometric mean improvement in the energy-delay product (EDP) when compared with the tensor core, across workloads with varying degrees of sparsity. Furthermore, SparSynergy significantly improves accuracy retention compared to state-of-the-art accelerators for representative DNNs. Jingkui Yang, Mei Wen, Junzhong Shen, Jianchao Yang, Yasong Cao, Minjin Tang, Zhaoyun Chen, Yang Shi 0008 |
DATE | 2 |
| 2025 | SmartBlock: Adaptive Block Floating Point Quantization for Efficient DNN AccelerationabstractDeep Neural Networks (DNNs) have achieved remarkable success as model sizes continue to grow, driving the need for optimizations in both computational and energy efficiency. Block Floating Point (BFP) quantization has emerged as an effective model compression technique, offering a favorable trade-off between model accuracy and hardware cost. However, the frequent use of floating-point (FP) accumulation across BFP blocks remains a significant bottleneck, limiting further improvements in energy efficiency. State-of-the-art (SotA) accelerators mitigate this issue by introducing low-overhead accumulators with a narrower dynamic range ahead of the FP accumulator to handle a small range of values. While this approach reduces the activation of power-hungry alignment and format conversion units, it increases the complexity of the processing elements (PEs), thereby limiting the overall energy savings. Xin Ju 0005, Jingkui Yang, Mei Wen, Minjin Tang, Zhaoyun Chen, Yang Shi 0008 |
ICPP | 3 |
| 2025 | Super Microscaling: Enhancing Precision and Hardware Efficiency in Deep Learning QuantizationabstractThe Microscaling (MX) data format, an state-of-the-art quantization technique for deep learning tensor operations, suffers from precision loss at low bit-widths and potential hardware overhead. To overcome these limitations, we propose the Super Microscaling (SMX) data format and its accompanying hardware engine. SMX ensures higher accuracy by optimizing the conversion from Floating Point (FP) to MX, particularly in low-bit width scenarios. We also present spatial-temporal reuse strategy and two-level dequantization hardware architecture, which boosts efficiency in terms of hardware area and energy consumption. Experimental results demonstrate that SMX outperforms MX, achieving average accuracy improvements of 55.86% for BERT, 40.89% for ResNet50, and 37.51% for VGG, while reducing hardware area and energy consumption by 22.74% and 25.46%, respectively. Zhuang Cao, Xin Ju 0005, Mei Wen, Yang Guo 0003 |
ISCAS | 4 |
| 2025 | CAMO: A High-Performance CIM-based Lightweight CNN Accelerator for Mobile DevicesabstractDigital Compute-in-Memory (CIM) macros revolutionize the Von Neumann architecture by significantly reducing data movements between CPU and memory. However, when dealing with lightweight CNNs with various convolution types, existing GEMM (general matrix multiplication)-oriented solutions suffer from underutilization and large activation traffic, leading to unsatisfied energy and area efficiencies. In this context, we propose CAMO, in which the key contributions are: (1) A novel convolution mapping mechanism suitable for CIM macros, that maximizes data reuse and reduces activation traffic. (2) A convolution-capable CIM macro, that also supports small-scale GEMM. (3) A CIM-based architecture that supports multiple computing modes. The experimental results show that CAMO achieves up to 31.49× performance speedup and 77.5% activation traffic reduction compared to the baseline architecture. Xin Ju 0005, Renyu Yang, Mei Wen, Junzhong Shen, Tianyu Wang 0009, Zhaoyan Shen, Zili Shao |
ISCAS | 3 |
| 2025 | A quantized network processor for mixed-precision deep learning models based on enhanced Microscaling format
Mei Wen, Xin Ju 0005, Zhuang Cao, Yang Guo 0003 |
J. Syst. Archit. | 2 |
| 2025 | SpikeFlow: A hardware-software co-designed systolic array for spiking neural networks
Yang Shi 0008, Zhaoyun Chen, Mei Wen |
J. Syst. Archit. | 4 |
| 2025 | ESCAN: Efficient GPU sharing for cascade neural network inference
Yang Shi 0008, Zhaoyun Chen, Mei Wen |
Neural Networks | 4 |
| 2025 | MAP-SIM: A DNN-Specific Mapping Optimization Framework for Shared-Memory CPU-Systolic Array ArchitecturesabstractAs performance demands continue to rise, Shared-Memory Heterogeneous Systems (SMHSs) have been widely adopted for their ability to enable efficient communication and data sharing between different heterogeneous cores. However, existing SMHS face challenges in uneven workload distribution among heterogeneous cores and suboptimal mapping schemes, preventing them from fully leveraging their architectural advantages. To address these issues, this paper proposes a mapping-aware framework for modeling SMHSs called MAP-SIM. By performing performance modeling for CPUs and Systolic Arrays (SAs), and considering rational schemes for the partition and mapping of computational tasks, MAP-SIM aims to evaluate and optimize the computational performance of heterogeneous multicore architectures. The experimental results show that compared to previous work, MAP-SIM can increase simulation speed by 14 to 67 times and can also enhance the computational performance of SMHS by 1.4 to 4.4 times. Mei Wen, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008, Tianyu Wang 0009, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | SPSA: Exploring Sparse-Packing Computation on Systolic Arrays From ScratchabstractSparse matrix-matrix multiplication (SpMM) and Generalized SpMM (SpGEMM) are essential computational kernels in domains, such as graph analytics and scientific computation. While systolic arrays have traditionally been employed as specialized architectures for complex computing problems like matrix multiplication, they exhibit inefficiency when dealing with sparse matrices. This inefficiency arises from the unnecessary operations performed by processing elements (PEs) that contain zero-valued entries, which do not contribute to the final result. To address this issue, we propose SPSA, a framework that leverages a sparse-packing algorithm suitable for systolic arrays to accelerate sparse matrix computations. Our approach achieves significant reduction of zero-valued items and improves matrix density by packing the rows or columns of the sparse matrix. Furthermore, we have introduced for the first time a data representation format tailored to systolic arrays, called CSXD, which further enhances storage and computational efficiency. Importantly, our adaptation scheme enables acceleration benefits even with limited resources. Through sparse packing, SPSA achieved a$5.2\times $performance improvement compared to the dense baseline, and further reached a$6.4\times $enhancement via CSXD. Simultaneously, CSXD realized an average storage efficiency improvement of$15.0\times $. Through extensive evaluations, SPSA outperforms previous designs on CPU, GPU, and ASIC platforms. Finally, in end-to-end evaluations, SPSA achieved a performance improvement of 3.9 times across the workloads of BERT, VGG19, and ResNet50. Minjin Tang, Mei Wen, Jianchao Yang, Zeyu Xue, Junzhong Shen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | SSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor CoresabstractSparse-Dense Matrix-Matrix Multiplication (SpMM) has emerged as a foundational primitive in HPC and AI. Recent advancements have aimed to accelerate SpMM by harnessing the powerful Tensor Cores found in modern GPUs. However, despite these efforts, existing methods frequently encounter performance degradation when ported across different Tensor Core architectures. Recognizing that scalable SpMM across multiple generations of Tensor Cores relies on the effective use of general-purpose instructions, we have meticulously developed a SpMM library named SSpMM. However, a significant conflict exists between granularity and performance in current Tensor Core instructions. To resolve this, we introduce the innovative Transpose Mapping Scheme, which elegantly implements fine-grained kernels using coarse-grained instructions. Additionally, we propose the Register Shuffle Method to further enhance performance. Finally, we introduce Sparse Vector Compression, a technique that ensures our kernels are scalable with both structured and unstructured sparsity. Our experimental results, conducted on four generations of Tensor Core GPUs using over 3,000 sparse matrices from well established matrix collections, demonstrate that SSpMM achieves an average speedup of 2.04×, 2.81×, 2.07×, and 1.87×, respectively, over the state-of-the-art SpMM solution. Furthermore, we have integrated SSpMM into PyTorch, achieving a 1.81× speedup in end-to-end Transformer inference compared to cuDNN. Zeyu Xue, Mei Wen, Jianchao Yang, Minjin Tang, Zhongdi Luo, Yang Shi 0008, Zhaoyun Chen, Junzhong Shen, Johannes Langguth |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | BitShare: An Efficient Precision-Scalable Accelerator with Combining-Like-Terms GEMMabstractNarrow-precision fixed-point (INT) computation is a significant approach for reducing memory requirements and enhancing the performance of accelerators for Deep Neural Networks (DNNs). Different DNNs, as well as different layers within the DNNs, may exhibit varying numerical distributions, necessitating INT formats with different minimum bit-widths. Therefore, DNN accelerators need to support multi-precision INT computations to strike a better balance between DNN inference accuracy and performance. However, existing precision-scalable accelerators face challenges such as low bandwidth utilization, insufficient utilization of computing resources across different precision modes, and complex circuit structures with associated overhead. In this paper, we propose (1) a hardware-friendly Combining-Like-Terms GEMM (CLT-GEMM) scheme that supports multiple computing modes of 2/4/8 bits and their combinations to align with the various bit-width settings of DNNs; (2) and subsequently design an efficient systolic accelerator with scalable precision, named BitShare, which features DataMap module and Multi-mode adder-tree-based accumulators. Compared to the state-of-the-art precision-scalable design, BitBlade, our accelerator achieves a 57.25% reduction in bandwidth requirement and exhibits an improvement of$1.14\times$and$1.12\times$in area and power efficiency$(2\mathbf{b}\times 2\mathbf{b})$, respectively. Yasong Cao, Mei Wen, Junzhong Shen, Zhongxing Li |
ASAP | 2 |
| 2024 | MAP-SIM: A Performance Model for Shared-Memory Heterogeneous Systems with Mapping Awareness
Mei Wen, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008 |
ICA3PP (1) | 2 |
| 2024 | MSA2: An Efficient Sparsity-Aware Accelerator for Matrix Multiplication with Multi-core Systolic Arrays
Minjin Tang, Mei Wen, Junzhong Shen, Jingkui Yang, Zeyu Xue, Zili Shao |
ICA3PP (3) | 2 |
| 2024 | Enhancing the PE Utilization for Multi-Precision Systolic Array via Optimizing Computation LatencyabstractSystolic array (SA) architectures are widely recognized as the optimal choice for Convolutional Neural Networks (CNNs). However, existing SAs suffer from reduced computational efficiency when confronted with an inadequate workload scale. Furthermore, this issue becomes even more pronounced in the accelerators that support multiple precisions. In this paper, by analyzing the under-utilization of processing element (PE), we propose a SA accelerator that optimizes computation latency for multi-precision scenarios. Considering dynamic changes in data precision, we incorporate a switching strategy to further enhance computational efficiency. Experimental results demonstrate that our proposed design only incurs a 1.203% increase in area compared to the classic approach, while achieving an average performance improvement of 20% on small-scale CNN models. Mei Wen, Xin Ju 0005, Junzhong Shen, Yang Guo 0003 |
ISCAS | 2 |
| 2024 | HyFiSS: A Hybrid Fidelity Stall-Aware Simulator for GPGPUsabstractThe widespread adoption of GPUs has driven the development of GPU simulators, which, in turn, lead advancements in both GPU architectures and software optimization. Trace-driven cycle-accurate Cycle-accurate simulators, which provide detailed microarchitectural models and clock-level precision, come at the cost of extended simulation times and require high computational resources. Their scalability has become a bottleneck. A growing trend is the adoption of cycle-approximate simulators, which introduce mathematical modeling of partial hardware units and utilize sampling to accelerate simulation. However, this approach faces challenges regarding the accuracy of performance predictions. To address these limitations, we introduce HyFiSS, a hybrid fidelity stall-aware GPU simulator. HyFiSS features fine-grained stall events tracking and attribution by constructing a detailed execution pipeline model for various stall events on Streaming Multiprocessors (SMs). It accurately emulates the thread block scheduler behavior using real-time scheduling logs and utilizes sampling based on thread block sets to minimize the precision loss due to fine-grained sampling points on the microarchitectural state. We achieve a balance between reliability, speed, and the level of simulation detail, especially regarding bottlenecks. By evaluating a diverse set of benchmarks, HyFiSS achieves a mean absolute percentage error in predicting active cycles that is comparable to the state-of-the-art cycle-accurate simulator Accel-Sim. Moreover, HyFiSS achieves a substantial 12.8 × speedup in the simulation efficiency compared to Accel-Sim. HyFiSS also requires at least 3.2 × less disk storage than both Accel-Sim and another state-of-the-art cycle-approximate simulator PPT-GPU due to its efficient SASS (Streaming Assembler) traces compression. With precise, per-cycle stall events statistics, HyFiSS can provide accurate GPU performance metrics and stall cause reporting. This significantly simplifies performance analysis, bottleneck identification, and performance optimization tasks for researchers, making it easier to enhance GPU performance effectively. Jianchao Yang, Mei Wen, Dong Chen 0015, Zhaoyun Chen, Zeyu Xue, Junzhong Shen, Yang Shi 0008 |
MICRO | 2 |
| 2024 | ESEN: Efficient GPU sharing of Ensemble Neural Networks
Yang Shi 0008, Zhaoyun Chen, Mei Wen |
Neurocomputing | 4 |
| 2024 | Optimizing VLIW Instruction Scheduling via a Two-Dimensional Constrained Dynamic ProgrammingabstractTypical embedded processors, such as Digital Signal Processors (DSPs), usually adopt Very Long Instruction Word (VLIW) architecture to improve computing efficiency. The performance of VLIW processors heavily relies on Instruction-Level Parallelism (ILP). Therefore, it is crucial to develop an efficient instruction scheduling algorithm to explore more ILP. While heuristic algorithms are widely used in modern compilers due to simple implementation and low computational cost, they have limitations in providing accurate solutions and are prone to local optima. On the other hand, exact algorithms can usually find the optimal solution, but their high time overhead makes them less suitable for large-scale problems. This article proposes a two-dimensional constrained dynamic programming (TDCDP) approach and a quantitative model for instruction scheduling. The TDCDP approach achieves near-optimal solutions within an acceptable time overhead. Furthermore, we integrate our TDCDP approach into mainstream compiler architecture, encompassing Pre- and Post-RA (register allocation) scheduling. We conduct a quantitative evaluation of TDCDP compared with four heuristic algorithms on a typical VLIW processor. Our approach achieves an efficiency improvement of up to 58.34% in final solutions compared with the heuristic algorithms. Additionally, the Post-RA Scheduling enhances programs with an average speedup of 14.04% than solely applying the Pre-RA Scheduling. Can Deng, Zhaoyun Chen, Yang Shi 0008, Yimin Ma, Mei Wen, Lei Luo 0002 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2024 | ABS: Accumulation Bit-Width Scaling Method for Designing Low-Precision Tensor CoreabstractA big gap exists between deep neural network (DNN) applications’ computational demand and the computing power of DNN accelerators. Low-precision floating-point (LP-FP) computation is one of the important means to improve the performance of DNN training and inference. However, the high-precision accumulators are typically applied to summating the dot products during general matrix multiplication (GEMM) in tensor cores (TCs). As the precision of data decreases, the accumulator becomes the main consumer of multiply-accumulate’s (MAC’s) area and power. Reducing the accumulators’ bit-width is of significant importance for improving the area- and energy-efficiency of TCs. There are two main challenges: 1) theoretical support on the floating-point (FP) formats with the lowest bit-width of TC’s accumulators and 2) how to integrate the LP-FP TC in the framework of DNN training and inference to evaluate its benefits. In this article, we propose accumulation bit-width scaling (ABS), a novel ABS method, to guide the design of LP-FP TCs. We 1) implement this method by constructing a novel variance retention ratio (VRR) model to predict the FP format with the minimum bit-width for TC’s accumulator; 2) provide a generator of DNN accelerator based on a systolic-array (SA) TC, supporting many low-precision configurations; and 3) design an LP-FP DNN executing framework that supports software-simulation mode and hardware-accelerator mode to run LP-FP DNN tasks. The experimental results show that the LP-FP TC guided by our ABS method has a maximum reduction of 76.47% and 75.60% in area and power consumption, respectively, compared with the advanced TCs. Yasong Cao, Mei Wen, Zhongdi Luo, Xin Ju 0005, Haolan Huang, Junzhong Shen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | Automatic End-to-End Joint Optimization for Kernel Compilation on DSPsabstractDigital signal processors (DSPs) commonly adopt VLIW-SIMD architecture and are extensively applied in most compute-heavy embedded sensing applications. The performances for DSP kernels rely heavily on compilations and handwritten optimizations. Hand-crafted methods suffer from heavy burden on programmers, while state-of-the-art automatic compilation methods always focus more on a certain aspect (tiling or auto-vectorization), lacking of global and sequential vision on the intact compilation optimization process. It still requires empirical adjustments by programmers in the actual scenario.In order to release programmers from kernel tuning, we propose JOKer, an automatic end-to-end multi-level code generator for kernel joint optimization on DSPs. JOKer integrates means of optimizations in compiling process and provides an end-to-end workflow for performance tuning. It explores compilation configurations through a reinforcement learning based agent for global optimal solution and generates high performance kernel codes for DSPs automatically. Zhaoyun Chen, Yang Shi 0008, Mei Wen, Chunyun Zhang |
DAC | 4 |
| 2023 | Releasing the Potential of Tensor Core for Unstructured SpMM using Tiled-CSR FormatabstractThe GPU has become a popular platform for AI applications, thanks in part to its Tensor Cores that address performance issues. However, the Sparse Matrix Multiplication (SpMM) kernel has remained a bottleneck despite significant advances in computing power. Due to the hardware mechanism of the Tensor Core, its programming granularity does not match SpMM. In this paper, we analyze the reasons why the unstructured SpMM kernel is not suitable for the Tensor Core, and propose the Tiled Compressed Sparse Row (Tiled-CSR) compression format. To address the issue of low non-zero rates in Tiled-CSR format, we exploit the row shuffle algorithm to improve the utilization of Tensor Cores and enhance computing density. We also utilize adaptive memory access modes and 3D-Grid tiling for the SpMM kernel to reduce memory access latency. The experimental results on NVIDIA A100 GPU with matrices in the Deep Learning Matrix Collection (DLMC) demonstrate that the Tiled-CSR format improves the utilization of Tensor Cores under different sparsity, with a maximum of 3.89× at 50% sparsity and a minimum of 1.82× at 90% sparsity compared to the SR-BCRS format. Additionally, our kernel achieves an average speedup of 1.54×(up to 2.12×) over Magicube. Zeyu Xue, Mei Wen, Zhaoyun Chen, Yang Shi 0008, Minjin Tang, Jianchao Yang, Zhongdi Luo |
ICCD | 2 |
| 2022 | Exploring ILP for VLIW Architecture by Quantified Modeling and Dynamic Programming-Based Instruction SchedulingabstractExploring the instruction-level parallelism (ILP) of Very Long Instruction Word (VLIW) architecture relies on instruction scheduling. List scheduling (LS) algorithms, which are most adopted in modern compilers, have limitations in searching for optimal solutions. This paper proposes a quantifiable model for instruction scheduling and a dynamic programming-based strategy (DPS). We evaluate DPS on a specified platform and realize high efficiency. The results suggest that the DPS achieves an efficiency improvement of up to 44.72% within acceptable time overhead. Can Deng, Zhaoyun Chen, Yang Shi 0008, Xichang Kong, Mei Wen |
ASP-DAC | 5 |
| 2022 | Light: A Component Enhances Faster and More Accurate Traffic Measurement*abstractThe greatest challenge when designing an online sketch method for data flow measurement is to reduce the storage cost of sketches with little loss of accuracy and obtain a higher bandwidth. To address this issue, we proposed Light component. By storing elephant and mice flows separately, the accuracy of the Light-enhanced sketches is substantially improved, and the processing speed is significantly faster due to a significant reduction in the required computational overhead. We implement Light-enhanced sketches alongside several existing sketch methods on CPU and compare their performance. Experiments show that, under the same storage conditions, Light-enhanced sketches can greatly reduce the Average Relative Error by 1.80 to 5.07 times, as well as increase the average processing speed of each packet by 4.96 to 9.05 times compared with their original structure. Our approach also achieves a stable performance in the measurement of traffic in different traffic distributions. More importantly, Light component can be deployed to different sketch methods, which demonstrates its reusability. Jianchao Yang, Mei Wen, Yang Shi 0008 |
ICC | 2 |
| 2022 | BP-Im2col: Implicit Im2col Supporting AI Backpropagation on Systolic ArraysabstractState-of-the-art systolic array-based accelerators adopt the traditional im2col algorithm to accelerate the inference of convolutional layers. However, traditional im2col cannot efficiently support AI backpropagation. Backpropagation in convolutional layers involves performing transposed convolution and dilated convolution, which usually introduces plenty of zero-spaces into the feature map or kernel. The zero-space data reorganization interfere with the continuity of training and incur additional and non-negligible overhead in terms of off- and on-chip storage, access and performance. Since countermeasures for backpropagation are rarely proposed, we propose BP-im2col, a novel im2col algorithm for AI backpropagation, and implement it in RTL on a TPU-like accelerator. Experiments on TPU-like accelerator indicate that BP-im2col reduces the backpropagation runtime by 34.9% on average, and reduces the bandwidth of off-chip memory and on-chip buffers by at least 22.7% and 70.6% respectively, over a baseline accelerator adopting the traditional im2col. It further reduces the additional storage overhead in the backpropagation process by at least 74.78%. Jianchao Yang, Mei Wen, Junzhong Shen, Yasong Cao, Minjin Tang, Renyu Yang, Jiawei Fei, Chunyuan Zhang |
ICCD | 2 |
| 2022 | Mentha: Enabling Sparse-Packing Computation on Systolic ArraysabstractGeneralized Sparse Matrix-Matrix Multiplication (SpGEMM) is a critical kernel in domains like graph analytic and scientific computation. As a kind of classical special-purpose architecture, systolic arrays were first used for complex computing problems, e.g., matrix multiplication. However, classical systolic arrays are not efficient enough when handling sparse matrices due to the fact that the PEs containing zero-valued entries perform unnecessary operations that do not contribute to the result. Accordingly, in this paper, we propose Mentha, a framework that enables systolic arrays to accelerate sparse matrix computation by employing a sparse-packing algorithm suitable for various dataflow of systolic array. Firstly, Mentha supports both online and offline methods. By packing the rows or columns of the sparse matrix, the zero-valued items in the matrix are significantly reduced and the density of the matrix is improved. In addition, acceleration benefits can be obtained by the adaptation scheme even with limited resources. Moreover, we reconfigure PEs in systolic arrays at a low cost (1.28x in area, 1.21x in power) and find that our method outperforms TPU-like systolic arrays by 1.2~3.3x in terms of SpMM and 1.3~4.4x in terms of SpGEMM when dealing with moderately sparse matrices (sparsity < 0.9), while its performance is at least 9.7x better than cuSPARSE. Furthermore, experimental results show a FLOPs reduction of roughly 3.4x in the neural network. Minjin Tang, Mei Wen, Yasong Cao, Junzhong Shen, Jianchao Yang, Jiawei Fei, Yang Guo 0003, Sheng Liu 0001 |
ICPP | 2 |
| 2022 | S-SIM: A Simulator for Systolic Array-based DNN Accelerators with Tile Access AwarenessabstractAs NN accelerators emerging, many analytical models are presented to help designers to carry out hardware design space exploration. However, these models cannot accurately simulate the systolic array-based NN accelerator due to their pervasiveness or abstraction. In this paper, we propose a compute-centric simulator driven by the execution of events from the tiles of the mapping matrix, which can accurately model the systolic array-based accelerator. The simulator focuses on the conflicts when the tile is used for data access, or various interruptions caused by hardware resource limitations. Experimental results show that the proposed simulator achieves more than 95% accuracy compared to the real scenes. Mei Wen, Renyu Yang, Junzhong Shen, Yasong Cao |
ISCAS | 2 |
| 2022 | TILE-SIM: A Systematic Approach to Systolic Array-based Accelerator EvaluationabstractThe systolic array provides extremely high efficiency for running matrix multiplication, and is one of the mainstream architectures of today’s deep learning accelerators. In order to develop efficient accelerators, people usually employ simulators to make design trade-offs. However, current simulators suffer from coarse-grained modeling methods and ideal assumptions, which limits their ability of describing structural characteristics of systolic arrays. In addition, they do not support the exploration of microarchitecture. This paper presents TILE-SIM, a computing-centric systematic method for evaluating systolic array accelerators by using an event-driven method. TILE-SIM can obtain accurate results and provide the best mapping scheme for different workload due to its fine-grained modeling technique and deny of ideal assumption. Experimental results show that TILE-SIM plays a significant role in design trade-offs and outperforms state-of-the-art simulators, with an accuracy of more than 95%. Mei Wen, Jiawei Fei, Junzhong Shen, Yasong Cao |
ISPASS | 2 |
| 2021 | sRouting: Towards a Better Flow Size Estimation Performance through Routing and Sketch ConfigurationabstractFlow size estimation is highly important and beneficial to various applications, including traffic engineering and anomaly detection. Considering the resource constraints and performance requirements in place, sketches are widely used to accomplish this task. However, when sketches are applied in real-life networks, the measurement performance is often decreased due to many practical reasons, including partial deployment of sketches, unique traffic characteristics, etc. In this paper, we present sRouting, a practical framework that aims to better utilize the deployed sketches. sRouting improves the measurement performance by optimizing which flows are monitored and where. We first investigate the relationship between the accuracy of a given sketch and the total number of packets, then formulate the offline problem in sRouting as an integer linear programming problem. To solve this problem efficiently, we devise a rounding-based algorithm and provide its performance guarantees. Furthermore, to handle dynamic changes in the network, we design an online adjustment algorithm capable of responding appropriately to these changes. Through experiments using real traces and typologies, we demonstrate that sRouting can significantly improve the volume of monitored traffic and reduce the measurement error without negatively impacting the network throughput. Yang Shi 0008, Mei Wen |
ICPP | 2 |
| 2021 | Automatic mapping and code optimization for OpenCL kernels on FT-matrix architecture (WIP paper)abstractFT-Matrix is a typical vector-SIMD architecture that refines the cooperation between scalar and vector units. This approach is widely used in digital signal processing, high-performance computing, and artificial intelligence, among other fields. FT-Matrix currently adopts C vector extension as the main programming model, improving the utilization efficiency of SIMD by providing explicit vector extension API. Moreover, it is difficult to efficiently transplant parallel programs (OpenCL, CUDA) adopted by users. This paper proposes an automatic mapping and code optimization method for OpenCL kernels on FT-Matrix architecture. The proposed approach solves these challenges by means of work item coalescing, slicing and rotation, and instruction-level code optimization. Preliminary results show that our method can achieve high performance and good hardware utilization for OpenCL kernels, as well as decreasing the programming difficulty on FT-Matrix. Mei Wen, Zhaoyun Chen, Yang Shi 0008, Chunyuan Zhang |
LCTES | 2 |
| 2020 | Towards Memory-Efficient Streaming Processing with Counter-Cascading Sketching on FPGAabstractObtaining item frequencies in data streams with limited space is a well-recognized and challenging problem in a wide range of applications. Sketch-based solutions have been widely used to address this challenge due to their ability to accurately record the data streams at a low memory cost. However, most sketches suffer from low memory utilization due to the adoption of a fixed counter size. Accordingly, in this work, we propose a counter-cascading scheduling algorithm to maximize the memory utilization of sketches without incurring any accuracy loss. In addition, we propose an FPGA-based system design that supports sketch parameter learning, counter-cascading record and online query. We implement our designs on Xilinx VCU118, and conduct evaluations on real-world traces, thereby demonstrating that our design can achieve higher accuracy with lower storage; the performance achieved is 10× ~ 20× better than that of state-of-the-art sketches. Minjin Tang, Mei Wen, Junzhong Shen, Chunyuan Zhang |
DAC | 2 |
| 2020 | Scalable FPGA-based Architecture for High-Performance Per-Flow Traffic MeasurementabstractPer-flow traffic measurement has emerged as a critical but challenging task in data center in recent years in the face of massive network traffic. Many approximate methods have been proposed to resolve the existing resource-accuracy trade-off in per-flow traffic measurement, one of which is the sketch-based method. However, sketches are affected by their high computational cost and low throughput; moreover, their measurement accuracy is hard to guarantee under the conditions of changing network bandwidth or flow size distribution. Recently, FPGA platforms have been widely deployed in data centers, as they demonstrate a good fit for high-speed network processing. In this work, we propose a scalable pipelined architecture for high high-throughput per-flow traffic measurement on FPGA. We adopts memory-friendly D-left hashing in our design, which guarantees high space utilization that successfully addressing the challenge of tracking high speed data stream under limit memory resource on FPGA. Comparisons with state-of-the-art sketch-based solutions show that our design outperforms state-of-the-art sketch-based methods in terms of throughput by over 80x. Junzhong Shen, Mei Wen, Minjin Tang, Chunyuan Zhang |
FPGA | 2 |
| 2020 | Towards a Deep-Pipelined Architecture for Accelerating Deep GCN on a Multi-FPGA Platform
Qixuan Cheng, Mei Wen, Junzhong Shen, Chunyuan Zhang |
ICA3PP (1) | 2 |
| 2020 | Optimized HybridSketch: More Efficient with Analysis and Algorithm
Mei Wen, Minjin Tang, Qun Huang 0001, Chunyuan Zhang |
ICA3PP (1) | 2 |
| 2020 | HybridSketch: A Memory-centric Precise Approach for Flow MeasurementabstractAs network bandwidth has rapidly developed, due to the high occupancy of memory and bandwidth required, the Sketch structure is favored by some researchers due to its limited memory usage and simple operation. But the accuracy will decrease when the Sketch system occupies less memory space. Traditional sketch algorithms and some other specially designed algorithms and structures are striving to improve accuracy. However, with the flow rate rapidly increasing, the on-chip memory will be the bottleneck of the system. Our network measurement system achieve good results focusing more on the memory usage. We proposes a hybrid method, HybridSketch, which focuses on the memory and precision of the system with mixing two measurement methods by quantitatively analyzing, modeling and allocating appropriate memory space to each method to achieve better results. Experimental results show that our method can provide 10× improvement in terms of precision, moreover, HybridSketch can provide the same level of precision with achieving 24× improvement in terms of memory size. Mei Wen, Minjin Tang, Qun Huang 0001, Chunyuan Zhang |
ICC | 2 |
| 2020 | Towards High-Efficiency Data Centers via Job-Aware Network SchedulingabstractDistributed jobs typically facing competition for multiple resources in modern data centers, especially for network. Without effective network scheduling, this competition can cause low efficiency of the data center. Previous work on network scheduling has been focused on reducing flow completion time or improving per-flow fairness. Yet, its effect on improving jobs’ performance is limited by the unawareness of relationships between communication and computation. In this paper, we focus on the problem of scheduling network resources for multiple jobs, with the specific objective to reduce the job completion time (JCT), which also makes the datacenter more efficient. With an in-depth investigation of communication and computation, we identify an opportunity for accelerating jobs in a way that occupies less bandwidth for DAG-based complicated modern jobs. Accordingly, this paper proposes JIT, a job-aware network scheduler that leverages the computational graph to accelerate jobs effectively. To cater to the goal of JIT, we first develop a mathematical model and formulate the scheduling problem as an integer linear programming (ILP) problem. We further prove that it has an equivalent linear programming (LP) problem through rigorous theoretical analysis in order to solve this ILP problem efficiently. Some reasonable simplifications are also adopted to reduce the solving time of JIT to only 1 second. The proposed JIT is simulated and compared against some state-of-the-art designs, and the simulation results demonstrate that JIT can achieve an acceleration of up to 1.55 × , which successfully improves the efficiency of the data center. Yang Shi 0008, Mei Wen, Chunyuan Zhang |
ICPP | 2 |
| 2020 | Incremental Deployment of Programmable Switches for Sketch-based Network MeasurementabstractThe emergence of programmable switches has boosted lots of research around many network aspects: mea-surements, security, quality of services. To explore the ad-vantages of programmable data planes while preserving the legacy networking systems, deploying programmable switches incrementally may be a more practical solution. In this paper, we deal with the programmable switch deploy problem for sketch-based network measurement, which has been overlooked before. We first analyze the desired properties of a good deployment for sketch-based network measurement with some examples. Based on summarized lessons, we then develop two Integer Linear Programming (ILP) models, namely TraceILP and TopoILP, to solve the deployment problem. If historical traffic traces are provided, TraceILP generates better deployment with historical information. Even if no traces are provided, TopoILP can still make a reasonable strategy according to the network topology. Evaluations on real ISP and datacenter topologies show that pro-posed models guarantee a promising measurement performance with only about 40% devices upgraded to programmable ones. Yang Shi 0008, Mei Wen, Chunyuan Zhang |
ISCC | 2 |
| 2020 | Toward an Efficient Deep Pipelined Template-Based Architecture for Accelerating the Entire 2-D and 3-D CNNs on FPGAabstract3-D convolutional neural networks (3-D CNNs) are used efficiently in many computer vision applications. Most previous work in this area has concentrated only on design and optimization of accelerators for 2-D CNNs, with few attempts having been made to accelerate 3-D CNNs on FPGA. We find the acceleration of 3-D CNNs on FPGA to be challenging due to their high computational complexity and storage demands. More importantly, although the computational patterns of 2-D and 3-D CNNs are analogous, the conventional approaches that have been adopted for acceleration of 2-D CNNs may be unfit for 3-D CNN acceleration. In this paper, in order to accelerate 2-D and 3-D CNNs using a uniform framework, we first propose a uniform template-based architecture that uses templates based on the Winograd algorithm to ensure the rapid development of 2-D and 3-D CNN accelerators. Then, with the aim of efficiently mapping all layers of 2-D/3-D CNNs onto a pipelined accelerator, techniques are developed to improve the throughput and computational efficiency of the accelerator, including layer fusion, layer clustering, and workload-balancing scheme. Finally, we demonstrate the effectiveness of the deep pipelined architecture by accelerating real-life 2-D and 3-D CNNs on the state-of-the-art FPGA platform. On VCU118, we achieve 3.7 TOPS for VGG-16, which outperforms state-of-the-art FPGA-based CNN accelerators. Comparisons with CPU and GPU solutions demonstrate that our implementation of 3-D CNN achieves gains of up to 17.8× and 64.2× in performance and energy relative to a CPU solution, and a 5.0× energy efficiency gain over a GPU solution. Junzhong Shen, You Huang, Mei Wen, Chunyuan Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Deep Learning Research and Development Platform: Characterizing and Scheduling with QoS Guarantees on GPU ClustersabstractDeep learning (DL) has been widely adopted in various domains of artificial intelligence (AI), achieving dramatic developments in industry and academia. Besides giant AI companies, numerous small and medium-sized enterprises, institutes, and universities (EIUs) have focused on the research and development (R&D) of DL. Considering the high cost of datacenters and high performance computing (HPC) systems, EIUs prefer adopting off-the-shelf GPU clusters as a DL R&D platform for multiple users and developers to process diverse DL workloads. In such scenarios, the scheduling of multiple DL tasks on a shared GPU cluster is both significant and challenging in terms of efficiently utilizing limited resources. Existing schedulers cannot predict the resource requirements of diverse DL workloads, leading to the under-utilization of computing resources and a decline in user satisfaction. This paper proposes GENIE, a QoS-aware dynamic scheduling framework for a shared GPU cluster, which achieves users' QoS guarantee and high system utilization. In accordance with an exhaustive characterization, GENIE analyzes the key factors that affect the performance of DL tasks and proposes a prediction model derived from lightweight profiling to estimate the processing rate and response latency for diverse DL workloads. Based on the prediction models, we propose a QoS-aware scheduling algorithm to identify the best placements for DL tasks and schedule them on the shared cluster. Experiments on a GPU cluster and large-scale simulations demonstrate that GENIE achieves a QoS-guarantee percentage improvement of up to 67.4 percent and a makespan reduction of up to 28.2 percent, compared to other baseline schedulers. Zhaoyun Chen, Wei Quan 0004, Mei Wen, Jianbin Fang, Jie Yu 0008, Chunyuan Zhang, Lei Luo 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Scale-out Acceleration for 3D CNN-based Lung Nodule Segmentation on a Multi-FPGA SystemabstractThree-dimensional convolutional neural networks (3D CNNs) have become a promising method in lung nodule segmentation. The high computational complexity and memory requirements of 3D CNNs make it challenging to accelerate 3D CNNs on a single FPGA. In this work, we focus on accelerating the 3D CNN-based lung nodule segmentation on a multi-FPGA platform by proposing an efficient mapping scheme that takes advantage of the massive parallelism provided by the platform, as well as maximizing the computational efficiency of the accelerators. Experimental results show that our system integrating with four Xilinx VCU118 can achieve state-of-the-art performance of 14.5 TOPS, in addition with a 29.4x performance gain over CPU and 10.5x more energy efficiency over GPU. Junzhong Shen, You Huang, Mei Wen, Chunyuan Zhang |
DAC | 4 |
| 2019 | GENIE: QoS-guided Dynamic Scheduling for CNN-based Tasks on SME ClustersabstractConvolutional Neural Network (CNN) has achieved dramatic developments in emerging Machine Learning (ML) services. Compared to online ML services, offline ML services that are full of diverse CNN workloads are common in small and medium-sized enterprises (SMEs), research institutes and universities. Efficient scheduling and processing of multiple CNN-based tasks on SME clusters is both significant and challenging. Existing schedulers cannot predict the resource requirements of CNN-based tasks. In this paper, we propose GENIE, a QoS-guided dynamic scheduling framework for SME clusters that achieves users' QoS guarantee and high system utilization. Based on a prediction model derived from lightweight profiling, a QoS-guided scheduling strategy is proposed to identify the best placements for CNN-based tasks. We implement GENIE as a plugin of Tensorflow and experiment with real SME clusters and large-scale simulations. The results of the experiments demonstrate that the QoS-guided strategy outperforms other baseline schedulers by up to 67.4% and 28.2% in terms of QoS-guarantee percentage and makespan. Zhaoyun Chen, Lei Luo 0002, Haoduo Yang, Jie Yu 0008, Mei Wen, Chunyuan Zhang |
DATE | 5 |
| 2019 | Accelerating 3D CNN-based Lung Nodule Segmentation on a Multi-FPGA SystemabstractLung nodule segmentation is one of the most significant steps in many Computer Aided Detection (CAD) systems used for lung nodule identification and classification. Three-dimensional convolutional neural networks (3D CNNs) have become a promising method in lung nodule segmentation, as this method can achieve higher detection accuracy than conventional methods. It has been proven that FPGAs can provide the most energy-efficient solution for CNN acceleration. However, the high computational complexity and memory requirements of 3D CNNs make it challenging to accelerate 3D CNNs on a single FPGA, as this will further bottleneck the performance of a 3D CNN-based CAD system. Accordingly, in this work, we focus on accelerating the 3D CNN-based lung nodule segmentation on a multi-FPGA platform by proposing an efficient mapping scheme that takes advantage of the massive parallelism provided by the platform, as well as maximizing the computational efficiency of the accelerators. Experimental results show that our system is able to achieve high computational efficiency and thereby a state-of-the-art performance of 14.5 TOPS at 200 MHz. Comparisons with CPU and GPU solutions demonstrate that our system achieves a 29.4x performance gain over CPU and a 10.5x energy efficiency improvement over GPU. Junzhong Shen, You Huang, Mei Wen, Chunyuan Zhang |
FPGA | 4 |
| 2019 | An Efficient Design Flow for Accelerating Complicated-connected CNNs on a Multi-FPGA PlatformabstractConvolutional Neural Networks (CNNs) have achieved impressive performance on various computer vision tasks. To facilitate better performance, some complicated-connected CNN models (e.g., GoogLeNet and DenseNet) have recently been proposed, and have achieved state-of-the-art performance in the fields of image classification and segmentation. However, CNNs are computation- and memory-intensive. Thus, it is significant to develop hardware accelerators in order to accelerate the inference and training processes of CNNs. Due to the high-performance, reconfigurable and energy-efficient nature of Field-Programmable Gate Arrays (FPGAs), many FPGA-based accelerators have been proposed to implement CNNs and have achieved higher throughput and energy efficiency. However, the large number of parameters involved in complicated-connected CNN models have exceeded the limited hardware resources of single FPGA board, which are unable to meet the memory and computation resource demands associated with mapping entire CNN models. Accordingly, in this paper, we propose a complete design flow to accelerate the inference of complicated-connected CNNs on a multi-FPGA platform, including DAG abstraction, mapping scheme generation and design space exploration. In addition, a multi-FPGA system with flexible inter-FPGA communications is proposed to efficiently support our design flow. Experimental results on representative models illustrate that the proposed multi-FPGA system design can achieve a throughput acceleration of up to 145.2× and 2.5× compared to CPU and GPU solutions, as well as an energy efficiency improvement of up to 139.1× and 4.8× compared to multi-core CPU and GPU solutions. Junzhong Shen, Mei Wen, Chunyuan Zhang |
ICPP | 3 |
| 2019 | Towards a Uniform Architecture for the Efficient Implementation of 2D and 3D Deconvolutional Neural Networks on FPGAsabstractThree-dimensional deconvolution is widely used in many computer vision applications. However, most previous works have only focused on accelerating 2D deconvolutional neural networks (DCNNs) on FPGAs, while the acceleration of 3D DCNNs has not been studied in depth as they have higher computational complexity and sparsity than 2D DCNNs. In this paper, we focus on the acceleration of both 2D and 3D DCNNs on FPGAs by proposing efficient schemes for mapping 2D and 3D DCNNs on a uniform architecture. By implementing our design on the Xilinx VC709 platform for four real-life 2D and 3D DCNNs, we can achieve up to 3.0 TOPS with high hardware efficiency. Comparisons with CPU and GPU solutions demonstrate that we can achieve an improvement of up to 63.3 × in throughput relative to a CPU solution and an improvement of up to 8.3 × in energy efficiency compared to a GPU solution. Junzhong Shen, Mei Wen, Chunyuan Zhang |
ISCAS | 3 |
| 2019 | KVSwitch: An In-network Load Balancer for Key-Value StoresabstractToday's cloud-based online services are underpinned by distributed key-value stores (KVSs). Keys and values are distributed across back-end servers in such scale-out systems. One primary real-life performance bottleneck occurs when storage servers suffer from load imbalance under skewed workloads. In this paper, we present KVSwitch, a centralized self-managing load balancer that leverages the power and flexibility of emerging programmable switches. The balance is achieved through dynamically predicting the hot items and creating replication strategies according to KVS loading. To overcome the challenges in realizing KVSwitch given the limitations of the switch hardware, we decompose KVSwitch's functions and carefully design them for the heterogeneous processors inside the switch. We prototype KVSwitch in a Tofino switch. Experimental results show that our solution can effectively keep the KVS servers balanced even under highly skewed workloads. Furthermore, KVSwitch only replicates 70% of hot items and consumes 9.88% of server memory rather than simply replicating all hot items to each server. Yang Shi 0008, Jiawei Fei, Mei Wen, Chunyuan Zhang |
ISCC | 3 |
| 2018 | Towards a Uniform Template-based Architecture for Accelerating 2D and 3D CNNs on FPGAabstractThree-dimensional convolutional neural networks (3D CNNs) are used efficiently in many computer vision applications. Most previous work in this area has concentrated only on designing and optimizing accelerators for 2D CNN, with few attempts made to accelerate 3D CNN on FPGA. We find accelerating 3D CNNs on FPGA to be challenge due to their high computational complexity and storage demands. More importantly, although the computation patterns of 2D and 3D CNNs are analogous, the conventional approaches adopted for accelerating 2D CNNs may be unfit for 3D CNN acceleration. In this paper, in order to accelerate 2D and 3D CNNs using a uniform framework, we propose a uniform template-based architecture that uses templates based on the Winograd algorithm to ensure fast development of 2D and 3D CNN accelerators. Furthermore, we also develop a uniform analytical model to facilitate efficient design space explorations of 2D and 3D CNN accelerators based on our architecture. Finally, we demonstrate the effectiveness of the template-based architecture by implementing accelerators for real-life 2D and 3D CNNs (VGG16 and C3D) on multiple FPGA platforms. On S2C VUS440, we achieve up to 1.13 TOPS and 1.11 TOPS under low resource utilization for VGG16 and C3D, respectively. End-to-end comparisons with CPU and GPU solutions demonstrate that our implementation of C3D achieves gains of up to 13x and 60x in performance and energy relative to a CPU solution, and a 6.4x energy efficiency gain over a GPU solution. Junzhong Shen, You Huang, Yuran Qiao, Mei Wen, Chunyuan Zhang |
FPGA | 5 |
| 2018 | Towards a Multi-array Architecture for Accelerating Large-scale Matrix Multiplication on FPGAsabstractLarge-scale floating-point matrix multiplication is a fundamental kernel in many scientific and engineering applications. Most existing work only focus on accelerating matrix multiplication on FPGA by adopting a linear systolic array. This paper towards the extension of this architecture by proposing a scalable and highly configurable multi-array architecture. In addition, we propose a work-stealing scheme to ensure the equality in the workload partition among multiple linear arrays. Furthermore, an analytical model is developed to determine the optimal design parameters. Experiments on a real-life convolutional neural network (CNN) show that we can obtain the optimal extension of the linear array architecture. Junzhong Shen, Yuran Qiao, You Huang, Mei Wen, Chunyuan Zhang |
ISCAS | 4 |
| 2017 | RVNet: A fast and high energy efficiency network packet processing system on RISC-VabstractRISC-V is a new open-source general-purpose instruction set architecture (ISA) developed by the University of California, Berkeley. It allows everyone to design their hardware circuits based on application characteristics and can be used in embedded devices, desktop computer and high-performance servers. In this paper, we use the RISC-V processor to design a fast network packet processing system. It aims to use less power and lower price to provide a faster network data processing capability for upper-layer applications in SDN and NFV. According to the results in our prototype on Field Programmable Gate Array (FPGA), our system has a comparable performance with DPDK, one of the fastest packet processing frameworks on the ×86 platform. It is worth mentioning that our system has higher (about 7.75 times) network packets processing energy efficiency than DPDK. Mei Wen, Chunyuan Zhang |
ASAP | 2 |
| 2017 | Optimizing OpenCL Implementation of Deep Convolutional Neural Network on FPGA
Yuran Qiao, Junzhong Shen, Dafei Huang, Qianming Yang, Mei Wen, Chunyuan Zhang |
NPC | 5 |
| 2017 | FPGA-accelerated deep convolutional neural networks for high throughput and energy efficiencyabstractSummary Recent breakthroughs in the deep convolutional neural networks (CNNs) have led to great improvements in the accuracy of both vision and auditory systems. Characterized by their deep structures and large numbers of parameters, deep CNNs challenge the computational performance of today. Hardware specialization in the form of field‐programmable gate array offers a promising path towards major leaps in computational performance while achieving high‐energy efficiency. In this paper, we focus on accelerating deep CNNs using the Xilinx Zynq‐zq7045 FPGA SoC. As most of the computational workload can be converted to matrix multiplications, we adopt a matrix multiplier‐based accelerator architecture. Dedicated units are designed to eliminate the conversion overhead. We also design a customized memory system according to the memory access pattern of CNNs. To make the accelerator easily usable by application developers, our accelerator supports Caffe, which is a widely used software framework of deep CNN. Different CNN models can be adopted by our accelerator, with good performance portability. The experimental results show that for a typical application of CNN, image classification, an average throughout of 77.8 GFLOPS is achieved, while the energy efficiency is 4.7× better than an Nvidia K20 GPGPU. © 2016 The Authors. Concurrency and Computation: Practice and Experience Published by John Wiley & Sons Ltd Yuran Qiao, Junzhong Shen, Qianming Yang, Mei Wen, Chunyuan Zhang |
Concurr. Comput. Pract. Exp. | 5 |
| 2017 | Applying Detection Proposals to Visual Tracking for Scale and Aspect Ratio Adaptability
Dafei Huang, Lei Luo 0002, Zhaoyun Chen, Mei Wen, Chunyuan Zhang |
Int. J. Comput. Vis. | 4 |
| 2017 | Exploiting a depth context model in visual tracking with correlation filterabstractRecently correlation filter based trackers have attracted considerable attention for their high computational efficiency. However, they cannot handle occlusion and scale variation well enough. This paper aims at preventing the tracker from failure in these two situations by integrating the depth information into a correlation filter based tracker. By using RGB-D data, we construct a depth context model to reveal the spatial correlation between the target and its surrounding regions. Furthermore, we adopt a region growing method to make our tracker robust to occlusion and scale variation. Additional optimizations such as a model updating scheme are applied to improve the performance for longer video sequences. Both qualitative and quantitative evaluations on challenging benchmark image sequences demonstrate that the proposed tracker performs favourably against state-of-the-art algorithms. Zhaoyun Chen, Lei Luo 0002, Dafei Huang, Mei Wen, Chunyuan Zhang |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2016 | Enabling Tissue-Scale Cardiac Simulations Using Heterogeneous Computing on Tianhe-2abstractWe develop a simulator for 3D tissue of the human cardiac ventricle with a physiologically realistic cell model and deploy it on the supercomputer Tianhe-2. In order to attain the full performance of the heterogeneous CPU-Xeon Phi design, we use carefully optimized codes for both devices and combine them to obtain suitable load balancing. Using a large number of nodes, we are able to perform tissue-scale simulations of the electrical activity and calcium handling in millions of cells, at a level of detail that tracks the states of trillions of ryanodine receptors. We can thus simulate arrythmogenic spiral waves and other complex arrhythmogenic patterns which arise from calcium handling deficiencies in human cardiac ventricle tissue. Due to extensive code tuning and parallelization via OpenMP, MPI, and SCIF/COI, large scale simulations of 10 heartbeats can be performed in a matter of hours. Test results indicate excellent scalability, thus paving the way for detailed whole-heart simulations in future generations of leadership class supercomputers. Johannes Langguth, Qiang Lan, Namit Gaur, Xing Cai, Mei Wen, Chunyuan Zhang |
ICPADS | 5 |
| 2015 | Enable Scale and Aspect Ratio Adaptability in Visual Tracking with Detection ProposalsabstractAmong increasingly complicated trackers in visual tracking area, recently proposed correlation filter based trackers have achieved appealing performance despite their great simplicity and superior speed. However, the filter input is a bounding box of fixed size, so they are not born with the adaptability to target’s scale and aspect ratio changes. Although scaleadaptive variants have been proposed, they are not flexible enough due to pre-defined scale sampling manners. Moreover, to the best of our knowledge, no correlation filter variant has been proposed to handle aspect ratio variation. To tackle this problem, this paper integrates the class-agnostic detection proposal method, which is widely adopted in object detection area, into a correlation filter tracker, and presents KCFDP tracker. The correlation filter part of KCFDP is based on KCF[2] with some modifications. We extend the HOG feature in KCF to a combination of HOG, intensity, and color naming by simply concatenating the three features, resulting in 42 feature channels. The model updating scheme in KCF, which is simple linear interpolation, is substituted with a more robust scheme presented in [1]. EdgeBoxes[4] is adopted to generate flexible detection proposals and enable the scale and aspect ratio adaptability of our tracker. It traverses the whole image in a sliding window manner, and scores every sampled bounding box according to the number of contours that are wholly enclosed. To accelerate EdgeBoxes and produce less unnecessary proposals, we set the minimum proposal area and aspect ratio range dynamically in sliding window sampling according to the current target size. In the tracking pipeline, KCF is firstly performed to estimate the preliminary target location ld . Within a patch zd extracted from current frame, KCF locates the target center according to the location of the maximum element in f : f(zd) = kxz d · α, (1) Dafei Huang, Lei Luo 0002, Mei Wen, Zhaoyun Chen, Chunyuan Zhang |
BMVC | 3 |
| 2015 | Unified Virtual Memory Support for Deep CNN Accelerator on SoC FPGA
Yuran Qiao, Junzhong Shen, Qianming Yang, Mei Wen |
ICA3PP (1) | 5 |
| 2015 | Fast tracking via context depth model learningabstractVisual tracking is one of the challenging tasks in computer vision. In this paper, we propose a fast and robust visual tracking algorithm which is directly extended from STC [1]. By exploring RGB-D data, we construct a context depth model to record spatial correlation between the low-level features from the target and its surrounding regions. According to the continuity and stability of target in depth image, we adopt region growing method and a model updating schema for scaling and occlusion detection. Both qualitative and quantitative evaluations on challenging benchmark image sequences demonstrate that the proposed tracker performs favorably against several state-of-the-art algorithms. Zhaoyun Chen, Lei Luo 0002, Mei Wen, Chunyuan Zhang |
ICIP | 3 |
| 2015 | Communication-hiding programming for clusters with multi-coprocessor nodesabstractSummary Future exascale systems are expected to adopt compute nodes that incorporate many accelerators. To shed some light on the upcoming software challenge, this paper investigates the particular topic of programming clusters that have multiple Xeon Phi coprocessors in each compute node. A new offload approach is considered for intra‐node communication, which combines Intel's APIs of coprocessor offload infrastructure (COI) and symmetric communication interface (SCIF) for achieving low latency. While the conventional pragma‐based offload approach allows simpler programming, the COI‐SCIF approach has three advantages in (1) lower overhead associated with launching offloaded code, (2) higher data transfer bandwidths, and (3) more advanced asynchrony between computation and data movement. The low‐level COI‐SCIF approach is also shown to have benefits over the MPI‐OpenMP counterpart, which belongs to the symmetric usage mode. Moreover, a hybird programming strategy based on COI‐SCIF is presented for joining the computational force of all CPUs and coprocessors, while realizing communication hiding. All the programming approaches are tested by a real‐world 3D application, for which the COI‐SCIF‐based approach shows a performance advantage on Tianhe‐2. Copyright © 2015 John Wiley & Sons, Ltd. Xinnan Dong, Mei Wen, Jun Chai, Xing Cai, Mandan Zhao, Chunyuan Zhang |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Improving performance portability for GPU-specific OpenCL kernels on multi-core/many-core CPUs by analysis-based transformationsabstractOpenCL is an open heterogeneous programming framework. Although OpenCL programs are functionally portable, they do not provide performance portability, so code transformation often plays an irreplaceable role. When adapting GPU-specific OpenCL kernels to run on multi-core/many-core CPUs, coarsening the thread granularity is necessary and thus has been extensively used. However, locality concerns exposed in GPU-specific OpenCL code are usually inherited without analysis, which may give side-effects on the CPU performance. Typically, the use of OpenCL’s local memory on multi-core/many-core CPUs may lead to an opposite performance effect, because local-memory arrays no longer match well with the hardware and the associated synchronizations are costly. To solve this dilemma, we actively analyze the memory access patterns using array-access descriptors derived from GPU-specific kernels, which can thus be adapted for CPUs by (1) removing all the unwanted local-memory arrays together with the obsolete barrier statements and (2) optimizing the coalesced kernel code with vectorization and locality re-exploitation. Moreover, we have developed an automated tool chain that makes this transformation of GPU-specific OpenCL kernels into a CPU-friendly form, which is accompanied with a scheduler that forms a new OpenCL runtime. Experiments show that the automated transformation can improve OpenCL kernel performance on a multi-core CPU by an average factor of 3.24. Satisfactory performance improvements are also achieved on Intel’s many-integrated-core coprocessor. The resultant performance on both architectures is better than or comparable with the corresponding OpenMP performance. Mei Wen, Dafei Huang, Changqing Xun, Dong Chen 0015 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2015 | An analytical GPU performance model for 3D stencil computations from the angle of data traffic
Huayou Su, Xing Cai, Mei Wen, Chunyuan Zhang |
J. Supercomput. | 3 |
| 2014 | Automated Transformation of GPU-Specific OpenCL Kernels Targeting Performance Portability on Multi-Core/Many-Core CPUs
Dafei Huang, Mei Wen, Changqing Xun, Dong Chen 0015, Xing Cai, Yuran Qiao, Nan Wu 0003, Chunyuan Zhang |
Euro-Par | 2 |
| 2014 | Utilizing Multiple Xeon Phi Coprocessors on One Compute Node
Xinnan Dong, Jun Chai, Mei Wen, Nan Wu 0003, Xing Cai, Chunyuan Zhang, Zhaoyun Chen |
ICA3PP (2) | 4 |
| 2013 | ACF: Networks-on-Chip Deadlock Recovery with Accurate Detection and Elastic Credit
Nan Wu 0003, Yuran Qiao, Mei Wen, Chunyuan Zhang |
APPT | 3 |
| 2013 | On the GPU-CPU Performance Portability of OpenCL for 3D Stencil ComputationsabstractAlthough OpenCL programming provides full code portability between different hardware platforms, performance portability can be far from satisfactory. In this work, we use a set of representative 3D stencil computations to study OpenCL's performance portability between GPUs and CPUs. For each stencil computation, we have devised different implementations of the computational kernel function, all being 100% code-portable between the two architectures. The most straightforward and compact implementation gives satisfactory CPU performance but performs poorly on GPUs, because such an implementation hampers effective use of the GPU hardware. By injecting code complexity into the involved loop nests, we can create kernel functions that still have full code portability but with increased performance portability. It is found that spatial data blocking and register reuse can be beneficial for performance on both GPUs and CPUs, whereas use of OpenCL's local memory (and subsequent temporal blocking) may only have positive effects on GPUs. Huayou Su, Nan Wu 0003, Mei Wen, Chunyuan Zhang, Xing Cai |
ICPADS | 3 |
| 2013 | Resource-efficient utilization of CPU/GPU-based heterogeneous supercomputers for Bayesian phylogenetic inference
Jun Chai, Huayou Su, Mei Wen, Xing Cai, Nan Wu 0003, Chunyuan Zhang |
J. Supercomput. | 3 |
| 2013 | Accelerating thread-intensive and explicit memory management programs with dynamic partial reconfiguration
Qianming Yang, Mei Wen, Nan Wu 0003, Chunyuan Zhang |
J. Supercomput. | 2 |
| 2012 | Using 1000+ GPUs and 10000+ CPUs for Sedimentary Basin SimulationsabstractIn cutting-edge CPU/GPU hybrid clusters, such as Tianhe-1A, the aggregate CPU computing capability may amount to up to 1/3 of the aggregate GPU computing capability. It thus goes without saying that the CPUs and GPUs should jointly carry out the computational work. However, to effectively and simultaneously use both the hardware components requires great care when developing the parallel implementations. The challenges include (1) finding a balanced division of the workload between the CPU and GPU sides, and (2) hiding various overheads by overlapping computations with CPU-GPU data transfers and/or MPI communications. We study these issues in the context of real-world sedimentary basin simulations. Numerical experiments show that an appropriately devised CPU-GPU hybrid implementation is able to handle a global mesh resolution of 131,072*131,072, and a double-precision rate of 62 TFlops is achieved by using 1024 GPUs and 12288 CPU cores on Tianhe-1A. Such an extreme computing capability will be of great importance for carrying out high-resolution and continental-scale stratigraphic simulations in future. Mei Wen, Huayou Su, Wenjie Wei, Nan Wu 0003, Xing Cai, Chunyuan Zhang |
CLUSTER | 1 |
| 2012 | The masala machine: accelerating thread-intensive and explicit memory management programs with dynamically reconfigurable FPGAs (abstract only)abstractA uniform FPGA-based architecture, an efficient programming model and a simple mapping method are paramount for PPGA technology to be more widely accepted. This paper presents MASALA, a dynamically reconfigurable FPGA-based accelerator specifically for parallel programs written in thread-intensive and explicit memory management (TEMM) programming models. The system uses TEMM programming model to parallelize the demanding application, including decomposing the application into separate thread blocks, decoupling compute and data load/store etc. Hardware engines are included into the MASALA by using partial dynamic reconfigure modules, each of which encapsulates Thread Process Engine implementing the thread functionality in hardware. A data dispatching scheme is also included in MASALA to enable the explicit communication among multiple memory hierarchies such as between inter-hardware engines, the host processor and hardware engines. At last, the paper illustrates a Multi-FPGA prototype system of the presented architecture: MASALA-SX. A large synthetic aperture radar (SAR) image formatting experiment shows that the MASALA architecture facilitates the construction of a TEMM program accelerator by providing it with greater performance and less power consumption than current CPU platforms, but without sacrificing programmability, flexibility and scalability. Mei Wen, Nan Wu 0003, Qianming Yang, Chunyuan Zhang |
FPGA | 1 |
| 2012 | Extending BORPH for shared memory reconfigurable computersabstractWe extend BORPH for shared memory reconfigurable computers in this paper. BORPH is an operating system designed for FPGA based reconfigurable computers. BORPH introduced the concept of hardware process in contrast to software process. With our extension, hardware processes are supported to communicate with other processes based on shared memory. In our system, the program of hardware process is not just hardware design, but the software program running on embedded processor in FPGA. Our experiment shows the overhead of shared memory segments management is acceptable. And with independent virtual memory access, bandwidth of repeated shared memory access is high. Changqing Xun, Mei Wen, Nan Wu 0003, Chunyuan Zhang, Hayden Kwok-Hay So |
FPL | 2 |
| 2012 | Parallelization Design of Irregular Algorithms of Video Processing on GPUsabstractIn this paper, we present the parallelization design consideration for irregular algorithms of video processing on GPUs. Enrich parallelism can be exploited by scheduling the processing order or making a tradeoff between performance and parallelism for irregular algorithms (such as CAVLC and deblocking filter). We implement a component-oriented CAVLC encoder and a direction-oriented deblocking filter on GPUs. The experiment results show that, compared with the implementation on CPU, the optimized parallel methods achieve high performance in term of speedup ratio from 63 to 44, relatively for deblocking filter and CAVLC. It shows that the rich parallelism is one of the most important factors to gain high performance for irregular algorithms based on GPUs. In addition, it seems that for some irregular kernels, the number of SM of GPU is more important to the performance than the computation capability. Huayou Su, Jun Chai, Mei Wen, Ju Ren 0002, Chunyuan Zhang |
ICME | 3 |
| 2012 | A Parallel H.264 Encoder with CUDA: Mapping and EvaluationabstractEfficient mapping of a real-time HD video application to graphics hardware is challenging. Developers face the challenges of choosing the right parallelism model, balancing thread's process granularity between massive computing resources on the GPU, and partitioning tasks between the CPU and GPU. The paper illustrated the mapping approaches by a case of HD H.264 encoder based on X264 reference code and then evaluating it on state-of-the-art CPU and GPUs in depth. In the paper, we first split most of the computing task into Single-Instruction Multiple-Thread (SIMT) kernels, which are then chained intocertaininput/output data stream. Then we implementeda completed H.264 encoding on the computer unified device architecture (CUDA) platform. Finally, we present methods for exploiting multi-level parallelism and memory efficiency when mapping H.264 code, which we use to increase the efficiency of the execution on GPUs. Our experimental results show that computation efficiency of GPU and then real-time encoding performance are achieved with CUDA. Nan Wu 0003, Mei Wen, Huayou Su, Ju Ren 0002, Chunyuan Zhang |
ICPADS | 2 |
| 2012 | Improving Performance of GPU Specific OpenCL Program on CPUsabstractOpenCL provides unified programming interface for various parallel computing platforms. The OpenCL framework manifests good functional portability, the programs can be run on platforms supporting OpenCL programming without any modification. However, most of the OpenCL programs are optimized for massively parallel processors, such as GPU, it's hard to achieve good performance on general multi-core processors without sophisticate modification to the GPU specific OpenCL programs. The major reason is the immense gap between CPU and GPU architecture. In this paper, we evaluate the performance portability of OpenCL programs between CPU and GPU, and analyse the reasons why GPU specific OpenCL programs are not fit for CPU. Based on the profiling, we proposed three optimization strategies for improving performance of GPU specific OpenCL programs on CPU, including increasing the granularity of task partition, optimizing the usage of memory hierarchy and block-based data accessing. In addition, we applied the proposed techniques on several benchmarks. The experimental results show that the performance of the optimized OpenCL programs achieve high performance in terms of speedup ratio from 2 to 4 on CPUs, when compared with their corresponding GPU specific ones. Qiang Lan, Changqing Xun, Mei Wen, Huayou Su, Chunyuan Zhang |
PDCAT | 3 |
| 2011 | A Multilevel Parallel Intra Coding for H.264/AVC Based on CUDAabstractIn this paper, we propose a multilevel parallel intra coding for H.264/AVC based on computed unified device architecture (CUDA). The proposed parallel algorithm improves the parallelism between 4×4 blocks within a macro block (MB) by throwing off some inappreciable prediction modes. By partitioning a frame into multi-slice, the parallelism between MBs can be exploited. In addition, a scalable parallel method for kernels is introduced to improve the performance of the proposed intra coding. Experimental results show that, more than 20 times speedup can be achieved with the assistance of GPU. Moreover, the entire encoder can meet the real-time processing requirement for HDTV. Huayou Su, Nan Wu 0003, Chunyuan Zhang, Mei Wen, Ju Ren 0002 |
ICIG | 4 |
| 2011 | High-efficient software parallel CAVLC encoder based on programmable stream processorabstractThis article presents an efficient software parallel CAVLC encoder based on programmable stream processors (Storm- SP16 and GPU). For static processor Storm SP16, a block-based 16 ways parallel CAVLC is presented with streaming processing. A component-oriented CAVLC encoder is proposed aiming at dynamic stream processor GPU. Experiments results show that, compared to the CPU version, more than 70 times of speedup can be obtained for the CAVLC based on Storm and over 50 times for GPU-based component-oriented CAVLC encoder. The throughput of the presented CAVLC encoder is more than 10 times higher over that of published software CAVLC encoders on DSP and multi-core platforms. Huayou Su, Chunyuan Zhang, Jun Chai, Mei Wen, Nan Wu 0003, Ju Ren 0002 |
ACM Multimedia | 4 |
| 2010 | Software Managed Instruction Scratchpad Memory Optimization in Stream Architecture Based on Hot Code Analysis of KernelsabstractStream processors, such as Imagine, GPGPUs, FT64 and MASA, typically uses software managed scratchpad instruction memory which improves performance and significantly reduces energy consumption. In this paper, we build a kernel-storage model to analyze the hot spot of kernels in stream programs. Based on the analysis, we define Kernel Hot Code and prove that scratchpad instruction memory should focus on the access efficiency of it. A methodology for finding Kernel Hot Code in the kernels of different structures is presented as well. In accordance with this method, we develop HOIS for Stream Architecture, which adopts a software managed scratchpad memory to store Kernel Hot Code, and uses a small hardware managed victim cache to store the Kernel Cool Code. HOIS is evaluated by measuring the performance of six applications on the MASA_S simulation platform. The results show that HOIS can achieve high efficiency in predictable applications with little performance loss. Yi He 0008, Ju Ren 0002, Mei Wen, Qianming Yang, Nan Wu 0003, Chunyuan Zhang |
DSD | 3 |
| 2009 | Cache streamization for high performance stream processorabstractDue to high bandwidth demand on memory system of stream applications, most of stream processors use software-managed streaming memory. However, this memory disadvantages ease of programming, compatibility, and supporting irregular stream access, which hinder the usage of stream processor in broader application domains. Meanwhile, hardware-managed coherent caches overcome these shortcomings of software-managed streaming memory with side-effect due to lack of supporting stream. For this problem, this paper developed a streamization cache whose performance is comparable to streaming memory but is more easy to use. The paper presents the motivation and details of our proposed design, including three stream-specific techniques for cache on data fetch policy, replacement policy and multi-client access. Moreover, a streamization cache instance is implemented in FT64, a 64-bit high performance stream processor. Based on a set of streaming application benchmark, the paper estimates the performance, power consumption and the area cost of the proposed architecture. Results show that these streamization techniques for cache are worthwhile. Nan Wu 0003, Mei Wen, Ju Ren 0002, Yi He 0008, Changqing Xun, Chunyuan Zhang |
HiPC | 2 |
| 2009 | Streaming HD H.264 encoder on programmable processorsabstractProgrammable processors have great advantage over dedicated ASIC design under intense time-to-market pressure. However, real-time encoding of high-definition (HD) H.264 video (up to 1080p) is a challenge to most existing programmable processors. On the other hand, model-based design is widely accepted in developing complex media program. Stream model, an emerging model-based programming method, shows surprising efficiency on many compute-intensive domains especially for media processing. On the basis, this paper proposes a set of streaming techniques for H.264 encoding, and then develops all of the code based on the X264 reference code. Our streaming H.264 encoder is a pure software implementation completely written in high-level language without special hardware/algorithm support. Real execution results show that our encoder achieves significant speedup over the original X264 encoder on various programmable architectures: on X86 CoreTM2 E8200 the speedup is 1.8x, on MIPS 4KEc the speedup is 3.7x, on TMS320 C6416 DSP the speedup is 5.5x, on stream processor STORM-SP16 G220 the speedup is 6.1x. Especially, on STORM processor, the streaming encoder achieves the performance of 30.6 frames per second for a 1080P HD sequence, satisfying the real-time requirement. These indicate that streaming is extremely efficient for this kind of media workload. Our work is also applicable for other media processing applications, and provides architecture insights into dedicated ASIC or FPGA HD H.264 encoders. Nan Wu 0003, Mei Wen, Ju Ren 0002, Huayou Su, Changqing Xun, Chunyuan Zhang |
ACM Multimedia | 2 |
| 2008 | Load scheduling: Reducing pressure on distributed register files for freeabstractIn this paper we describe load scheduling, a novel method that balances load among register files by residual resources. Load scheduling can reduce register pressure for clustered VLIW processors with distributed register files while not increasing VLIW scheduling length. We have implemented load scheduling in compiler for Imagine and FT64 stream processors. The result shows that the proposed technique effectively reduces the number of variables spilled to memory, and can even eliminate it. The algorithm presented in this paper is extremely efficient in embedded processor with limited register resource because it can improve registers utilization instead of increasing the requirement for the number of registers. Mei Wen, Nan Wu 0003, Maolin Guan, Chunyuan Zhang |
ASP-DAC | 1 |
| 2007 | FT64: Scientific Computing with Streams
Mei Wen, Nan Wu 0003, Chunyuan Zhang, Qianming Yang, Changqing Xun |
HiPC | 1 |
| 2005 | Multiple-Morphs Adaptive Stream Architecture
Mei Wen, Nan Wu 0003, Chunyuan Zhang |
J. Comput. Sci. Technol. | 1 |
| 2004 | A Parallel Reed-Solomon Decoder on the Imagine Stream Processor
Mei Wen, Chunyuan Zhang, Nan Wu 0003, Li Li 0005 |
ISPA | 1 |