EDBT 2026 Demo / reviewers in the wild / expert
Chunyuan Zhang
dblp:03/2728
· DBLP profile ↗
75ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Computer networks · 4Software engineering, systems software and programming languages · 3 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NeuroPDE: A Neuromorphic PDE Solver Based on Spintronic and Ferroelectric DevicesabstractIn recent years, new methods for solving partial differential equations (PDEs) such as Monte Carlo random walk methods have gained considerable attention. However, due to the lack of hardware-intrinsic randomness in the conventional von Neumann architecture, the performance of PDE solvers is limited. In this paper, we introduce NeuroPDE, a hardware design for neuromorphic PDE solvers that utilizes emerging spintronic and ferroelectric devices. NeuroPDE incorporates spin neurons that are capable of probabilistic transmission to emulate random walks, along with ferroelectric synapses that store continuous weights non-volatilely. The proposed NeuroPDE achieves a squared error of less than 1e-2 compared to analytical solutions when solving diffus3.48× to 315× speedup in execution time and an energy consumption advantage of 2.7× to 29.8× over advanced CMOS-based neuromorphic chips. By leveraging the inherent physical stochasticity of emerging devices, this study paves the way for future probabilistic neuromorphic computing systems. Siqing Fu, Lizhou Wu, Chunyuan Zhang, Sheng Ma, Yuhan Tang, Jixuan Tang |
ICCAD | 4 |
| 2024 | Cell DINO: End-to-End Cell Segmentation and Tracking with TransformerabstractPrecise cell segmentation and tracking are essential in biomedical research, but current methods are often complex and inefficient. We propose Cell DINO, an extension of the Transformer-based Mask DINO, designed for cell segmentation and tracking. By introducing rotated bounding boxes, track queries, and mitosis queries, our model detects, segments, and tracks cells simultaneously. Benchmarked on the Cell Tracking Challenge, Cell DINO ranked first on DIC-C2DH-HeLa and second on Fluo-N2DH-GOWT1 datasets, demonstrating its effectiveness and simplicity. Lei Luo 0002, Chunyuan Zhang |
BIBM | 4 |
| 2024 | AFMA-Track: Adaptive Fusion of Motion and Appearance for Robust Multi-object Tracking
Chunyuan Zhang |
ICPR (16) | 3 |
| 2024 | SparGD: A Sparse GEMM Accelerator with Dynamic DataflowabstractDeep learning has become a highly popular research field, and previously deep learning algorithms ran primarily on CPUs and GPUs. However, with the rapid development of deep learning, it was discovered that existing processors could not meet the specific large-scale computing requirements of deep learning, and custom deep learning accelerators have become popular. The majority of the primary workloads in deep learning are general matrix-matrix multiplications (GEMMs), and emerging GEMMs are highly sparse and irregular. The TPU and SIGMA are typical GEMM accelerators in recent years, but the TPU does not support sparsity, and both the TPU and SIGMA have insufficient utilization rates of the Processing Element (PE). We design and implement SparGD, a sparse GEMM accelerator with dynamic dataflow. SparGD has specific PE structures, flexible distribution networks and reduction networks, and a simple dataflow switching module. When running sparse and irregular GEMMs, SparGD can maintain high PE utilization while utilizing sparsity, and can switch to the optimal dataflow according to the computing environment. For sparse, irregular GEMMs, our experimental results show that SparGD outperforms systolic arrays by 30 times and SIGMA by 3.6 times. Bo Wang 0159, Sheng Ma, Shengbai Luo, Lizhou Wu, Chunyuan Zhang |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2023 | Recursive least squares method for training and pruning convolutional neural networksabstractAbstract Convolutional neural networks (CNNs) have shown good performance in many practical applications. However, their high computational and storage requirements make them difficult to deploy on resource-constrained devices. To address this issue, in this paper, we propose a novel iterative structured pruning algorithm for CNNs based on the recursive least squares (RLS) optimization. Our algorithm combines inverse input autocorrelation matrices with weight matrices to evaluate and prune unimportant input channels or nodes in each CNN layer and performs the next pruning operation when the testing loss is tuned down to the last unpruned level. Our algorithm can be used to prune feedforward neural networks (FNNs) as well. The fast convergence speed of the RLS optimization allows our algorithm to prune CNNs and FNNs multiple times in a small number of epochs. We validate its effectiveness in pruning VGG-16 and ResNet-50 on CIFAR-10 and CIFAR-100 and pruning a three-layer FNN on MNIST. Compared with four popular pruning algorithms, our algorithm can adaptively prune CNNs according to the learning task difficulty and can effectively prune CNNs and FNNs with a small or even no reduction in accuracy. In addition, our algorithm can prune the original sample features in the input layer. Tianzong Yu, Chunyuan Zhang |
Appl. Intell. | 2 |
| 2023 | RHS-TRNG: A Resilient High-Speed True Random Number Generator Based on STT-MTJ DeviceabstractHigh-quality random numbers are very critical to many fields such as cryptography, finance, and scientific simulation, which calls for the design of reliable true random number generators (TRNGs). Limited by entropy source, throughput, reliability, and system integration, existing TRNG designs are difficult to be deployed in real computing systems to greatly accelerate target applications. This study proposes a TRNG circuit named resilient high-speed (RHS)-TRNG based on spin-transfer torque magnetic tunnel junction (STT-MTJ). RHS-TRNG generates resilient and high-speed random bit sequences exploiting the stochastic switching characteristics of STT-MTJ. By circuit/system codesign, we integrate RHS-TRNG into a reduced instruction set computer-V (RISC-V) processor as an acceleration component, which is driven by customized random number generation instructions. Our experimental results show that a single cell of RHS-TRNG has a random bit generation speed of up to 303 Mb/s, which is the highest among existing MTJ-based TRNGs. Higher throughput can be achieved by exploiting cell-level parallelism. RHS-TRNG also shows strong resilience against PVT variations thanks to our designs using bidirectional switching currents and dual generator units. In addition, our system evaluation results using gem5 simulator suggest that the system equipped with RHS-TRNG can achieve 3.4–$12\times $higher performance in speeding up option pricing programs than software implementations of random number generation. Siqing Fu, Chunyuan Zhang, Hanqing Li, Sheng Ma, Lizhou Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | BP-Im2col: Implicit Im2col Supporting AI Backpropagation on Systolic ArraysabstractState-of-the-art systolic array-based accelerators adopt the traditional im2col algorithm to accelerate the inference of convolutional layers. However, traditional im2col cannot efficiently support AI backpropagation. Backpropagation in convolutional layers involves performing transposed convolution and dilated convolution, which usually introduces plenty of zero-spaces into the feature map or kernel. The zero-space data reorganization interfere with the continuity of training and incur additional and non-negligible overhead in terms of off- and on-chip storage, access and performance. Since countermeasures for backpropagation are rarely proposed, we propose BP-im2col, a novel im2col algorithm for AI backpropagation, and implement it in RTL on a TPU-like accelerator. Experiments on TPU-like accelerator indicate that BP-im2col reduces the backpropagation runtime by 34.9% on average, and reduces the bandwidth of off-chip memory and on-chip buffers by at least 22.7% and 70.6% respectively, over a baseline accelerator adopting the traditional im2col. It further reduces the additional storage overhead in the backpropagation process by at least 74.78%. Jianchao Yang, Mei Wen, Junzhong Shen, Yasong Cao, Minjin Tang, Renyu Yang, Jiawei Fei, Chunyuan Zhang |
ICCD | 8 |
| 2021 | Automatic mapping and code optimization for OpenCL kernels on FT-matrix architecture (WIP paper)abstractFT-Matrix is a typical vector-SIMD architecture that refines the cooperation between scalar and vector units. This approach is widely used in digital signal processing, high-performance computing, and artificial intelligence, among other fields. FT-Matrix currently adopts C vector extension as the main programming model, improving the utilization efficiency of SIMD by providing explicit vector extension API. Moreover, it is difficult to efficiently transplant parallel programs (OpenCL, CUDA) adopted by users. This paper proposes an automatic mapping and code optimization method for OpenCL kernels on FT-Matrix architecture. The proposed approach solves these challenges by means of work item coalescing, slicing and rotation, and instruction-level code optimization. Preliminary results show that our method can achieve high performance and good hardware utilization for OpenCL kernels, as well as decreasing the programming difficulty on FT-Matrix. Mei Wen, Zhaoyun Chen, Yang Shi 0008, Chunyuan Zhang |
LCTES | 5 |
| 2020 | Towards Memory-Efficient Streaming Processing with Counter-Cascading Sketching on FPGAabstractObtaining item frequencies in data streams with limited space is a well-recognized and challenging problem in a wide range of applications. Sketch-based solutions have been widely used to address this challenge due to their ability to accurately record the data streams at a low memory cost. However, most sketches suffer from low memory utilization due to the adoption of a fixed counter size. Accordingly, in this work, we propose a counter-cascading scheduling algorithm to maximize the memory utilization of sketches without incurring any accuracy loss. In addition, we propose an FPGA-based system design that supports sketch parameter learning, counter-cascading record and online query. We implement our designs on Xilinx VCU118, and conduct evaluations on real-world traces, thereby demonstrating that our design can achieve higher accuracy with lower storage; the performance achieved is 10× ~ 20× better than that of state-of-the-art sketches. Minjin Tang, Mei Wen, Junzhong Shen, Chunyuan Zhang |
DAC | 5 |
| 2020 | Scalable FPGA-based Architecture for High-Performance Per-Flow Traffic MeasurementabstractPer-flow traffic measurement has emerged as a critical but challenging task in data center in recent years in the face of massive network traffic. Many approximate methods have been proposed to resolve the existing resource-accuracy trade-off in per-flow traffic measurement, one of which is the sketch-based method. However, sketches are affected by their high computational cost and low throughput; moreover, their measurement accuracy is hard to guarantee under the conditions of changing network bandwidth or flow size distribution. Recently, FPGA platforms have been widely deployed in data centers, as they demonstrate a good fit for high-speed network processing. In this work, we propose a scalable pipelined architecture for high high-throughput per-flow traffic measurement on FPGA. We adopts memory-friendly D-left hashing in our design, which guarantees high space utilization that successfully addressing the challenge of tracking high speed data stream under limit memory resource on FPGA. Comparisons with state-of-the-art sketch-based solutions show that our design outperforms state-of-the-art sketch-based methods in terms of throughput by over 80x. Junzhong Shen, Mei Wen, Minjin Tang, Chunyuan Zhang |
FPGA | 5 |
| 2020 | Towards a Deep-Pipelined Architecture for Accelerating Deep GCN on a Multi-FPGA Platform
Qixuan Cheng, Mei Wen, Junzhong Shen, Chunyuan Zhang |
ICA3PP (1) | 5 |
| 2020 | Optimized HybridSketch: More Efficient with Analysis and Algorithm
Mei Wen, Minjin Tang, Qun Huang 0001, Chunyuan Zhang |
ICA3PP (1) | 5 |
| 2020 | HybridSketch: A Memory-centric Precise Approach for Flow MeasurementabstractAs network bandwidth has rapidly developed, due to the high occupancy of memory and bandwidth required, the Sketch structure is favored by some researchers due to its limited memory usage and simple operation. But the accuracy will decrease when the Sketch system occupies less memory space. Traditional sketch algorithms and some other specially designed algorithms and structures are striving to improve accuracy. However, with the flow rate rapidly increasing, the on-chip memory will be the bottleneck of the system. Our network measurement system achieve good results focusing more on the memory usage. We proposes a hybrid method, HybridSketch, which focuses on the memory and precision of the system with mixing two measurement methods by quantitatively analyzing, modeling and allocating appropriate memory space to each method to achieve better results. Experimental results show that our method can provide 10× improvement in terms of precision, moreover, HybridSketch can provide the same level of precision with achieving 24× improvement in terms of memory size. Mei Wen, Minjin Tang, Qun Huang 0001, Chunyuan Zhang |
ICC | 5 |
| 2020 | Towards High-Efficiency Data Centers via Job-Aware Network SchedulingabstractDistributed jobs typically facing competition for multiple resources in modern data centers, especially for network. Without effective network scheduling, this competition can cause low efficiency of the data center. Previous work on network scheduling has been focused on reducing flow completion time or improving per-flow fairness. Yet, its effect on improving jobs’ performance is limited by the unawareness of relationships between communication and computation. In this paper, we focus on the problem of scheduling network resources for multiple jobs, with the specific objective to reduce the job completion time (JCT), which also makes the datacenter more efficient. With an in-depth investigation of communication and computation, we identify an opportunity for accelerating jobs in a way that occupies less bandwidth for DAG-based complicated modern jobs. Accordingly, this paper proposes JIT, a job-aware network scheduler that leverages the computational graph to accelerate jobs effectively. To cater to the goal of JIT, we first develop a mathematical model and formulate the scheduling problem as an integer linear programming (ILP) problem. We further prove that it has an equivalent linear programming (LP) problem through rigorous theoretical analysis in order to solve this ILP problem efficiently. Some reasonable simplifications are also adopted to reduce the solving time of JIT to only 1 second. The proposed JIT is simulated and compared against some state-of-the-art designs, and the simulation results demonstrate that JIT can achieve an acceleration of up to 1.55 × , which successfully improves the efficiency of the data center. Yang Shi 0008, Mei Wen, Chunyuan Zhang |
ICPP | 3 |
| 2020 | Incremental Deployment of Programmable Switches for Sketch-based Network MeasurementabstractThe emergence of programmable switches has boosted lots of research around many network aspects: mea-surements, security, quality of services. To explore the ad-vantages of programmable data planes while preserving the legacy networking systems, deploying programmable switches incrementally may be a more practical solution. In this paper, we deal with the programmable switch deploy problem for sketch-based network measurement, which has been overlooked before. We first analyze the desired properties of a good deployment for sketch-based network measurement with some examples. Based on summarized lessons, we then develop two Integer Linear Programming (ILP) models, namely TraceILP and TopoILP, to solve the deployment problem. If historical traffic traces are provided, TraceILP generates better deployment with historical information. Even if no traces are provided, TopoILP can still make a reasonable strategy according to the network topology. Evaluations on real ISP and datacenter topologies show that pro-posed models guarantee a promising measurement performance with only about 40% devices upgraded to programmable ones. Yang Shi 0008, Mei Wen, Chunyuan Zhang |
ISCC | 3 |
| 2020 | Estimation of the parameters of a weighted nuclear norm model and its application in image denoising
Hongyao Deng, Jinsong Tao, Xiuli Song, Chunyuan Zhang |
Inf. Sci. | 4 |
| 2020 | Toward an Efficient Deep Pipelined Template-Based Architecture for Accelerating the Entire 2-D and 3-D CNNs on FPGAabstract3-D convolutional neural networks (3-D CNNs) are used efficiently in many computer vision applications. Most previous work in this area has concentrated only on design and optimization of accelerators for 2-D CNNs, with few attempts having been made to accelerate 3-D CNNs on FPGA. We find the acceleration of 3-D CNNs on FPGA to be challenging due to their high computational complexity and storage demands. More importantly, although the computational patterns of 2-D and 3-D CNNs are analogous, the conventional approaches that have been adopted for acceleration of 2-D CNNs may be unfit for 3-D CNN acceleration. In this paper, in order to accelerate 2-D and 3-D CNNs using a uniform framework, we first propose a uniform template-based architecture that uses templates based on the Winograd algorithm to ensure the rapid development of 2-D and 3-D CNN accelerators. Then, with the aim of efficiently mapping all layers of 2-D/3-D CNNs onto a pipelined accelerator, techniques are developed to improve the throughput and computational efficiency of the accelerator, including layer fusion, layer clustering, and workload-balancing scheme. Finally, we demonstrate the effectiveness of the deep pipelined architecture by accelerating real-life 2-D and 3-D CNNs on the state-of-the-art FPGA platform. On VCU118, we achieve 3.7 TOPS for VGG-16, which outperforms state-of-the-art FPGA-based CNN accelerators. Comparisons with CPU and GPU solutions demonstrate that our implementation of 3-D CNN achieves gains of up to 17.8× and 64.2× in performance and energy relative to a CPU solution, and a 5.0× energy efficiency gain over a GPU solution. Junzhong Shen, You Huang, Mei Wen, Chunyuan Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Deep Learning Research and Development Platform: Characterizing and Scheduling with QoS Guarantees on GPU ClustersabstractDeep learning (DL) has been widely adopted in various domains of artificial intelligence (AI), achieving dramatic developments in industry and academia. Besides giant AI companies, numerous small and medium-sized enterprises, institutes, and universities (EIUs) have focused on the research and development (R&D) of DL. Considering the high cost of datacenters and high performance computing (HPC) systems, EIUs prefer adopting off-the-shelf GPU clusters as a DL R&D platform for multiple users and developers to process diverse DL workloads. In such scenarios, the scheduling of multiple DL tasks on a shared GPU cluster is both significant and challenging in terms of efficiently utilizing limited resources. Existing schedulers cannot predict the resource requirements of diverse DL workloads, leading to the under-utilization of computing resources and a decline in user satisfaction. This paper proposes GENIE, a QoS-aware dynamic scheduling framework for a shared GPU cluster, which achieves users' QoS guarantee and high system utilization. In accordance with an exhaustive characterization, GENIE analyzes the key factors that affect the performance of DL tasks and proposes a prediction model derived from lightweight profiling to estimate the processing rate and response latency for diverse DL workloads. Based on the prediction models, we propose a QoS-aware scheduling algorithm to identify the best placements for DL tasks and schedule them on the shared cluster. Experiments on a GPU cluster and large-scale simulations demonstrate that GENIE achieves a QoS-guarantee percentage improvement of up to 67.4 percent and a makespan reduction of up to 28.2 percent, compared to other baseline schedulers. Zhaoyun Chen, Wei Quan 0004, Mei Wen, Jianbin Fang, Jie Yu 0008, Chunyuan Zhang, Lei Luo 0002 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2019 | Scale-out Acceleration for 3D CNN-based Lung Nodule Segmentation on a Multi-FPGA SystemabstractThree-dimensional convolutional neural networks (3D CNNs) have become a promising method in lung nodule segmentation. The high computational complexity and memory requirements of 3D CNNs make it challenging to accelerate 3D CNNs on a single FPGA. In this work, we focus on accelerating the 3D CNN-based lung nodule segmentation on a multi-FPGA platform by proposing an efficient mapping scheme that takes advantage of the massive parallelism provided by the platform, as well as maximizing the computational efficiency of the accelerators. Experimental results show that our system integrating with four Xilinx VCU118 can achieve state-of-the-art performance of 14.5 TOPS, in addition with a 29.4x performance gain over CPU and 10.5x more energy efficiency over GPU. Junzhong Shen, You Huang, Mei Wen, Chunyuan Zhang |
DAC | 5 |
| 2019 | GENIE: QoS-guided Dynamic Scheduling for CNN-based Tasks on SME ClustersabstractConvolutional Neural Network (CNN) has achieved dramatic developments in emerging Machine Learning (ML) services. Compared to online ML services, offline ML services that are full of diverse CNN workloads are common in small and medium-sized enterprises (SMEs), research institutes and universities. Efficient scheduling and processing of multiple CNN-based tasks on SME clusters is both significant and challenging. Existing schedulers cannot predict the resource requirements of CNN-based tasks. In this paper, we propose GENIE, a QoS-guided dynamic scheduling framework for SME clusters that achieves users' QoS guarantee and high system utilization. Based on a prediction model derived from lightweight profiling, a QoS-guided scheduling strategy is proposed to identify the best placements for CNN-based tasks. We implement GENIE as a plugin of Tensorflow and experiment with real SME clusters and large-scale simulations. The results of the experiments demonstrate that the QoS-guided strategy outperforms other baseline schedulers by up to 67.4% and 28.2% in terms of QoS-guarantee percentage and makespan. Zhaoyun Chen, Lei Luo 0002, Haoduo Yang, Jie Yu 0008, Mei Wen, Chunyuan Zhang |
DATE | 6 |
| 2019 | Accelerating 3D CNN-based Lung Nodule Segmentation on a Multi-FPGA SystemabstractLung nodule segmentation is one of the most significant steps in many Computer Aided Detection (CAD) systems used for lung nodule identification and classification. Three-dimensional convolutional neural networks (3D CNNs) have become a promising method in lung nodule segmentation, as this method can achieve higher detection accuracy than conventional methods. It has been proven that FPGAs can provide the most energy-efficient solution for CNN acceleration. However, the high computational complexity and memory requirements of 3D CNNs make it challenging to accelerate 3D CNNs on a single FPGA, as this will further bottleneck the performance of a 3D CNN-based CAD system. Accordingly, in this work, we focus on accelerating the 3D CNN-based lung nodule segmentation on a multi-FPGA platform by proposing an efficient mapping scheme that takes advantage of the massive parallelism provided by the platform, as well as maximizing the computational efficiency of the accelerators. Experimental results show that our system is able to achieve high computational efficiency and thereby a state-of-the-art performance of 14.5 TOPS at 200 MHz. Comparisons with CPU and GPU solutions demonstrate that our system achieves a 29.4x performance gain over CPU and a 10.5x energy efficiency improvement over GPU. Junzhong Shen, You Huang, Mei Wen, Chunyuan Zhang |
FPGA | 5 |
| 2019 | An Efficient Design Flow for Accelerating Complicated-connected CNNs on a Multi-FPGA PlatformabstractConvolutional Neural Networks (CNNs) have achieved impressive performance on various computer vision tasks. To facilitate better performance, some complicated-connected CNN models (e.g., GoogLeNet and DenseNet) have recently been proposed, and have achieved state-of-the-art performance in the fields of image classification and segmentation. However, CNNs are computation- and memory-intensive. Thus, it is significant to develop hardware accelerators in order to accelerate the inference and training processes of CNNs. Due to the high-performance, reconfigurable and energy-efficient nature of Field-Programmable Gate Arrays (FPGAs), many FPGA-based accelerators have been proposed to implement CNNs and have achieved higher throughput and energy efficiency. However, the large number of parameters involved in complicated-connected CNN models have exceeded the limited hardware resources of single FPGA board, which are unable to meet the memory and computation resource demands associated with mapping entire CNN models. Accordingly, in this paper, we propose a complete design flow to accelerate the inference of complicated-connected CNNs on a multi-FPGA platform, including DAG abstraction, mapping scheme generation and design space exploration. In addition, a multi-FPGA system with flexible inter-FPGA communications is proposed to efficiently support our design flow. Experimental results on representative models illustrate that the proposed multi-FPGA system design can achieve a throughput acceleration of up to 145.2× and 2.5× compared to CPU and GPU solutions, as well as an energy efficiency improvement of up to 139.1× and 4.8× compared to multi-core CPU and GPU solutions. Junzhong Shen, Mei Wen, Chunyuan Zhang |
ICPP | 4 |
| 2019 | Towards a Uniform Architecture for the Efficient Implementation of 2D and 3D Deconvolutional Neural Networks on FPGAsabstractThree-dimensional deconvolution is widely used in many computer vision applications. However, most previous works have only focused on accelerating 2D deconvolutional neural networks (DCNNs) on FPGAs, while the acceleration of 3D DCNNs has not been studied in depth as they have higher computational complexity and sparsity than 2D DCNNs. In this paper, we focus on the acceleration of both 2D and 3D DCNNs on FPGAs by proposing efficient schemes for mapping 2D and 3D DCNNs on a uniform architecture. By implementing our design on the Xilinx VC709 platform for four real-life 2D and 3D DCNNs, we can achieve up to 3.0 TOPS with high hardware efficiency. Comparisons with CPU and GPU solutions demonstrate that we can achieve an improvement of up to 63.3 × in throughput relative to a CPU solution and an improvement of up to 8.3 × in energy efficiency compared to a GPU solution. Junzhong Shen, Mei Wen, Chunyuan Zhang |
ISCAS | 4 |
| 2019 | KVSwitch: An In-network Load Balancer for Key-Value StoresabstractToday's cloud-based online services are underpinned by distributed key-value stores (KVSs). Keys and values are distributed across back-end servers in such scale-out systems. One primary real-life performance bottleneck occurs when storage servers suffer from load imbalance under skewed workloads. In this paper, we present KVSwitch, a centralized self-managing load balancer that leverages the power and flexibility of emerging programmable switches. The balance is achieved through dynamically predicting the hot items and creating replication strategies according to KVS loading. To overcome the challenges in realizing KVSwitch given the limitations of the switch hardware, we decompose KVSwitch's functions and carefully design them for the heterogeneous processors inside the switch. We prototype KVSwitch in a Tofino switch. Experimental results show that our solution can effectively keep the KVS servers balanced even under highly skewed workloads. Furthermore, KVSwitch only replicates 70% of hot items and consumes 9.88% of server memory rather than simply replicating all hot items to each server. Yang Shi 0008, Jiawei Fei, Mei Wen, Chunyuan Zhang |
ISCC | 4 |
| 2018 | Towards a Uniform Template-based Architecture for Accelerating 2D and 3D CNNs on FPGAabstractThree-dimensional convolutional neural networks (3D CNNs) are used efficiently in many computer vision applications. Most previous work in this area has concentrated only on designing and optimizing accelerators for 2D CNN, with few attempts made to accelerate 3D CNN on FPGA. We find accelerating 3D CNNs on FPGA to be challenge due to their high computational complexity and storage demands. More importantly, although the computation patterns of 2D and 3D CNNs are analogous, the conventional approaches adopted for accelerating 2D CNNs may be unfit for 3D CNN acceleration. In this paper, in order to accelerate 2D and 3D CNNs using a uniform framework, we propose a uniform template-based architecture that uses templates based on the Winograd algorithm to ensure fast development of 2D and 3D CNN accelerators. Furthermore, we also develop a uniform analytical model to facilitate efficient design space explorations of 2D and 3D CNN accelerators based on our architecture. Finally, we demonstrate the effectiveness of the template-based architecture by implementing accelerators for real-life 2D and 3D CNNs (VGG16 and C3D) on multiple FPGA platforms. On S2C VUS440, we achieve up to 1.13 TOPS and 1.11 TOPS under low resource utilization for VGG16 and C3D, respectively. End-to-end comparisons with CPU and GPU solutions demonstrate that our implementation of C3D achieves gains of up to 13x and 60x in performance and energy relative to a CPU solution, and a 6.4x energy efficiency gain over a GPU solution. Junzhong Shen, You Huang, Yuran Qiao, Mei Wen, Chunyuan Zhang |
FPGA | 6 |
| 2018 | Towards a Multi-array Architecture for Accelerating Large-scale Matrix Multiplication on FPGAsabstractLarge-scale floating-point matrix multiplication is a fundamental kernel in many scientific and engineering applications. Most existing work only focus on accelerating matrix multiplication on FPGA by adopting a linear systolic array. This paper towards the extension of this architecture by proposing a scalable and highly configurable multi-array architecture. In addition, we propose a work-stealing scheme to ensure the equality in the workload partition among multiple linear arrays. Furthermore, an analytical model is developed to determine the optimal design parameters. Experiments on a real-life convolutional neural network (CNN) show that we can obtain the optimal extension of the linear array architecture. Junzhong Shen, Yuran Qiao, You Huang, Mei Wen, Chunyuan Zhang |
ISCAS | 5 |
| 2018 | Design of Practical Experiences to Improve Student Understanding of Efficiency and Scalability Issues in High Performance Computing: (Abstract Only)abstractWith the increasing demand of big data technology, there has been a growing interest of introducing high performance computing in computer science curriculum. One challenge in helping students understand the nature of efficiency and scalability issues in high performance computing is the lack of opportunities for them to be engaged in large-scale applications that run on supercomputer system architecture. This poster presents a collection of example projects that have been used in a parallel computing course in multiple universities in China, including National University of Defense Technology, Sun Yat-sen University and Hunan University. These projects were adopted from a wide range of scientific computing applications such as CFD, text mining of biomedical literature and so on. The large-scale computing resource for courses is supported by two National Supercomputing Centers, one in Guangzhou and the other in Changsha. The poster describes the background, objective, structure, task, practice process and outcome for each project. It also discusses the impact on student understanding all kinds of key topics and major challenges related to computational efficiency and scalability. Such projects build a positive practical environment to make students indulge in doing all kinds of interesting and helpful trials to validate their assumptions, especially when they have different perspectives or results for one problem. The poster presents our design evaluation rubric to reflect the effectiveness of our practice, as well as the statistics about the students" achievements for the last three semesters. Juan Chen 0001, Li Shen 0007, Jianping Yin, Chunyuan Zhang |
SIGCSE | 4 |
| 2017 | RVNet: A fast and high energy efficiency network packet processing system on RISC-VabstractRISC-V is a new open-source general-purpose instruction set architecture (ISA) developed by the University of California, Berkeley. It allows everyone to design their hardware circuits based on application characteristics and can be used in embedded devices, desktop computer and high-performance servers. In this paper, we use the RISC-V processor to design a fast network packet processing system. It aims to use less power and lower price to provide a faster network data processing capability for upper-layer applications in SDN and NFV. According to the results in our prototype on Field Programmable Gate Array (FPGA), our system has a comparable performance with DPDK, one of the fastest packet processing frameworks on the ×86 platform. It is worth mentioning that our system has higher (about 7.75 times) network packets processing energy efficiency than DPDK. Mei Wen, Chunyuan Zhang |
ASAP | 3 |
| 2017 | Winograd Algorithm for 3D Convolution Neural Networks
Qiang Lan, Hongjun He, Chunyuan Zhang |
ICANN (2) | 4 |
| 2017 | Optimizing OpenCL Implementation of Deep Convolutional Neural Network on FPGA
Yuran Qiao, Junzhong Shen, Dafei Huang, Qianming Yang, Mei Wen, Chunyuan Zhang |
NPC | 6 |
| 2017 | FPGA-accelerated deep convolutional neural networks for high throughput and energy efficiencyabstractSummary Recent breakthroughs in the deep convolutional neural networks (CNNs) have led to great improvements in the accuracy of both vision and auditory systems. Characterized by their deep structures and large numbers of parameters, deep CNNs challenge the computational performance of today. Hardware specialization in the form of field‐programmable gate array offers a promising path towards major leaps in computational performance while achieving high‐energy efficiency. In this paper, we focus on accelerating deep CNNs using the Xilinx Zynq‐zq7045 FPGA SoC. As most of the computational workload can be converted to matrix multiplications, we adopt a matrix multiplier‐based accelerator architecture. Dedicated units are designed to eliminate the conversion overhead. We also design a customized memory system according to the memory access pattern of CNNs. To make the accelerator easily usable by application developers, our accelerator supports Caffe, which is a widely used software framework of deep CNN. Different CNN models can be adopted by our accelerator, with good performance portability. The experimental results show that for a typical application of CNN, image classification, an average throughout of 77.8 GFLOPS is achieved, while the energy efficiency is 4.7× better than an Nvidia K20 GPGPU. © 2016 The Authors. Concurrency and Computation: Practice and Experience Published by John Wiley & Sons Ltd Yuran Qiao, Junzhong Shen, Qianming Yang, Mei Wen, Chunyuan Zhang |
Concurr. Comput. Pract. Exp. | 6 |
| 2017 | Applying Detection Proposals to Visual Tracking for Scale and Aspect Ratio Adaptability
Dafei Huang, Lei Luo 0002, Zhaoyun Chen, Mei Wen, Chunyuan Zhang |
Int. J. Comput. Vis. | 5 |
| 2017 | Exploiting a depth context model in visual tracking with correlation filterabstractRecently correlation filter based trackers have attracted considerable attention for their high computational efficiency. However, they cannot handle occlusion and scale variation well enough. This paper aims at preventing the tracker from failure in these two situations by integrating the depth information into a correlation filter based tracker. By using RGB-D data, we construct a depth context model to reveal the spatial correlation between the target and its surrounding regions. Furthermore, we adopt a region growing method to make our tracker robust to occlusion and scale variation. Additional optimizations such as a model updating scheme are applied to improve the performance for longer video sequences. Both qualitative and quantitative evaluations on challenging benchmark image sequences demonstrate that the proposed tracker performs favourably against state-of-the-art algorithms. Zhaoyun Chen, Lei Luo 0002, Dafei Huang, Mei Wen, Chunyuan Zhang |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2016 | Multikernel Recursive Least-Squares Temporal Difference Learning
Chunyuan Zhang, Qingxin Zhu, Xinzheng Niu |
ICIC (3) | 1 |
| 2016 | Enabling Tissue-Scale Cardiac Simulations Using Heterogeneous Computing on Tianhe-2abstractWe develop a simulator for 3D tissue of the human cardiac ventricle with a physiologically realistic cell model and deploy it on the supercomputer Tianhe-2. In order to attain the full performance of the heterogeneous CPU-Xeon Phi design, we use carefully optimized codes for both devices and combine them to obtain suitable load balancing. Using a large number of nodes, we are able to perform tissue-scale simulations of the electrical activity and calcium handling in millions of cells, at a level of detail that tracks the states of trillions of ryanodine receptors. We can thus simulate arrythmogenic spiral waves and other complex arrhythmogenic patterns which arise from calcium handling deficiencies in human cardiac ventricle tissue. Due to extensive code tuning and parallelization via OpenMP, MPI, and SCIF/COI, large scale simulations of 10 heartbeats can be performed in a matter of hours. Test results indicate excellent scalability, thus paving the way for detailed whole-heart simulations in future generations of leadership class supercomputers. Johannes Langguth, Qiang Lan, Namit Gaur, Xing Cai, Mei Wen, Chunyuan Zhang |
ICPADS | 6 |
| 2015 | Enable Scale and Aspect Ratio Adaptability in Visual Tracking with Detection ProposalsabstractAmong increasingly complicated trackers in visual tracking area, recently proposed correlation filter based trackers have achieved appealing performance despite their great simplicity and superior speed. However, the filter input is a bounding box of fixed size, so they are not born with the adaptability to target’s scale and aspect ratio changes. Although scaleadaptive variants have been proposed, they are not flexible enough due to pre-defined scale sampling manners. Moreover, to the best of our knowledge, no correlation filter variant has been proposed to handle aspect ratio variation. To tackle this problem, this paper integrates the class-agnostic detection proposal method, which is widely adopted in object detection area, into a correlation filter tracker, and presents KCFDP tracker. The correlation filter part of KCFDP is based on KCF[2] with some modifications. We extend the HOG feature in KCF to a combination of HOG, intensity, and color naming by simply concatenating the three features, resulting in 42 feature channels. The model updating scheme in KCF, which is simple linear interpolation, is substituted with a more robust scheme presented in [1]. EdgeBoxes[4] is adopted to generate flexible detection proposals and enable the scale and aspect ratio adaptability of our tracker. It traverses the whole image in a sliding window manner, and scores every sampled bounding box according to the number of contours that are wholly enclosed. To accelerate EdgeBoxes and produce less unnecessary proposals, we set the minimum proposal area and aspect ratio range dynamically in sliding window sampling according to the current target size. In the tracking pipeline, KCF is firstly performed to estimate the preliminary target location ld . Within a patch zd extracted from current frame, KCF locates the target center according to the location of the maximum element in f : f(zd) = kxz d · α, (1) Dafei Huang, Lei Luo 0002, Mei Wen, Zhaoyun Chen, Chunyuan Zhang |
BMVC | 5 |
| 2015 | Fast tracking via context depth model learningabstractVisual tracking is one of the challenging tasks in computer vision. In this paper, we propose a fast and robust visual tracking algorithm which is directly extended from STC [1]. By exploring RGB-D data, we construct a context depth model to record spatial correlation between the low-level features from the target and its surrounding regions. According to the continuity and stability of target in depth image, we adopt region growing method and a model updating schema for scaling and occlusion detection. Both qualitative and quantitative evaluations on challenging benchmark image sequences demonstrate that the proposed tracker performs favorably against several state-of-the-art algorithms. Zhaoyun Chen, Lei Luo 0002, Mei Wen, Chunyuan Zhang |
ICIP | 4 |
| 2015 | Communication-hiding programming for clusters with multi-coprocessor nodesabstractSummary Future exascale systems are expected to adopt compute nodes that incorporate many accelerators. To shed some light on the upcoming software challenge, this paper investigates the particular topic of programming clusters that have multiple Xeon Phi coprocessors in each compute node. A new offload approach is considered for intra‐node communication, which combines Intel's APIs of coprocessor offload infrastructure (COI) and symmetric communication interface (SCIF) for achieving low latency. While the conventional pragma‐based offload approach allows simpler programming, the COI‐SCIF approach has three advantages in (1) lower overhead associated with launching offloaded code, (2) higher data transfer bandwidths, and (3) more advanced asynchrony between computation and data movement. The low‐level COI‐SCIF approach is also shown to have benefits over the MPI‐OpenMP counterpart, which belongs to the symmetric usage mode. Moreover, a hybird programming strategy based on COI‐SCIF is presented for joining the computational force of all CPUs and coprocessors, while realizing communication hiding. All the programming approaches are tested by a real‐world 3D application, for which the COI‐SCIF‐based approach shows a performance advantage on Tianhe‐2. Copyright © 2015 John Wiley & Sons, Ltd. Xinnan Dong, Mei Wen, Jun Chai, Xing Cai, Mandan Zhao, Chunyuan Zhang |
Concurr. Comput. Pract. Exp. | 6 |
| 2015 | A Computational Model of the Short-Cut Rule for 2D Shape DecompositionabstractWe propose a new 2D shape decomposition method based on the short-cut rule. The short-cut rule originates from cognition research, and states that the human visual system prefers to partition an object into parts using the shortest possible cuts. We propose and implement a computational model for the short-cut rule and apply it to the problem of shape decomposition. The model we proposed generates a set of cut hypotheses passing through the points on the silhouette, which represent the negative minima of curvature. We then show that most part-cut hypotheses can be eliminated by analysis of local properties of each. Finally, the remaining hypotheses are evaluated in ascending length order, which guarantees that of any pair of conflicting cuts only the shortest will be accepted. We demonstrate that, compared with state-of-the-art shape decomposition methods, the proposed approach achieves decomposition results, which better correspond to human intuition as revealed in psychological experiments. Lei Luo 0002, Chunhua Shen, Xinwang Liu 0002, Chunyuan Zhang |
IEEE Trans. Image Process. | 4 |
| 2015 | An analytical GPU performance model for 3D stencil computations from the angle of data traffic
Huayou Su, Xing Cai, Mei Wen, Chunyuan Zhang |
J. Supercomput. | 4 |
| 2014 | Rethread: A Low-Cost Transient Fault Recovery Scheme for Multithreaded ProcessorsabstractTransient fault recovery is important in processor availability. However, significant silicon or performance over-heads are characteristics of existing techniques. We uncover an opportunity to reduce the overheads dramatically in modern processors that appears as a side-effect of introducing hardware multithreading to improve performance. We observe that threads are usually short code sequences with no branches and few memory side-effects, which means that the number of checkpoints is small and constant. In addition, the state structures of a thread already presented in hardware can be reused to provide check pointing. In this paper, we demonstrate this principle of using a hardware/software co-design called Rethread, which features compiler-generated code annotations and automatic recovery in hardware by restarting threads. This approach provides the ability to recover from transient faults without dedicated hardware. Moreover, results show performance degradation under both fault-free condition (less than 5%) and as a function of fault rate. Qiang Yang 0006, Raphael 'kena' Poss, Chris R. Jesshope, Chunyuan Zhang |
ARES | 5 |
| 2014 | A fault detection mechanism in a Data-flow scheduled Multithreaded processorabstractThis paper designs and implements the Redundant Multi-Threading (RMT) in a Data-flow scheduled MultiThreaded (DMT) multicore processor, called Data-flow scheduled Redundant Multi-Threading (DRMT). Meanwhile, It presents Asynchronous Output Comparison (AOC) for RMT techniques to avoid fault detection related inter-core communication and alleviate the performance and hardware overheads induced by output comparison. Results show that the performance overhead of DRMT is less than 60% even when the number of threads is four times the number of processing elements. Also the performance and hardware overheads of AOC are insignificant. Qiang Yang 0006, Raphael 'kena' Poss, Chris R. Jesshope, Chunyuan Zhang |
DATE | 5 |
| 2014 | Automated Transformation of GPU-Specific OpenCL Kernels Targeting Performance Portability on Multi-Core/Many-Core CPUs
Dafei Huang, Mei Wen, Changqing Xun, Dong Chen 0015, Xing Cai, Yuran Qiao, Nan Wu 0003, Chunyuan Zhang |
Euro-Par | 8 |
| 2014 | Utilizing Multiple Xeon Phi Coprocessors on One Compute Node
Xinnan Dong, Jun Chai, Mei Wen, Nan Wu 0003, Xing Cai, Chunyuan Zhang, Zhaoyun Chen |
ICA3PP (2) | 7 |
| 2013 | ACF: Networks-on-Chip Deadlock Recovery with Accurate Detection and Elastic Credit
Nan Wu 0003, Yuran Qiao, Mei Wen, Chunyuan Zhang |
APPT | 4 |
| 2013 | On the GPU-CPU Performance Portability of OpenCL for 3D Stencil ComputationsabstractAlthough OpenCL programming provides full code portability between different hardware platforms, performance portability can be far from satisfactory. In this work, we use a set of representative 3D stencil computations to study OpenCL's performance portability between GPUs and CPUs. For each stencil computation, we have devised different implementations of the computational kernel function, all being 100% code-portable between the two architectures. The most straightforward and compact implementation gives satisfactory CPU performance but performs poorly on GPUs, because such an implementation hampers effective use of the GPU hardware. By injecting code complexity into the involved loop nests, we can create kernel functions that still have full code portability but with increased performance portability. It is found that spatial data blocking and register reuse can be beneficial for performance on both GPUs and CPUs, whereas use of OpenCL's local memory (and subsequent temporal blocking) may only have positive effects on GPUs. Huayou Su, Nan Wu 0003, Mei Wen, Chunyuan Zhang, Xing Cai |
ICPADS | 4 |
| 2013 | Efficient fine-grained shared buffer management for multiple OpenCL devicesabstractOpenCL programming provides full code portability between different hardware platforms, and can serve as a good programming candidate for heterogeneous systems, which typically consist of a host processor and several accelerators. However, to make full use of the computing capacity of such a system, programmers are requested to manage diverse OpenCL-enabled devices explicitly, including distributing the workload between different devices and managing data transfer between multiple devices. All these tedious jobs pose a huge challenge for programmers. In this paper, a distributed shared OpenCL memory (DSOM) is presented, which relieves users of having to manage data transfer explicitly, by supporting shared buffers across devices. DSOM allocates shared buffers in the system memory and treats the on-device memory as a software managed virtual cache buffer. To support fine-grained shared buffer management, we designed a kernel parser in DSOM for buffer access range analysis. A basic modified, shared, invalid cache coherency is implemented for DSOM to maintain coherency for cache buffers. In addition, we propose a novel strategy to minimize communication cost between devices by launching each necessary data transfer as early as possible. This strategy enables overlap of data transfer with kernel execution. Our experimental results show that the applicability of our method for buffer access range analysis is good, and the efficiency of DSOM is high. Changqing Xun, Dong Chen 0015, Qiang Lan, Chunyuan Zhang |
J. Zhejiang Univ. Sci. C | 4 |
| 2013 | Resource-efficient utilization of CPU/GPU-based heterogeneous supercomputers for Bayesian phylogenetic inference
Jun Chai, Huayou Su, Mei Wen, Xing Cai, Nan Wu 0003, Chunyuan Zhang |
J. Supercomput. | 6 |
| 2013 | Accelerating thread-intensive and explicit memory management programs with dynamic partial reconfiguration
Qianming Yang, Mei Wen, Nan Wu 0003, Chunyuan Zhang |
J. Supercomput. | 4 |
| 2013 | Shape Similarity Analysis by Self-Tuning Locally Constrained Mixed-DiffusionabstractSimilarity analysis is a powerful tool for shape matching/retrieval and other computer vision tasks. In the literature, various shape (dis)similarity measures have been introduced. Different measures specialize on different aspects of the data. In this paper, we consider the problem of improving retrieval accuracy by systematically fusing several different measures. To this end, we propose the locally constrained mixed-diffusion method, which partly fuses the given measures into one and propagates on the resulted locally dense data space. Furthermore, we advocate the use of self-adaptive neighborhoods to automatically determine the appropriate size of the neighborhoods in the diffusion process, with which the retrieval performance is comparable to the best manually tuned kNNs. The superiority of our approach is empirically demonstrated on both shape and image datasets. Our approach achieves a score of 100% in the bull's eye test on the MPEG-7 shape dataset, which is the best reported result to date. Lei Luo 0002, Chunhua Shen, Chunyuan Zhang, Anton van den Hengel |
IEEE Trans. Multim. | 3 |
| 2012 | Using 1000+ GPUs and 10000+ CPUs for Sedimentary Basin SimulationsabstractIn cutting-edge CPU/GPU hybrid clusters, such as Tianhe-1A, the aggregate CPU computing capability may amount to up to 1/3 of the aggregate GPU computing capability. It thus goes without saying that the CPUs and GPUs should jointly carry out the computational work. However, to effectively and simultaneously use both the hardware components requires great care when developing the parallel implementations. The challenges include (1) finding a balanced division of the workload between the CPU and GPU sides, and (2) hiding various overheads by overlapping computations with CPU-GPU data transfers and/or MPI communications. We study these issues in the context of real-world sedimentary basin simulations. Numerical experiments show that an appropriately devised CPU-GPU hybrid implementation is able to handle a global mesh resolution of 131,072*131,072, and a double-precision rate of 62 TFlops is achieved by using 1024 GPUs and 12288 CPU cores on Tianhe-1A. Such an extreme computing capability will be of great importance for carrying out high-resolution and continental-scale stratigraphic simulations in future. Mei Wen, Huayou Su, Wenjie Wei, Nan Wu 0003, Xing Cai, Chunyuan Zhang |
CLUSTER | 6 |
| 2012 | The masala machine: accelerating thread-intensive and explicit memory management programs with dynamically reconfigurable FPGAs (abstract only)abstractA uniform FPGA-based architecture, an efficient programming model and a simple mapping method are paramount for PPGA technology to be more widely accepted. This paper presents MASALA, a dynamically reconfigurable FPGA-based accelerator specifically for parallel programs written in thread-intensive and explicit memory management (TEMM) programming models. The system uses TEMM programming model to parallelize the demanding application, including decomposing the application into separate thread blocks, decoupling compute and data load/store etc. Hardware engines are included into the MASALA by using partial dynamic reconfigure modules, each of which encapsulates Thread Process Engine implementing the thread functionality in hardware. A data dispatching scheme is also included in MASALA to enable the explicit communication among multiple memory hierarchies such as between inter-hardware engines, the host processor and hardware engines. At last, the paper illustrates a Multi-FPGA prototype system of the presented architecture: MASALA-SX. A large synthetic aperture radar (SAR) image formatting experiment shows that the MASALA architecture facilitates the construction of a TEMM program accelerator by providing it with greater performance and less power consumption than current CPU platforms, but without sacrificing programmability, flexibility and scalability. Mei Wen, Nan Wu 0003, Qianming Yang, Chunyuan Zhang |
FPGA | 4 |
| 2012 | Extending BORPH for shared memory reconfigurable computersabstractWe extend BORPH for shared memory reconfigurable computers in this paper. BORPH is an operating system designed for FPGA based reconfigurable computers. BORPH introduced the concept of hardware process in contrast to software process. With our extension, hardware processes are supported to communicate with other processes based on shared memory. In our system, the program of hardware process is not just hardware design, but the software program running on embedded processor in FPGA. Our experiment shows the overhead of shared memory segments management is acceptable. And with independent virtual memory access, bandwidth of repeated shared memory access is high. Changqing Xun, Mei Wen, Nan Wu 0003, Chunyuan Zhang, Hayden Kwok-Hay So |
FPL | 4 |
| 2012 | Parallelization Design of Irregular Algorithms of Video Processing on GPUsabstractIn this paper, we present the parallelization design consideration for irregular algorithms of video processing on GPUs. Enrich parallelism can be exploited by scheduling the processing order or making a tradeoff between performance and parallelism for irregular algorithms (such as CAVLC and deblocking filter). We implement a component-oriented CAVLC encoder and a direction-oriented deblocking filter on GPUs. The experiment results show that, compared with the implementation on CPU, the optimized parallel methods achieve high performance in term of speedup ratio from 63 to 44, relatively for deblocking filter and CAVLC. It shows that the rich parallelism is one of the most important factors to gain high performance for irregular algorithms based on GPUs. In addition, it seems that for some irregular kernels, the number of SM of GPU is more important to the performance than the computation capability. Huayou Su, Jun Chai, Mei Wen, Ju Ren 0002, Chunyuan Zhang |
ICME | 5 |
| 2012 | A Parallel H.264 Encoder with CUDA: Mapping and EvaluationabstractEfficient mapping of a real-time HD video application to graphics hardware is challenging. Developers face the challenges of choosing the right parallelism model, balancing thread's process granularity between massive computing resources on the GPU, and partitioning tasks between the CPU and GPU. The paper illustrated the mapping approaches by a case of HD H.264 encoder based on X264 reference code and then evaluating it on state-of-the-art CPU and GPUs in depth. In the paper, we first split most of the computing task into Single-Instruction Multiple-Thread (SIMT) kernels, which are then chained intocertaininput/output data stream. Then we implementeda completed H.264 encoding on the computer unified device architecture (CUDA) platform. Finally, we present methods for exploiting multi-level parallelism and memory efficiency when mapping H.264 code, which we use to increase the efficiency of the execution on GPUs. Our experimental results show that computation efficiency of GPU and then real-time encoding performance are achieved with CUDA. Nan Wu 0003, Mei Wen, Huayou Su, Ju Ren 0002, Chunyuan Zhang |
ICPADS | 5 |
| 2012 | Improving Performance of GPU Specific OpenCL Program on CPUsabstractOpenCL provides unified programming interface for various parallel computing platforms. The OpenCL framework manifests good functional portability, the programs can be run on platforms supporting OpenCL programming without any modification. However, most of the OpenCL programs are optimized for massively parallel processors, such as GPU, it's hard to achieve good performance on general multi-core processors without sophisticate modification to the GPU specific OpenCL programs. The major reason is the immense gap between CPU and GPU architecture. In this paper, we evaluate the performance portability of OpenCL programs between CPU and GPU, and analyse the reasons why GPU specific OpenCL programs are not fit for CPU. Based on the profiling, we proposed three optimization strategies for improving performance of GPU specific OpenCL programs on CPU, including increasing the granularity of task partition, optimizing the usage of memory hierarchy and block-based data accessing. In addition, we applied the proposed techniques on several benchmarks. The experimental results show that the performance of the optimized OpenCL programs achieve high performance in terms of speedup ratio from 2 to 4 on CPUs, when compared with their corresponding GPU specific ones. Qiang Lan, Changqing Xun, Mei Wen, Huayou Su, Chunyuan Zhang |
PDCAT | 6 |
| 2011 | A Multilevel Parallel Intra Coding for H.264/AVC Based on CUDAabstractIn this paper, we propose a multilevel parallel intra coding for H.264/AVC based on computed unified device architecture (CUDA). The proposed parallel algorithm improves the parallelism between 4×4 blocks within a macro block (MB) by throwing off some inappreciable prediction modes. By partitioning a frame into multi-slice, the parallelism between MBs can be exploited. In addition, a scalable parallel method for kernels is introduced to improve the performance of the proposed intra coding. Experimental results show that, more than 20 times speedup can be achieved with the assistance of GPU. Moreover, the entire encoder can meet the real-time processing requirement for HDTV. Huayou Su, Nan Wu 0003, Chunyuan Zhang, Mei Wen, Ju Ren 0002 |
ICIG | 3 |
| 2011 | High-efficient software parallel CAVLC encoder based on programmable stream processorabstractThis article presents an efficient software parallel CAVLC encoder based on programmable stream processors (Storm- SP16 and GPU). For static processor Storm SP16, a block-based 16 ways parallel CAVLC is presented with streaming processing. A component-oriented CAVLC encoder is proposed aiming at dynamic stream processor GPU. Experiments results show that, compared to the CPU version, more than 70 times of speedup can be obtained for the CAVLC based on Storm and over 50 times for GPU-based component-oriented CAVLC encoder. The throughput of the presented CAVLC encoder is more than 10 times higher over that of published software CAVLC encoders on DSP and multi-core platforms. Huayou Su, Chunyuan Zhang, Jun Chai, Mei Wen, Nan Wu 0003, Ju Ren 0002 |
ACM Multimedia | 2 |
| 2010 | Software Managed Instruction Scratchpad Memory Optimization in Stream Architecture Based on Hot Code Analysis of KernelsabstractStream processors, such as Imagine, GPGPUs, FT64 and MASA, typically uses software managed scratchpad instruction memory which improves performance and significantly reduces energy consumption. In this paper, we build a kernel-storage model to analyze the hot spot of kernels in stream programs. Based on the analysis, we define Kernel Hot Code and prove that scratchpad instruction memory should focus on the access efficiency of it. A methodology for finding Kernel Hot Code in the kernels of different structures is presented as well. In accordance with this method, we develop HOIS for Stream Architecture, which adopts a software managed scratchpad memory to store Kernel Hot Code, and uses a small hardware managed victim cache to store the Kernel Cool Code. HOIS is evaluated by measuring the performance of six applications on the MASA_S simulation platform. The results show that HOIS can achieve high efficiency in predictable applications with little performance loss. Yi He 0008, Ju Ren 0002, Mei Wen, Qianming Yang, Nan Wu 0003, Chunyuan Zhang |
DSD | 6 |
| 2009 | Joint Channel Width Adaptation, Topology Control, and Routing for Multi-Radio Multi-Channel Wireless Mesh NetworksabstractCapacity limitation is one of the fundamental issues in wireless mesh networks. The aggregate capacity can be increased by equipping each mesh router with multiple radios tuned into distinct frequency channels. However, most past research efforts that attempt to exploit multiple channels assume orthogonal channels of fixed pre-determined width, which prohibits the further effective use of spectrum resource. In this paper, we use mathematical programming to address how to optimally adapt channel width to make full use of the spectrum resource. We formulate the channel width adaptation, topology control and routing as a joint mixed 0-1 integer linear optimization problem. Simulation results show that our algorithm can significantly improve spectrum use efficiency and network performance. Chunyuan Zhang |
CCNC | 2 |
| 2009 | Cache streamization for high performance stream processorabstractDue to high bandwidth demand on memory system of stream applications, most of stream processors use software-managed streaming memory. However, this memory disadvantages ease of programming, compatibility, and supporting irregular stream access, which hinder the usage of stream processor in broader application domains. Meanwhile, hardware-managed coherent caches overcome these shortcomings of software-managed streaming memory with side-effect due to lack of supporting stream. For this problem, this paper developed a streamization cache whose performance is comparable to streaming memory but is more easy to use. The paper presents the motivation and details of our proposed design, including three stream-specific techniques for cache on data fetch policy, replacement policy and multi-client access. Moreover, a streamization cache instance is implemented in FT64, a 64-bit high performance stream processor. Based on a set of streaming application benchmark, the paper estimates the performance, power consumption and the area cost of the proposed architecture. Results show that these streamization techniques for cache are worthwhile. Nan Wu 0003, Mei Wen, Ju Ren 0002, Yi He 0008, Changqing Xun, Chunyuan Zhang |
HiPC | 7 |
| 2009 | Streaming HD H.264 encoder on programmable processorsabstractProgrammable processors have great advantage over dedicated ASIC design under intense time-to-market pressure. However, real-time encoding of high-definition (HD) H.264 video (up to 1080p) is a challenge to most existing programmable processors. On the other hand, model-based design is widely accepted in developing complex media program. Stream model, an emerging model-based programming method, shows surprising efficiency on many compute-intensive domains especially for media processing. On the basis, this paper proposes a set of streaming techniques for H.264 encoding, and then develops all of the code based on the X264 reference code. Our streaming H.264 encoder is a pure software implementation completely written in high-level language without special hardware/algorithm support. Real execution results show that our encoder achieves significant speedup over the original X264 encoder on various programmable architectures: on X86 CoreTM2 E8200 the speedup is 1.8x, on MIPS 4KEc the speedup is 3.7x, on TMS320 C6416 DSP the speedup is 5.5x, on stream processor STORM-SP16 G220 the speedup is 6.1x. Especially, on STORM processor, the streaming encoder achieves the performance of 30.6 frames per second for a 1080P HD sequence, satisfying the real-time requirement. These indicate that streaming is extremely efficient for this kind of media workload. Our work is also applicable for other media processing applications, and provides architecture insights into dedicated ASIC or FPGA HD H.264 encoders. Nan Wu 0003, Mei Wen, Ju Ren 0002, Huayou Su, Changqing Xun, Chunyuan Zhang |
ACM Multimedia | 7 |
| 2009 | Interference-aware Broadcast Routing and Channel Assignment for Multi-Radio Wireless Mesh NetworksabstractAn important problem in multi-radio multi-channel wireless mesh networks is how to perform efficient network-wide broadcasting. While route discovery or energy efficiency is the major concern for broadcasting in mobile ad-hoc networks or sensor networks, in WMNs, more attention should be paid to the high-throughput schemes. In this paper, we propose an interference-aware broadcast routing and channel assignment scheme for IEEE802.11-based multi-radio multi-channel mesh networks. We first formulate the problem as a mixed integer linear programming which jointly consider the broadcast routing and channel assignment. And then we propose heuristic suboptimal algorithms. Our schemes enable nodes to operate with minimum interference while the channel diversity can be fully exploited. Simulation results show that out schemes can significantly improve the broadcast performance compared with previous work. Li Li 0005, Chunyuan Zhang |
VTC Fall | 3 |
| 2009 | QoS-aware on-demand channel width adaptation protocols for multi-radio ad-hoc networksabstractHow to efficiently use the spectrum resource to provide quality-of-service (QoS) support is a challenging task for wireless communication systems. In this paper, a resource reservation-based on-demand spectrum assignment and routing protocol is proposed to provide QoS support for IEEE 802.11-based multiradio multi-channel ad-hoc networks. Our work distinguishes from prior ones in that we don't treat the spectrum as the set of discrete orthogonal channels but the continuous resource, i.e. we use channel width adaptation when allocating spectrum resource. We develop distributed resource assignment protocols that utilize the AODV routing to perform admission control and resource reservation. Simulation results show that our protocol can efficiently utilize the network resources to provide QoS support. Li Li 0005, Chunyuan Zhang |
WCNC | 2 |
| 2008 | Load scheduling: Reducing pressure on distributed register files for freeabstractIn this paper we describe load scheduling, a novel method that balances load among register files by residual resources. Load scheduling can reduce register pressure for clustered VLIW processors with distributed register files while not increasing VLIW scheduling length. We have implemented load scheduling in compiler for Imagine and FT64 stream processors. The result shows that the proposed technique effectively reduces the number of variables spilled to memory, and can even eliminate it. The algorithm presented in this paper is extremely efficient in embedded processor with limited register resource because it can improve registers utilization instead of increasing the requirement for the number of registers. Mei Wen, Nan Wu 0003, Maolin Guan, Chunyuan Zhang |
ASP-DAC | 4 |
| 2008 | Transform coding on programmable stream processors
Chunyuan Zhang, Li Li 0005, Ju Ren 0002 |
J. Supercomput. | 2 |
| 2007 | The Design on SEU-Tolerant Information Processing System of the On-Board-Computer
Chunyuan Zhang, Dong Liu 0022, Sheng-xin Weng |
APPT | 2 |
| 2007 | FT64: Scientific Computing with Streams
Mei Wen, Nan Wu 0003, Chunyuan Zhang, Qianming Yang, Changqing Xun |
HiPC | 3 |
| 2007 | Efficient Broadcasting in Multi-radio Multi-channel and Multi-hop Wireless Networks Based on Self-pruning
Li Li 0005, Chunyuan Zhang |
HPCC | 3 |
| 2007 | Quantification of Cut Sequence Set for Fault Tree Analysis
Dong Liu 0022, Chunyuan Zhang, Weiyan Xing, Rui Li 0051 |
HPCC | 2 |
| 2006 | A Streaming Implementation of Transform and Quantization in H.264
Chunyuan Zhang, Li Li 0005, Ming Pang |
HPCC | 2 |
| 2006 | Prediction-Table Based Fault-Tolerant Real-Time Scheduling AlgorithmabstractIn order to predict accurately whether primary versions of real-time tasks is executable in software fault-tolerant module, a new algorithm, PTBA, prediction-table based algorithm, is presented. PTBA uses prediction-table to predict whether a host primary can meet its pre-deadline. Prediction-table contains the pre-assignment information of tasks between the current time and the alternates' notification time. If the prediction result shows that host primary has not enough time to execute, it will be aborted. Otherwise, prediction-table is referenced to schedule tasks with low overhead. The novelty of PTBA is that it schedules primaries according to their corresponding alternates' notification time and has no extra scheduling overhead in prediction-table mode. Simulation results show that PTBA allows more execution time for primaries and wastes less processor time than the well-known similar algorithms. PTBA is appropriate to the situation where the periods of tasks are short and software fault probability is low Dong Liu 0022, Chunyuan Zhang, Rui Li 0051 |
PDCAT | 2 |
| 2005 | Multiple-Morphs Adaptive Stream Architecture
Mei Wen, Nan Wu 0003, Chunyuan Zhang |
J. Comput. Sci. Technol. | 4 |
| 2004 | A Case of SCMP with TLS
Jianzhuang Lu, Chunyuan Zhang |
ISPA | 2 |
| 2004 | A Parallel Reed-Solomon Decoder on the Imagine Stream Processor
Mei Wen, Chunyuan Zhang, Nan Wu 0003, Li Li 0005 |
ISPA | 2 |