VLDB 2026 Research / reviewers in the wild / expert
Weifeng Zhang 0003
dblp:19/949-3
· DBLP profile ↗
22ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-4529-1679ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-DesignabstractEfficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing dynamic schedulers struggle with large-model scheduling due to the mismatch between static parallelism (SP)-aware scheduling and AP-based execution, leading to cluster inefficiencies such as degraded throughput and prolonged job queuing. This paper presents Arena, a large-model training system that co-designs dynamic scheduling and adaptive parallelism to achieve high cluster efficiency. To reduce scheduling costs while improving decision quality, Arena designs low-cost, disaggregated profiling and AP-tailored, load-aware performance estimation, while unifying them by sharding the joint scheduling-parallelism optimization space via a grid abstraction. Building on this, Arena dynamically schedules profiled jobs in elasticity and heterogeneity dimensions, and executes them using efficient AP with pruned search space. Evaluated on heterogeneous testbeds and production workloads, Arena reduces job completion time by up to 49.3% and improves cluster throughput by up to 1.6X. Chunyu Xue, Weihao Cui, Quan Chen 0002, Chen Chen 0067, Han Zhao 0005, Shulai Zhang, Linmei Wang, Limin Xiao 0001, Weifeng Zhang 0003, Jing Yang 0017, Bingsheng He, Minyi Guo |
EuroSys | 10 |
| 2026 | NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
Chen Nie, Limin Xiao 0001, Weifeng Zhang 0003, Zhezhi He |
ISCA | 8 |
| 2026 | EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu 0003, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Zhan Wang 0003, Yuchao Zhang 0004, Yang Liu 0038, Xiangrui Yang 0002, Xiaohe Hu, Limin Xiao 0001, Weifeng Zhang 0003, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu |
SIGCOMM | 25 |
| 2026 | APU: Accelerate Point Cloud Neural Networks via Unified Processing-in-SRAM ArchitectureabstractRecent advances in deep learning have expanded point cloud applications by point-based neural networks (PNNs). However, the escalating complexity and computational demands of PNNs overwhelm conventional computers. Specialized PNN accelerators have emerged, significantly outperforming modern CPUs and GPUs. Nevertheless, existing designs remain inefficient when handling performance-critical mapping kernels of PNNs, involving diverse arithmetic functions (e.g., add, multiply, sort) across separate hardware modules. This fragmentation restricts hardware sharing and data locality, leading to area overhead, redundant data movements, and under-utilization. Therefore, a unified and efficient micro-architecture for mapping kernels is needed to enhance performance and reduce data transfers. This paper presents APU, an efficient processing-in-memory (PIM) architecture for PNN acceleration. We introduce the first unified SRAM-PIM micro-architecture that supports all mapping kernels in mainstream PNNs. Data movement is reduced through extensive on-chip memory and maximized data locality viain-situcomputing approach. At the algorithmic level, we introduce mask grouping and aggregation to eliminate costly sorting operations, enabled by hardware support for in-memory vector max-search. This refined strategy reduces computational overhead and data transfers while improving inference accuracy.We further enhance performance by exploiting parallelism across PNN operations and applying mixed-precision quantization. Evaluated on real-world PNN workloads, APU outperforms the state-of-the-art accelerator by 2.54× in speedup and 4.54× in energy saving. Chen Nie, Kang You, Yu Feng 0007, Limin Xiao 0002, Weifeng Zhang 0003, Zhezhi He |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2025 | PICK: An SRAM-based Processing-in-Memory Accelerator for K-Nearest-Neighbor Search in Point CloudsabstractK-nearest neighbor (kNN) search is a fundamental operation in various point cloud applications, such as autonomous driving. However, the heavy computational intensity and memory demands of kNN search pose significant challenges for efficient implementation, especially in resource-constrained scenarios. To address these challenges, we propose PICK, a processing-in-memory (PIM) architecture designed to accelerate kNN search in point cloud applications. PICK leverages bit-serial-based PIM (BS-PIM) and customized circuits to efficiently handle key operations of kNN search: distance calculation and top-k selection. The run-time off-chip access is eliminated thanks to the large on-chip memory. For distance calculation, we introduce a bit-width clipping technique to reduce the latency of bit-serial execution with negligible accuracy degradation, providing flexible trade-offs between performance and precision. Besides, we propose a filtering-and-selection strategy that realizes approximately constant time complexity for arbitrary values of k. Furthermore, a two-stage pipeline is implemented to parallelize distance calculation and top-k search, effectively hiding latency and improving throughput. According to our experiments, PICK achieves $4.17 \times$ speedup and a $4.42 \times$ energy saving over the state-of-the-art design. Chen Nie, Liming Xiao, Weifeng Zhang 0003, Zhezhi He |
DAC | 4 |
| 2025 | Boosting Deep Vector Quantization with Progressive Distribution TransformationabstractVector quantization (VQ) is an effective technique for data compression and is widely used in approximate nearest neighbor search (ANNs). Based on the multi-codebook quantization (MCQ), most of existing VQ methods focus on the refinement of codebook structure to enhance performance. However, the uneven distribution of raw vectors poses challenges for effective codebook learning. Previous attempts to learn distribution transformations for raw vectors have struggled with minimal accuracy improvements and limited scalability due to poor convergence. To address these problems, we propose a Progressive Distribution Transformation (PDT) algorithm based on denoising autoencoder for VQ optimization. Our approach employs multi-head distribution transformation (MHDT) blocks to progressively learn transformation offsets, allowing raw vectors to be transformed into a quantization-friendly latent space step-by-step. Notably, the proposed PDT algorithm can be seamlessly integrated with diverse existing deep MCQ methods, facilitating end-to-end training and efficient inference. Furthermore, it also enables easy integration with VQ-based ANNs methods. Experimental results show that our method outperforms state-of-the-art MCQ methods by 20.6% and significantly boosts existing deep MCQ techniques by 58.5% in search accuracy across various similarity search datasets, while also exhibiting improved scalability and efficiency. The code has been available at https://github.com/Lenovo-ICI/PDT-VQ. Weikang Wang 0002, Weifeng Zhang 0003 |
KDD (2) | 4 |
| 2024 | VSPIM: SRAM Processing-in-Memory DNN Acceleration via Vector-Scalar OperationsabstractProcessing-in-Memory (PIM) has been widely explored for accelerating data-intensive machine learning computation that mainly consists of general-matrix-multiplication (GEMM), by mitigating the burden of data movements and exploiting the ultra-high memory parallelism. The two mainstreams of PIM, the analog- and digital-type, have both been exploited in accelerating machine learning workloads by numerous outstanding prior works. Currently, the digital-PIM is increasingly favored due to the broader computing support and the avoidance of errors caused by intrinsic non-idealities, e.g., process variation. Nevertheless, it still lacks further optimization considering the characteristics of the GEMM computation, including better efficient data layout and scheduling, and the ability to handle the sparsity of activations at the bit-level. To boost the performance and efficiency of digital SRAM PIM, we propose the architecture called VSPIM that performs the computation in a bit-serial fashion, with unique support of vector-scalar computing pattern. The novelties of the VSPIM can be concluded as follows: 1) support bit-serial based scalar-vector computing via ingenious parallel bit-broadcasting; 2) refine the GEMM mapping strategy and computing pattern to enhance performance and efficiency; 3) powered by the introduced scalar-vector operation, the bit-sparsity of activation is leveraged to halt unnecessary computation to maximize efficiency and throughput. Our comprehensive evaluation shows that, compared to the state-of-the-art SRAM-based digital-PIM design (Neural Cache), VSPIM can significantly boost the performance and energy efficiency by up to$8.87\times$and$4.81\times$respectively, with negligible area overhead, upon multiple representative neural networks. Chen Nie, Chenyu Tang, Jie Lin 0004, Chenyang Lv, Ting Cao 0007, Weifeng Zhang 0003, Li Jiang 0002, Xiaoyao Liang, Weikang Qian, Yanan Sun 0003, Zhezhi He |
IEEE Trans. Computers | 7 |
| 2024 | A Classical Architecture for Digital Quantum ComputersabstractScaling bottlenecks the making of digital quantum computers, posing challenges from both the quantum and the classical components. We present a classical architecture to cope with a comprehensive list of the latter challenges all at once , and implement it fully in an end-to-end system by integrating a multi-core RISC-V CPU with our in-house control electronics. Our architecture enables scalable, high-precision control of large quantum processors and accommodates evolving requirements of quantum hardware. A central feature is a microarchitecture executing quantum operations in parallel on arbitrary predefined qubit groups. Another key feature is a reconfigurable quantum instruction set that supports easy qubit re-grouping and instructions extensions. As a demonstration, we implement the widely-studied surface code quantum computing workflow, which is instructive for being demanding on both the controllers and the integrated classical computation. Our design, for the first time, reduces instruction issuing and transmission costs to constants, which do not scale with the number of qubits, without adding any overheads in decoding or dispatching. Our system uses a dedicated general-purpose CPU for both qubit control and classical computation, including syndrome decoding. Implementing recent theoretical proposals as decoding firmware that parallelizes general inner decoders, we can achieve unprecedented decoding capabilities of up to distances 47 and 67 with the currently available systems-on-chips for physical error rate p = 0.001 and p = 0.0001, respectively, all in just 1 μs. Rui Chao, Cupjin Huang, Linghang Kong, Guoyang Chen, Dawei Ding 0002, Haishan Feng, Yihuai Gao, Xiaotong Ni, Liwei Qiu, Yueming Yang, Yaoyun Shi, Weifeng Zhang 0003, Peng Zhou 0030 |
ACM Trans. Quantum Comput. | 16 |
| 2023 | GIM: Versatile GNN Acceleration with Reconfigurable Processing-in-MemoryabstractRecent boost of deep learning has revolutionized many machine learning tasks, including the graph neural networks (GNNs) that are specifically designed for non-Euclidean graph data. GNNs have been widely adopted in numerous real-world applications, such as the recommendation system. However, with increasingly enlarged graph size and complexity, GNN performance on conventional computers has been severely hindered by the memory bottleneck. The challenge attracts wide investigations, and the processing-in-memory (PIM) architecture arises as one of the most promising solutions. Prior works have leveraged the ReRAM crossbars as analog dot-product engines to accelerate the vector-matrix multiplications in GNN, and achieve prominent performance improvements over modern CPUs and GPUs. Nevertheless, analog computing is known to be variation-vulnerable, which hampers the inference accuracy of GNN. Besides, the mixed-signal peripherals (e.g., ADC) are hardware-expensive and specialize in dense computations, which makes the analog crossbar-based PIM not the ideal candidate for GNN inference whose computation is of great sparsity.In this work, we propose a novel digital-PIM architecture for GNN acceleration, namely GIM. Our compact yet efficient digital computing paradigm can greatly boost computing parallelism with a minimum budget. GIM integrates dedicated optimizations on both operand- and bit-sparsity, to eliminate sparse computations thus significantly boost the performance. Meanwhile, at the software level, we implement data-layout optimizations to minimize the inter-memory communications and maximize computing parallelism. Our design derives prominent performance improvements over the modern CPU, GPU, and state-of-the-art PIM-based accelerators. Compared to modern CPU and GPU, GIM averagely achieves 24485× and 778× of speedup, and 78480× and 8906× of energy reduction. Compared to the state-of-the-art PIM-based GNN accelerators ReFlip and PIMGCN, GIM averagely achieves 9.0× and 73.4× of throughput boost with 15.2× and 95.6× of efficiency improvements. Chen Nie, Guoyang Chen, Weifeng Zhang 0003, Zhezhi He |
ICCD | 3 |
| 2022 | N3H-Core: Neuron-designed Neural Network Accelerator via FPGA-based Heterogeneous Computing CoresabstractAccelerating the neural network inference by FPGA has emerged as a popular option, since the reconfigurability and high performance computing capability of FPGA intrinsically satisfies the computation demand of the fast-evolving neural algorithms. However, the popular neural accelerators on FPGA (e.g., Xilinx DPU) mainly utilize the DSP resources for constructing their processing units, while the rich LUT resources are not well exploited. Via the software-hardware co-design approach, in this work, we develop an FPGA-based heterogeneous computing system for neural network acceleration. From the hardware perspective, the proposed accelerator consists of DSP- and LUT-based GEneral Matrix-Multiplication (GEMM) computing cores, which forms the entire computing system in a heterogeneous fashion. The DSP- and LUT-based GEMM cores are computed w.r.t a unified Instruction Set Architecture (ISA) and unified buffers. Along the data flow of the neural network inference path, the computation of the convolution/fully-connected layer is split into two portions, handled by the DSP- and LUT-based GEMM cores asynchronously. From the software perspective, we mathematically and systematically model the latency and resource utilization of the proposed heterogeneous accelerator, regarding varying system design configurations. Through leveraging the reinforcement learning technique, we construct a framework to achieve end-to-end selection and optimization of the design specification of target heterogeneous accelerator, including workload split strategy, mixed-precision quantization scheme, and resource allocation of DSP- and LUT-core. In virtue of the proposed design framework and heterogeneous computing system, our design outperforms the state-of-the-art Mix&Match design with latency reduced by 1.12-1.32x with higher inference accuracy. The N3H-core is open-sourced at: https://github.com/elliothe/N3H_Core. Zhihan Xu, Zhezhi He, Weifeng Zhang 0003, Xiaobing Tu, Xiaoyao Liang, Li Jiang 0002 |
FPGA | 4 |
| 2021 | PIM-DL: Boosting DNN Inference on Digital Processing In-Memory Architectures via Data Layout OptimizationsabstractDigital processing in-memory (DPIM) provides very low overhead, highly parallel computation in conventional memory, which significantly accelerates data-intensive workloads like deep neural networks (DNNs). DPIM-based DNN accelerators require that data be properly laid out to make the best use of the available in-memory operations. However, existing DPIM accelerators tend to optimize for a particular DNN dataflow, neglecting the large design space of data layout. This work systematically investigates the data layout for DPIM DNN acceleration. We propose a mapping framework to represent the whole design space of DPIM data layout for general DNN models. Our investigation shows that an exhaustive exploration on the whole design space for mapping a DNN application to DPIM architecture is not computationally tractable. Therefore, we propose a compiler-level optimization, PIM-DL, that finds highly efficient data layouts for DPIM DNN acceleration using a two-level dynamic programming algorithm and a heuristic-based search. Our experiments show that DNN DPIM solutions created by our PIM-DL provide 3.7× and 4.3× better performance and energy efficiency as compared to the state of the art under the same hardware constraints. Minxuan Zhou, Guoyang Chen, Mohsen Imani, Saransh Gupta, Weifeng Zhang 0003, Tajana Rosing |
PACT | 5 |
| 2021 | Simple Augmentation Goes a Long Way: ADRL for DNN Quantization
Lin Ning 0001, Guoyang Chen, Weifeng Zhang 0003, Xipeng Shen |
ICLR | 3 |
| 2021 | Enabling energy-efficient DNN training on hybrid GPU-FPGA acceleratorsabstractDNN training consumes orders of magnitude more energy than inference and requires innovative use of accelerators to improve energy-efficiency. However, despite having complementary features, GPUs and FPGAs have been mostly used independently for the entire training process, thus neglecting the opportunity in assigning individual but distinct operations to the most suitable hardware. In this paper, we take the initiative to explore new opportunities and viable solutions in enabling energy-efficient DNN training on hybrid accelerators. To overcome fundamental challenges including avoiding training throughput loss, enabling fast design space exploration, and efficient scheduling, we propose a comprehensive framework, Hype-training, that utilizes a combination of offline characterization, performance modeling, and online scheduling of individual operations. Experimental tests using NVIDIA V100 GPUs and Intel Stratix 10 FPGAs show that, Hype-training is able to exploit a mixture of GPUs and FPGAs at a fine granularity to achieve significant energy reduction, by 44.3% on average and up to 59.7%, without any loss in training throughput. Hype-training can also enforce power caps more effectively than state-of-the-art power management mechanisms on GPUs. Xin He 0054, Hao Chen 0002, Guoyang Chen, Weifeng Zhang 0003, Dong Li 0001 |
ICS | 6 |
| 2021 | EGEMM-TC: accelerating scientific computing on tensor cores with extended precisionabstractNvidia Tensor Cores achieve high performance with half-precision matrix inputs tailored towards deep learning workloads. However, this limits the application of Tensor Cores especially in the area of scientific computing with high precision requirements. In this paper, we build Emulated GEMM on Tensor Cores (EGEMM-TC) to extend the usage of Tensor Cores to accelerate scientific computing applications without compromising the precision requirements. First, EGEMM-TC employs an extendable workflow of hardware profiling and operation design to generate a lightweight emulation algorithm on Tensor Cores with extended-precision. Second, EGEMM-TC exploits a set of Tensor Core kernel optimizations to achieve high performance, including the highly-efficient tensorization to exploit the Tensor Core memory architecture and the instruction-level optimizations to coordinate the emulation computation and memory access. Third, EGEMM-TC incorporates a hardware-aware analytic model to offer large flexibility for automatic performance tuning across various scientific computing workloads and input datasets. Extensive evaluations show that EGEMM-TC can achieve on average 3.13× and 11.18× speedup over the cuBLAS kernels and the CUDA-SDK kernels on CUDA Cores, respectively. Our case study on several scientific computing applications further confirms that EGEMM-TC can generalize the usage of Tensor Cores and achieve about 1.8× speedup compared to the hand-tuned, highly-optimized implementations running on CUDA Cores. Boyuan Feng, Guoyang Chen, Weifeng Zhang 0003, Yuan Xie 0001, Yufei Ding 0001 |
PPoPP | 4 |
| 2021 | Software-Defined Design Space Exploration for an Efficient DNN Accelerator ArchitectureabstractDeep neural networks (DNNs) have been shown to outperform conventional machine learning algorithms across a wide range of applications, e.g., image recognition, object detection, robotics, and natural language processing. However, the high computational complexity of DNNs often necessitates extremely fast and efficient hardware. The problem gets worse as the size of neural networks grows exponentially. As a result, customized hardware accelerators have been developed to accelerate DNN processing without sacrificing model accuracy. However, previous accelerator design studies have not fully considered the characteristics of the target applications, which may lead to sub-optimal architecture designs. On the other hand, new DNN models have been developed for better accuracy, but their compatibility with the underlying hardware accelerator is often overlooked. In this article, we propose an application-driven framework for architectural design space exploration of DNN accelerators. This framework is based on a hardware analytical model of individual DNN operations. It models the accelerator design task as a multi-dimensional optimization problem. We demonstrate that it can be efficaciously used in application-driven accelerator architecture design: we use the framework to optimize the accelerator configurations for eight representative DNNs and select the configuration with the highest geometric mean performance. The geometric mean performance improvement of the selected DNN configuration relative to the architectural configuration optimized only for each individual DNN ranges from 12.0 to 117.9 percent. Given a target DNN, the framework can generate efficient accelerator design solutions with optimized performance and area. Furthermore, we explore the opportunity to use the framework for accelerator configuration optimization under simultaneous diverse DNN applications. The framework is also capable of improving neural network models to best fit the underlying hardware resources. We demonstrate that it can be used to analyze the relationship between the operations of the target DNNs and the corresponding accelerator configurations, based on which the DNNs can be tuned for better processing efficiency on the given accelerator without sacrificing accuracy. Ye Yu 0003, Yingmin Li, Shuai Che, Niraj K. Jha, Weifeng Zhang 0003 |
IEEE Trans. Computers | 5 |
| 2020 | Regularized Training and Tight Certification for Randomized Smoothed Classifier with Provable RobustnessabstractRecently smoothing deep neural network based classifiers via isotropic Gaussian perturbation is shown to be an effective and scalable way to provide state-of-the-art probabilistic robustness guarantee against ℓ2 norm bounded adversarial perturbations. However, how to train a good base classifier that is accurate and robust when smoothed has not been fully investigated. In this work, we derive a new regularized risk, in which the regularizer can adaptively encourage the accuracy and robustness of the smoothed counterpart when training the base classifier. It is computationally efficient and can be implemented in parallel with other empirical defense methods. We discuss how to implement it under both standard (non-adversarial) and adversarial training scheme. At the same time, we also design a new certification algorithm, which can leverage the regularization effect to provide tighter robustness lower bound that holds with high probability. Our extensive experimentation demonstrates the effectiveness of the proposed training and certification approaches on CIFAR-10 and ImageNet datasets. Huijie Feng, Chunpeng Wu, Guoyang Chen, Weifeng Zhang 0003, Yang Ning |
AAAI | 4 |
| 2020 | iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank ArchitectureabstractImage processing is becoming an increasingly important domain for many applications on workstations and the datacenter that require accelerators for high performance and energy efficiency. GPU, which is the state-of-the-art accelerator for image processing, suffers from the memory bandwidth bottleneck. To tackle this bottleneck, near-bank architecture provides a promising solution due to its enormous bank-internal bandwidth and low-energy memory access. However, previous work lacks hardware programmability, while image processing workloads contain numerous heterogeneous pipeline stages with diverse computation and memory access patterns. Enabling programmable near-bank architecture with low hardware overhead remains challenging.This work proposes iPIM, the first programmable in-memory image processing accelerator using near-bank architecture. We first design a decoupled control-execution architecture to provide lightweight programmability support. Second, we propose the SIMB (Single-Instruction-Multiple-Bank) ISA to enable flexible control flow and data access. Third, we present an end-to-end compilation flow based on Halide that supports a wide range of image processing applications and maps them to our SIMB ISA. We further develop iPIM-aware compiler optimizations, including register allocation, instruction reordering, and memory order enforcement to improve performance. We evaluate a set of representative image processing applications on iPIM and demonstrate that on average iPIM obtains 11.02× acceleration and 79.49% energy saving over an NVIDIA Tesla V100 GPU. Further analysis shows that our compiler optimizations contribute 3.19× speedup over the unoptimized baseline. Peng Gu 0008, Xinfeng Xie, Yufei Ding 0001, Guoyang Chen, Weifeng Zhang 0003, Dimin Niu, Yuan Xie 0001 |
ISCA | 5 |
| 2019 | Parallel Training via Computation Graph TransformationabstractParallel training can speed up the convergence of machine learning models via splitting the workload into multiple accelerators by the wide array of possible parallel paradigms (e.g., data parallelism, model parallelism, attribute parallelism, and pipelining parallelism). However, most machine learning frameworks lack sufficient support for these flexible and sometimes complex parallel training schemes (e.g., TensorFlow does not provide convenient APIs for any paradigm other than data parallelism), and the engineering effort to support all parallelisms in all machine learning frameworks seems gigantic. In this paper, we demonstrate that most parallel training designs/paradigms can be abstracted as a computation graph transformation problem, so that they are realized via computation graph duplication, splitting, augmentation, and assignment to different accelerators, which are then connected by send/recv channels for tensor communications. Furthermore, conducting such computation graph transformations in a por table I R allows the engineering efforts of parallel training to be widely applied across machine learning frameworks. We propose an extensible parallel training search space which describes parallel training schemes in a declarative fashion. We then implement a computation graph transformation compiler that can instantiate the parallel schemes into explicit execution plans, which are readily executable on modern machine learning frameworks (such as TensorFlow). We maximize code reuse by handling parallel configurations and computation graph transformations in extended ONNX, which can be ported to machine learning frameworks by adapting their existing ONNX frontend/backend implementations. Our design reflects a few good themes in machine learning frameworks, including code reuse via powerful IR (as in MLIR) and separation of declaration and realization (as in Halide/TVM). Fei Wang 0046, Guoyang Chen, Weifeng Zhang 0003, Tiark Rompf |
IEEE BigData | 3 |
| 2019 | Energy-Efficient and Quality-Assured Approximate Computing Framework Using a Co-Training MethodabstractApproximate computing is a promising design paradigm that introduces a new dimension—error—into the original design space. By allowing the inexact computation in error-tolerance applications, approximate computing can gain both performance and energy efficiency. A neural network (NN) is a universal approximator in theory and possesses a high level of parallelism. The emerging deep neural network accelerators deployed with NN-based approximator is thereby a promising candidate for approximate computing. Nevertheless, the approximation result must satisfy the users’ requirement, and the approximation result varies across different applications. We normally deploy an NN-based classifier to ensure the approximation quality. Only the inputs predicted to meet the quality requirement can be executed by the approximator. The potential of these two NNs, however, is fully explored; the involving of two NNs in approximate computing imposes critical optimization questions, such as two NNs’ distinct views of the input data space, how to train the two correlated NNs, and what are their topologies. In this article, we propose a novel NN-based approximate computing framework with quality insurance. We advocate a co-training approach that trains the classifier and the approximator alternately to maximize the agreement of the two NNs on the input space. In each iteration, we coordinate the training of the two NNs with a judicious selection of training data. Next, we explore different selection policies and propose to select training data from multiple iterations, which can enhance the invocation of the approximate accelerator. In addition, we optimize the classifier by integrating a dynamic threshold tuning algorithm to improve the invocation of the approximate accelerator further. The increased invocation of accelerator leads to higher energy efficiency under the same quality requirement. We propose two efficient algorithms to explore the smallest topology of the NN-based approximator and the classifier to achieve the quality requirement. The first algorithm straightforward searches the minimum topology using a greedy strategy. However, the first algorithm incurs too much training overhead. To solve this issue, the second one gradually grows the topology of NNs to match the quality requirement by transferring the learned parameters. Experimental results show significant improvement on the quality and the energy efficiency compared to the existing NN-based approximate computing frameworks. Li Jiang 0002, Zhuoran Song, Haiyue Song, Chengwen Xu, Qiang Xu 0001, Naifeng Jing, Weifeng Zhang 0003, Xiaoyao Liang |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2007 | Accelerating and Adapting Precomputation Threads for Effcient PrefetchingabstractSpeculative precomputation enables effective cache prefetching for even irregular memory access behavior, by using an alternate thread on a multithreaded or multi-core architecture. This paper describes a system that constructs and runs precomputation based prefetching threads via event-driven dynamic optimization. Precomputation threads are dynamically constructed by a runtime compiler from the program's frequently executed hot traces, and are adapted to the memory behavior automatically. Both construction and execution of the prefetching threads happen in another thread, imposing little overhead on the main thread. This paper also presents several techniques to accelerate the precomputation threads, including colocation of p-threads with hot traces, dynamic stride prediction, and automatic adaptation of runahead and jumpstart distance. The adaptive prefetching achieves 42% speedup, a 17% improvement over existing p-thread prefetching schemes Weifeng Zhang 0003, Dean M. Tullsen, Brad Calder |
HPCA | 1 |
| 2006 | A Self-Repairing Prefetcher in an Event-Driven Dynamic Optimization FrameworkabstractSoftware prefetching has been demonstrated as a powerful technique to tolerate long load latencies. However, to be effective, prefetching must target the most critical (frequently missing) loads, and prefetch them sufficiently far in advance. This is difficult to do correctly with a static optimizer, because locality characteristics and cache latencies vary across data inputs and across different machines. This paper presents a mechanism that dynamically inserts prefetch instructions into frequently executed hot traces. Hot traces are dynamically analyzed to identify delinquent loads and the appropriate prefetch distance for those loads. Those prefetches are then inserted into the hot trace. The low overhead of the event-driven dynamic optimization system allows the optimizer to continuously monitor the performance of the software prefetches. This is done to find an accurate and stable prefetch distance and to adapt to changes in program behavior using what we call self-repairing prefetching. Relative to the baseline hardware stride prefetching, we find a total 23% improvement when we use the self-repairing mechanism to adoptively discover the best prefetch distance for each load, which is 12% better performance than dynamic prefetching techniques without adaptive repairing. Weifeng Zhang 0003, Brad Calder, Dean M. Tullsen |
CGO | 1 |
| 2006 | Speculative Code Value Specialization Using the Trace Cache Fill UnitabstractValue specialization is a technique which can improve a program's performance when its code frequently takes the same values. In this paper, speculative value specialization is applied dynamically by utilizing the trace cache hardware. We implement a small, efficient hardware profiler to identify loads that have semi-invariant runtime values. A specialization engine off the program's critical path generates highly optimized traces using these values, which reside in the trace cache. Specialized traces are dynamically verified during execution, and mis-specialization is recovered automatically without new hardware overhead. Our simulation shows that dynamic value specialization in the trace cache achieves a 17% speedup, even over a system with support for hardware value prediction. When combined with other techniques aimed at tolerating memory latencies, this technique still performs well -this technique combined with an aggressive hardware prefetcher achieves 24% better performance than prefetching alone. Weifeng Zhang 0003, Brad Calder, Dean M. Tullsen, Steve Checkoway |
ICCD | 1 |