Size Zheng 0001

dblp:254/6617-1 · DBLP profile ↗
← Back
30ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0002-9471-1780ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 7 first-author · 25 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 SwiftSpec: Disaggregated Speculative Decoding and Fused Kernels for Low-Latency LLM Inference
abstract
Low-latency, single-request decoding of large language models is critical for interactive systems with tight SLA demands. Prior work reduces latency through speculative decoding (combining a small draft model with a larger target model), but the draft model remains on the critical path, and communication overhead limits scaling across GPUs due to the small batch size associated with single-request decoding. To address these limitations, this paper introduces SwiftSpec: a system architecture that disaggregates draft and target models across homogeneous GPUs within a single node and utilizes NCCL-low-latency primitives directly to improve the performance of core GEMM and attention kernels. Our implementation includes 3k lines of custom CUDA for fused kernels and an evolving tree cache for KV-cache consistency and maximized reuse between draft and target models. On a single 8×H800 GPU node, SwiftSpec achieves 347 tokens/s for Llama-3-70B---1.3× faster than NVIDIA's own benchmarks on a higher-performance 8×H200 setup---and averages 1.75× faster decoding than state-of-the-art speculative decoding across five model families and six datasets. Specifically, we find that for Llama-3-70B SwiftSpec is significantly faster across all 480 tested queries, showing 1.7× speedup over the best open-source baseline for 95th percentile requests. Code for SwiftSpec will be available at https://github.com/ByteDance-Seed/SwiftSpec
Ziheng Jiang, Chengquan Jiang, Menghan Yu, Size Zheng 0001, Haibin Lin, Xin Liu 0086, Henry Hoffmann
ASPLOS (2)5
2026 LATIAS: A General Architecture-Operator Model for Spatial Accelerators with Complex Topology and Memory Hierarchy
abstract
Spatial accelerators are widely deployed for deep neural networks, but their architectural diversity—from hierarchical to dataflow designs—makes accurate architecture–operator modeling difficult, limiting operator optimization and hardware utilization. Existing models abstract hardware as hierarchical chains and operators as loop trees, which cannot capture essential features of modern dataflow accelerators, including heterogeneous processing elements (PEs), uni-directional interconnects, and cross-PE memory hierarchies, leading to inaccurate latency prediction. We propose LATIAS, a unified framework that introduces (1) an architecture graph with uni-directional edges to represent arbitrary topologies, and (2) a dataflow-aware tile-centric notation that augments loop trees with transfer nodes to model diverse dataflows. Building on these, LATIAS further provides a graph-guided tree analysis that accurately resolves tensor residency and latency under hardware constraints. Experiments on representative operators (GEMM, vector, fused vector) and operator shapes extracted from DNNs (BERT, ViT, T5) on Huawei Ascend 910B3 show that LATIAS achieves over 0.99 correlation with runtime measurements—substantially outperforming prior models—and provides actionable insights for architectural design.
Chengrui Zhang, Liancheng Jia, Renze Chen, Xiuping Cui, Size Zheng 0001, Shengen Yan, Yu Wang 0002, Yun Liang 0001
DATE7
2026 DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
abstract
As the Mix-of-Experts (MoE) architecture increases the number of parameters in large models, there is an even greater need for model quantization. However, existing quantization methods overlook the expert dynamics of MoE across multiple datasets. Moreover, the existing static quantization cannot adapt MoE to various data change scenarios. In this paper, we perform a multi-level analysis to reveal MoE dynamics and define the significance of each channel/each expert. Based on the analysis results, we propose DynaMo, an end-to-end MoE quantization framework. DynaMo adopts an expert-level mixed-precision baseline quantization strategy, which ensures the quantized MoEs are compatible with multiple existing datasets. Furthermore, DynaMo incorporates a channel-level dynamic switching mechanism to adapt these quantized MoE models to novel datasets. Experiments show that DynaMo achieves a 2.78~4.54 PPL decrease and a 1.85%~3.77% accuracy improvement in various datasets, with ~3× inference speedup and negligible overhead.
Xiuping Cui, Size Zheng 0001, Maoliang Li, Yun Liang 0001
DATE3
2026 MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
abstract
We present MegaScale-MoE, a production system tailored for the efficient training of large-scale mixture-of-experts (MoE) models. MoE emerges as a promising architecture to scale large language models (LLMs) to unprecedented sizes, thereby enhancing model performance. However, existing MoE training systems experience a degradation in training efficiency, exacerbated by the escalating scale of MoE models and the continuous evolution of hardware.
Chao Jin 0007, Ziheng Jiang, Zhihao Bai, Juncai Liu, Xiang Li 0067, Ningxin Zheng, Qi Huang 0001, Wen Heng, Yiyuan Ma, Wenlei Bao, Size Zheng 0001, Xuegui Zheng, Yanghua Peng, Haibin Lin, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086
EuroSys14
2026 TENET-v2: Applying Relation-Centric Notation to Model and Optimize Data Swizzle in the Cache of Modern NPU
abstract
Swizzle is a data access pattern optimization technique by reorganizing the execution order of computational tasks to improve the cache locality in modern NPUs. Existing analysis and optimization techniques lack support for swizzleaware modeling on NPUs and fail to effectively capture cache behavior across diverse swizzle configurations. To this end, we propose TENET-v2, a framework for modeling and optimizing swizzle. We introduce a relation-centric notation to characterize different cache access patterns, thus exploring wider swizzle space. Then, we propose a hybrid performance model for cache analysis. The proposed performance model uses an analytical approach to quantify cache miss behavior under unsaturated cache conditions (non-saturated misses), and employs a simulation method combined with an early exiting mechanism to rapidly model cache behavior under saturated cache conditions (saturated misses). Experimental evaluations demonstrate that TENET-v2 achieves an average absolute error of 1.05 % in read hit rate compared to real-world hardware. Evaluation on a variety of DNNs shows that TENET-v2 outperforms existing tensor program optimizers by up to$\mathbf{1. 5} \times$on A100 GPUs. We also demonstrate NPU cache size optimization based on TENET-v2.
Fangxu Guo, Liqiang Lu, Jinghan Zhang 0015, Jie Zhang 0177, Chenli Xue, Chengpeng Wu, Yun Liang 0001, Size Zheng 0001, Jianwei Yin
HPCA13
2026 UniEP: Unified Expert-Parallel MegaKernel MoE for LLM Training
abstract
As LLM training grows increasingly resource-intensive and expert parallelism (EP) becomes essential for scaling MoE models, EP optimizations are widely adopted in production frameworks like Megatron-LM. Existing solutions often rely on ad-hoc, complex kernels that lack adaptability across diverse optimization configurations and frequently neglect numerical stability, failing to meet the strict precision requirements of large-scale training.
Size Zheng 0001, Xuegui Zheng, Li-Wen Chang, Jidong Zhai
HPDC1
2026 Tetris: Efficient Long-context LLM Serving with Chunkwise Dynamic Sequence Parallelism
Xuegui Zheng, Yijin Guan, Size Zheng 0001, Li-Wen Chang, Shufan Liu, Xin Liu 0086, Guangyu Sun 0003
ISCA6
2026 PipeComm: Maximizing Link Utilization Through Pipeline-Aware Collective Communication Synthesis
Ruifan Xu, Yuze Luo, Yuhao Meng, Size Zheng 0001, Meng Li 0004, Yun Liang 0001
ISCA4
2026 MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware
abstract
Large language models (LLMs) are prominent for their superior ability in language understanding and generation. However, a notorious problem for LLM inference is low computational utilization caused by the memory bottleneck, since it typically requires large memory capacity and high bandwidth to process neural weights. By integrating processing cores into memory, Processing-In-Memory (PIM) architecture excels at alleviating memory bottleneck; with the recent release of the first commodity near-bank PIM hardware (NBP), PIM becomes off-the-shelf and shows great potential for accelerating LLM inference practically. However, simply shoehorning LLM inference on NBP can not achieve satisfactory performance due to its inherent limitations: weak compute performance, frequent cache misses caused by the limited working memory capacity, and poor inter-PIM-core communication bandwidth. To address these limitations, we propose MI-LLM, an efficient system deploying LLM inference on NBP hardware. Its key idea is to build NBP-aware Lookup Tables (LUTs) and completely replace multiplications with lookups on LUTs, thereby mitigating the limitation of weak compute performance. 1) To reduce the model accuracy drop caused by the use of LUT, MI-LLM tailors a learning-based LUT construction method to maintain the model accuracy. 2) To cope with frequent cache misses caused by LUT sizes far exceeding PIM working memory capacity, MI-LLM introduces the design of PIM-aware linear kernel, with the optimization of intra-row and inter-row reordering enabled, to enhance LUT lookup locality. 3) MI-LLM further proposes a model partitioning scheme to minimize inter-PIM-core communication. Kernel-level benchmarks reveal that MI-LLM achieves a 9% throughput improvement and an 11% increase in energy efficiency over GPU implementations. Compared to FP8 quantization, MI-LLM incurs only a 0.24 times increase in perplexity, demonstrating minimal accuracy degradation. Moreover, in our end-to-end evaluation, MI-LLM requires 80% fewer ALU operation ticks per output token than the GPU baseline.
Puyun Hu, Minhui Xie, Linjiang Li, Kuiyaohui Zhang, Erge Xiang, Jing Wang 0055, Size Zheng 0001, Xiao Zhang 0001, Yunpeng Chai
IEEE Trans. Computers7
2025 QRAMsim: Efficiently Simulating, Analyzing, and Optimizing Large-Scale Quantum Random Access Memory
Chenning Tao, Yujie Ji, Liqiang Lu, Size Zheng 0001, Jianwei Yin
APPT4
2025 DyREM: Dynamically Mitigating Quantum Readout Error with Embedded Accelerator
abstract
Quantum readout error is the most significant source of error, substantially reducing the measurement fidelity. Tensor-product-based readout error mitigation has been proposed to address this issue by approximating the mitigation matrix. However, this method inevitably encounters the dynamic generation of the mitigation matrix, leading to long latency. In this paper, we propose DyREM, a software-hardware codesign approach that mitigates readout errors with an embedded accelerator. The main innovation lies in leveraging the inherent sparsity in the nonzero probability distribution of quantum states and calculating the tensor product on an embedded accelerator. Specifically, using the output sparsity, our dataflow dynamically downsamples the original mitigation matrix, which dramatically reduces the memory requirement. Then, we design DyREM architecture that can flexibly gate the redundant computation of nonzero quantum states. Experiments demonstrate that DyREM achieves an average speedup of $9.6 \times \sim 2000 \times$ and fidelity improvements of $1.03 \times \sim 1.15 \times$ compared to state-of-the-art readout error mitigation methods.
Kaiwen Zhou 0003, Liqiang Lu, Debin Xiang, Chenning Tao, Xinkui Zhao, Size Zheng 0001, Jianwei Yin
DAC7
2025 MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
abstract
Mixture-of-Experts (MoE) models face deployment challenges due to their large parameter counts and computational demands. We explore quantization for MoE models and highlight two key insights: 1) linear blocks exhibit varying quantization sensitivity, and 2) divergent expert activation frequencies create heterogeneous computational characteristics. Based on these observations, we introduce MxMoE, a mixed-precision optimization framework for MoE models that considers both algorithmic and system perspectives. MxMoE navigates the design space defined by parameter sensitivity, expert activation dynamics, and hardware resources to derive efficient mixed-precision configurations. Additionally, MxMoE automatically generates optimized mixed-precision GroupGEMM kernels, enabling parallel execution of GEMMs with different precisions. Evaluations show that MxMoE outperforms existing methods, achieving 2.4 lower Wikitext-2 perplexity than GPTQ at 2.25-bit and delivering up to 3.4x speedup over full precision, as well as up to 29.4% speedup over uniform quantization at equivalent accuracy with 5-bit weight-activation quantization. Our code is available at https://github.com/cat538/MxMoE.
Haojie Duanmu, Zhihang Yuan, Size Zheng 0001, Jiangfei Duan, Xingcheng Zhang, Dahua Lin
ICML4
2025 ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
abstract
With the widespread deployment of long-context large language models (LLMs), there has been a growing demand for efficient support of high-throughput inference. However, as the key-value (KV) cache expands with the sequence length, the increasing memory footprint and the need to access it for decoding both result in low throughput when serving long-context LLMs. While various dynamic sparse attention methods have been proposed to accelerate inference while maintaining generation quality, they either fail to sufficiently reduce GPU memory usage or introduce significant decoding latency by offloading the KV cache to the CPU. We present ShadowKV, a high-throughput long-context LLM inference system that stores the low-rank key cache and offloads the value cache to reduce the memory footprint for larger batch sizes and longer sequences. To minimize decoding latency, ShadowKV employs an accurate KV selection strategy that reconstructs minimal sparse KV pairs on-the-fly. By evaluating ShadowKV on benchmarks like RULER, LongBench, and models such as Llama-3.1-8B and GLM-4-9B-1M, we demonstrate that it achieves up to 6$\times$ larger batch sizes and 3.04$\times$ higher throughput on an A100 GPU without sacrificing accuracy, even surpassing the performance achievable with infinite batch size under the assumption of infinite GPU memory.
Hanshi Sun, Li-Wen Chang, Wenlei Bao, Size Zheng 0001, Ningxin Zheng, Xin Liu 0086, Harry Dong, Yuejie Chi, Beidi Chen
ICML4
2025 Qtenon: Towards Low-Latency Architecture Integration for Accelerating Hybrid Quantum-Classical Computing
abstract
Hybrid quantum-classical algorithms have shown great promise in leveraging the computational potential of quantum systems.However, the efficiency of these algorithms is severely constrained by the limitations of current quantum hardware architectures.These architectures, which typically feature a decoupled design, lack both hardware support for low-latency communication and software support for fine-grained optimization.In this paper, we propose Qtenon, a tightly coupled system for efficient hybrid quantum-classical algorithm acceleration.Qtenon is composed of both hardware part and software part.To enable efficient communication and computation, the hardware part provides a unified memory hierarchy, an efficient quantum controller, as well as a multi-stage processing pipeline.The unified memory hierarchy functions as a communication buffer between host and quantum accelerators, with dedicated data paths and interfaces provided by the quantum controller.The multi-stage pipeline leverages hardware pipelines to fully exploit parallelism.To program hybrid quantum-classical algorithms on the hardware, our software part provides a set of instructions for data communication and computation.The instructions also enable fine-grained synchronization and efficient scheduling for quantum-host interaction.We design Qtenon as a RISC-V extended chip and implement it using Chisel.In evaluation, we achieve up to 14.9× end-to-end speedup compared to state-of-the-art work for hybrid quantum-classical algorithms.
Chenning Tao, Liqiang Lu, Size Zheng 0001, Li-Wen Chang, Minghua Shen, Fangxin Liu, Kaiwen Zhou 0003, Jianwei Yin
ISCA3
2024 MAGIS: Memory Optimization via Coordinated Graph Transformation and Scheduling for DNN
abstract
Recently, memory consumption of Deep Neural Network (DNN) rapidly increases, mainly due to long lifetimes and large shapes of tensors. Graph scheduling has emerged as an effective memory optimization technique, which determines the optimal execution, re-computation, swap-out, and swap-in timings for each operator/tensor. However, it often hurts performance significantly and can only manipulate tensors' lifetimes but not shapes, limiting the optimization space. We find that graph transformation, which can change the tensor shapes and graph structure, creates a new trade-off space between memory and performance. Nevertheless, graph transformation are applied separately so far, with primary focus on optimizing performance and not memory.
Renze Chen, Zijian Ding, Size Zheng 0001, Chengrui Zhang, Jingwen Leng, Xuanzhe Liu, Yun Liang 0001
ASPLOS (3)3
2024 SpecPIM: Accelerating Speculative Inference on PIM-Enabled System via Architecture-Dataflow Co-Exploration
abstract
Generative large language models' (LLMs) inference suffers from inefficiency because of the token dependency brought by autoregressive decoding. Recently, speculative inference has been proposed to alleviate this problem, which introduces small language models to generate draft tokens and adopts the original large language model to conduct verification. Although speculative inference can enhance the efficiency of the decoding procedure, we find that it presents variable resource demands due to the distinct computation patterns of the models used in speculative inference. This variability impedes the full realization of speculative inference's acceleration potential in current systems.
Cong Li 0008, Zhe Zhou 0002, Size Zheng 0001, Jiaxi Zhang 0001, Yun Liang 0001, Guangyu Sun 0003
ASPLOS (3)3
2024 MoteNN: Memory Optimization via Fine-grained Scheduling for Deep Neural Networks on Tiny Devices
abstract
There has been a growing trend in deploying deep neural networks (DNNs) on tiny devices. However, deploying DNNs on such devices poses significant challenges due to the contradiction between DNNs' substantial memory requirements and the stringent memory constraints of tiny devices. Some prior works incur large latency overhead to save memory and target only simple CNNs, while others employ coarse-grained scheduling for complicated networks, leading to limited memory footprint reduction. This paper proposes MoteNN that performs fine-grained scheduling via operator partitioning on arbitrary DNNs to dramatically reduce peak memory usage with little latency overhead. MoteNN presents a graph representation named Axis Connecting Graph (ACG) to perform operator partition at graph-level efficiently. MoteNN further proposes an algorithm that finds the partition and schedule guided by memory bottlenecks. We evaluate MoteNN using various popular networks and show that MoteNN achieves up to 80% of peak memory usage reduction compared to the state-of-art works with nearly no latency overhead on tiny devices.
Renze Chen, Zijian Ding, Size Zheng 0001, Meng Li 0004, Yun Liang 0001
DAC3
2024 SpREM: Exploiting Hamming Sparsity for Fast Quantum Readout Error Mitigation
abstract
The current Noisy Intermediate-Scale Quantum (NISQ) era suffers from high quantum readout error that severely reduces the measurement fidelity. Matrix-based error mitigation has been demonstrated as a promising software-level technique, which performs matrix-vector multiplication to calibrate the probability distribution with noise. However, this approach shows poor scalability and limited fidelity improvement as the matrix size exponentially increases with the number of qubits. In this paper, we propose SpREM to exploit the inherent sparsity in the mitigation matrix. Inspired by the interaction mechanism between qubits, we identify structured sparsity patterns using Hamming distance. With this insight, we propose the Hamming-Distance Sparse Row (HDSR) compression method and its format, which can achieve higher sparsity than threshold-based pruning meanwhile exhibiting great fidelity improvement. Finally, we propose the computational dataflow of the HDSR format and implement it on hardware. Experiments demonstrate that SpREM achieves 98.9% sparsity and a 27.3× reduction in fidelity loss on the real-world quantum device, compared to threshold-based pruning. It achieves an average 11.2× ~ 36.4× speedup compared to Xilinx Vitis SPARSE library and NVIDIA A100 GPU implementations.
Liqiang Lu, Siwei Tan, Size Zheng 0001, Jianwei Yin
DAC4
2024 ArkVale: Efficient Generative LLM Inference with Recallable Key-Value Eviction
abstract
Large Language Models (LLMs) are widely used in today's tasks of natural language processing. To support applications like multi-turn chats, document understanding, and content generation, models with long context lengths are growing in importance. However, managing long contexts brings substantial challenges due to the expansion of key-value cache (KV cache). Longer KV cache requires larger memory, limiting the batch-size thus decreasing throughput. Also, computing attention over long KV cache incurs more memory access, hurting the end-to-end latency. Prior works find that it is sufficient to use only the recent and high-impact tokens for attention computation, allowing the eviction of less vital tokens to shrink cache size. Nonetheless, we observe a dynamic shift in token importance across different decoding steps. Tokens initially evicted might regain importance after certain decoding steps. To address this, we propose ArkVale, a page-based KV cache manager that can recognize and recall currently important tokens evicted before. We asynchronously copy the filled page into external memory (e.g., CPU memory) as backup and summarize it into a much smaller digest by constructing the bounding-volume of its keys. Before attention computation, we measure all pages' importance based on their digests, recall the important ones, evict the unimportant ones, and select the top-ranked pages for attention computation. Experiment results show that ArkVale performs well on various long context tasks with negligible accuracy loss under 2k$\sim$4k cache budget and can improve decoding latency to $2.2\times$ and batching throughput to $4.6\times$ because it applies attention on only a small subset of pages and reduce per-sample memory usage of KV cache.
Renze Chen, Zhuofeng Wang, Beiquan Cao, Size Zheng 0001, Xuechao Wei, Shengen Yan, Meng Li 0004, Yun Liang 0001
NeurIPS5
2024 Rubick: A Unified Infrastructure for Analyzing, Exploring, and Implementing Spatial Architectures via Dataflow Decomposition
abstract
The fast-growing tensor applications expose tremendous dataflow alternatives when implemented on spatial architectures that feature large PE arrays and abundant interconnection resources. Prior works develop various notations and performance models for dataflows. Though these notations are very useful for understanding the reuse, bandwidth, and performance of dataflows, they do not define the underlying hardware implementation. Due to the semantic gap, analysis based on these notations cannot capture the detailed architectural features between different dataflows, leading to inefficient design space exploration and suboptimal designs. To address these issues, we propose Rubick, a unified infrastructure for analyzing, exploring, and implementing spatial architectures. The main innovation of Rubick is it decomposes the dataflow into two low-level intermediate representations: access entry and data layout. Access entry specifies how data enter into the PE arrays from memory, while data layout specifies how data are arranged and accessed. These two representations allow us to infer the hardware implementation details such as PE interconnection and memory structure, which are amenable for structural analysis and systematic exploration. Based on this decomposition analysis, Rubick provides opportunities for micro-architecture optimization and efficient design space exploration. Our experiments demonstrate that Rubick can reduce 82.4% of wire resources with only a 2.7% latency increase by optimizing access entry IR, and achieve 70.8% memory overhead reduction by optimizing data layout IR. Rubick also accelerates the DSE time of dataflows by up to 1.1×105X, saving the time from several days to minutes. The source code of Rubick is publically available on (https://link-omitted-for-blind-review).
Liqiang Lu, Zizhang Luo, Size Zheng 0001, Jieming Yin, Jason Cong, Yun Liang 0001, Jianwei Yin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Rubick: A Synthesis Framework for Spatial Architectures via Dataflow Decomposition
abstract
Dataflows are critical for spatial architectures designed for tensor applications. Prior works develop various notations and hardware generation frameworks for dataflows. However, due to the semantic gap between notations and low-level details, analysis based on these notations cannot capture the detailed architectural features between different dataflows, so these works failed to provide architectural optimization and efficient design space exploration (DSE) at the same time.We propose Rubick, a synthesis framework for spatial architecture. Rubick decomposes the dataflow into two low-level intermediate representations including access entry and data layout. Access entry specifies how data enter into the PE arrays from memory, while data layout specifies how data are arranged and accessed. Based on this decomposition, Rubick provides efficient DSE and generates optimized hardware. Experiments show that the DSE time is accelerated by up to 1.1×105X and performance on FPGA is improved by 13%.
Zizhang Luo, Liqiang Lu, Size Zheng 0001, Jieming Yin, Jason Cong, Jianwei Yin, Yun Liang 0001
DAC3
2023 Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoC
abstract
The DNN models are now pervasively used for various applications. Meanwhile, the computing hardware has shifted towards heterogeneous system composed of various accelerators. The intertwined complexity of DNN models and hardware makes it challenging for mapping DNN models. Existing mapping frameworks suffer from inefficiencies due to under utilization of computation and bandwidth in heterogeneous SoC. In this paper, we propose COMB, a mapping framework that coordinates the memory and computation and data transfer overhead of heterogeneous accelerators to achieve latency improvement and energy efficiency with two optimizations: dataflow grouping and accelerator mapping. Dataflow grouping maps multiple independent DNN layers to the same accelerator at the same time to spatially share the hardware resources; accelerator mapping finds the optimized placement of the layer groups to accelerators to reduce data transfer overhead. These two optimizations provide a huge design space for heterogeneous DNN mapping. To explore the space efficiently, we present a hybrid scheduling algorithm by combining greedy algorithm and genetic algorithm. In evaluation, COMB achieves 1.28× and 1.37× speedup for latency compared to MAGMA and H2H; COMB also reduces 22.7% and 29.2% energy consumption compared to MAGMA and H2H.
Size Zheng 0001, Siyuan Chen 0007, Yun Liang 0001
DAC1
2023 Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion
abstract
Machine learning models with various tensor operators are becoming ubiquitous in recent years. There are two types of operators in machine learning: compute-intensive operators (e.g., GEMM and convolution) and memory-intensive operators (e.g., ReLU and softmax). In emerging machine learning models, compute-intensive operators are usually organized in a chain structure. With the continual specialization of hardware, the gap between computing performance and memory bandwidth has become more prominent. Consequently, the implementations of many compute-intensive operator chains are bounded by memory bandwidth, and generating fused kernels to improve locality for these compute-intensive operators becomes necessary. But in existing machine learning compilers, there lack both precise analysis and efficient optimization for compute-intensive operator chains on different accelerators. As a result, they usually produce sub-optimal performance for these operator chains.In this paper, we propose Chimera, an optimizing framework that can efficiently improve the locality of compute-intensive operator chains on different hardware accelerators. In Chimera, each compute-intensive operator is composed of a series of computation blocks. To generate efficient fused kernels for the operator chains, optimizations for both inter-block and intra-block are required. For inter-block optimization, Chimera decides the optimized block execution order by minimizing the data movement volume among blocks using an analytical model. For intra-block optimization, Chimera uses unified replaceable micro kernels to apply hardware-specific optimizations for different accelerators. Finally, Chimera generates fused kernels for compute-intensive operator chains. Evaluation of batch GEMM chains and convolution chains on CPU, GPU, and NPU shows that Chimera achieves up to 2.87×, 2.29×, and 2.39× speedups to hand-tuned libraries. Compared to state-of-the-art compilers, the speedups are up to 2.29×, 1.64×, and 1.14× for CPU, GPU, and NPU.
Size Zheng 0001, Siyuan Chen 0007, Peidi Song, Renze Chen, Shengen Yan, Dahua Lin, Jingwen Leng, Yun Liang 0001
HPCA1
2023 ARES: A Mapping Framework of DNNs Towards Diverse PIMs with General Abstractions
abstract
Numerous architectures based on processing-in-memory (PIM) have recently emerged, exhibiting diversity in memory types, compute functions, memory mapping constraints, etc. To effectively utilize PIM hardware for deploying deep neural networks (DNNs), programmers face the challenge of mapping computations and data across multiple memory arrays, scheduling computation and data transfers, while adhering to various hardware constraints. Existing mapping approaches, however, are tailored to specific architectures and lack a general formulation for mapping optimization, limiting their applicability and performance. In this paper, we present ARES, a comprehensive mapping framework designed for diverse PIM architectures. The core of the framework is hardware abstractions for PIMs, which is inspired by the fact that DNNs on PIM hardware can be represented by a tensorized compute function and data layout constraints in the memory array. This abstraction forms the basis for constructing a mapping space that encompasses both compute and memory constraints. Through exploration of this mapping space, we derive efficient mapping strategies tailored to different PIM hardware configurations. Experimental evaluation conducted on four distinct hardware architectures demonstrates that compared to state-of-the-art mapping methods, ARES yields up to a 70% speed improvement for single operator mapping and a 50% speedup for overall network mapping.
Xiuping Cui, Size Zheng 0001, Le Ye, Yun Liang 0001
ICCAD2
2023 TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based Analysis
abstract
With the increasing size of DNN models and the growing discrepancy between compute performance and memory bandwidth, fusing multiple layers together to reduce off-chip memory access has become a popular approach in dataflow design. However, designing such dataflows requires flexible and accurate performance models to facilitate evaluation, architecture analysis, and design space exploration. Unfortunately, current state-of-the-art performance models are limited to the dataflows of single operator acceleration, making them inapplicable to operator fusion dataflows.
Size Zheng 0001, Siyuan Chen 0007, Liancheng Jia, Guangyu Sun 0003, Runsheng Wang, Yun Liang 0001
MICRO1
2022 AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction
abstract
Hardware specialization is a promising trend to sustain performance growth. Spatial hardware accelerators that employ specialized and hierarchical computation and memory resources have recently shown high performance gains for tensor applications such as deep learning, scientific computing, and data mining. To harness the power of these hardware accelerators, programmers have to use specialized instructions with certain hardware constraints. However, these hardware accelerators and instructions are quite new and there is a lack of understanding of the hardware abstraction, performance optimization space, and automatic methodologies to explore the space. Existing compilers use hand-tuned computation implementations and optimization templates, resulting in sub-optimal performance and heavy development costs.
Size Zheng 0001, Renze Chen, Anjiang Wei, Yicheng Jin, Qin Han, Liqiang Lu, Bingyang Wu, Shengen Yan, Yun Liang 0001
ISCA1
2022 NeoFlow: A Flexible Framework for Enabling Efficient Compilation for High Performance DNN Training
abstract
Deep neural networks (DNNs) are increasingly deployed in various image recognition and natural language processing applications. The continuous demand for accuracy and high performance has led to innovations in DNN design and a proliferation of new operators. However, existing DNN training frameworks such as PyTorch and TensorFlow only support a limited range of operators and rely on hand-optimized libraries to provide efficient implementations for these operators. To evaluate novel neural networks with new operators, the programmers have to either replace the holistic new operators with existing operators or provide low-level implementations manually. Therefore, a critical requirement for DNN training frameworks is to provide high-performance implementations for the neural networks containing new operators automatically in the absence of efficient library support. In this paper, we introduce NeoFlow, which is a flexible framework for enabling efficient compilation for high-performance DNN training. NeoFlow allows the programmers to directly write customized expressions as new operators to be mapped to graph representation and low-level implementations automatically, providing both high programming productivity and high performance. First, NeoFlow provides expression-based automatic differentiation to support customized model definitions with new operators. Then, NeoFlow proposes an efficient compilation system that partitions the neural network graph into subgraphs, explores optimized schedules, and generates high-performance libraries for subgraphs automatically. Finally, NeoFlow develops an efficient runtime system to combine the compilation and training as a whole by overlapping their execution. In the experiments, we examine the numerical accuracy and performance of NeoFlow. The results show that NeoFlow can achieve similar or even better performance at the operator and whole graph level for DNNs compared to deep learning frameworks. Especially, for novel networks training, the geometric mean speedups of NeoFlow to PyTorch, TensorFlow, and CuDNN are 3.16X, 2.43X, and 1.92X, respectively.
Size Zheng 0001, Renze Chen, Yicheng Jin, Anjiang Wei, Bingyang Wu, Shengen Yan, Yun Liang 0001
IEEE Trans. Parallel Distributed Syst.1
2021 HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation
abstract
Tensor computations overwhelm traditional general-purpose computing devices due to the large amounts of data and operations of the computations. They call for a holistic solution composed of both hardware acceleration and software mapping. Hardware/software (HW/SW) co-design optimizes the hardware and software in concert and produces high-quality solutions. There are two main challenges in the co-design flow. First, multiple methods exist to partition tensor computation and have different impacts on performance and energy efficiency. Besides, the hardware part must be implemented by the intrinsic functions of spatial accelerators. It is hard for programmers to identify and analyze the partitioning methods manually. Second, the overall design space composed of HW/SW partitioning, hardware optimization, and software optimization is huge. The design space needs to be efficiently explored. To this end, we propose an agile co-design approach HASCO that provides an efficient HW/SW solution to dense tensor computation. We use tensor syntax trees as the unified IR, based on which we develop a two-step approach to identify partitioning methods. For each method, HASCO explores the hardware and software design spaces. We propose different algorithms for the explorations, as they have distinct objectives and evaluation costs. Concretely, we develop a multi-objective Bayesian optimization algorithm to explore hardware optimization. For software optimization, we use heuristic and Q-learning algorithms. Experiments demonstrate that HASCO achieves a 1.25X to 1.44X latency reduction through HW/SW co-design compared with developing the hardware and software separately.
Qingcheng Xiao, Size Zheng 0001, Bingzhe Wu, Pengcheng Xu 0005, Xuehai Qian, Yun Liang 0001
ISCA2
2020 FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous System
abstract
Tensor computation plays a paramount role in a broad range of domains, including machine learning, data analytics, and scientific computing. The wide adoption of tensor computation and its huge computation cost has led to high demand for flexible, portable, and high-performance library implementation on heterogeneous hardware accelerators such as GPUs and FPGAs. However, the current tensor library implementation mainly requires programmers to manually design low-level implementation and optimize from the algorithm, architecture, and compilation perspectives. Such a manual development process often takes months or even years, which falls far behind the rapid evolution of the application algorithms.
Size Zheng 0001, Yun Liang 0001, Shuo Wang 0009, Renze Chen, Kaiwen Sheng
ASPLOS1
2020 SuSy: A Programming Model for Productive Construction of High-Performance Systolic Arrays on FPGAs
abstract
Systolic algorithms are one of the killer applications on spatial architectures such as FPGAs and CGRAs. However, it requires a tremendous amount of human effort to design and implement a high-performance systolic array for a given algorithm using the traditional RTL-based methodology. On the other hand, existing high-level synthesis (HLS) tools either (1) force the programmers to do "micro-coding" where too many optimizations must be carried out through tedious code restructuring and insertion of vendor-specific pragmas, or (2) give them too little control to influence a push-button compilation flow to achieve high quality of results.
Yi-Hsiang Lai, Hongbo Rong, Size Zheng 0001, Xiuping Cui, Yunshan Jia, Jie Wang 0022, Brendan Sullivan, Zhiru Zhang, Yun Liang 0001, Youhui Zhang, Jason Cong, Nithin George, Christopher J. Hughes, Pradeep Dubey
ICCAD3