Zitao Mo

dblp:249/5473 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0001-8623-5465ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 6 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Boosting the Performance of Tree-Based Speculative Decoding of LLMs on FPGAs
abstract
As an efficient alternative to autoregressive decoding, tree-based speculative decoding (SD) has been widely adopted to accelerate LLM inference on GPUs. However, due to the notable disparity in compute power and memory bandwidth, we observe that a specific target-draft model pair with a proper decoding configuration, despite demonstrating significant performance gains on GPUs, often fails to maintain its efficacy on FPGAs, and may even underperform the standard autoregressive decoding approachIn this paper, we propose an analytical framework to revive the performance of tree-based speculative decoding on FPGAs. We introduce effective performance, a roofline-based metric designed to: 1) assess whether a specific target-draft model pair can benefit from SD for the given FPGA platform, and 2) determine the optimal decoding configuration to achieve peak performance when SD is applicable. We also propose a prior-score-based search strategy to identify the optimal tree structure for a preset number of nodes, further enhancing the performance. We evaluate our method on AMD FPGA platforms using two state-of-the-art SD algorithms: LongSpec and EAGLE-3. Our approach demonstrates a speedup of 2.54-3.89× over autoregressive decoding.
Tielong Liu, Gang Li 0015, Zitao Mo, Minnan Pei, Jian Cheng 0001
DATE3
2026 APEX: Integer-only Non-linear Function Approximation for Efficient Cross-Modal Inference
Peihuan Ni, Zitao Mo, Tielong Liu, Hongli Wen, Minnan Pei, Junwen Si, Weifan Guan, Peisong Wang 0001, Qinghao Hu 0001, Gang Li 0015, Jian Cheng 0001
DATE2
2026 MATA: A Memory-Efficient Attention Accelerator for LLMs Exploiting Look-Back KV Cache Pruning
abstract
Transformer-based Large Language Models (LLMs) have sparked a new wave of AI applications. However, their large computational complexity and memory footprint pose significant challenges for real-world deployment. Although dedicated transformer accelerators have been widely explored, we observe that they are unefficient for decoder-only LLMs that feature autoregressive computations with KV Cache. Our in-depth analysis reveals that DRAM accesses induced by the KV Cache dominate the overall attention process. To address this issue, we propose aMemory-efficientATtentionAccelerator (MATA) for LLMs through algorithm and hardware co-design. Specifically,at the algorithm level, to mitigate the overhead caused by the linear increase of KV Cache, we propose a post-training Look-Back pruning method. It dynamically discards unimportant tokens through a comprehensive scoring scheme, thereby restricting KV Cache to a constant volume.At the hardware level, to identify important tokens with low latency, we design a Slice Top-K (STK) engine that can complete top-k-based sorting withO(N) time complexity. Moreover, we present the Adaptive Dataflow, which adaptively performs different inference phases of LLMs, thus significantly enhancing the PE array utilization. On average, our MATA can achieve speedups of 3.56×, 2.23× and 2.04×, 1.56× energy savings over two state-of-the-art transformer accelerators SpAtten and FACT, respectively.
Gang Li 0015, Tielong Liu, Zitao Mo, Xiaoyao Liang, Jian Cheng 0001
IEEE Trans. Computers4
2025 LISLLM: Long Context Inference of Large Language Models with Short KV Cache
abstract
Large Language Models (LLMs) have demonstrated remarkable performance across various language tasks. However, in the process of generating long texts, the linearly increasing keyvalue (KV) cache imposes a large volume of memory footprint on HBM, resulting in significant time and energy consumption during inference. By retaining only the initial and recent tokens, existing methods ignore the intermediate tokens and the effect of the distance between current tokens and those in the KV cache, which causes significant accuracy degradation and limits the potential for KV cache compression. To overcome the above problems, we propose LISLLM, a KV cache compression mechanism that takes both the difference in importance and the distance between tokens into consideration for efficient inference of LLMs in long-context settings. Compared to the state-of-theart method, our method achieves compression ratios of up to$1.91 \times$and speedup of$1.20 \times$with better accuracy.
Tielong Liu, Gang Li 0015, Zitao Mo, Xingting Yao, Jian Cheng 0001
ICPADS4
2025 GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing
abstract
3D Gaussian Splatting (3DGS) has emerged as a leading neural rendering technique for high-fidelity view synthesis, prompting the development of dedicated 3DGS accelerators for resource-constrained platforms.The conventional decoupled preprocessing-rendering dataflow in existing accelerators has two major limitations: 1) a
Minnan Pei, Gang Li 0015, Junwen Si, Zitao Mo, Peisong Wang 0001, Zhuoran Song, Xiaoyao Liang, Jian Cheng 0001
MICRO5
2024 MEGA: A Memory-Efficient GNN Accelerator Exploiting Degree-Aware Mixed-Precision Quantization
abstract
Graph Neural Networks (GNNs) are becoming a promising technique in various domains due to their excellent capabilities in modeling non-Euclidean data. Although a spectrum of accelerators has been proposed to accelerate the inference of GNNs, our analysis demonstrates that the latency and energy consumption induced by DRAM access still significantly impedes the improvement of performance and energy efficiency. To address this issue, we propose a Memory - Efficient GNN Accelerator (MEGA) through algorithm and hardware co-design in this work. Specifically, at the algorithm level, through an in-depth analysis of the node property, we observe that the data-independent quantization in previous works is not optimal in terms of accuracy and memory efficiency. This motivates us to propose the Degree-Aware mixed-precision quantization method, in which a proper bitwidth is learned and allocated to a node according to its in-degree to compress GNNs as much as possible while maintaining accuracy. At the hardware level, we employ a heterogeneous architecture design in which the aggregation and combination phases are implemented separately with different dataflows. In order to boost the performance and energy efficiency, we also present an Adaptive-Package format to alleviate the storage overhead caused by the fine-grained bitwidth and diverse sparsity, and a Condense-Edge scheduling method to enhance the data locality and further alleviate the access irregularity induced by the extremely sparse adjacency matrix in the graph. We implement our MEGA accelerator in a 28nm technology node. Extensive experiments demonstrate that MEGA can achieve an average speedup of 38.3 ×, 7.1 ×, 4.0 ×, 3.6× and 47.6 ×, 7.2 ×, 5.4 ×, 4.5 × energy savings over four state-of-the-art GNN accelerators, HyGCN, GCNAX, GROW, and SGCN, respectively, while retaining task accuracy.
Fanrong Li, Gang Li 0015, Zejian Liu, Zitao Mo, Qinghao Hu 0001, Xiaoyao Liang, Jian Cheng 0001
HPCA5
2023 $\rm A^2Q$: Aggregation-Aware Quantization for Graph Neural Networks
Fanrong Li, Zitao Mo, Qinghao Hu 0001, Gang Li 0015, Zejian Liu, Xiaoyao Liang, Jian Cheng 0001
ICLR3
2022 GLIF: A Unified Gated Leaky Integrate-and-Fire Neuron for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have been studied over decades to incorporate their biological plausibility and leverage their promising energy efficiency. Throughout existing SNNs, the leaky integrate-and-fire (LIF) model is commonly adopted to formulate the spiking neuron and evolves into numerous variants with different biological features. However, most LIF-based neurons support only single biological feature in different neuronal behaviors, limiting their expressiveness and neuronal dynamic diversity. In this paper, we propose GLIF, a unified spiking neuron, to fuse different bio-features in different neuronal behaviors, enlarging the representation space of spiking neurons. In GLIF, gating factors, which are exploited to determine the proportion of the fused bio-features, are learnable during training. Combining all learnable membrane-related parameters, our method can make spiking neurons different and constantly changing, thus increasing the heterogeneity and adaptivity of spiking neurons. Extensive experiments on a variety of datasets demonstrate that our method obtains superior performance compared with other SNNs by simply changing their neuronal formulations to GLIF. In particular, we train a spiking ResNet-19 with GLIF and achieve $77.35\%$ top-1 accuracy with six time steps on CIFAR-100, which has advanced the state-of-the-art. Codes are available at https://github.com/Ikarosy/Gated-LIF.
Xingting Yao, Fanrong Li, Zitao Mo, Jian Cheng 0001
NeurIPS3
2020 ProxyBNN: Learning Binarized Neural Networks via Proxy Matrices
Zitao Mo, Ke Cheng 0002, Qinghao Hu 0001, Peisong Wang 0001, Qingshan Liu 0001, Jian Cheng 0001
ECCV (3)2
2020 Ladder Pyramid Networks For Single Image Super-Resolution
abstract
Benefiting from the powerful representation capability of convolutional neural networks, the performance of single image super-resolution (SISR) has been substantially improved in recent years. However, many current CNN-based methods are computation-intensive because of large-size intermediate feature maps and inefficient convolutions. To resolve these problems, we propose Ladder Pyramid Network (LPN) for single image super-resolution. Firstly, we use strided convolution to reduce the size of the intermediate feature maps and thus reducing computation burden. In order to better balance the effectiveness and efficiency, we propose Ladder Pyramid Module to gradually fuse hierarchical features to enhance performance. Secondly, lightweight convolution block similar to Inverted Residual Module of Mobilenet-v2 was introduced into SISR, with which we build the network backbone and ladder feature pyramid. Experimental results demonstrate that the proposed Ladder Pyramid Network can achieve comparable or better performance than previous lightweight networks while reducing the amount of computation.
Zitao Mo, Gang Li 0015, Jian Chen 0001
ICIP1
2020 FSA: A Fine-Grained Systolic Accelerator for Sparse CNNs
abstract
Sparsity, as an intrinsic property of convolutional neural networks (CNNs), has been widely employed for hardware acceleration, and many customized accelerators tailored for sparse weights or activations have been proposed in these years. However, the irregular sparse patterns introduced by both weights and activations are much more challenging for efficient computation. For example, due to the issues of access contention, workload imbalance, and tile fragmentation, the state-of-the-art sparse accelerator SCNN fails to fully leverage the benefits of sparsity, leading to nonoptimal results for both speedup and energy efficiency. In this article, we propose an efficient sparse CNN accelerator for both weights and activations, namely finegrained systolic accelerator (FSA), which jointly optimizes both hardware dataflow and software partitioning and scheduling strategy. Specifically, to deal with the access contentions problem, we present a fine-grained systolic dataflow, in which the activations move rhythmically along the horizontal processing element array while the weights are fed into the array in a fine-grained order. We then propose a hybrid network partitioning strategy that sets different partitioning strategies for different layers to balance the workload and alleviate the fragmentation problem caused by both sparse weights and activations. Finally, we present a scheduling search strategy to find the optimized schedules for neural networks, which can further improve energy efficiency. Extensive evaluations show that the proposed FSA consistently outperforms SCNN over AlexNet, VGGNet, GoogLeNet, and ResNet with an average speedup of 1.74× and up to 13.86× energy efficiency.
Fanrong Li, Gang Li 0015, Zitao Mo, Jian Cheng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 ODE-Inspired Network Design for Single Image Super-Resolution
abstract
Single image super-resolution, as a high dimensional structured prediction problem, aims to characterize fine-grain information given a low-resolution sample. Recent advances in convolutional neural networks are introduced into super-resolution and push forward progress in this field. Current studies have achieved impressive performance by manually designing deep residual neural networks but overly relies on practical experience. In this paper, we propose to adopt an ordinary differential equation (ODE)-inspired design scheme for single image super-resolution, which have brought us a new understanding of ResNet in classification problems. Not only is it interpretable for super-resolution but it provides a reliable guideline on network designs. By casting the numerical schemes in ODE as blueprints, we derive two types of network structures: LF-block and RK-block, which correspond to the Leapfrog method and Runge-Kutta method in numerical ordinary differential equations. We evaluate our models on benchmark datasets, and the results show that our methods surpass the state-of-the-arts while keeping comparable parameters and operations.
Zitao Mo, Peisong Wang 0001, Yang Liu 0021, Mingyuan Yang, Jian Cheng 0001
CVPR2