Tielong Liu

dblp:357/3668 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Boosting the Performance of Tree-Based Speculative Decoding of LLMs on FPGAs
abstract
As an efficient alternative to autoregressive decoding, tree-based speculative decoding (SD) has been widely adopted to accelerate LLM inference on GPUs. However, due to the notable disparity in compute power and memory bandwidth, we observe that a specific target-draft model pair with a proper decoding configuration, despite demonstrating significant performance gains on GPUs, often fails to maintain its efficacy on FPGAs, and may even underperform the standard autoregressive decoding approachIn this paper, we propose an analytical framework to revive the performance of tree-based speculative decoding on FPGAs. We introduce effective performance, a roofline-based metric designed to: 1) assess whether a specific target-draft model pair can benefit from SD for the given FPGA platform, and 2) determine the optimal decoding configuration to achieve peak performance when SD is applicable. We also propose a prior-score-based search strategy to identify the optimal tree structure for a preset number of nodes, further enhancing the performance. We evaluate our method on AMD FPGA platforms using two state-of-the-art SD algorithms: LongSpec and EAGLE-3. Our approach demonstrates a speedup of 2.54-3.89× over autoregressive decoding.
Tielong Liu, Gang Li 0015, Zitao Mo, Minnan Pei, Jian Cheng 0001
DATE1
2026 APEX: Integer-only Non-linear Function Approximation for Efficient Cross-Modal Inference
Peihuan Ni, Zitao Mo, Tielong Liu, Hongli Wen, Minnan Pei, Junwen Si, Weifan Guan, Peisong Wang 0001, Qinghao Hu 0001, Gang Li 0015, Jian Cheng 0001
DATE3
2026 Towards efficient and accurate spiking neural networks via adaptive bit allocation
abstract
Multi-bit spiking neural networks (SNNs) have recently become a heated research spot, pursuing energy-efficient and high-accurate AI. However, with more bits involved, the associated memory and computation demands escalate to the point where the performance improvements become disproportionate. Based on the insight that different layers demonstrate different importance and extra bits could be wasted and interfering, this paper presents an adaptive bit allocation strategy for direct-trained SNNs, achieving fine-grained layer-wise allocation of memory and computation resources. Thus, SNN's efficiency and accuracy can be improved. Specifically, we parametrize the temporal lengths and the bit widths of weights and spikes, and make them learnable and controllable through gradients. To address the challenges caused by changeable bit widths and temporal lengths, we propose the refined spiking neuron, which can handle different temporal lengths, enable the derivation of gradients for temporal lengths, and suit spike quantization better. In addition, we theoretically formulate the step-size mismatch problem of learnable bit widths, which may incur severe quantization errors to SNN, and accordingly propose the step-size renewal mechanism to alleviate this issue. Experiments on various datasets, including the static CIFAR and ImageNet datasets and the dynamic CIFAR-DVS and DVS-GESTURE datasets, demonstrate that our methods can reduce the overall memory and computation cost while achieving higher accuracy. Particularly, our SEWResNet-34 can achieve a 2.69 % accuracy gain and 4.16 × lower bit budgets over the advanced baseline work on ImageNet. This work is open-sourced at this link.
Xingting Yao, Qinghao Hu 0001, Tielong Liu, Gang Li 0015, Peisong Wang 0001, Jian Cheng 0001
Neural Networks4
2026 MATA: A Memory-Efficient Attention Accelerator for LLMs Exploiting Look-Back KV Cache Pruning
abstract
Transformer-based Large Language Models (LLMs) have sparked a new wave of AI applications. However, their large computational complexity and memory footprint pose significant challenges for real-world deployment. Although dedicated transformer accelerators have been widely explored, we observe that they are unefficient for decoder-only LLMs that feature autoregressive computations with KV Cache. Our in-depth analysis reveals that DRAM accesses induced by the KV Cache dominate the overall attention process. To address this issue, we propose aMemory-efficientATtentionAccelerator (MATA) for LLMs through algorithm and hardware co-design. Specifically,at the algorithm level, to mitigate the overhead caused by the linear increase of KV Cache, we propose a post-training Look-Back pruning method. It dynamically discards unimportant tokens through a comprehensive scoring scheme, thereby restricting KV Cache to a constant volume.At the hardware level, to identify important tokens with low latency, we design a Slice Top-K (STK) engine that can complete top-k-based sorting withO(N) time complexity. Moreover, we present the Adaptive Dataflow, which adaptively performs different inference phases of LLMs, thus significantly enhancing the PE array utilization. On average, our MATA can achieve speedups of 3.56×, 2.23× and 2.04×, 1.56× energy savings over two state-of-the-art transformer accelerators SpAtten and FACT, respectively.
Gang Li 0015, Tielong Liu, Zitao Mo, Xiaoyao Liang, Jian Cheng 0001
IEEE Trans. Computers3
2025 LISLLM: Long Context Inference of Large Language Models with Short KV Cache
abstract
Large Language Models (LLMs) have demonstrated remarkable performance across various language tasks. However, in the process of generating long texts, the linearly increasing keyvalue (KV) cache imposes a large volume of memory footprint on HBM, resulting in significant time and energy consumption during inference. By retaining only the initial and recent tokens, existing methods ignore the intermediate tokens and the effect of the distance between current tokens and those in the KV cache, which causes significant accuracy degradation and limits the potential for KV cache compression. To overcome the above problems, we propose LISLLM, a KV cache compression mechanism that takes both the difference in importance and the distance between tokens into consideration for efficient inference of LLMs in long-context settings. Compared to the state-of-theart method, our method achieves compression ratios of up to$1.91 \times$and speedup of$1.20 \times$with better accuracy.
Tielong Liu, Gang Li 0015, Zitao Mo, Xingting Yao, Jian Cheng 0001
ICPADS1