Longkai Cheng

dblp:392/9996 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0000-1679-8254ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 79% Deep learning architectures and training · 16% Language models and text generation · 5%
Network and information security
1 paper
Hardware security and side channels · 100%
Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
1.922026
EDSD: Entropy-Driven Design for Faster Speculative Decoding · ACL (1) 2026
HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference acceleration · EMNLP 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
1.012026
EDSD: Entropy-Driven Design for Faster Speculative Decoding · ACL (1) 2026
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference acceleration · EMNLP 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference acceleration · EMNLP 2025
Hardware security and side channels › trusted execution environments
ARM TrustZone
0.912025
SmartZone: Runtime Support for Secure and Efficient On-Device Inference on ARM TrustZone · IEEE Trans. Computers 2025
Hardware security and side channels
trusted execution environments
0.912025
SmartZone: Runtime Support for Secure and Efficient On-Device Inference on ARM TrustZone · IEEE Trans. Computers 2025
Machine learning › Efficient and distributed learning
draft model training
0.312026
EDSD: Entropy-Driven Design for Faster Speculative Decoding · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model
0.312025
SmartZone: Runtime Support for Secure and Efficient On-Device Inference on ARM TrustZone · IEEE Trans. Computers 2025
Machine learning › Efficient and distributed learning
model inference
0.312025
SmartZone: Runtime Support for Secure and Efficient On-Device Inference on ARM TrustZone · IEEE Trans. Computers 2025

Methods — techniques the papers use, named apart from their topics

secure memory management · 2.6multithreading · 1.7speculative decoding · 1.0entropy-driven alignment · 1.0top-k routing · 0.9post-training calibration · 0.9multi-threading · 0.9
YearPublicationVenuePosition
2026 EDSD: Entropy-Driven Design for Faster Speculative Decoding
abstract
Speculative decoding has emerged as a promising paradigm for accelerating large language model inference by leveraging a lightweight draft model to generate multiple candidate tokens.However, existing methods often incur substantial training overhead to mitigate information misalignment between autoregressive draft model training and decoding.To address this challenge, we propose EDSD, an Entropy-Driven Speculative Decoding framework that uses entropy as a unified, interpretable signal for both draft model training and architectural design.EDSD drives the draft model to progressively align with the target model in an easy-to-hard manner while establishing tokenlevel alignment as a dominant design principle.Extensive experiments on seven LLMs demonstrate that EDSD improves training efficiency by 24.8%, increases the average acceptance length by 4.0%, and achieves a 4.1% speedup compared to state-of-the-art methods.Furthermore, EDSD improves robustness to system prompt variations by more than 5×.Our findings establish entropy-driven alignment as an effective and principled foundation for efficient speculative decoding.We make our draft model weights available at https://github.com/KerwinKai/EDSD.
Longkai Cheng, Ximing Wang, Jiangcai Zhu, Kailai Shao, Haixiang Hu
ACL (1)1
2025 HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference acceleration
abstract
Mixture of Experts (MoE) architectures have emerged as a promising paradigm for scaling model capacity through top-k routing mechanisms.Although reducing the number of activated experts inherently enables inference acceleration, this efficiency gain typically comes at the cost of significant performance degradation.To address this trade-off between efficiency and performance, we propose Hook-MoE, a plug-and-play single-layer compensation framework that effectively restores performance using only a small post-training calibration set.Our method strategically inserts a lightweight trainable Hook module immediately preceding selected transformer blocks.Comprehensive evaluations on four popular MoE models, with an average performance degradation of only 2.5% across various benchmarks, our method reduces the number of activated experts by more than 50% and achieves a 1.42× inference speed-up during the prefill stage.Through systematic analysis, we further reveal that the upper layers require fewer active experts, offering actionable insights for refining dynamic expert selection strategies and enhancing the overall efficiency of MoE models.We make our code available at https://github.com/KerwinKai/HookMoE.
Longkai Cheng, Along He, Mulin Li, Xueshuo Xie, Tao Li 0022
EMNLP1
2025 SmartZone: Runtime Support for Secure and Efficient On-Device Inference on ARM TrustZone
abstract
On-device inference is a burgeoning paradigm that performs model inference locally on end devices, allowing private data to remain local. ARM TrustZone as a widely supported trusted execution environment has been applied to provide confidentiality protection for on-device inference. However, with the rise of large-scale models like large language models (LLMs), TrustZone-based on-device inference faces challenges in migration difficulties and inefficient execution. The rudimentary TEE OS on TrustZone lacks both the inference runtime needed for building models and the parallel support necessary to accelerate inference. Moreover, the limited secure memory resources on end devices further constrain the model size and degrade performance. In this paper, we propose SmartZone to provide runtime support for secure and efficient on-device inference on TrustZone. SmartZone consists three main components: (1) a trusted inference-oriented operator set, providing the underlying mechanisms adapted to the TrustZone’s execution mode for trusted inference of DNN models and LLMs. (2) the proactive multi-threading parallel support, which increases the number of CPU cores in the secure state via cross-world thread collaboration to achieve parallelism, and (3) the on-demand secure memory management method, which statically allocates the appropriate secure memory size based on pre-execution resource analysis. We implement a prototype of SmartZone on the Raspberry Pi 3B+ board and evaluate it on four well-known DNN models and llama2 LLM. Extensive experimental results show that SmartZone provides end-to-end protection for on-device inference while maintaining excellent performance. Compared to the origin trusted inference, SmartZone accelerates the inference speed by up to 4.26× and reduces energy consumption by 65.81%.
Zhaolong Jian, Qiankun Dong, Longkai Cheng, Xueshuo Xie, Tao Li 0022
IEEE Trans. Computers4