EDBT 2026 Demo / reviewers in the wild / expert
Jianyu Wei
dblp:133/8724
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling LLM Test-Time Compute with Mobile NPU on SmartphonesabstractDeploying Large Language Models (LLMs) on mobile devices faces the challenge of insufficient performance in smaller models and excessive resource consumption in larger ones. This paper highlights that mobile Neural Processing Units (NPUs) have underutilized computational resources, particularly their matrix multiplication units, during typical LLM inference. To leverage this wasted compute capacity, we propose applying parallel test-time scaling techniques on mobile NPUs to enhance the performance of smaller LLMs. However, this approach confronts inherent NPU challenges, including inadequate hardware support for fine-grained quantization and low efficiency in general-purpose computations. To overcome these, we introduce two key techniques: a hardware-aware tile quantization scheme that aligns group quantization with NPU memory access patterns, and efficient LUT-based replacements for complex operations such as Softmax and dequantization. We design and implement an end-to-end inference system that leverages the NPU's compute capability to support test-time scaling on Qualcomm Snapdragon platforms. Experiments show our approach brings significant speedups: up to 19.0× for mixed-precision GEMM and 2.2× for Softmax. More importantly, we demonstrate that smaller models using test-time scaling can match or exceed the accuracy of larger models, achieving a new performance-cost Pareto frontier. Zixu Hao, Jianyu Wei, Tuowei Wang, Minxing Huang, Huiqiang Jiang, Shiqi Jiang 0002, Ting Cao 0003, Ju Ren 0001 |
EuroSys | 2 |
| 2026 | Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge DevicesabstractLarge language models (LLMs) are increasingly deployed on edge devices. To meet strict resource constraints, real-world deployment has pushed LLM quantization from 8-bit to 4-bit, 2-bit, and now 1.58-bit. Combined with lookup table (LUT)-based inference, CPUs run these ultra-low-bit LLMs even faster than NPUs, opening new opportunities for ubiquitous on-device intelligence. Weijun Wang 0001, Jianyu Wei, Ting Cao 0003, Yunxin Liu 0001 |
MobiSys | 4 |
| 2026 | Cognitive Jamming-Aided UAV Multi-User Covert CommunicationabstractIntroducing friendly jammer to unmanned aerial vehicle (UAV) covert communication could further enhance the covert performance. However, due to the additional communication overhead required for information synchronization (IS) among cooperating parties, the security risk is further elevated. It is noteworthy that cognitive jamming (CJ) could sense the cooperating node signals without IS links. Therefore, in this paper, we propose the cooperative UAVs covert communication scheme assisted by the CJ that maximize the minimum average covert rate (MACR). Specifically, the UAV (Alice) tries to transmit its private message to legitimate users under the surveillance of the Wardens. Since the Warden is in passive surveillance, the practical location uncertainty of the Warden is considered, and we further analyze the worst-case situation that derives the minimal detection error probability of the Wardens. Moreover, we maximize the MACR by alternatively optimizing the sense time ratio and trajectory of CJ, the transmit power and trajectory of the Alice. To expedite algorithmic convergence, we further introduce a principled initialization scheme. The simulation results reveal that the covert performance of the proposed algorithm outperforms the other schemes, especially, in the multi-warden scenario. Furthermore, compared with the full-time jammer and probabilistic jammer, the CJ can not only achieve a higher covert rate with less average jamming power, but also require no information synchronization link with Alice based on spectrum sensing, which promises considerable application prospects in real-world systems. Jianyu Wei, Yan Guo 0002, Haichao Wang 0001, Jiangchun Gu, Jiteng Liu, Guoru Ding |
IEEE J. Sel. Areas Commun. | 1 |
| 2026 | Fluid Antenna Array-Enabled AAV Covert Communications Against Active WardenabstractAutonomous aerial vehicle (AAV) covert communication could further enhance the quality and coverage of covert channels. However, in complex low-altitude environments, the air-to-ground (A2G) links are susceptible to the fading effects, which may degrade the performance of covert communication. In this paper, we investigate the fluid antenna (FA) enabled AAV covert communication, where a AAV equipped with FA serves multiple ground users in the presence of an active Warden. We aim to maximize the minimum average covert rate by jointly optimizing the beamforming vectors, AAV trajectory and fluid antenna positions. First, we derive closed-form expressions for the minimum detection error probability (MDEP) and the optimal detection threshold, accounting for uncertainties in both the noise variance and the self-interference channel coefficient. Secondly, we propose an alternating optimization algorithm subject to the covertness constraint, power constraint, and the Warden’s position uncertainty. Specifically, the original nonconvex problem is decomposed into tractable subproblems via the block coordinate descent, which could be solved successively by successive convex approximation, semidefinite relaxation, and Dinkelbach transformation. What’s more, a low-complexity algorithm is developed for the single-user scenario to improve the practical applicability of the proposed framework. Finally, simulation results validate the effectiveness of the proposed FA-AAV covert communication scheme. Moreover, compared to the fixed position antenna scheme, the FA-AAV could improve the covert performance, especially in strong channel fading environments, which is beneficial for practical application. Jianyu Wei, Yan Guo 0002, Haichao Wang 0001, Jiangchun Gu, Yunyang Zhang, Jiawei Yi, Xinliang Chen, Guoru Ding |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | Bitnet.cpp: Efficient Edge Inference for Ternary LLMsabstractJinheng Wang, Hansong Zhou, Ting Song, Shijie Cao, Yan Xia, Ting Cao, Jianyu Wei, Shuming Ma, Hongyu Wang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jinheng Wang, Hansong Zhou, Shijie Cao, Yan Xia 0005, Jianyu Wei, Shuming Ma, Furu Wei |
ACL (1) | 7 |
| 2025 | T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on EdgeabstractThe deployment of Large Language Models (LLMs) on edge devices is increasingly important to enhance on-device intelligence. Weight quantization is crucial for reducing the memory footprint of LLMs on devices. However, low-bit LLMs necessitate mixed precision matrix multiplication (mpGEMM) of low precision weights and high precision activations during inference. Existing systems, lacking native support for mpGEMM, resort to dequantize weights for high precision computation. Such an indirect way can lead to a significant inference overhead. Jianyu Wei, Shijie Cao, Ting Cao 0003, Lingxiao Ma, Lei Wang 0222, Yanyong Zhang, Mao Yang 0004 |
EuroSys | 1 |
| 2025 | LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceabstractLarge Language Model (LLM) inference becomes resource-intensive, prompting a shift toward low-bit model weights to reduce the memory footprint and improve efficiency.Such low-bit LLMs necessitate the mixed-precision matrix multiplication (mpGEMM), an important yet underexplored operation involving the multiplication of lower-precision weights with higher-precision activations.Off-theshelf hardware does not support this operation natively, leading to indirect, thus inefficient, dequantization-based implementations.In this paper, we study the lookup table (LUT)-based approach for mpGEMM and find that a conventional LUT implementation fails to achieve the promised gains.To unlock the full potential of LUT-based mpGEMM, we propose LUT Tensor Core, a softwarehardware co-design for low-bit LLM inference.LUT Tensor Core differentiates itself from conventional LUT designs through: 1) * Work is done during internship at Microsoft Research. Zhiwen Mo, Lei Wang 0222, Jianyu Wei, Zhichen Zeng 0002, Shijie Cao, Lingxiao Ma, Naifeng Jing, Ting Cao 0003, Jilong Xue, Fan Yang 0024, Mao Yang 0004 |
ISCA | 3 |
| 2025 | Demo: EdgeMind-OS: A Plug-and-Play Embodied Intelligence System for Real-Time On-Device DeploymentabstractBuilding an always-on, contextual AI assistant that proactively supports humans remains a central goal in Embodied AI—yet cloud-based pipelines struggle to meet due to delay, bandwidth, and privacy constraints. This demo presents EdgeMind-OS, a fully on-device intelligence system designed for embodied agents operating in real-world scenarios. Edge-Mind-OS features a hierarchical architecture combining a real-time StreamBrain, modular skill experts, and a dynamic scene-episode memory. Achieving up to 7.3× faster local processing, it enables low-latency, privacy-preserving, and plug-and-play deployment across tasks such as semantic navigation, spatial memory recall, and multimodal interaction. We demonstrate how EdgeMind-OS empowers a mobile robot with only basic locomotion capabilities to perform realtime, free-form user-robot interaction through autonomous perception, reasoning and action —without reliance on external cloud infrastructure. Jianyu Wei, Fucheng Jia, Liang Mi, Ruofei Ju, Xianye Wang, Yikai Zheng, Weijun Wang 0001, Shiqi Jiang 0002, Yunxin Liu 0001, Ting Cao 0003 |
MobiCom | 2 |
| 2024 | Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert InferenceabstractLarge language models (LLMs) based on transformers have made significant strides in recent years, the success of which is driven by scaling up their model size. Despite their high algorithmic performance, the computational and memory requirements of LLMs present unprecedented challenges. To tackle the high compute requirements of LLMs, the Mixture-ofExperts (MoE) architecture was introduced which is able to scale its model size without proportionally scaling up its computational requirements. Unfortunately, MoE’s high memory demands and dynamic activation of sparse experts restrict its applicability to real-world problems. Previous solutions that offload MoE’s memory-hungry expert parameters to CPU memory fall short because the latency to migrate activated experts from CPU to GPU incurs high performance overhead. Our proposed Pre-gated MoE system effectively tackles the compute and memory challenges of conventional MoE architectures using our algorithm-system codesign. Pre-gated MoE employs our novel pre-gating function which alleviates the dynamic nature of sparse expert activation, allowing our proposed system to address the large memory footprint of MoEs while also achieving high performance. We demonstrate that Pre-gated MoE is able to improve performance, reduce GPU memory consumption, while also maintaining the same level of model quality. These features allow our Pre-gated MoE system to cost-effectively deploy large-scale LLMs using just a single GPU with high performance. Ranggi Hwang, Jianyu Wei, Shijie Cao, Changho Hwang, Xiaohu Tang 0003, Ting Cao 0003, Mao Yang 0004 |
ISCA | 2 |
| 2023 | NN-Stretch: Automatic Neural Network Branching for Parallel Inference on Heterogeneous Multi-ProcessorsabstractMobile devices are increasingly equipped with heterogeneous multiprocessors, e.g., CPU + GPU + DSP. Yet existing Neural Network (NN) inference fails to fully utilize the computing power of the heterogeneous multi-processors due to the sequential structures of NN models. Towards this end, this paper proposes NN-Stretch, a new model adaption strategy, as well as the supporting system. It automatically branches a given model according to the processor architecture characteristics. Compared to other popular model adaption techniques such as model pruning that often sacrifices accuracy, NN-Stretch accelerates inference while preserving accuracy. Jianyu Wei, Ting Cao 0003, Shijie Cao, Shiqi Jiang 0002, Shaowei Fu, Mao Yang 0004, Yanyong Zhang, Yunxin Liu 0001 |
MobiSys | 1 |
| 2023 | Autonomous confrontation strategy learning evolution mechanism of unmanned system group under actual combat in the loop
Shiguang Hu, Binghan Lei, Jianyu Wei |
Comput. Commun. | 7 |
| 2021 | nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devicesabstractWith the recent trend of on-device deep learning, inference latency has become a crucial metric in running Deep Neural Network (DNN) models on various mobile and edge devices. To this end, latency prediction of DNN model inference is highly desirable for many tasks where measuring the latency on real devices is infeasible or too costly, such as searching for efficient DNN models with latency constraints from a huge model-design space. Yet it is very challenging and existing approaches fail to achieve a high accuracy of prediction, due to the varying model-inference latency caused by the runtime optimizations on diverse edge devices. Li Lyna Zhang, Shihao Han, Jianyu Wei, Ningxin Zheng, Ting Cao 0003, Yuqing Yang 0001, Yunxin Liu 0001 |
MobiSys | 3 |