Tuowei Wang

dblp:326/0527 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0006-8272-8151ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
abstract
Deploying Large Language Models (LLMs) on mobile devices faces the challenge of insufficient performance in smaller models and excessive resource consumption in larger ones. This paper highlights that mobile Neural Processing Units (NPUs) have underutilized computational resources, particularly their matrix multiplication units, during typical LLM inference. To leverage this wasted compute capacity, we propose applying parallel test-time scaling techniques on mobile NPUs to enhance the performance of smaller LLMs. However, this approach confronts inherent NPU challenges, including inadequate hardware support for fine-grained quantization and low efficiency in general-purpose computations. To overcome these, we introduce two key techniques: a hardware-aware tile quantization scheme that aligns group quantization with NPU memory access patterns, and efficient LUT-based replacements for complex operations such as Softmax and dequantization. We design and implement an end-to-end inference system that leverages the NPU's compute capability to support test-time scaling on Qualcomm Snapdragon platforms. Experiments show our approach brings significant speedups: up to 19.0× for mixed-precision GEMM and 2.2× for Softmax. More importantly, we demonstrate that smaller models using test-time scaling can match or exceed the accuracy of larger models, achieving a new performance-cost Pareto frontier.
Zixu Hao, Jianyu Wei, Tuowei Wang, Minxing Huang, Huiqiang Jiang, Shiqi Jiang 0002, Ting Cao 0003, Ju Ren 0001
EuroSys3
2025 Neuralink: Fast on-Device LLM Inference with Neuron Co-Activation Linking
abstract
Large Language Models (LLMs) have achieved remarkable success across various domains, yet deploying them on mobile devices remains an arduous challenge due to their extensive computational and memory demands.While lightweight LLMs have been developed to fit mobile environments, they suffer from degraded model accuracy.In contrast, sparsitybased techniques minimize DRAM usage by selectively transferring only relevant neurons to DRAM while retaining the full model in external storage, such as flash.However, such approaches are critically limited by numerous I/O operations, particularly on smartphones with severe IOPS constraints.In this paper, we propose Neuralink, a novel approach that accelerates LLM inference on smartphones by optimizing neuron placement in flash memory.Neuralink leverages the concept of Neuron Co-Activation, where neurons frequently activated together are linked to facilitate continuous read access and optimize I/O efficiency.Our approach incorporates a two-stage solution: an offline stage that reorganizes neuron placement based on co-activation patterns, and an online stage that employs tailored data access and caching strategies to align well with hardware characteristics.Evaluations conducted on a variety of smartphones and LLMs demonstrate that Neuralink achieves on average 1.49× improvements in end-to-end latency compared to the state-of-the-art.As the first solution to optimize storage placement under sparsity, Neuralink explores a new * Both authors contributed equally to this research.
Tuowei Wang, Ruwen Fan, Minxing Huang, Zixu Hao, Kun Li 0016, Ting Cao 0003, Youyou Lu, Yaoxue Zhang, Ju Ren 0001
ASPLOS (3)1
2025 FedAF: Alignment-Augmented Fusion for Federated Multimodal Learning with Small Labels
abstract
Federated multimodal learning is an emerging advancement in artificial intelligence, enabling the integration of data from diverse modalities while preserving data privacy. However, limited labeled data and modality heterogeneity on the clients pose significant challenges for effective federated multimodal model training. To address these challenges, this paper introduces FedAF, a novel alignment-augmented fusion framework tailored for federated multimodal learning. FedAF extracts unbiased and complementary information from multiple modalities with small data, enabling effective modality fusion and feature alignment for improving system performance. The framework introduces a three-stage strategy. First, FedAF utilizes labeled data to create unbiased anchor points, addressing disparities in client feature distributions. Second, FedAF employs a weighted enhancement contrast fusion scheme to improve feature clustering and reduce feature overlap. Finally, a multimodal semisupervised algorithm mitigates data heterogeneity and overfitting. Extensive experiments demonstrate that FedAF significantly outperforms baseline methods, showcasing its effectiveness in federated multimodal learning scenarios.
Guanbo Wang, Yongheng Deng, Yingjun Wu, Xinyi Li 0005, Tuowei Wang, Yaoxue Zhang, Ju Ren 0001
IWQoS7
2025 JENGA: Enhancing LLM Long-Context Fine-tuning with Contextual Token Sparsity
Tuowei Wang, Kun Li 0016, Ting Cao 0003, Ju Ren 0001, Yaoxue Zhang
USENIX ATC1
2024 Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
abstract
The adaptation of pre-trained large language models (LLMs) to diverse downstream tasks via fine-tuning is critical for numerous applications. However, the inefficiency of parameterefficient fine-tuning (PEFT) techniques presents significant challenges in terms of time investments and operational costs. In this paper, we first introduce a nuanced form of sparsity, termed Shadowy Sparsity, which is distinctive in fine-tuning and has not been adequately addressed for acceleration. Under Shadowy Sparsity, we propose Long Exposure1, an efficient system to accelerate PEFT for LLMs. Long Exposure comprises three key components: Shadowy-sparsity Exposer employs a prolonged sensing range to capture more sparsity details under shadowy sparsity; Sequence-oriented Predictor provides efficient yet accurate predictions to handle large sequence inputs and constantly-evolving parameters; and Dynamic-aware Operator facilitates more structured computational patterns and coalesced memory accesses, addressing dynamic sparse operations. Extensive evaluations show that Long Exposure outperforms state-of-the-arts with up to a $2.49 \times$ speedup in end-to-end fine-tuning, offering promising advancements in accelerating PEFT for LLMs.1Long Exposure is available at https://github.com/HPHEX/LongExposure.
Tuowei Wang, Kun Li 0016, Zixu Hao, Donglin Bai, Ju Ren 0001, Yaoxue Zhang, Ting Cao 0003, Mao Yang 0004
SC1
2024 GAuV: A Graph-Based Automated Verification Framework for Perfect Semi-Honest Security of Multiparty Computation Protocols
abstract
Proving the security of a Multiparty Computation (MPC) protocol is a difficult task. Under the current simulation-based definition of MPC, a security proof consists of a simulator, which is usually specific to the concrete protocol and requires to be manually constructed, together with a theoretical analysis of the output distribution of the simulator and corrupted parties’ views in the real world. This presents an obstacle in verifying the security of a given MPC protocol. Moreover, an instance of a secure MPC protocol can easily lose its security guarantee due to careless implementation, and such a security issue is hard to detect in practice.(p)(/p)In this work, we propose a general automated framework to verify the perfect security of instances of MPC protocols against the semi-honest adversary. Our framework has perfect soundness: any protocol that is proven secure under our framework is also secure under the simulation-based definition of MPC. We demonstrate the completeness of our framework by showing that for any instance of the well-known BGW protocol, our framework can prove its security for every corrupted party set with polynomial time. Unlike prior work that only focuses on black-box privacy which requires the outputs of corrupted parties to contain no information about the inputs of the honest parties, our framework may potentially be used to prove the security of arbitrary MPC protocols. (p)(/p)We implement our framework as a prototype. The evaluation shows that our prototype automatically proves the perfect semi-honest security of BGW protocols and B2A (binary to arithmetic) conversion protocols in reasonable durations.
Xingyu Xie, Tuowei Wang, Shizhen Xu
SP4
2023 EINNET: Optimizing Tensor Programs with Derivation-Based Transformations
Liyan Zheng 0001, Haojie Wang 0004, Jidong Zhai, Muyan Hu, Zixuan Ma, Tuowei Wang, Shuhong Huang, Xupeng Miao, Shizhi Tang, Kezhao Huang
OSDI6
2023 FedINC: An Exemplar-Free Continual Federated Learning Framework with Small Labeled Data
abstract
Federated learning (FL) has shown great promise for privacy-preserving learning by enabling collaborative training on decentralized clients. However, in realistic FL scenarios, clients often collect new data continuously, join or exit learning dynamically. As a result, the global model tends to forget old knowledge while learning new knowledge. Meanwhile, labeling the continuously arriving data in real-time is usually challenging. Therefore, the catastrophic forgetting problem intertwined with the label deficiency issue poses significant challenges for both learning new knowledge and consolidating old knowledge. To address these challenges, we develop a novel exemplar-free continual federated learning framework named FedINC, to learn a global incremental model with limited labeled data. We begin by excavating the cause of catastrophic forgetting via in-depth empirical studies. Based on that, we introduce targeted mechanisms for FedINC, including a hybrid contrastive learning mechanism to efficiently learn new knowledge with limited labeled data, a plastic feature regularization mechanism to preserve old task's representation space, a prototype-guided regularization mechanism to mitigate feature overlap between old and new classes while aligning the features of non-iid clients, and a prototype evolution mechanism for flexible and efficient incremental classification. Extensive experiments demonstrate the superior performance of FedINC in terms of both convergence speed and accuracy of the global model.
Yongheng Deng, Sheng Yue 0001, Tuowei Wang, Guanbo Wang, Ju Ren 0001, Yaoxue Zhang
SenSys3
2023 Optimizing DNNs With Partially Equivalent Transformations and Automated Corrections
abstract
Deep neural network (DNN) applications are typically represented by tensor programs. To boost the performance of DNN computations, existing works adopt fully equivalent transformations for tensor program optimization by guaranteeing the equivalence on each element of tensors. However, as there are thousands of elements in a tensor, such optimization misses the opportunities that allow the in-equivalence of minority elements. In this work, we proposePet, the first work that introduces partially equivalent transformations to optimize tensor programs. To maintain the functional equivalence of tensor programs,Petautomatically finds and corrects the in-equivalent positions by leveraging the multi-linearity of DNN computations.Petfurther uses a mutation manager to improve search efficiency. Evaluation results show thatPetcan achieve up to 1.98$\times$and 2.20$\times$speedups on NVIDIA Tesla A100 and V100 respectively compared with existing DNN frameworks by introducing new optimization opportunities of partially equivalent transformations.
Haojie Wang 0004, Jidong Zhai, Mingyu Gao 0001, Feng Zhang 0007, Tuowei Wang, Zixuan Ma, Shizhi Tang, Liyan Zheng 0001, Kaiyuan Rong, Yuanyong Chen
IEEE Trans. Computers5