Yang Wang 0053

dblp:w/YangWang53 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-7322-4062ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 9 since 2021Computer networks · 7 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021
YearPublicationVenuePosition
2026 LUT-LLM: Efficient Language Model Inference with Memory-based Computations on FPGAs
abstract
The rapid development of large language models (LLM) has greatly enhanced everyday applications. While many FPGA-based accelerators, with flexibility for fine-grained data control, exhibit superior speed and energy efficiency compared to GPUs, recent GPU-specific optimizations have diminished this advantage. When limited to arithmetic-based computation, FPGAs often underperform GPUs due to their comparatively fewer computational resources. To address this challenge, we exploit a key advantage of FPGAs over GPUs: abundant distributed on-chip memory embedded among computational units. We believe that shifting LLM inference from arithmetic-based to memory-based computations through table lookups can improve the efficiency on FPGAs to compete with GPUs. However, existing methods are inefficient or unable to scale and deploy language models due to algorithm and architecture design limitations. This paper introduces LUT-LLM1, the first FPGA accelerator that deploys 1B+ language model with memory-based computation, leveraging vector quantization. We construct a performance model, evaluate multiple quantization schemes, and identify the activation-weight vector co-quantization as the most effective approach. To support this scheme, LUT-LLM features (1) a bandwidth-aware parallel centroid search to reduce decoding latency, (2) efficient 2D table lookups, and (3) a spatial-temporal hybrid design to reduce data caching for a higher throughput table lookup. We develop a training recipe that converts existing models to support table lookups with high accuracy and prototype LUT-LLM for Qwen 3 1.7B model on the AMD V80 FPGA, reducing arithmetic operations by 4× and achieving a 1.10∼3.29× faster generation speed and a 3.05∼6.60× higher energy efficiency than GPUs.
Zifan He, Shengyu Ye, Rui Ma 0021, Yang Wang 0053, Jason Cong
FCCM4
2026 SuperBench: A Proactive Validation System for Improving Reliability of Cloud AI Infrastructure
abstract
Reliability in cloud AI infrastructure is crucial for cloud service providers, prompting the widespread use of hardware redundancies. However, these redundancies can inadvertently lead to hidden degradation, known as “gray failure”, for AI workloads, significantly affecting end-to-end performance and concealing performance issues, which complicates root cause analysis for failures and regressions. We introduce SuperBench, a proactive validation system for AI infrastructure that mitigates hidden degradation caused by hardware redundancies and enhances overall reliability. SuperBench features a comprehensive benchmark suite, capable of evaluating individual hardware components and representing most real AI workloads. It comprises a Validator that learns benchmark criteria to pinpoint defective components clearly. Additionally, SuperBench incorporates a Selector to balance validation time and issue-related penalties, enabling optimal timing for validation execution with a tailored subset of benchmarks. Through testbed evaluation and simulation, we demonstrate that SuperBench can increase the mean time between incidents by up to 22.61×. SuperBench has been successfully deployed in Azure production, validating hundreds of thousands of GPUs every year.
Yifan Xiong 0001, Ziyue Yang 0002, Guoshuai Zhao 0001, Dong Zhong, Boris Pinzur, Jie Zhang 0048, Yang Wang 0053, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng 0005, Yongqiang Xiong, Lidong Zhou
ACM Trans. Comput. Syst.10
2025 Miniature: Fast AI Supercomputer Networks Simulation on FPGAs
Yicheng Qian, Ran Shu 0001, Rui Ma 0021, Yang Wang 0053, Derek Chiou, Nadeen Gebara, Luca Piccolboni, Miriam Leeser, Yongqiang Xiong
APNet4
2025 BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
abstract
Large language models (LLMs) have demonstrated remarkable performance across various machine learning tasks. Yet the substantial memory footprint of LLMs significantly hinders their deployment. In this paper, we improve the accessibility of LLMs through BitMoD1, an algorithm-hardware co-design solution that enables efficient LLM acceleration at low weight precision. On the algorithm side, BitMoD introduces fine-grained data type adaptation that uses a different numerical data type to quantize a group of (e.g., 128) weights. Through the careful design of these new data types, BitMoD is able to quantize LLM weights to very low precision (e.g., 4 bits and 3 bits) while maintaining high accuracy. On the hardware side, BitMoD employs a bitserial processing element to easily support multiple numerical precisions and data types; our hardware design includes two key innovations: First, it employs a unified representation to process different weight data types, thus reducing the hardware cost. Second, it adopts a bit-serial dequantization unit to rescale the per-group partial sum with minimal hardware overhead. Our evaluation on six representative LLMs demonstrates that BitMoD significantly outperforms state-of-the-art LLM quantization and acceleration methods. For discriminative tasks, BitMoD can quantize LLM weights to 4 -bit with1Code is available at: https://github.com/yc2367/BitMoD-HPCA-25
Yuzong Chen 0001, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang 0053, Marta Andronic, George A. Constantinides, Mohamed S. Abdelfattah
HPCA4
2025 LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
abstract
The emergence of neural network capabilities invariably leads to a significant surge in computational demands due to expanding model sizes and increased computational complexity. To reduce model size and lower inference costs, recent research has focused on simplifying models and designing hardware accelerators using low-bit quantization. However, due to numerical representation limits, scalar quantization cannot reduce bit width lower than 1-bit, diminishing its benefits. To break through these limitations, we introduce LUT-DLA, a Look-Up Table (LUT) Deep Learning Accelerator Framework that utilizes vector quantization to convert neural network models into LUTs, achieving extreme low-bit quantization. The LUT-DLA framework facilitates efficient and cost-effective hardware accelerator designs and supports the LUTBoost algorithm, which helps to transform various DNN models into LUT-based models via multistage training, drastically cutting both computational and hardware overhead. Additionally, through co-design space exploration, LUT-DLA assesses the impact of various model and hardware parameters to fine-tune hardware configurations for different application scenarios, optimizing performance and efficiency. Our comprehensive experiments show that LUT-DLA achieves improvements in power efficiency and area efficiency with gains of 1.4~7.0× and 1.5~146.1×, respectively, while maintaining only a modest accuracy drop. For CNNs, accuracy decreases by 0.1%~3.1% using the L2distance similarity, 0.1%~3.4% with the L1distance similarity, and 0.1%~3.8% when employing the Chebyshev distance similarity. For transformer-based models, the accuracy drop ranges from 1.4% to 3.0%.
Shengyu Ye, Chunyun Chen, Yang Wang 0053, Fan Yang 0024, Ting Cao 0003, Cheng Liu 0008, Mohamed M. Sabry, Mao Yang 0004
HPCA4
2025 PipeThreader: Software-Defined Pipelining for Efficient DNN Execution
Yu Cheng 0030, Lei Wang 0222, Yining Shi 0001, Yuqing Xia, Lingxiao Ma, Jilong Xue, Yang Wang 0053, Zhiwen Mo, Fan Yang 0024, Mao Yang 0004, Zhi Yang 0001
OSDI7
2025 DSTC: Dual-Side Sparse Tensor Core for DNNs Acceleration on Modern GPU Architectures
abstract
Leveraging sparsity in deep neural network (DNN) models holds significant promise for accelerating model inference. However, current GPUs can only harness sparsity in model weights, leaving activations unutilized due to their dynamic and unpredictable nature, which poses a considerable challenge for exploitation. In our research, we introduce a novel architectural approach aimed at effectively leveraging dual-side sparsity, encompassing both weight and activation sparsity. Our methodology involves a systematic examination of previous sparsity-related architectures, and culminating in the proposal of an uncharted paradigm that combines outer-product computation primitive and bitmap-based encoding format. Our approach showcases feasibility through minimal modifications to existing production-scale inner-product-based Tensor Cores. We introduce a set of innovative ISA extensions and carefully co-design matrix-matrix multiplication and convolution algorithms, the two predominant computation patterns in contemporary DNN models, to exploit our novel dual-side sparse Tensor Core. Our evaluation demonstrates the efficacy of our design, unlocking the full potential of dual-side DNN sparsity and delivering performance enhancements of up to an order of magnitude while incurring only modest hardware overhead.
Chen Zhang 0001, Yang Wang 0053, Cong Guo 0003, Yunxin Liu 0001, Jingwen Leng, Zhigang Ji, Yuan Xie 0001, Ru Huang 0001
IEEE Trans. Computers2
2024 PIM-DL: Expanding the Applicability of Commodity DRAM-PIMs for Deep Learning via Algorithm-System Co-Optimization
abstract
DRAM-based processing-in-memory (DRAM-PIM) has gained commercial prominence in recent years. However, their integration for deep learning acceleration poses inherent challenges. Existing DRAM-PIMs are limited in computational capabilities, primarily applicable for element-wise and GEMV operators. Unfortunately, these operators contribute only a small portion of the execution time in most DNN workloads. Current systems still necessitate powerful hosts to handle a significant portion of compute-heavy operators.
Cong Li 0008, Zhe Zhou 0002, Yang Wang 0053, Fan Yang 0093, Ting Cao 0003, Mao Yang 0004, Yun Liang 0001, Guangyu Sun 0003
ASPLOS (2)3
2024 VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models
abstract
Scaling model size significantly challenges the deployment and inference of Large Language Models (LLMs).Due to the redundancy in LLM weights, recent research has focused on pushing weight-only quantization to extremely low-bit (even down to 2 bits).It reduces memory requirements, optimizes storage costs, and * Contribution during internship at Microsoft Research
Jicheng Wen, Yang Wang 0053, Shengyu Ye, Li Lyna Zhang, Ting Cao 0003, Cheng Li 0001, Mao Yang 0004
EMNLP3
2024 NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
abstract
The Compute Express Link (CXL) interconnect makes it feasible to integrate diverse types of memory into servers via its byte-addressable SerDes links. Considering the various access latency, harnessing the full potential of CXL-based heterogeneous memory systems requires efficient memory tiering. However, prior work can hardly make a fundamental progress owing to low-resolution and high-overhead memory access profiling techniques. To address this critical challenge, we propose a novel memory tiering solution called NeoMem, which features a hardware/software co-design. NeoMem offloads memory profiling functions to CXL device-side controllers, integrating a dedicated hardware unit called NeoProf. NeoProf readily monitors memory accesses and provides the OS with crucial page hotness statistics and other useful system state information. On the OS kernel side, we design a revamped memory-tiering strategy, enabling accurate and timely hot page promotion based on NeoProf statistics. We implement NeoMem on a real FPGA-based CXL memory platform and Linux kernel v6.3. Comprehensive evaluations demonstrate that NeoMem achieves 32% ~ 67% geomean speedup over several existing memory tiering solutions.
Zhe Zhou 0002, Tao Zhang 0032, Yang Wang 0053, Ran Shu 0001, Shuotao Xu, Peng Cheng 0005, Yongqiang Xiong, Jie Zhang 0048, Guangyu Sun 0003
MICRO4
2024 SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
Yifan Xiong 0001, Ziyue Yang 0002, Guoshuai Zhao 0001, Dong Zhong, Boris Pinzur, Jie Zhang 0048, Yang Wang 0053, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng 0005, Yongqiang Xiong, Lidong Zhou
USENIX ATC10
2023 LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table Lookup
abstract
On-device Deep Neural Network (DNN) inference consumes significant computing resources and development efforts. To alleviate that, we propose LUT-NN, the first system to empower inference by table lookup, to reduce inference cost. LUT-NN learns the typical features for each operator, named centroid, and precompute the results for these centroids to save in lookup tables. During inference, the results of the closest centroids with the inputs can be read directly from the table, as the approximated outputs without computations.
Xiaohu Tang 0003, Yang Wang 0053, Ting Cao 0003, Li Lyna Zhang, Qi Chen 0009, Deng Cai 0001, Yunxin Liu 0001, Mao Yang 0004
MobiCom2
2022 Romou: rapidly generate high-performance tensor kernels for mobile GPUs
abstract
Mobile GPU, as a ubiquitous and powerful accelerator, plays an important role in accelerating on-device DNN (Deep Neural Network) inference. The frequent-upgrade and diversity of mobile GPUs require automatic kernel generation to empower fast DNN deployment. However, current generated kernels have poor performance.
Rendong Liang, Ting Cao 0003, Jicheng Wen, Manni Wang, Yang Wang 0053, Jianhua Zou, Yunxin Liu 0001
MobiCom5
2022 SparTA: Deep-Learning Model Sparsity via Tensor-with-Sparsity-Attribute
Ningxin Zheng, Quanlu Zhang, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Yang Wang 0053, Mao Yang 0004, Lidong Zhou
OSDI7
2022 FlexMon: A flexible and fine-grained traffic monitor for programmable networks
Yang Wang 0053, Xiong Wang 0001, Shizhong Xu, Ci He, Jing Ren 0002, Shui Yu 0001
J. Netw. Comput. Appl.1
2021 NeuralMon: Graph Neural Network for Flow Measurement Allocation
abstract
Fine-grained and accurate network flow measurements are essential for various network management tasks. In recent years, the evolution of programmable networks enables flow measurement on the switch. However, limited hardware resources on programmable switches drive the shift of measurement from a single switch to network-wide coordinations. This paper aims to optimize the allocation strategy of flow measurement among switches under the objective of measurement coverage and accuracy in network-wide measurement scenarios. We design a Graph Neural Network model, NeuralMon, that can model and solve the above problem precisely. NeuralMon converts network topologies and network flows into a hypergraph and transforms the flow measurement task allocation problem into a node classification problem. NeuralMon is effective in learning the task allocation solution from the network topologies and flows directly. Even on untrained real-world network topologies, NeuralMon still provides excellent performance.
Yang Wang 0053, Xiong Wang 0001, Zhuobin Huang, Ci He, Shizhong Xu
GLOBECOM1
2021 Dual-side Sparse Tensor Core
abstract
Leveraging sparsity in deep neural network (DNN) models is promising for accelerating model inference. Yet existing GPUs can only leverage the sparsity from weights but not activations, which are dynamic, unpredictable, and hence challenging to exploit. In this work, we propose a novel architecture to efficiently harness the dual-side sparsity (i.e., weight and activation sparsity). We take a systematic approach to understand the (dis)advantages of previous sparsity-related architectures and propose a novel, unexplored paradigm that combines outer-product computation primitive and bitmap-based encoding format. We demonstrate the feasibility of our design with minimal changes to the existing production-scale inner-product-based Tensor Core. We propose a set of novel ISA extensions and co-design the matrix-matrix multiplication and convolution algorithms, which are the two dominant computation patterns in today’s DNN models, to exploit our new dual-side sparse Tensor Core. Our evaluation shows that our design can fully unleash the dual-side DNN sparsity and improve the performance by up to one order of magnitude with small hardware overhead.
Yang Wang 0053, Chen Zhang 0001, Cong Guo 0003, Yunxin Liu 0001, Jingwen Leng
ISCA1
2020 LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression
abstract
BERT is a cutting-edge language representation model pre-trained by a large corpus, which achieves superior performances on various natural language understanding tasks. However, a major blocking issue of applying BERT to online services is that it is memory-intensive and leads to unsatisfactory latency of user requests, raising the necessity of model compression. Existing solutions leverage the knowledge distillation framework to learn a smaller model that imitates the behaviors of BERT. However, the training procedure of knowledge distillation is expensive itself as it requires sufficient training data to imitate the teacher model. In this paper, we address this issue by proposing a tailored solution named LadaBERT (Lightweight adaptation of BERT through hybrid model compression), which combines the advantages of different model compression methods, including weight pruning, matrix factorization and knowledge distillation. LadaBERT achieves state-of-the-art accuracy on various public datasets while the training overheads can be reduced by an order of magnitude.
Yihuan Mao, Yujing Wang 0002, Chufan Wu, Chen Zhang 0001, Yang Wang 0053, Quanlu Zhang, Yaming Yang 0001, Yunhai Tong, Jing Bai 0010
COLING5
2018 MOSC: a method to assign the outsourcing of service function chain across multiple clouds
Xiong Wang 0001, Yangming Zhao, Tongyu Song, Yang Wang 0053, Shizhong Xu, Lemin Li
Comput. Networks5
2016 Towards optimal outsourcing of service function chain across multiple clouds
abstract
As Network Function Virtualization (NFV) becomes reality and cloud computing offers a scalable pay-as-you-go charging model, more network operators would like to outsource their Service Function Chains (SFC) to the public clouds in order to reduce the operational cost. However, how to minimize the operational cost with Quality of Service (QoS) guarantee when outsourcing SFC is still an open problem. In this paper, we are to study this problem when there are large number of candidate cloud providers with diverse pricing schemes of network functions. In addition, extra delay is introduced as the result of outsourcing SFCs. Firstly, we formulate this problem as an Integer Linear Programming (ILP) model. Then we design an efficient heuristic algorithm named QoS-Guaranteed SFC Outsourcing algorithm (QGSO) based on Hidden Markov Model (HMM). The extensive simulations show that QGSO saves up to 75.8% cost compared with that of deploying network functions in local network. QGSO also achieves up to 42.6% cost savings compared with the result of first-fit based optimization algorithm.
Shizhong Xu, Xiong Wang 0001, Yangming Zhao, Ke Li 0001, Yang Wang 0053, Wei Wang 0171, Lemin Li
ICC6
2009 Products of Mealy-type fuzzy finite state machines
Zhiwen Mo, Dong Qiu, Yang Wang 0053
Fuzzy Sets Syst.4