VLDB 2026 Research / reviewers in the wild / expert
Ting Cao 0003
dblp:03/8334-3
· DBLP profile ↗
49ranked-venue papers
0as first author
49since 2021 · last 2026
0000-0002-9107-013XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 21 · 21 since 2021Systems, architecture and hardware · 16 · 16 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling LLM Test-Time Compute with Mobile NPU on SmartphonesabstractDeploying Large Language Models (LLMs) on mobile devices faces the challenge of insufficient performance in smaller models and excessive resource consumption in larger ones. This paper highlights that mobile Neural Processing Units (NPUs) have underutilized computational resources, particularly their matrix multiplication units, during typical LLM inference. To leverage this wasted compute capacity, we propose applying parallel test-time scaling techniques on mobile NPUs to enhance the performance of smaller LLMs. However, this approach confronts inherent NPU challenges, including inadequate hardware support for fine-grained quantization and low efficiency in general-purpose computations. To overcome these, we introduce two key techniques: a hardware-aware tile quantization scheme that aligns group quantization with NPU memory access patterns, and efficient LUT-based replacements for complex operations such as Softmax and dequantization. We design and implement an end-to-end inference system that leverages the NPU's compute capability to support test-time scaling on Qualcomm Snapdragon platforms. Experiments show our approach brings significant speedups: up to 19.0× for mixed-precision GEMM and 2.2× for Softmax. More importantly, we demonstrate that smaller models using test-time scaling can match or exceed the accuracy of larger models, achieving a new performance-cost Pareto frontier. Zixu Hao, Jianyu Wei, Tuowei Wang, Minxing Huang, Huiqiang Jiang, Shiqi Jiang 0002, Ting Cao 0003, Ju Ren 0001 |
EuroSys | 7 |
| 2026 | BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV CacheabstractThe rise of long-context Large Language Models (LLMs) amplifies memory and bandwidth demands during autoregressive decoding, as the Key-Value (KV) cache grows with each generated token. Low-bit KV-cache quantization (e.g., 4-bit or 2-bit) can reduce memory footprint while preserving accuracy, but existing systems suffer from slow decoding due to their exclusive reliance on CUDA cores, neglecting Tensor Cores—the primary source of compute on modern GPUs. We present BitDecoding, a new long-context LLMs inference system with low-bit KV cache. BitDecoding enables efficient low-bit KV cache decoding by cooperatively leveraging CUDA Cores and Tensor Cores. It introduces methods for automatically inducing optimized layouts to exploit Tensor Cores, along with novel warp-level parallelization strategies for dequantization. For unified system support, BitDecoding includes a query transformation module supporting diverse attention variants, a quantization kernel to support both tensor-wise and channelwise scaling used in various quantization algorithms with high performance, and a dequantization kernel with a softwaredefined pipeline to coordinate CUDA and Tensor Cores execution for mix-precision operations. In addition, architecture-specific optimizations leverage Hopper's warpgroup tensor instructions and Blackwell's native low-precision tensor formats to maximize decoding throughput on the latest GPU generations. Evaluated on Blackwell, Hopper, Ada, and Ampere architectures, BitDecoding attains on average a 7.5× decoding speedup over FP16 FlashDecoding-v2, and further reaches up to 8.6× with native MXFP4 formats on Blackwell, while surpassing the state-of-the-art low-bit system QServe by up to 4.3×. On LLaMA-3.1-8B with a 128K context, BitDecoding reduces singlebatch decoding latency by 3×, demonstrating substantial improvements for long-context generation, and is open sourced at https://github.com/OpenBitSys/BitDecoding. Dayou Du, Shijie Cao, Jianyi Cheng, Luo Mai, Ting Cao 0003, Mao Yang 0004 |
HPCA | 5 |
| 2026 | Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge DevicesabstractLarge language models (LLMs) are increasingly deployed on edge devices. To meet strict resource constraints, real-world deployment has pushed LLM quantization from 8-bit to 4-bit, 2-bit, and now 1.58-bit. Combined with lookup table (LUT)-based inference, CPUs run these ultra-low-bit LLMs even faster than NPUs, opening new opportunities for ubiquitous on-device intelligence. Weijun Wang 0001, Jianyu Wei, Ting Cao 0003, Yunxin Liu 0001 |
MobiSys | 5 |
| 2026 | AVA: Towards Agentic Video Analytics with Vision Language Models
Yuxuan Yan, Shiqi Jiang 0002, Ting Cao 0003, Yifan Yang 0004, Qianqian Yang 0002, Yuanchao Shu, Yuqing Yang 0001, Lili Qiu |
NSDI | 3 |
| 2026 | Efficient Remote KV Cache Reuse with GPU-native Video Codec
Liang Mi, Weijun Wang 0001, Jinghan Chen, Ting Cao 0003, Haipeng Dai 0001, Yunxin Liu 0001 |
SIGCOMM | 4 |
| 2025 | Neuralink: Fast on-Device LLM Inference with Neuron Co-Activation LinkingabstractLarge Language Models (LLMs) have achieved remarkable success across various domains, yet deploying them on mobile devices remains an arduous challenge due to their extensive computational and memory demands.While lightweight LLMs have been developed to fit mobile environments, they suffer from degraded model accuracy.In contrast, sparsitybased techniques minimize DRAM usage by selectively transferring only relevant neurons to DRAM while retaining the full model in external storage, such as flash.However, such approaches are critically limited by numerous I/O operations, particularly on smartphones with severe IOPS constraints.In this paper, we propose Neuralink, a novel approach that accelerates LLM inference on smartphones by optimizing neuron placement in flash memory.Neuralink leverages the concept of Neuron Co-Activation, where neurons frequently activated together are linked to facilitate continuous read access and optimize I/O efficiency.Our approach incorporates a two-stage solution: an offline stage that reorganizes neuron placement based on co-activation patterns, and an online stage that employs tailored data access and caching strategies to align well with hardware characteristics.Evaluations conducted on a variety of smartphones and LLMs demonstrate that Neuralink achieves on average 1.49× improvements in end-to-end latency compared to the state-of-the-art.As the first solution to optimize storage placement under sparsity, Neuralink explores a new * Both authors contributed equally to this research. Tuowei Wang, Ruwen Fan, Minxing Huang, Zixu Hao, Kun Li 0016, Ting Cao 0003, Youyou Lu, Yaoxue Zhang, Ju Ren 0001 |
ASPLOS (3) | 6 |
| 2025 | T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on EdgeabstractThe deployment of Large Language Models (LLMs) on edge devices is increasingly important to enhance on-device intelligence. Weight quantization is crucial for reducing the memory footprint of LLMs on devices. However, low-bit LLMs necessitate mixed precision matrix multiplication (mpGEMM) of low precision weights and high precision activations during inference. Existing systems, lacking native support for mpGEMM, resort to dequantize weights for high precision computation. Such an indirect way can lead to a significant inference overhead. Jianyu Wei, Shijie Cao, Ting Cao 0003, Lingxiao Ma, Lei Wang 0222, Yanyong Zhang, Mao Yang 0004 |
EuroSys | 3 |
| 2025 | LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning AcceleratorabstractThe emergence of neural network capabilities invariably leads to a significant surge in computational demands due to expanding model sizes and increased computational complexity. To reduce model size and lower inference costs, recent research has focused on simplifying models and designing hardware accelerators using low-bit quantization. However, due to numerical representation limits, scalar quantization cannot reduce bit width lower than 1-bit, diminishing its benefits. To break through these limitations, we introduce LUT-DLA, a Look-Up Table (LUT) Deep Learning Accelerator Framework that utilizes vector quantization to convert neural network models into LUTs, achieving extreme low-bit quantization. The LUT-DLA framework facilitates efficient and cost-effective hardware accelerator designs and supports the LUTBoost algorithm, which helps to transform various DNN models into LUT-based models via multistage training, drastically cutting both computational and hardware overhead. Additionally, through co-design space exploration, LUT-DLA assesses the impact of various model and hardware parameters to fine-tune hardware configurations for different application scenarios, optimizing performance and efficiency. Our comprehensive experiments show that LUT-DLA achieves improvements in power efficiency and area efficiency with gains of 1.4~7.0× and 1.5~146.1×, respectively, while maintaining only a modest accuracy drop. For CNNs, accuracy decreases by 0.1%~3.1% using the L2distance similarity, 0.1%~3.4% with the L1distance similarity, and 0.1%~3.8% when employing the Chebyshev distance similarity. For transformer-based models, the accuracy drop ranges from 1.4% to 3.0%. Shengyu Ye, Chunyun Chen, Yang Wang 0053, Fan Yang 0024, Ting Cao 0003, Cheng Liu 0008, Mohamed M. Sabry, Mao Yang 0004 |
HPCA | 6 |
| 2025 | StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
Hao Wu 0067, Yifan Yang 0004, Shiqi Jiang 0002, Qianxi Zhang, Donglin Bai, Zhibo Chen 0001, Ting Cao 0003 |
ICCV | 8 |
| 2025 | LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceabstractLarge Language Model (LLM) inference becomes resource-intensive, prompting a shift toward low-bit model weights to reduce the memory footprint and improve efficiency.Such low-bit LLMs necessitate the mixed-precision matrix multiplication (mpGEMM), an important yet underexplored operation involving the multiplication of lower-precision weights with higher-precision activations.Off-theshelf hardware does not support this operation natively, leading to indirect, thus inefficient, dequantization-based implementations.In this paper, we study the lookup table (LUT)-based approach for mpGEMM and find that a conventional LUT implementation fails to achieve the promised gains.To unlock the full potential of LUT-based mpGEMM, we propose LUT Tensor Core, a softwarehardware co-design for low-bit LLM inference.LUT Tensor Core differentiates itself from conventional LUT designs through: 1) * Work is done during internship at Microsoft Research. Zhiwen Mo, Lei Wang 0222, Jianyu Wei, Zhichen Zeng 0002, Shijie Cao, Lingxiao Ma, Naifeng Jing, Ting Cao 0003, Jilong Xue, Fan Yang 0024, Mao Yang 0004 |
ISCA | 8 |
| 2025 | Demo: EdgeMind-OS: A Plug-and-Play Embodied Intelligence System for Real-Time On-Device DeploymentabstractBuilding an always-on, contextual AI assistant that proactively supports humans remains a central goal in Embodied AI—yet cloud-based pipelines struggle to meet due to delay, bandwidth, and privacy constraints. This demo presents EdgeMind-OS, a fully on-device intelligence system designed for embodied agents operating in real-world scenarios. Edge-Mind-OS features a hierarchical architecture combining a real-time StreamBrain, modular skill experts, and a dynamic scene-episode memory. Achieving up to 7.3× faster local processing, it enables low-latency, privacy-preserving, and plug-and-play deployment across tasks such as semantic navigation, spatial memory recall, and multimodal interaction. We demonstrate how EdgeMind-OS empowers a mobile robot with only basic locomotion capabilities to perform realtime, free-form user-robot interaction through autonomous perception, reasoning and action —without reliance on external cloud infrastructure. Jianyu Wei, Fucheng Jia, Liang Mi, Ruofei Ju, Xianye Wang, Yikai Zheng, Weijun Wang 0001, Shiqi Jiang 0002, Yunxin Liu 0001, Ting Cao 0003 |
MobiCom | 12 |
| 2025 | SeerAttention: Self-distilled Attention Gating for Efficient Long-context PrefillingabstractAttention is the cornerstone of modern Large Language Models (LLMs). Yet its quadratic complexity hinders efficiency and scalability, especially for long-context processing. A promising approach is to leverage sparsity in attention. However, existing sparsity-based solutions predominantly rely on predefined patterns or heuristics at the attention head level, struggling to adapt dynamically to different contexts efficiently. We propose SeerAttention, a simple yet effective attention mechanism that directly learns the block-level attention sparsity from the LLM itself. Inspired by the gating mechanism in Mixture of Experts (MoE), SeerAttention augments the conventional attention with a **learnable gate** that **selectively activates important blocks** within the attention map. Specifically, the gate first pools the query (Q) and key (K) tensors along the sequence dimension and processes them through learnable linear layers. The resulting matrices are then multiplied together to produce the gating scores, which are used to predict block-level attention sparsity. Combined with our block-sparse FlashAttention kernel, SeerAttention can achieve significant speedup on GPUs. When applied to pre-trained LLMs, SeerAttention only requires training the gate parameters in a lightweight self-distillation manner, allowing rapid convergence. Our evaluation results demonstrate that SeerAttention achieves better model accuracy and lower latency for long-context pre-filling compared to prior methods. Code is available at: https://github.com/microsoft/SeerAttention. Yizhao Gao 0002, Zhichen Zeng 0002, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao 0003, Fan Yang 0024, Mao Yang 0004 |
NeurIPS | 9 |
| 2025 | FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core UnitsabstractWhile Tensor Core Units (TCUs) excel in AI tasks, their application to HPC algorithms like stencil computations faces significant challenges due to sparsity, which leads to underutilization and exacerbates memory-bound limitations. This paper introduces FlashFFTStencil1, a memory-efficient stencil computing system designed to bridge FFT to fully-dense stencil computations on TCUs. Aimed at bound shifting, FlashFFTStencil comprises three key techniques: Kernel Tailoring on HBM fuses distinct kernels to enhance parallelism while reducing memory transfer and footprint; Architecture Aligning on SMEM restructures FFT-based stencil computations into dense matrix multiplications tailored for shared memory architecture; Computation Streamlining on TCU optimizes TCU utilization and thread parallelism by minimizing pipeline stalls and maximizing register reuse. Notably, a distinctive extension is FlashFFTStencil's ability to enable theoretically unrestricted temporal fusion by FFT. Results show that FlashFFTStencil achieves effective sparsity-free bound shifting, with an average speedup of 2.57x over the state-of-the-art. FlashFFTStencil pioneers a new era in unifying computational patterns within the HPC landscape and bridges them with cutting-edge AI-driven hardware innovations like TCUs. Haozhi Han, Kun Li 0016, Donglin Bai, Yiwei Zhang 0009, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
PPoPP | 9 |
| 2025 | Jigsaw: Toward Conflict-free Vectorized Stencil Computation by Tessellating Swizzled RegistersabstractStencil computation plays a pivotal role in numerous scientific and engineering applications. Previous studies have extensively investigated vectorization techniques to enhance in-core parallelism; however, the performance bottleneck caused by data alignment conflicts (DAC) has not been effectively resolved in all dimensions. This paper proposes Jigsaw, a conflict-free vectorization method to reduce DAC across all dimensions by tessellating swizzled finest-grained lanes. Jigsaw comprises three key components: Lane-based Butterfly Vectorization, SVD-based Dimension Flattening, and Iteration-based Temporal Merging. These components effectively address DAC across spatial and temporal dimensions. Experimental results on different machines demonstrate that Jigsaw could achieve a significant improvement compared to the state-of-the-art techniques, with an average speedup of 2.31x on various stencil kernels. Yiwei Zhang 0009, Kun Li 0016, Haozhi Han, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
PPoPP | 6 |
| 2025 | Matrix Is All You Need: Rearchitecting Quantum Chemistry to Scale on AI AcceleratorsabstractScientific computing remains fundamentally misaligned with the execution paradigm of modern AI accelerators, which rely on structured, low-precision matrix operations for performance and scalability. Quantum chemistry exemplifies this gap through three core scalability limits: irregular computational patterns, fragmented hardware utilization, and limited scientific reach. Haozhi Han, Kun Li 0016, Fusong Ju, Qi Li 0039, Hong An, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
SC | 8 |
| 2025 | SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity TransformationabstractSparse Tensor Cores offer exceptional performance gains for AI workloads by exploiting structured 2:4 sparsity. However, their potential remains untapped for core scientific workloads such as stencil computations, which exhibit irregular sparsity patterns. Qi Li 0039, Kun Li 0016, Haozhi Han, Yunquan Zhang, Junshi Chen 0003, Hong An, Ting Cao 0003, Mao Yang 0004 |
SC | 9 |
| 2025 | Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality AlignmentabstractThis paper presents Babel, the expandable modality alignment model, specially designed for multi-modal sensing. While there has been considerable work on multi-modality alignment, they all struggle to effectively incorporate multiple sensing modalities due to the data scarcity constraints. How to utilize multi-modal data with partial pairings in sensing remains an unresolved challenge. Shenghong Dai, Shiqi Jiang 0002, Yifan Yang 0004, Ting Cao 0003, Mo Li 0001, Suman Banerjee 0001, Lili Qiu |
SenSys | 4 |
| 2025 | JENGA: Enhancing LLM Long-Context Fine-tuning with Contextual Token Sparsity
Tuowei Wang, Kun Li 0016, Ting Cao 0003, Ju Ren 0001, Yaoxue Zhang |
USENIX ATC | 4 |
| 2025 | Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile DevicesabstractDiffusion models have revolutionized image synthesis applications. Many studies focus on using approximate computation such as model quantization to reduce inference costs on mobile devices. However, due to their extensive model parameters and autoregressive inference fashion, the overhead of diffusion models remains high, which is challenging for mobile devices to handle. To reduce the inference overhead of diffusion models on mobile devices, we proposeLUT-Diff, an algorithm-system co-design specifically tailored for mobile device diffusion model inference optimization.LUT-Diffoptimizes using lookup tables and can efficiently generate a series of lookup table candidates for diffusion models without end-to-end training. During inference,LUT-Diffadaptively selects the best inference strategy based on the application/user's latency budget. Additionally,LUT-Diffincludes a parallel inference engine that rapidly completes model inference through CPU-GPU co-scheduling. Extensive experiments demonstrate thatLUT-Diffcan generate images comparable to the original model, with an up to 0.012 MSE in generated images.LUT-Diffcan also achieve up to 9.1× inference acceleration and reduce the inference memory footprint by up to 70.9% compared to baseline methods. Moreover,LUT-Diffcan save at least 3281× the learning cost of lookup tables. Qipeng Wang 0001, Shiqi Jiang 0002, Yifan Yang 0004, Ruiqi Liu 0001, Yuanchun Li 0003, Ting Cao 0003, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Anatomizing Deep Learning Inference in Web BrowsersabstractWeb applications have increasingly adopted Deep Learning (DL) through in-browser inference , wherein DL inference performs directly within Web browsers. The actual performance of in-browser inference and its impacts on the Quality of Experience ( QoE ) remain unexplored, and urgently require new QoE measurements beyond traditional ones, e.g., mainly focusing on page load time. To bridge this gap, we make the first comprehensive performance measurement of in-browser inference to date. Our approach proposes new metrics to measure in-browser inference: responsiveness, smoothness, and inference accuracy. Our extensive analysis involves 9 representative DL models across Web browsers of 50 popular PC devices and 20 mobile devices. The results reveal that in-browser inference exhibits a substantial latency gap, averaging 16.9 times slower on CPU and 4.9 times slower on GPU compared to native inference on PC devices. The gap on mobile CPU and mobile GPU is 15.8 times and 7.8 times, respectively. Furthermore, we identify contributing factors to such latency gap, including underutilized hardware instruction sets, inherent overhead in the runtime environment, resource contention within the browser, and inefficiencies in software libraries and GPU abstractions. Additionally, in-browser inference imposes significant memory demands, at times exceeding 334.6 times the size of the DL models themselves, partly attributable to suboptimal memory management. We also observe that in-browser inference leads to a significant 67.2% increase in the time it takes for GUI components to render within Web browsers, significantly affecting the overall user QoE of Web applications reliant on this technology. Qipeng Wang 0001, Shiqi Jiang 0002, Zhenpeng Chen 0001, Yuanchun Li 0003, Aoyu Li, Yun Ma 0002, Ting Cao 0003, Xuanzhe Liu |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2024 | BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-DistillationabstractThe upscaling of Large Language Models (LLMs) has yielded impressive advances in natural language processing, yet it also poses significant deployment challenges.Weight quantization has emerged as a widely embraced solution to reduce memory and computational demands.This paper introduces BitDistiller, a framework that synergizes Quantization-Aware Training (QAT) with Knowledge Distillation (KD) to boost the performance of LLMs at ultra-low precisions (sub-4-bit).Specifically, BitDistiller first incorporates a tailored asymmetric quantization and clipping technique to maximally preserve the fidelity of quantized weights, and then proposes a novel Confidence-Aware Kullback-Leibler Divergence (CAKLD) objective, which is employed in a self-distillation manner to enable faster convergence and superior model performance.Empirical evaluations demonstrate that BitDistiller significantly surpasses existing methods in both 3-bit and 2-bit configurations on general language understanding and complex reasoning benchmarks.Notably, Bit-Distiller is shown to be more cost-effective, demanding fewer data and training resources. Dayou Du, Shijie Cao, Ting Cao 0003, Xiaowen Chu 0001, Ningyi Xu |
ACL (1) | 5 |
| 2024 | PIM-DL: Expanding the Applicability of Commodity DRAM-PIMs for Deep Learning via Algorithm-System Co-OptimizationabstractDRAM-based processing-in-memory (DRAM-PIM) has gained commercial prominence in recent years. However, their integration for deep learning acceleration poses inherent challenges. Existing DRAM-PIMs are limited in computational capabilities, primarily applicable for element-wise and GEMV operators. Unfortunately, these operators contribute only a small portion of the execution time in most DNN workloads. Current systems still necessitate powerful hosts to handle a significant portion of compute-heavy operators. Cong Li 0008, Zhe Zhou 0002, Yang Wang 0053, Fan Yang 0093, Ting Cao 0003, Mao Yang 0004, Yun Liang 0001, Guangyu Sun 0003 |
ASPLOS (2) | 5 |
| 2024 | VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language ModelsabstractScaling model size significantly challenges the deployment and inference of Large Language Models (LLMs).Due to the redundancy in LLM weights, recent research has focused on pushing weight-only quantization to extremely low-bit (even down to 2 bits).It reduces memory requirements, optimizes storage costs, and * Contribution during internship at Microsoft Research Jicheng Wen, Yang Wang 0053, Shengyu Ye, Li Lyna Zhang, Ting Cao 0003, Cheng Li 0001, Mao Yang 0004 |
EMNLP | 6 |
| 2024 | Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language ModelsabstractEfficient deployment of Large Language Models (LLMs) requires low-bit quantization to reduce model size and inference cost. Besides low-bit integer formats (e.g., INT8/INT4) used in previous quantization works, emerging low-bit floating-point formats (e.g., FP8/FP4) supported by advanced hardware like NVIDIA’s H100 GPU offer an alternative. Our study finds that introducing floating-point formats significantly improves LLMs quantization. We also discover that the optimal quantization format varies across layers. Therefore, we select the optimal format for each layer, which we call the Mixture of Formats Quantization (MoFQ) method. Our MoFQ method achieves better or comparable results over current methods in weight-only (W-only) and weight-activation (WA) post-training quantization scenarios across various tasks, with no additional hardware overhead. Lingran Zhao, Shijie Cao, Ting Cao 0003, Fan Yang 0024, Mao Yang 0004, Shanghang Zhang, Ningyi Xu |
ICME | 6 |
| 2024 | Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert InferenceabstractLarge language models (LLMs) based on transformers have made significant strides in recent years, the success of which is driven by scaling up their model size. Despite their high algorithmic performance, the computational and memory requirements of LLMs present unprecedented challenges. To tackle the high compute requirements of LLMs, the Mixture-ofExperts (MoE) architecture was introduced which is able to scale its model size without proportionally scaling up its computational requirements. Unfortunately, MoE’s high memory demands and dynamic activation of sparse experts restrict its applicability to real-world problems. Previous solutions that offload MoE’s memory-hungry expert parameters to CPU memory fall short because the latency to migrate activated experts from CPU to GPU incurs high performance overhead. Our proposed Pre-gated MoE system effectively tackles the compute and memory challenges of conventional MoE architectures using our algorithm-system codesign. Pre-gated MoE employs our novel pre-gating function which alleviates the dynamic nature of sparse expert activation, allowing our proposed system to address the large memory footprint of MoEs while also achieving high performance. We demonstrate that Pre-gated MoE is able to improve performance, reduce GPU memory consumption, while also maintaining the same level of model quality. These features allow our Pre-gated MoE system to cost-effectively deploy large-scale LLMs using just a single GPU with high performance. Ranggi Hwang, Jianyu Wei, Shijie Cao, Changho Hwang, Xiaohu Tang 0003, Ting Cao 0003, Mao Yang 0004 |
ISCA | 6 |
| 2024 | FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge DevicesabstractDue to the popularity of deep neural networks (DNNs) and considerations over network overhead, data privacy, and inference latency, there is a growing interest in deploying DNNs to edge devices in recent years. However, the limited memory becomes a major bottleneck for on-device DNN deployment, making it crucial to reduce the memory footprint of DNN. The mainstream model customization solutions require intensive deployment efforts and may lead to severe accuracy degradation, and existing deep learning (DL) frameworks don't take memory as a priority. Besides, recent works to enhance the memory management scheme cannot be directly applied because of several challenges, including the unbalanced memory footprint across layers, the inevitable overhead of memory management, and the memory budget dynamicity. To tackle these challenges, we introduce FlexNN, an efficient and adaptive memory management framework for DNN inference on memory-constrained devices. FlexNN uses a slicing-loading-computing joint planning approach, to achieve optimal memory utilization and minimal memory management overhead. We implemented FlexNN atop NCNN, and conducted comprehensive evaluations with common model architectures on various devices. The results have shown that our approach is able to adapt to different memory constraints with optimal latency-memory trade-offs. For example, FlexNN can reduce the memory consumption by 93.81% with only a 3.64% increase in latency, as compared with the original NCNN on smartphones. Yuanchun Li 0003, Yuanzhe Li 0001, Ting Cao 0003, Yunxin Liu 0001 |
MobiCom | 4 |
| 2024 | Empowering In-Browser Deep Learning Inference on Edge Through Just-In-Time Kernel OptimizationabstractWeb is increasingly becoming the primary platform to deliver AI services onto edge devices, making in-browser deep learning (DL) inference more prominent. Nevertheless, the heterogeneity of edge devices, combined with the underdeveloped state of Web hardware acceleration practices, hinders current in-browser inference from achieving its full performance potential on target devices. Fucheng Jia, Shiqi Jiang 0002, Ting Cao 0003, Tianrui Xia, Yuanchun Li 0003, Qipeng Wang 0001, Ju Ren 0001, Yunxin Liu 0001, Lili Qiu, Mao Yang 0004 |
MobiSys | 3 |
| 2024 | Poster: Design of Elastic Deep Neural Network Candidate Spaces for Inference on Diverse DevicesabstractDeep Neural Network (DNN) inference on edge devices is now a common practice. However, tailoring a model for multiple devices involves a lot of time and effort. While elastic models, also known as weight-sharing models, have been proposed as an efficient solution to create high-accuracy, low-latency modules, the design of an elastic model's candidate space (search space) has been underexplored. We identified a new characteristic in candidate spaces, which we named sensitivity, made a design rationale for candidate spaces based on it, and then built a preliminary algorithm to generate candidate spaces. Results show that we can get a range of models (with a 2.75× FLOPs range) on the Pareto frontier of the space by training only once. Jeongho Won, Ting Cao 0003, Huiqiang Jiang, Junehwa Song |
MobiSys | 2 |
| 2024 | LitePred: Transferable and Scalable Latency Prediction for Hardware-Aware Neural Architecture Search
Chengquan Feng, Li Lyna Zhang, Yuanchi Liu, Jiahang Xu, Chengruidong Zhang, Ting Cao 0003, Mao Yang 0004, Haisheng Tan |
NSDI | 7 |
| 2024 | Ladder: Enabling Efficient Low-Precision Deep Learning Computing through Hardware-aware Tensor Transformation
Lei Wang 0222, Lingxiao Ma, Shijie Cao, Quanlu Zhang, Jilong Xue, Yining Shi 0001, Ningxin Zheng, Ziming Miao, Fan Yang 0024, Ting Cao 0003, Yuqing Yang 0001, Mao Yang 0004 |
OSDI | 10 |
| 2024 | ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor CoresabstractTensor Core Unit (TCU) is increasingly integrated into modern high-performance processors to enhance matrix multiplication performance. However, constrained to its over-specification, its potential for improving other critical scientific operations like stencil computations remains untapped. Yuetao Chen, Kun Li 0016, Donglin Bai, Lei Wang 0222, Lingxiao Ma, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
PPoPP | 9 |
| 2024 | Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy SparsityabstractThe adaptation of pre-trained large language models (LLMs) to diverse downstream tasks via fine-tuning is critical for numerous applications. However, the inefficiency of parameterefficient fine-tuning (PEFT) techniques presents significant challenges in terms of time investments and operational costs. In this paper, we first introduce a nuanced form of sparsity, termed Shadowy Sparsity, which is distinctive in fine-tuning and has not been adequately addressed for acceleration. Under Shadowy Sparsity, we propose Long Exposure1, an efficient system to accelerate PEFT for LLMs. Long Exposure comprises three key components: Shadowy-sparsity Exposer employs a prolonged sensing range to capture more sparsity details under shadowy sparsity; Sequence-oriented Predictor provides efficient yet accurate predictions to handle large sequence inputs and constantly-evolving parameters; and Dynamic-aware Operator facilitates more structured computational patterns and coalesced memory accesses, addressing dynamic sparse operations. Extensive evaluations show that Long Exposure outperforms state-of-the-arts with up to a $2.49 \times$ speedup in end-to-end fine-tuning, offering promising advancements in accelerating PEFT for LLMs.1Long Exposure is available at https://github.com/HPHEX/LongExposure. Tuowei Wang, Kun Li 0016, Zixu Hao, Donglin Bai, Ju Ren 0001, Yaoxue Zhang, Ting Cao 0003, Mao Yang 0004 |
SC | 7 |
| 2024 | LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresabstractStencil computations play a pivotal role in numerous scientific and industrial applications, yet their efficient execution on specialized hardware accelerators like Tensor Core Units (TCUs) remains a challenge. This paper introduces LoRAStencil1, a novel stencil computing system designed to mitigate memory access redundancies on TCUs through low-rank adaptation. We first identify a nuanced form of this redundancy, dimension residue, specific to TCUs. Then LoRAStencil leverages orchestrated mathematical transformations to decompose stencil weight matrices into smaller rank-1 matrices, facilitating efficient data gathering along residual dimensions. It comprises three key components: memory-efficient Residual Dimension Gathering to facilitate more data reuse, compute-saving Pyramidal Matrix Adaptation to exploit the inherent low-rank characteristics, and performance-boosting Butterfly Vector Swapping to circumvent all data shuffles. Comprehensive evaluations demonstrate that LoRAStencil address dimension residues effectively, which outperforms state-of-the-arts with up to a 2.16x speedup, offering promising advancements for efficient tensorized stencil computation on TCUs by Low-Rank Adaptation. Yiwei Zhang 0009, Kun Li 0016, Jiawen Cheng, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
SC | 6 |
| 2024 | HiMoDepth: Efficient Training-Free High-Resolution On-Device Depth PerceptionabstractHigh-resolution depth estimation, with a minimum resolution of$1280\times 960$, is essential for achieving more immersive experiences in on-device 3D vision applications. However, implementing high-resolution solutions on resource-limited mobile devices presents significant challenges, such as the need for additional expensive depth sensors, computation-intensive machine learning models requiring large-scale datasets, or the need for device motion while the target object remains stationary. In this study, we propose HiMoDepth, an efficient training-free high-resolution depth estimation system that utilizes widely-available on-device dual cameras. HiMoDepth consists of two modules: 1) homogenizing the on-device heterogeneous cameras by iteratively cropping the Field-of-Views to make the focal length of the cameras equal and filtering out the out-of-sync frames based on time stamps, and 2) designing a hierarchical mobile GPU-friendly stereo matching method that effectively reduces the latency of stereo matching with high-resolution depth maps by using efficient data layout, reducing the number of memory accesses, and searching the corresponding pixel over a coarse-to-fine hierarchy. We implement HiMoDepth on multiple commodity mobile devices and conduct comprehensive evaluations. Experimental results show that HiMoDepth significantly outperforms the baselines in both accuracy and running speed on mobile devices that support high-resolution depth maps. Ju Ren 0001, Bangwen He, Youngki Lee 0001, Ting Cao 0003, Yuanchun Li 0003, Yaoxue Zhang, Yunxin Liu 0001 |
IEEE Trans. Mob. Comput. | 7 |
| 2023 | Adam Accumulation to Reduce Memory Footprints of Both Activations and Gradients for Large-Scale DNN TrainingabstractRunning out of GPU memory has become a main bottleneck for large-scale DNN training. How to reduce the memory footprint during training has received intensive research attention. We find that previous gradient accumulation reduces activation memory but fails to be compatible with gradient memory reduction due to a contradiction between preserving gradients and releasing gradients. To address this issue, we propose a novel optimizer accumulation method for Adam, named Adam Accumulation (AdamA), which enables reducing both activation and gradient memory. Specifically, AdamA directly integrates gradients into optimizer states and accumulates optimizer states over micro-batches, so that gradients can be released immediately after use. We mathematically and experimentally demonstrate AdamA yields the same convergence properties as Adam. Evaluated on transformer-based models, AdamA achieves up to 23% memory reduction compared to gradient accumulation with less than 2% degradation in training throughput. Notably, AdamA can work together with memory reduction methods for optimizer states to fit 1.26×~3.14× larger models over PyTorch and DeepSpeed baseline on GPUs with different memory capacities. Yibo Han, Shijie Cao, Guohao Dai 0001, Youshan Miao, Ting Cao 0003, Fan Yang 0024, Ningyi Xu |
ECAI | 6 |
| 2023 | ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile DevicesabstractNeural Architecture Search (NAS) has shown promising performance in the automatic design of vision transformers (ViT) exceeding 1G FLOPs. However, designing lightweight and low-latency ViT models for diverse mobile devices remains a big challenge. In this work, we propose ElasticViT, a two-stage NAS approach that trains a high-quality ViT supernet over a very large search space for covering a wide range of mobile devices, and then searches an optimal sub-network (subnet) for direct deployment. However, current supernet training methods that rely on uniform sampling suffer from the gradient conflict issue: the sampled subnets can have vastly different model sizes (e.g., 50M vs. 2G FLOPs), leading to different optimization directions and inferior performance. To address this challenge, we propose two novel sampling techniques: complexity-aware sampling and performance-aware sampling. Complexity-aware sampling limits the FLOPs difference among the subnets sampled across adjacent training steps, while covering different-sized subnets in the search space. Performance-aware sampling further selects subnets that have good accuracy, which can reduce gradient conflicts and improve supernet quality. Our discovered models, ElasticViT models, achieve top-1 accuracy from 67.2% to 80.0% on ImageNet from 60M to 800M FLOPs without extra retraining, outperforming all prior CNNs and ViTs in terms of accuracy and latency. Our tiny and small models are also the first ViT models that surpass state-of-the-art CNNs with significantly lower latency on mobile devices. For instance, ElasticViT-S1 runs 2.62× faster than EfficientNet-B0 with 0.1% higher accuracy. Li Lyna Zhang, Huiqiang Jiang, Jiahang Xu, Ting Cao 0003, Quanlu Zhang, Yuqing Yang 0001, Mao Yang 0004 |
ICCV | 5 |
| 2023 | SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 InferenceabstractThe combination of Neural Architecture Search (NAS) and quantization has proven successful in automatically designing low-FLOPs INT8 quantized neural networks (QNN). However, directly applying NAS to design accurate QNN models that achieve low latency on real-world devices leads to inferior performance. In this work, we identify that the poor INT8 latency is due to the quantization-unfriendly issue: the operator and configuration (e.g., channel width) choices in prior art search spaces lead to diverse quantization efficiency and can slow down the INT8 inference speed. To address this challenge, we propose SpaceEvo, an automatic method for designing a dedicated, quantization-friendly search space for each target hardware. The key idea of SpaceEvo is to automatically search hardware-preferred operators and configurations to construct the search space, guided by a metric called Q-T score to quantify how quantization-friendly a candidate search space is. We further train a quantized-for-all supernet over our discovered search space, enabling the searched models to be directly deployed without extra retraining or quantization. Our discovered models, SEQnet, establish new SOTA INT8 quantized accuracy under various latency constraints, achieving up to 10.1% accuracy improvement on ImageNet than prior art CNNs under the same latency. Extensive experiments on real devices show that SpaceEvo consistently outperforms manually-designed search spaces with up to 2.5× faster speed while achieving the same accuracy. Li Lyna Zhang, Jiahang Xu, Quanlu Zhang, Yujing Wang 0002, Yuqing Yang 0001, Ningxin Zheng, Ting Cao 0003, Mao Yang 0004 |
ICCV | 8 |
| 2023 | Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer InferenceabstractDeploying pre-trained transformer models like BERT on downstream tasks in resource-constrained scenarios is challenging due to their high inference cost, which grows rapidly with input sequence length. In this work, we propose a constraint-aware and ranking-distilled token pruning method ToP, which selectively removes unnecessary tokens as input sequence passes through layers, allowing the model to improve online inference speed while preserving accuracy. ToP overcomes the limitation of inaccurate token importance ranking in the conventional self-attention mechanism through a ranking-distilled token distillation technique, which distills effective token rankings from the final layer of unpruned models to early layers of pruned models. Then, ToP introduces a coarse-to-fine pruning approach that automatically selects the optimal subset of transformer layers and optimizes token pruning decisions within these layers through improved L0 regularization. Extensive experiments on GLUE benchmark and SQuAD tasks demonstrate that ToP outperforms state-of-the-art token pruning and model compression methods with improved accuracy and speedups. ToP reduces the average FLOPs of BERT by 8.1X while achieving competitive accuracy on GLUE, and provides a real latency speedup of up to 7.4X on an Intel CPU. Code is available at https://github.com/microsoft/Moonlit/tree/main/ToP Li Lyna Zhang, Jiahang Xu, Yujing Wang 0002, Shaoguang Yan, Yunqing Xia, Yuqing Yang 0001, Ting Cao 0003, Hao Sun 0015, Qi Zhang 0066, Mao Yang 0004 |
KDD | 8 |
| 2023 | LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table LookupabstractOn-device Deep Neural Network (DNN) inference consumes significant computing resources and development efforts. To alleviate that, we propose LUT-NN, the first system to empower inference by table lookup, to reduce inference cost. LUT-NN learns the typical features for each operator, named centroid, and precompute the results for these centroids to save in lookup tables. During inference, the results of the closest centroids with the inputs can be read directly from the table, as the approximated outputs without computations. Xiaohu Tang 0003, Yang Wang 0053, Ting Cao 0003, Li Lyna Zhang, Qi Chen 0009, Deng Cai 0001, Yunxin Liu 0001, Mao Yang 0004 |
MobiCom | 3 |
| 2023 | NN-Stretch: Automatic Neural Network Branching for Parallel Inference on Heterogeneous Multi-ProcessorsabstractMobile devices are increasingly equipped with heterogeneous multiprocessors, e.g., CPU + GPU + DSP. Yet existing Neural Network (NN) inference fails to fully utilize the computing power of the heterogeneous multi-processors due to the sequential structures of NN models. Towards this end, this paper proposes NN-Stretch, a new model adaption strategy, as well as the supporting system. It automatically branches a given model according to the processor architecture characteristics. Compared to other popular model adaption techniques such as model pruning that often sacrifices accuracy, NN-Stretch accelerates inference while preserving accuracy. Jianyu Wei, Ting Cao 0003, Shijie Cao, Shiqi Jiang 0002, Shaowei Fu, Mao Yang 0004, Yanyong Zhang, Yunxin Liu 0001 |
MobiSys | 2 |
| 2023 | Boosting DNN Cold Inference on Edge DevicesabstractDNNs are ubiquitous on edge devices nowadays. With its increasing importance and use cases, it's not likely to pack all DNNs into device memory and expect that each inference has been warmed up. Therefore, cold inference, the process to read, initialize, and execute a DNN model, is becoming commonplace and its performance is urgently demanded to be optimized. To this end, we present NNV12, the first on-device inference engine optimizing cold inference. NNV12 is built atop three novel optimization knobs: selecting a proper kernel (i.e., operator implementation) for each DNN operator, bypassing the weights transformation process by caching the post-transformed weights on disk, and pipelined execution of many kernels on asymmetric processors. To tackle with the huge search space, NNV12 employs a heuristic-based scheme to obtain a near-optimal kernel scheduling plan. We fully implement a prototype of NNV12 and evaluate its performance across extensive experiments. It shows that NNV12 achieves up to 15.2× speedup compared to the state-of-the-art DNN engines on edge CPUs and 401.5× speedup on edge GPUs, respectively. Rongjie Yi, Ting Cao 0003, Ao Zhou 0001, Xiao Ma 0009, Shangguang Wang, Mengwei Xu 0001 |
MobiSys | 2 |
| 2022 | SwiftPruner: Reinforced Evolutionary Pruning for Efficient Ad RelevanceabstractAd relevance modeling plays a critical role in online advertising systems including Microsoft Bing. To leverage powerful transformers like BERT in this low-latency setting, many existing approaches perform ad-side computations offline. While efficient, these approaches are unable to serve cold start ads, resulting in poor relevance predictions for such ads. This work aims to design a new, low-latency BERT via structured pruning to empower real-time online inference for cold start ads relevance on a CPU platform. Our challenge is that previous methods typically prune all layers of the transformer to a high, uniform sparsity, thereby producing models which cannot achieve satisfactory inference speed with an acceptable accuracy. Li Lyna Zhang, Youkow Homma, Yujing Wang 0002, Mao Yang 0004, Ruofei Zhang, Ting Cao 0003 |
CIKM | 7 |
| 2022 | Romou: rapidly generate high-performance tensor kernels for mobile GPUsabstractMobile GPU, as a ubiquitous and powerful accelerator, plays an important role in accelerating on-device DNN (Deep Neural Network) inference. The frequent-upgrade and diversity of mobile GPUs require automatic kernel generation to empower fast DNN deployment. However, current generated kernels have poor performance. Rendong Liang, Ting Cao 0003, Jicheng Wen, Manni Wang, Yang Wang 0053, Jianhua Zou, Yunxin Liu 0001 |
MobiCom | 2 |
| 2022 | MobiDepth: real-time depth estimation using on-device dual camerasabstractReal-time depth estimation is critical for the increasingly popular augmented reality and virtual reality applications on mobile devices. Yet existing solutions are insufficient as they require expensive depth sensors or motion of the device, or have a high latency. We propose MobiDepth, a real-time depth estimation system using the widely-available on-device dual cameras. While binocular depth estimation is a mature technique, it is challenging to realize the technique on commodity mobile devices due to the different focal lengths and unsynchronized frame flows of the on-device dual cameras and the heavy stereo-matching algorithm. Ju Ren 0001, Bangwen He, Ting Cao 0003, Yuanchun Li 0003, Yaoxue Zhang, Yunxin Liu 0001 |
MobiCom | 6 |
| 2022 | CoDL: efficient CPU-GPU co-execution for deep learning inference on mobile devicesabstractConcurrent inference execution on heterogeneous processors is critical to improve the performance of increasingly heavy deep learning (DL) models. However, available inference frameworks can only use one processor at a time, or hardly achieve speedup by concurrent execution compared to using one processor. This is due to the challenges to 1) reduce data sharing overhead, and 2) properly partition each operator between processors. Fucheng Jia, Ting Cao 0003, Shiqi Jiang 0002, Yunxin Liu 0001, Ju Ren 0001, Yaoxue Zhang |
MobiSys | 3 |
| 2022 | Turbo: Opportunistic Enhancement for Edge Video AnalyticsabstractEdge computing is being widely used for video analytics. To alleviate the inherent tension between accuracy and cost, various video analytics pipelines have been proposed to optimize the usage of GPU on edge nodes. Nonetheless, we find that GPU compute resources provisioned for edge nodes are commonly under-utilized due to video content variations, subsampling and filtering at different places of a video analytics pipeline. As opposed to model and pipeline optimization, in this work, we study the problem of opportunistic data enhancement using the non-deterministic and fragmented idle GPU resources. In specific, we propose a task-specific discrimination and enhancement module, and a model-aware adversarial training mechanism, providing a way to exploit idle resources to identify and transform pipeline-specific, low-quality images in an accurate and efficient manner. A multi-exit enhancement model structure and a resource-aware scheduler is further developed to make online enhancement decisions and fine-grained inference execution under latency and GPU resource constraints. Experiments across multiple video analytics pipelines and datasets reveal that our system boosts DNN object detection accuracy by 7.27 -- 11.34% by judiciously allocating 15.81 -- 37.67% idle resources on frames that tend to yield greater marginal benefits from enhancement. Yan Lu 0006, Shiqi Jiang 0002, Ting Cao 0003, Yuanchao Shu |
SenSys | 3 |
| 2022 | Hyperion: A Generic and Distributed Mobile Offloading Framework on OpenCLabstractDespite the significant development of mobile device SoCs, they are still inefficient in computing computation-intensive workloads, such as high-resolution image processing and AR/VR applications. Offloading offers a promising way to leverage cloud or edge servers for acceleration, but existing offloading is limited to specific tasks or specific hardware/software platforms, resulting in significant engineering overhead. To address this problem, we focus on the underlying layer of these applications (i.e., OpenCL) and propose Hyperion, a generic and distributed mobile offloading framework built on OpenCL. To achieve high-performance distributed execution for Hyperion, we first take a deep insight into the OpenCL data structures and design regularity-aware kernel analyzer to analyze the data dependency of work-groups and identify the essential data to offload. Then, context-aware execution time predictor is proposed to estimate the computing time of a given partitioned kernel workload that is highly impacted by many runtime factors. These techniques are integrated into pipeline-enabled and network-adaptive scheduler to make scheduling decisions, which coordinates the kernel partition and workload scheduling to form pipeline processing between data transmission and distributed execution with flexible adaptability to network dynamics. Extensive experimental results demonstrate that Hyperion achieves superior performance with an average 3.80× speedup compared with the best baseline and flexible adaptation to dynamic network conditions and available computing resources. Ziyan Fu 0001, Ju Ren 0001, Yunxin Liu 0001, Ting Cao 0003, Yue-Zhi Zhou, Yaoxue Zhang |
SenSys | 4 |
| 2021 | AsyMo: scalable and efficient deep-learning inference on asymmetric mobile CPUsabstractOn-device deep learning (DL) inference has attracted vast interest. Mobile CPUs are the most common hardware for on-device inference and many inference frameworks have been developed for them. Yet, due to the hardware complexity, DL inference on mobile CPUs suffers from two common issues: the poor performance scalability on the asymmetric multiprocessor, and energy inefficiency. Manni Wang, Shaohua Ding, Ting Cao 0003, Yunxin Liu 0001, Fengyuan Xu |
MobiCom | 3 |
| 2021 | nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devicesabstractWith the recent trend of on-device deep learning, inference latency has become a crucial metric in running Deep Neural Network (DNN) models on various mobile and edge devices. To this end, latency prediction of DNN model inference is highly desirable for many tasks where measuring the latency on real devices is infeasible or too costly, such as searching for efficient DNN models with latency constraints from a huge model-design space. Yet it is very challenging and existing approaches fail to achieve a high accuracy of prediction, due to the varying model-inference latency caused by the runtime optimizations on diverse edge devices. Li Lyna Zhang, Shihao Han, Jianyu Wei, Ningxin Zheng, Ting Cao 0003, Yuqing Yang 0001, Yunxin Liu 0001 |
MobiSys | 5 |