Yitao Hu

dblp:148/1971 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
20since 2021 · last 2026
0009-0004-0458-0900ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 2 first-author · 16 since 2021Computer networks · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel
abstract
LLM serving is increasingly dominated by decode attention, which is a memory-bound operation due to massive KV cache loading from global memory. Meanwhile, real-world workloads exhibit substantial, hierarchical shared prefixes across requests (e.g., system prompts, tools/templates, RAG). Existing attention implementations fail to fully exploit prefix sharing: one-query-per-CTA execution repeatedly loads shared prefix KV cache, while one-size-fits-all tiling leaves on-chip resources idle and exacerbates bubbles for uneven KV lengths. These choices amplify memory bandwidth pressure and stall memory-bound decode attention.
Jinjun Yi, Yitao Hu, Hao Wang 0022, Laiping Zhao, Yuhao Zhang 0006, Wenxin Li 0001, Keqiu Li
ASPLOS (2)3
2026 PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
abstract
Modern deep neural network (DNN) and large language model (LLM) applications integrate multiple models into inference pipelines with stringent latency requirements for customized tasks. To mitigate extensive request timeouts caused by accumulation, systems for inference pipelines commonly drop a subset of requests so the remaining ones can satisfy latency constraints. Since it is commonly believed that request dropping adversely affects goodput, existing systems only drop requests when they have to, which we call reactive dropping. However, this reactive policy can not maintain high goodput, as it neither makes timely dropping decisions nor identifies the proper set of requests to drop, leading to issues of dropping requests too late or dropping the wrong set of requests.
Yitao Hu, Mingfang Ji, Wei Yang 0013, Yuhao Zhang 0006, Laiping Zhao, Wenxin Li 0001, Xiulong Liu 0001, Wenyu Qu, Hao Wang 0022
EuroSys2
2026 Accelerating ML Inference via Opportunistic Pre-Loading on Serverless Clusters
abstract
Serverless computing has emerged as a novel paradigm in cloud computing, characterized by its agile scalability, cost-effective pay-as-you-go billing, and user-friendly capabilities for Machine Learning (ML) inference tasks. Developers wrap their ML algorithms into serverless functions and run them in containers. However, the well-known cold-start problem significantly slows down the response time of functions. To address cold-starts, the technique of pre-warming, which proactively maintains containers in a warm state, has gained widespread adoption across both research and industry. Nevertheless, we observed that pre-warming does not address the distinct delays caused by the loading of ML artifacts. According to our analysis, in ML inference functions, the time required to load libraries and models significantly exceeds the time needed to warm containers. Thus, relying solely on pre-warming is insufficient for mitigating cold-starts. This paper presentsTyche, an opportunistic pre-loading approach designed to eliminate the latency associated with loading ML artifacts, enabling near-instant inference and minimizing function execution time.Tychefully leverages the idle memory in warmed containers and GPUs to pre-load required libraries and models, striking an optimal balance between acceleration and resource efficiency. Additionally,Tycheis tailored for large-scale serverless platforms, incorporating cluster-wide scheduling and lightweight locality-aware load balancing to enhance performance. We designTycheto be transparent to providers and compatible with existing pre-warming solutions. Experiments on OpenWhisk with real-world workloads show thatTychereduces up to 93% loading latency and achieves up to 8× speedup compared to state-of-the-art pre-warming solutions. Compared with the state-of-the-art serverless pre-loading solution,Tychealso achieves up to 1.9× speedup.
Yifan Sui, Hanfei Yu, Yitao Hu, Hao Wang 0022
IEEE Trans. Parallel Distributed Syst.3
2026 ViDA: Lossless VideoQA Acceleration via Selective Sparse Self-Speculation With Parallel Computational Load Management
abstract
Video large language models (VideoLLMs) have significantly advanced video question answering (VideoQA) applications, which demand both low latency and high accuracy. To meet the requirements, VideoLLMs are typically deployed on GPUs for parallel acceleration. However, the massive computational load from long video contexts often makes such acceleration insufficient. Existing techniques like token pruning and speculative decoding attempt to address this challenge by altering the computational load, but often fail to balance both speed and accuracy. Sparse self-speculation mitigates these limitations via selecting a subset of tokens on a specific token budget to draft the output and then verify it using all tokens. However, existing sparse self-speculation is designed for text-based scenarios and cannot be directly applied to VideoQA tasks, as it fails to account for discrepancies in critical multimodal tokens and the dynamic nature of optimal token budget in VideoQA, leading to suboptimal scale and inappropriate composition of parallel computational load. We argue that achieving both high accuracy and low latency in VideoQA tasks requires managing the computational load with awareness of these discrepancies and dynamics. To achieve this, we introduce ViDA, a selective sparse self-speculation inference system. It progressively searches token budgets by iteratively refining lower and upper bounds of search space derived from long contexts, aiming to find and allocate varying optimal budgets in real-time adaptively. Additionally, it leverages insights from discrepancies in critical multimodal tokens to perform a discrepancy-aware token selection approach for identifying critical tokens. Evaluations across various VideoQA workloads show that compared to state-of-the-art methods, ViDA preserves exact model outputs while reducing average time-per-outputtoken (TPOT) by 15% to 46% and average end-to-end latency by up to 28%, while decreasing the divergence from optimal token budget distribution by up to 93
Yitao Hu, Yuhao Zhang 0006, Laiping Zhao, Wenxin Li 0001, Keqiu Li
IEEE Trans. Parallel Distributed Syst.2
2025 SmartCache: Two-Dimensional KV-Cache Similarity for Efficient Long-Context LLM Decoding
abstract
Large language models (LLMs) achieve state-of-the-art performance in many NLP tasks but incur prohibitive memory-access and compute costs when processing very long contexts due to linearly growing KV Cache. Existing static sparsification methods rely on fixed heuristics, while dynamic schemes incur substantial runtime overhead. To address this trade-off, we propose SmartCache, a sparse inference system that exploits two-dimensional KV Cache similarity across adjacent decoding iterations and neighboring layers. SmartCache combines a similarity-driven dual-path selection algorithm, which adaptively reuses TopK KV entries from both the previous iteration and the preceding layer with a rolling-array cache index manager that reduces index storage complexity from$O(L \cdot k)$to$O(k)$. We analyze the layer and iterative sparse patterns of KV Cache in long context LLM decoding and show that SmartCache maintains semantic consistency while drastically reducing redundant computation and memory traffic. Extensive experiments on Llama-3-8B-Instruct-Gradient-1048k, Qwen2.5-7B-Instruct-1M, and glm-4-9b-chat-1m across four long-context benchmarks report up to$\mathbf{3 0. 5 \%}$end-to-end latency reduction and$\mathbf{1 5} \boldsymbol{\%} \mathbf{- 2 3 \%}$average latency reduction, with inference accuracy degradation constrained within 2% and occasional slight improvements. These results indicate that SmartCache offers a practical, high-accuracy solution for scalable long-sequence LLM inference.
Kaining Hui, Yitao Hu, Sheng Chen 0015, Xiulong Liu 0001, Keqiu Li
HPCC6
2025 SuperSpec: Enhanced Verification and Sampling for End-to-End LLM Speculative Decoding
abstract
Modern LLM decoding has the drawbacks of high cost and slow speed, and speculative decoding has been shown to be an effective solution to this problem. However, the inference latency still poses a significant challenge to maintaining service level objectives (SLOs) in systems that employ multiple draft models for speculative decoding. The verification phase in such systems if reliant on tree attention often constitutes a bottleneck especially when draft sequences lack common prefixes and substantially underutilizes GPU parallelism while increasing end-to-end latency. We introduce SuperSpec, an end-to-end speculative decoding system designed to co-optimize verification, sampling and draft generation. SuperSpec integrates three pivotal innovations: an Efficient Batch Verifier, which substitutes treebased flattening with batch parallel validation and layer-wise KV Cache replication; a Global Optimal Sampler, which assesses all candidate sequences within a batch to ascertain the longest valid path, thereby circumventing the local optima frequently encountered in tree-based rejection sampling; and a Dynamic Adaptive Multi-Drafter, which dynamically modulates the speculative length (K) for each drafter predicated on real-time idleness metrics and acceptance rates. Empirical evaluations of Qwen2.5-72B and the OPT-66B on various datasets show that SuperSpec improves average acceptance rate by 6.4% to 30.2%, and the end-to-end inference acceleration ratio by 7.12% to 62.06%, when compared to the state-of-the-art tree-based speculative decoding system SpecInfer. These improvements were achieved without compromising the quality of text generation, making SuperSpec an effective solution for accelerating LLM inference.
Yitao Hu, Sheng Chen 0015, Xiulong Liu 0001, Keqiu Li
HPCC6
2025 MoEoM: Joint Compute and Memory-Aware Balancing for Fast MoE Inference
abstract
Mixture-of-Experts (MoE) architectures have emerged as a scalable and efficient alternative to dense Transformer models by activating only a subset of experts per layer. However, deploying MoE models in multi-GPU environments faces severe challenges due to expert load imbalance and the resulting inefficient GPU utilization. Existing static replication strategies fail to adapt to dynamic token distributions, while dynamic rebalancing methods incur excessive communication and memory-access overheads, often outweighing the benefits of load balancing. This paper presents MoEoM, an inference system that enhances the efficiency of MoE models by innovatively taking memory-access costs into consideration. MoEoM integrates two complementary techniques: (i) a load-aware offline expert deployer, which symmetrically groups experts across GPUs and selectively replicates high-load experts, and (ii) an I/O-aware online token reallocator, which dynamically redistributes tokens among original and backup experts to minimize the maximum latency across GPUs. Experimental evaluations on state-of-theart MoE models demonstrate that MoEoM reduces end-toend inference latency by up to 16.6 %, improves throughput by 11–20 % in the prefill stage and 11–23 % in the decode stage, and decreases cross-GPU imbalance by 65–80 % (about 75 % on average), compared to prior MoE inference baselines. These results highlight the importance of incorporating both computation and memory-access overhead into expert placement and token scheduling for efficient large-scale MoE deployment.
Ziqi Gong, Yitao Hu, Sheng Chen 0015, Wenxin Li 0001, Keqiu Li
ICPADS2
2025 Lark: A Buffer-aware Building Block for Programmable Packet Scheduling in Datacenters
abstract
Programmable packet scheduling enables users to customize scheduling algorithms flexibly without designing new ASICs. Existing schemes prefre to approximate optimal Push-In First-Out (PIFO) using First-In First-Out (FIFO) queues in commodity programmable switches. Despite its availability, these schemes suffer performance degradation due to the unawareness of available switch buffer. To be specific, when the port buffer is drained, existing schemes discard all incoming packets, even though these packets have higher priorities than the enqueued packets. In this paper, we reveal that the problem's key culprit is the lack of coordination between buffer management and packet scheduling in the switch. To fill this gap, we present Lark, a buffer-aware building block for programmable scheduling schemes designed to solve the above problem. Its key idea is to proactively drop the low-priority packets when the allocated buffer is to be drained, thereby admitting the later-arriving high-priority packets. Lark contains two modules, a lightweight gradient-based online prediction module and a simple priority-based decision module. Lark relies the former module to identify whether the allocated buffer is to be drained and uses the later to determine whether to drop the incoming packet. We have integrated Lark into two representative schemes, SP-PIFO and AIFO. Our large-scale evaluations over three realistic workloads show that Lark can significantly optimize their key metrics without sacrificing throughnut.
Song Zhang 0008, Wenxin Li 0001, Yulong Li 0001, Lide Suo, Sheng Chen 0015, Yitao Hu, Laiping Zhao, Keqiu Li
INFOCOM7
2025 Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
abstract
Advances in deep neural networks (DNNs) have significantly contributed to the development of real-time video processing applications. Efficient scheduling of DNN workloads in cloud-hosted inference systems is crucial to minimizing serving costs while meeting application latency constraints. However, existing systems suffer from excessive module latency during request dispatching, low execution throughput during module scheduling, and wasted latency budget during latency splitting for multi-DNN applications, which undermines their capability to minimize the serving cost. In this paper, we design a DNN inference system called Harpagon, which minimizes the serving cost under latency constraints with a three-level design. It first maximizes the batch collection rate with a batch-aware request dispatch policy to minimize the module latency. It then maximizes the module throughput with multi-tuple configurations and proper amount of dummy requests. It also carefully splits the end-to-end latency into per-module latency budget to minimize the total serving cost for multi-DNN applications. Evaluation shows that Harpagon outperforms the state of the art by 1.49 to 2.37 times in serving cost while satisfying the latency objectives. Additionally, compared to the optimal solution using brute force search, Harpagon derives the lower bound of serving cost for 91.5% workloads with millisecond level runtime.
Yitao Hu, Ziqi Gong, Guotao Yang, Wenxin Li 0001, Xiulong Liu 0001, Keqiu Li, Hao Wang 0022
INFOCOM2
2025 TightLLM: Maximizing Throughput for LLM Inference via Adaptive Offloading Policy
abstract
Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks, largely due to their substantial model size. However, this also results in significant GPU memory demands during inference. To address these challenges on hardware with limited GPU memory, existing approaches employ offloading techniques that offload unused tensors to CPU memory, thereby reducing GPU memory usage. Since offloading involves data transfer between GPU and CPU, it introduces transfer overhead. To mitigate this, prior works typically overlap data transfer with GPU computation using a fixed pipelining strategy applied uniformly across all inference iterations, referred to asstaticoffloading. However, static offloading policies fail to maximize inference throughput because they cannot adapt to the dynamically changing transfer overhead during the inference process, leading to increasing GPU idleness and reduced inference throughput.We propose that offloading policies should beadaptiveto the varying transfer overhead across inference iterations to maximize inference throughput. To this end, we design and implement an adaptive offloading-based inference system called TightLLM with two key innovations. First, its key-value (KV) distributor employs atrade-compute-for-transferstrategy to address growing transfer overhead by dynamically recomputing portions of the KV cache, effectively overlapping data transfer with computation and minimizing GPU idleness. Second, TightLLM’s weight loader slices model weights and distributes the loading processacross multiple batches, amortizing the excessive weight loading overhead and significantly improving throughput. Evaluation across various combinations of GPU hardware and LLM models shows that TightLLM achieves 1.3 to 23 times higher throughput during the decoding phase and 1.2 to 22 times higher throughput in the prefill phase compared to state-of-the-art offloading systems. Due to the higher throughput in prefill and decoding phases, TightLLM can reduce the completion time for large-scale tasks, which involve processing and generating a substantial number of tokens, by 59.6% to 94.9%.
Yitao Hu, Xiulong Liu 0001, Guotao Yang, Sheng Chen 0015, Laiping Zhao, Wenxin Li 0001, Keqiu Li
IEEE Trans. Computers1
2025 SLOpt: Serving Real-Time Inference Pipeline With Strict Latency Constraint
abstract
The rise of Machine Learning as a Service (MLaaS) has driven the demand for complex and customized real-time inference tasks, often requiring cascading multiple deep neural network (DNN) models into inference pipelines. However, these pipelines pose significant challenges due to scheduling complexity, particularly in maintaining strict latency service level objectives (SLOs). Existing systems serve pipelines with model-independent scheduling policies, which ignore the unique workload characteristics introduced by model cascading in the inference pipeline, leading to SLO violations and resource inefficiencies. In this paper, we propose that the serving system should exploit the model-cascading nature and inter-model workload dependency of the inference pipeline to ensure strict latency SLO cost-effectively. Based on this, we design and implementSLOpt, a serving system optimized for real-time inference pipelines with a three-stage co-design of workload estimation, resource provisioning, and request execution.SLOptproposes cascade workload estimation and ahead-of-time tuning, which together address the challenge of cascade blocking and head-of-line blocking in workload estimation and resource provisioning.SLOptfurther implements an adaptive batch drop policy to mitigate latency amplification issues within the pipeline. These innovations enableSLOptto reduce the 99th percentile latency (P99 latency) by 1.4 to 2.5 times compared to the state of the arts while lowering serving costs by up to 29%. Moreover, to achieve comparable P99 latency,SLOptrequires up to 70% less cost than existing systems. Extensive evaluations on a 64-GPU cluster demonstrateSLOpt’s effectiveness in meeting strict P99 latency SLOs under diverse real-world workloads.
Yitao Hu, Guotao Yang, Ziqi Gong, Laiping Zhao, Wenxin Li 0001, Xiulong Liu 0001, Wenyu Qu
IEEE Trans. Computers2
2024 FUYAO: DPU-enabled Direct Data Transfer for Serverless Computing
abstract
Serverless computing typically relies on the third-party forwarding method to transmit data between functions. This method couples control flow and data flow together, resulting in significantly slow data transmission speeds. This challenge makes it difficult for the serverless computing paradigm to meet the low-latency requirements of web services.
Laiping Zhao, Zhaolin Duan, Sheng Chen 0015, Yitao Hu, Zhiyuan Su, Wenyu Qu
ASPLOS (3)6
2024 Pre-Warming is Not Enough: Accelerating Serverless Inference With Opportunistic Pre-Loading
abstract
Serverless computing has rapidly prospered as a new cloud computing paradigm with agile scalability, pay-as-you-go pricing, and ease-to-use features for Machine Learning (ML) inference tasks. Users package their ML code into lightweight serverless functions and execute them using containers. Unfortunately, a notorious problem, called cold-starts, hinders serverless computing from providing low-latency function executions. To mitigate cold-starts, pre-warming, which keeps containers warm predictively, has been widely accepted by academia and industry. However, pre-warming fails to eliminate the unique latency incurred by loading ML artifacts. We observed that for ML inference functions, the loading of libraries and models takes significantly more time than container warming. Consequently, pre-warming alone is not enough to mitigate the ML inference function's cold-starts.
Yifan Sui, Hanfei Yu, Yitao Hu, Hao Wang 0022
SoCC3
2024 WQEFC: A Scalable and Low-Latency RDMA Messages Scheduler for Mixed Messages
abstract
RDMA has been widely deployed to improve the performance of applications with frequently fine-grained remote access. However, restricted on-chip resources result in cache misses under high concurrency that significantly degrade network performance. QPC-aware solutions only focus on the number of concurrent QPs, ignoring the impact of WQE within QPs. SMART limits the number of WQEs in each QP with a credit-based scheme. Nevertheless, we find that equal treatment increases the tail latency of messages ranging from 32 bytes to 1024 bytes by 2× when mixing messages of different sizes. In this paper, we introduce WQEFC, a scalable RDMA message scheduler that provides lower latency and higher throughput for applications with heavily concurrent messages. Our key insight is that there is a significant difference in the sensitivity to cache miss between messages of different sizes. For messages smaller than 32 bytes, which are sensitive to cache misses, we combine the credit limiter and sub-message poller, limiting the number of concurrent wqes to avoid cache miss while ensuring optimal message completion latency. For other messages, which are insensitive to cache miss, we assign them a higher priority and use the sub-message poller to ensure message concurrency while reducing cache miss. We implement WQEFC as a middleware between the driver layer and the application layer for flexible deployment. WQEFC outperforms the state-of-the-art solution Smart by increasing system throughput by 61.3%, and reducing the tail latency of messages smaller than 32 bytes and larger than 32 bytes by 41.6% and 72.6%, respectively.
Yaozhen Li, Lide Suo, Xiancheng Meng, Yiren Pang, Wenxin Li 0001, Keqiu Li, Yitao Hu
HPCC8
2024 Efficient Disaggregated Memory Eviction with Glitter
abstract
Memory disaggregation, a promising technique allowing applications to use remote memory, is increasingly appealing in datacenters due to its high resource utilization. Operationally, the application’s host server constantly evicts unused data to remote to make room for memory allocation of new pages. Inefficient evictions allow memory usage to hit its limit, resulting in application blocking, which brings severe throughput degradation. However, most existing works neglect the importance of eviction. They offload the eviction to a background thread and set a fixed trigger timing, rendering a belated eviction. Worse still, they overlook the impact of network congestion on eviction efficiency, making their strategy flawed in large-scale scenarios. In this paper, we present Glitter, an adaptive, multi-level awareness eviction solution that accelerates applications by minimizing the overhead of application blocking from host and network aspects. For host, Glitter presents an adaptive eviction threshold adjustment to optimize the eviction timing, reducing the occurrence of application blocking. For network, Glitter adopts an eviction flow scheduling to address the hazards posed by flow contention at switches, decreasing the duration of each application blocking. Through comprehensive experiments, Glitter gives an average 1.4 throughput boost to Fastswap, a state-of-the-art disaggregated×memory system.
Linxuan Zhong, Wenxin Li 0001, Yulong Li 0001, Jiawen Shen, Song Zhang 0008, Wenyu Qu, Yitao Hu
HPCC7
2024 PPT: A Pragmatic Transport for Datacenters
abstract
This paper introduces PPT, a pragmatic transport that achieves comparable performance to proactive transports while maintaining good deployability as reactive transports. Our key idea is to run a low-priority control loop to leverage the available bandwidth left by the reactive transports. The main challenge is to send just enough packets to improve performance without harming the primary control loop. We combine two unconventional techniques: an intermittent loop initialization and an exponential window decrease, enabling us to dynamically identify and fill the spare bandwidth. We further complement PPT's design with a buffer-aware flow scheduling scheme to optimize the average FCT of small flows without prior knowledge of flow size information. We have implemented a PPT prototype in the Linux kernel with ~400 lines of code and demonstrated that compared to Homa, it delivers up to 46.3% lower overall average FCT and even 25%/55.5% lower average/tail FCT of small flows in an Memcached workload.
Lide Suo, Yiren Pang, Wenxin Li 0001, Renjie Pei, Keqiu Li, Xiulong Liu 0001, Xin He 0043, Yitao Hu, Guyue Liu
SIGCOMM8
2023 DeepLat: Achieving Minimum Worst Case Latency for DNN Inference with Batch-Aware Dispatching
Jiaheng Gao, Yitao Hu
ICA3PP (1)2
2023 High-throughput Sampling, Communicating and Training for Reinforcement Learning Systems
abstract
Reinforcement Learning (RL) algorithms require large amounts of computational resources and time to train due to the simultaneous model training and real-time interactions with simulation environments. This results in an RL algorithm's training time potentially exceeding days to months. We present HRL, a comprehensive optimization system designed to improve the system throughput of RL algorithms. HRL addresses the challenges of discovering and identifying bottlenecks in the sampling, training, or communication stages and promptly improving their performance. We propose a group-parallel pipeline method to improve the sampling efficiency and apply data quantization and multi-learner training to resolve the network and learning bottlenecks. The HRL system has been fully implemented and integrated with XingTian. The results of the experiments show that the HRL system can improve the throughput by 18.6%-90.6%.
Laiping Zhao, Xinan Dai, Yusong Xin, Yitao Hu, Keqiu Li
IWQoS5
2023 Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network
abstract
Container overlay network, though being widely adopted to enable communication between containers on different hosts, is a key downside for latency-sensitive applications. The state-of-the-art solution seeks to shorten the data path in packet processing by replacing overlay connection file descriptors with host namespace ones. While promising, it must block each overlay connection until the relevant host connection is set up, thus heavily influencing the request latency. In this paper, we present ShuntFlow, a systematic data delivery framework that seamlessly integrates the host and overlay networks to reduce the application's request-response latency. ShuntFlow first lets all connections flow in the overlay network directly. Then, it adopts a simple-yet-effective syscall-threshold-based mechanism to pick appropriate connections and switches their data delivery to the host network in a blocking-free way using a multi-threading technique. As such, unnecessary connection switches are prevented; yet, the pre-setup phase dilemma is eliminated. We have implemented a ShuntFlow prototype based on Linux and Docker and evaluated it extensively on a 40 Gbps testbed. The results show that ShuntFlow achieves 13%/72% and 19%/69% reductions, in average/tail request-response latency of a web server and an in-memory key-value store, respectively, while incurring less CPU overhead, compared to Slim.
Wenxin Li 0001, Yiren Pang, Renjie Pei, Yitao Hu, Lide Suo, Keqiu Li
IEEE Trans. Parallel Distributed Syst.5
2021 Scrooge: A Cost-Effective Deep Learning Inference System
abstract
Advances in deep learning (DL) have prompted the development of cloud-hosted DL-based media applications that process video and audio streams in real-time. Such applications must satisfy throughput and latency objectives and adapt to novel types of dynamics, while incurring minimal cost. Scrooge, a system that provides media applications as a service, achieves these objectives by packing computations efficiently into GPU-equipped cloud VMs, using an optimization formulation to find the lowest cost VM allocations that meet the performance objectives, and rapidly reacting to variations in input complexity (e.g., changes in participants in a video). Experiments show that Scrooge can save serving cost by 16-32% (which translate to tens of thousands of dollars per year) relative to the state-of-the-art while achieving latency objectives for over 98% under dynamic workloads.
Yitao Hu, Rajrup Ghosh, Ramesh Govindan
SoCC1
2018 Olympian: Scheduling GPU Usage in a Deep Neural Network Model Serving System
abstract
Deep neural networks (DNNs) are emerging as important drivers for GPU (Graphical Processing Unit) usage. Routinely, now, cloud offerings include GPU-capable VMs, and GPUs are used for training and testing DNNs. A popular way to run inference (or testing) tasks with DNNs is to use middleware called a serving system. Tensorflow-Serving (TF-Serving) is an example of a DNN serving system. In this paper, we consider the problem of carefully scheduling multiple concurrent DNNs in a serving system on a single GPU to achieve fairness or service differentiation objectives, a capability crucial to cloud-based TF-Serving offerings. In scheduling DNNs, we face two challenges: how to schedule, and switch between, different DNN jobs at low overhead; and, how to account for their usage. Our system, Olympian, extends TF-Serving to enable fair sharing of a GPU across multiple concurrent large DNNs at low overhead, a capability TF-Serving by itself is not able to achieve. Specifically, Olympian can run concurrent instances of several large DNN models such as Inception, ResNet, GoogLeNet, AlexNet and VGG, provide each with an equal share of the GPU, while interleaving them at timescales of 1-2 ms, and incurring an overhead of less than 2%. It achieves this by leveraging the predictability of GPU computations to profile GPU resource usage models offline, then using these to achieve low overhead switching between DNNs.
Yitao Hu, Swati Rallapalli, Bong Jun Ko, Ramesh Govindan
Middleware1
2016 ALPS: accurate landmark positioning at city scales
abstract
Context awareness is crucial for ubiquitous computing, and position is an important aspect of context. In an ideal world, every stationary object or entity in the built environment would be associated with position, so that applications can have precise spatial context about the environment surrounding a human. In this paper, we take a step towards this ideal: by analyzing images from Google Street View that cover different perspectives of a given object and triangulating the location of the object, our system, ALPS, can discover and localize common landmarks at the scale of a city accurately and with high coverage. ALPS contains several novel techniques that help improve the accuracy, coverage, and scalability of localization. Evaluations of ALPS on many cities in the United States show that it can localize storefronts with a coverage higher than 90% and a median error of 5 meters.
Yitao Hu, Suman Nath, Ramesh Govindan
UbiComp1
2015 Data Acquisition for Real-Time Decision-Making under Freshness Constraints
abstract
The paper describes a novel algorithm for timely sensor data retrieval in resource-poor environments under freshness constraints. Consider a civil unrest, national security, or disaster management scenario, where a dynamic situation evolves and a decision-maker must decide on a course of action in view of latest data. Since the situation changes, so is the best course of action. The scenario offers two interesting constraints. First, one should be able to successfully compute the course of action within some appropriate time window, which we call the decision deadline. Second, at the time the course of action is computed, the data it is based on must be fresh (i.e., within some corresponding validity interval). We call it the freshness constraint. These constraints create an interesting novel problem of timely data retrieval. We address this problem in resource-scarce environments, where network resource limitations require that data objects (e.g., pictures and other sensor measurements pertinent to the decision) generally remain at the sources. Hence, one must decide on (i) which objects to retrieve and (ii) in what order, such that the cost of deciding on a valid course of action is minimized while meeting data freshness and decision deadline constraints. Such an algorithm is reported in this paper. The algorithm is shown in simulation to reduce the cost of data retrieval compared to a host of baselines that consider time or resource constraints. It is applied in the context of minimizing cost of finding unobstructed routes between specified locations in a disaster zone by retrieving data on the health of individual route segments.
Shaohan Hu, Shuochao Yao, Haiming Jin, Yiran Zhao 0001, Yitao Hu, Nooreddin Naghibolhosseini, Shen Li 0002, Akash Kapoor, William Dron, Lu Su 0001, Amotz Bar-Noy, Pedro A. Szekely, Ramesh Govindan, Reginald L. Hobbs, Tarek F. Abdelzaher
RTSS5
2014 Critical sensing range for mobile heterogeneous camera sensor networks
abstract
In camera sensor networks (CSNs), full view coverage, in which any direction of any point in the operational region is covered by at least one camera sensor, is of great significance since image shot at the frontal viewpoint considerably increases the possibility to recognize the object. However, finding the critical condition to achieve full view coverage in mobile heterogeneous CSNs remains an open question. In this paper, we analyze both the static and mobile random deployed camera sensor networks. A centralized parameter - equivalent sensing radius (ESR) - is defined to evaluate the critical requirement for asymptotic full view coverage in heterogeneous CSNs. We derive the critical sensing range for full view coverage under static model, 2-dimensional random walk mobility model, 1-dimensional random walk mobility model and random rotating model. We then discuss the impact of various mobility patterns on sensing energy consumption and study the relationship between ESR and percentage of full view coverage, and show that random walk mobility model can decrease the sensing energy consumption under certain delay tolerance. To our knowledge, our work is the very first that derive the critical condition to achieve full view coverage in mobile heterogeneous CSNs.
Yitao Hu, Xinbing Wang, Xiaoying Gan
INFOCOM1