EDBT 2026 Demo / reviewers in the wild / expert
Yanning Yang
dblp:252/1235
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
Yanning Yang, Dong Du 0003, Zhigang Mao, Naifeng Jing, Yubin Xia, Haibo Chen 0001 |
ISCA | 4 |
| 2026 | SC-AFE: A channel-adaptive and feature-enhanced semantic communication for wireless image transmission
Chunhua Zhu, Yanning Yang, Bohui Wang |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | Resource-Efficient Orchestration for Heterogeneous Serverless Computing with Harmonized Effectiveness and PracticabilityabstractCurrent serverless platforms struggle to optimize resource utilization for both CPU and GPU functions due to their dynamic and fine-grained nature. Conventional techniques like overcommitment and autoscaling fall short, often sacrificing utilization for practicability or incurring performance tradeoffs. Overcommitment requires predicting performance to prevent QoS violation, introducing tradeoff between prediction accuracy and overheads. Autoscaling requires scaling instances in response to load fluctuations quickly to reduce resource wastage, but more frequent scaling also leads to more cold start overheads. The rich concurrency of GPU resources further complicates GPU instance orchestration, such as setting right batch sizes. This article introduces Jiagu to harmonize efficiency with practicability through the following novel techniques. First, pre-decision scheduling achieves accurate prediction while eliminating overheads by decoupling prediction and scheduling. Second, dual-staged scaling achieves frequent adjustment of instances with minimum overhead. Third, Jiagu conducts an in-depth analysis about the complexity of the relationship between GPU function configuration and execution. It then proposes batch-aware scaling that achieves optimal configurations for both batch size setting and autoscaling, addressing all the challenges according to the analysis. We have implemented a prototype and evaluated it using real-world applications and traces from the public cloud platform. Our evaluation shows an improvement in deployment density over commercial clouds (with Kubernetes) while maintaining QoS for both CPU and GPU functions (54.8% and 18% respectively), and 81.0%–93.7% lower scheduling costs and a 57.4%–69.3% reduction in cold start latency compared to existing QoS-aware schedulers. Yanning Yang, Dong Du 0003, Yubin Xia, Haibo Chen 0001 |
ACM Trans. Comput. Syst. | 2 |
| 2024 | On-demand and Parallel Checkpoint/Restore for GPU ApplicationsabstractLeveraging serverless computing for cloud-based machine learning services is on the rise, promising cost-efficiency and flexibility are crucial for ML applications relying on high-performance GPUs and substantial memory. However, despite modern serverless platforms handling diverse devices like GPUs seamlessly on a pay-as-you-go basis, a longstanding challenge remains: startup latency, a well-studied issue when serverless is CPU-centric. For example, initializing GPU apps with minor GPU models, like MobileNet, demands several seconds. For more intricate models such as GPT-2, startup latency can escalate to around 10 seconds, vastly overshadowing the short computation time for GPU-based inference. Prior solutions tailored for CPU serverless setups, like fork() and Checkpoint/Restore, cannot be directly and effectively applied due to differences between CPUs and GPUs. Yanning Yang, Dong Du 0003, Haitao Song 0001, Yubin Xia |
SoCC | 1 |
| 2024 | Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in Serverless Computing with Jiagu
Yanning Yang, Dong Du 0003, Yubin Xia, James R. Larus, Haibo Chen 0001 |
USENIX ATC | 2 |
| 2022 | ScalaRAID: optimizing linux software RAID system for next-generation storageabstractRAID has been widely adopted to enhance the performance, capacity, and reliability of the existing storage systems. However, we observe that the Linux software RAID (mdraid) suffers from its poor implementation of the lock mechanism. To address this, we propose ScalaRAID, which refines the role domain of locks and designs a new data structure to prevent different threads from preempting the RAID resources. By doing so, ScalaRAID can maximize the thread-level parallelism and reduce the time consumption of I/O request handling. Our evaluation results reveal that ScalaRAID can improve throughput by 89.4% while decreasing 99.99th percentile latency by 85.4% compared to mdraid. Shushu Yi, Yanning Yang, Yunxiao Tang, Chen Yue, Myoungsoo Jung, Jie Zhang 0048 |
HotStorage | 2 |
| 2019 | A Fog Computing Paradigm for Efficient Information Services in VANETabstractWith recent advances in wireless communications, vehicular networks have attracted great interests in both industry and academia. This work aims at proposing a novel vehicular fog computing paradigm including both the system architecture and the scheduling algorithm. Specifically, we present a hierarchical architecture, which integrates the paradigm of both fog computing and the software defined networking (SDN). Then, we formulate a novel problem called Cooperative Service in Vehicular Fog Computing (CS-VFC), which aims at maximizing the bandwidth efficiency by coordinating the service in both the fog layer and the cloud layer. We prove that CS-VFC is NP-hard. On this basis, we propose an on-line scheduling algorithm, which incorporates with the network coding and makes scheduling decisions at SDN controller. In particular, it will determine the coding policy for each cloud node, and then it will implement both the intra and inter cooperation strategies at the fog layer. Finally, we build the simulation model by implementing NS3 simulator and SUMO. A comprehensive simulation is carried out to demonstrate the superiority of the proposed system architecture and the solution. Ke Xiao 0001, Kai Liu 0001, Yanning Yang, Liang Feng 0001, Jingjing Cao, Victor C. S. Lee |
WCNC | 4 |