EDBT 2026 Demo / reviewers in the wild / expert
Yifan Wang 0005
dblp:47/6959-5
· DBLP profile ↗
15ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-3878-0343ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 7 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling High-Utilization and Low-Contention FaaS: A Request-Level Resource Provisioning ApproachabstractFunction-as-a-Service offers cost efficiency but often suffers from resource underutilization. This underutilization stems from the instance-level resource provisioning pattern, an issue that existing optimizations have failed to resolve fundamentally. The core problem is that static coarse-grained instance-level resource allocation cannot match the millisecond-level burstiness of dynamic requests. Consequently, it is difficult for current systems to achieve high resource utilization while maintaining high quality of service (QoS) guarantees. To address the problem, this paper advocates a shift to request-level resource provisioning, which redefines the individual request as the atomic unit for scheduling and resource management. We implement this approach in RRP, a scalable FaaS platform that enables efficient per-request resource allocation and release. RRP unifies instance placement and request routing with low-overhead, millisecond-level global visibility. Our evaluation shows that RRP significantly outperforms state-of-the-art instance-level platforms and algorithms. By matching resources to each request’s needs and isolating them from contention, RRP achieves low latency and high utilization. Specifically, on real-world Azure traces, RRP achieves speedups of 1.33 × –30.15 × for average end-to-end latency and 1.37 × –61.46 × for P99 latency, and raises CPU utilization from 44.80%–56.32% to 72.49% under bursty loads. Runfu Li, Zishu Yu, Yifan Wang 0005, Xiaohui Peng 0002, Ninghui Sun, Zhiwei Xu 0002 |
HPDC | 3 |
| 2026 | Group-TopK: Optimizing Distributed Training on Edge Devices via Communication Compression
Yifan Wang 0005, Xiaohui Peng 0002, Haohao Ma, Hui Sun 0002, Deke Guo, Boyu Diao |
HPDC | 1 |
| 2026 | LASS: Reducing Cold Startup Latency in Serverless Through Loaded Library SharingabstractIn serverless scenario, function invocation runs in an individual container. Lightweight container technology has significantly reduced the startup latency of container. The library loading process now becomes a critical performance bottleneck of serverless function cold startup. The state-of-the-art approaches leverage the process fork operation to reduce the cold startup latency in serverless computing by reusing the loaded libraries. However, the fork operation can only share libraries between parent process and forked process. For security, the libraries loaded by the parent process should be a subset of those required by the forked process, which limits opportunities to eliminate library loading overhead. To address this problem, we propose theLASSsystem, which enables multiple processes to share initialized libraries in a composable and efficient manner.LASSallows a process to securely reuse libraries loaded by multiple processes, thereby reducing library loading latency to the millisecond level. Compared to the state-of-the-art approaches,LASScan improve average library loading speed by more than 10.3×, and reduce 99thpercentile end-to-end latency by 34%–57%. Zishu Yu, Runfu Li, Yifan Wang 0005, Xiaohui Peng 0002, Zhiwei Xu 0002 |
IEEE Trans. Computers | 4 |
| 2026 | Breaking Cloud Dependencies: A Distributed Ledger Approach to IoT Device Usufruct Management
Xiaohui Peng 0002, Yifan Wang 0005, Deke Guo |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Tide: A Distributed Runtime Management Framework for Things-Edge-Cloud Computing ContinuumabstractThe increasing number of connected IoT devices produces massive amounts of sensed data at the network edge. A new computing paradigm, called Things-Edge-Cloud (TEC) collaboration, has been proposed to meet real-time and high-throughput requirements. Most of the existing work focuses on workload scheduling across the computing nodes in TEC with ad-hoc implementations using runtime and management frameworks designed for the cloud. In this paper, we propose RSEP to model the entities and their relationship in the TEC computing continuum (TEC3). We then design and implement Tide—a distributed runtime management framework for TEC3based on the RSEP model, which enables elastic resource allocation and seamless computation offloading. It employs runtime environment isolation and physical resource binding to enforce strong isolation without incurring performance penalties. To decouple runtime and framework, Tide provides a set of portable application interfaces that allow the management of variety runtimes. We implement Tide from scratch and compare its latency and throughput with KubeEdge, Ray, and bare-metal implementations. Experimental results show that Tide improves throughput by 2x and reduces average latency, 95th percentile latency, and latency standard deviation by${4 5. 8 2 \%, 4 8. 3 6 \%}$, and${2 8. 9 6 \%}$, respectively. Specifically, Tide achieves${8 7. 3 \%}$of the ideal goodput, exceeding other platforms more than 10x. Xiaohui Peng 0002, Wenkai Yan, Yifan Wang 0005, Shoujian Zheng, Zhiwei Xu 0002 |
IPDPS | 3 |
| 2025 | A Case Study on Benchmarking Distributed AI SystemsabstractThe rapid growth of artificial intelligence (AI), particularly in computer vision (CV), necessitates distributed computing for efficient model training. Existing benchmarks often lack adaptability to emerging scenarios or focus on limited applications. To address these gaps, this paper studies a case on a comprehensive benchmark suite for distributed AI training systems. We classifies AI tasks into four categories, LargeScale, Moderate Complexity, High Load, and High-Performance, based on single-load computation and load concurrency, with representative models evaluated on Ray and DeepSpeed across diverse hardware. The experiments reveal fragmented framework performance. DeepSpeed excels in stability and efficiency for Large-Scale and Moderate Complexity tasks, leveraging advanced memory optimization. Ray outperforms in High Load and High-Performance tasks due to its dynamic resource scheduling but shows greater variability. These results highlight the need for task-specific framework selection tailored to hardware and performance requirements. We provides valuable insights for optimizing distributed AI training and bridges limitations in current benchmarks. Future work aims to expand task categories and framework support to align with the evolving demands of distributed AI systems. Jianwei Gao, Xiaohui Peng 0002, Jiamu Liu, Yifan Wang 0005, Deke Guo |
IWQoS | 5 |
| 2025 | EdgeInferFlow: A Distributed Inference Acceleration Method for Deep Learning Chained Structure Models for Edge Devices
Hanfeng Zhai, Yifan Wang 0005, Xiaohui Peng 0002, Xueqi Li 0001 |
NPC (1) | 2 |
| 2025 | DSparse: A Distributed Training Method for Edge Clusters Based on Sparse Update
Xiaohui Peng 0002, Yixuan Sun, Zhenghui Zhang, Yifan Wang 0005 |
J. Comput. Sci. Technol. | 4 |
| 2024 | Snapipeline: Accelerating Snapshot Startup for FaaS ContainersabstractDue to the frequent starts and stops of numerous services in FaaS (Function as a Service), reducing cold start overhead is a core issue in improving the performance of container-based FaaS services. Snapshot and restore-based mechanisms effectively reduce the cold start time of containers by transforming container initialization overhead into restoration overhead. Consequently, this mechanism has become a research hotspot in accelerating the cold start of FaaS containers. Researchers introduce snapshot compression and decompress the snapshots to reduce the storage cost before starting instances. However, existing works have the following shortcomings: (1) File-mapped memory pages are not processed during snapshot compression, resulting in a significant amount of redundant data in memory; (2) The serial execution of snapshot decompression and instance restoration leads to high instance startup latency. Yuqiao Lan, Xiaohui Peng 0002, Yifan Wang 0005 |
SoCC | 3 |
| 2024 | CMS: A Computility Resource Status Management and Storage Framework
Xiaohui Peng 0002, Kuo Chang, Yifan Wang 0005 |
NPC (1) | 4 |
| 2024 | Hawk: An Efficient NALM System for Accurate Low-Power Appliance RecognitionabstractNon-intrusive Appliance Load Monitoring (NALM) aims to recognize individual appliance usage from the main meter without indoor sensors. However, existing systems struggle to balance dataset construction efficiency and event/state recognition accuracy, especially for low-power appliance recognition. This paper introduces Hawk, an efficient and accurate NALM system that operates in two stages: dataset construction and event recognition. In the data construction stage, we efficiently collect a balanced and diverse dataset, HawkDATA, based on balanced Gray code and enable automatic data annotations via a sampling synchronization strategy called shared perceptible time. During the event recognition stage, our algorithm pipeline integrates steady-state differential pre-processing and voting-based post-processing for accurate event recognition from the aggregate current. Experimental results show that HawkDATA takes only 1/71.5 of the collection time to collect 6.34x more appliance state combinations than the baseline. In HawkDATA and a widely used dataset, Hawk achieves an average F1 score of 93.94% for state recognition and 97.07% for event recognition, which is a 47.98% and 11.57% increase over SOTA algorithms. Furthermore, selected appliance subsets and the model trained from HawkDATA are deployed in two real-world scenarios with many unknown background appliances. The average F1 scores of event recognition are 96.02% and 94.76%. Hawk's source code and HawkDATA are accessible at https://github.com/WZiJ/SenSys24-Hawk. Xingzhou Zhang, Yifan Wang 0005, Xiaohui Peng 0002, Zhiwei Xu 0002 |
SenSys | 3 |
| 2020 | Fengyi: Trusted Data Sharing in VANETs with BlockchainabstractSuperior to traditional vehicles, intelligent vehicles (IV) can share data in Vehicular Ad-Hoc Networks (VANETs) to provide a more comfortable and safer driving experience, based on the assumption that the data sharing is accountable and reliable. However, trusted data sharing in VANETs is always a paramount concern. We define that the trusted data sharing includes the three properties, namely data sharing accountability, privacy preservation, and transmission confidentiality. To address the problem, we propose a comprehensive solution including the trusted ledger model (TLM) and implement the Fengyi system to verify it. The TLM is a model that ensures the consistency of multiple data resources in a low trust distributed computing environment. Then, distributed Fengyi ledgers based on the TLM are proposed to keep data sharing accountable and private in VANETs. The Fengyi system is designed and implemented to provide authentication and encrypted communication services with the ledgers. Finally, we deploy the Fengyi system on three different platforms, checking the effectiveness and efficiency of the Fengyi system. The results show that the system can ensure trusted data sharing in VANETs, and the time cost for the verification of data sharing on-road is average 253.33μs, 38% lower than that in recent research. Yifan Wang 0005, Xiaohui Peng 0002 |
PRDC | 2 |
| 2019 | OpenEI: An Open Framework for Edge IntelligenceabstractIn the last five years, edge computing has attracted tremendous attention from industry and academia due to its promise to reduce latency, save bandwidth, improve availability, and protect data privacy to keep data secure. At the same time, we have witnessed the proliferation of AI algorithms and models which accelerate the successful deployment of intelligence mainly in cloud services. These two trends, combined together, have created a new horizon: Edge Intelligence (EI). The development of EI requires much attention from both the computer systems research community and the AI community to meet these demands. However, existing computing techniques used in the cloud are not applicable to edge computing directly due to the diversity of computing sources and the distribution of data sources. We envision that there missing a framework that can be rapidly deployed on edge and enable edge AI capabilities. To address this challenge, in this paper we first present the definition and a systematic review of EI. Then, we introduce an Open Framework for Edge Intelligence (OpenEI), which is a lightweight software platform to equip edges with intelligent processing and data sharing capability. We analyze four fundamental EI techniques which are used to build OpenEI and identify several open problems based on potential research directions. Finally, four typical application scenarios enabled by OpenEI are presented. Xingzhou Zhang, Yifan Wang 0005, Sidi Lu, Liangkai Liu, Lanyu Xu, Weisong Shi |
ICDCS | 2 |
| 2019 | Edge Computing for Autonomous Driving: Opportunities and ChallengesabstractSafety is the most important requirement for autonomous vehicles; hence, the ultimate challenge of designing an edge computing ecosystem for autonomous vehicles is to deliver enough computing power, redundancy, and security so as to guarantee the safety of autonomous vehicles. Specifically, autonomous driving systems are extremely complex; they tightly integrate many technologies, including sensing, localization, perception, decision making, as well as the smooth interactions with cloud platforms for high-definition (HD) map generation and data storage. These complexities impose numerous challenges for the design of autonomous driving edge computing systems. First, edge computing systems for autonomous driving need to process an enormous amount of data in real time, and often the incoming data from different sensors are highly heterogeneous. Since autonomous driving edge computing systems are mobile, they often have very strict energy consumption restrictions. Thus, it is imperative to deliver sufficient computing power with reasonable energy consumption, to guarantee the safety of autonomous vehicles, even at high speed. Second, in addition to the edge system design, vehicle-to-everything (V2X) provides redundancy for autonomous driving workloads and alleviates stringent performance and energy constraints on the edge side. With V2X, more research is required to define how vehicles cooperate with each other and the infrastructure. Last, safety cannot be guaranteed when security is compromised. Thus, protecting autonomous driving edge computing systems against attacks at different layers of the sensing and computing stack is of paramount concern. In this paper, we review state-of-the-art approaches in these areas as well as explore potential solutions to address these challenges. Shaoshan Liu, Liangkai Liu, Jie Tang 0003, Bo Yu 0014, Yifan Wang 0005, Weisong Shi |
Proc. IEEE | 5 |
| 2018 | OpenVDAP: An Open Vehicular Data Analytics Platform for CAVsabstractIn this paper, we envision the future connected and autonomous vehicles (CAVs) as a sophisticated computer on wheels, with substantial on-board sensors as data sources and a variety of services running on top to support autonomous driving or other functions. In general, these services are computationally expensive, especially for the machine learning based applications (e.g., CNN-based object detection). Nevertheless, the on-board computation unit possess limited compute resources, raising a huge challenge to deploy these computation-intensive services on the vehicle. On the contrary, the cloud-based architecture conceptually with unconstrained resources suffers from unexpected extended latency that attributes to the large-scale Internet data transmission; thus, adversely affecting the services' real-time performance, quality of services and user experiences. To address this dilemma, inspired by the promising edge computing paradigm, we propose to build an Open Vehicular Data Analytics Platform (OpenVDAP) for CAVs, which is a full-stack edge based platform including an on-board computing/communication unit, an isolation-supported and security & privacy-preserved vehicle operation system, an edge-aware application library, as well as an optimal workload of?oading and scheduling strategy, allowing CAVs to dynamically detect each service's status, computation overhead and the optimal of?oading destination so that each service could be finished within an acceptable latency and limited bandwidth consumption. Most importantly, contrast to the proprietary platform, OpenVDAP is an open-source platform that offers free APIs and real-?eld vehicle data to the researchers and developers in the community, allowing them to deploy and evaluate applications on the real environment. Qingyang Zhang 0001, Yifan Wang 0005, Xingzhou Zhang, Liangkai Liu, Xiaopei Wu, Weisong Shi, Hong Zhong 0001 |
ICDCS | 2 |