Mingyu Wu 0001

dblp:91/7977-1 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-4270-4124ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 8 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 miniK8s: A Pedagogical Cloud-Native System
abstract
The rapid adoption of cloud-native technologies, particularly containerization and orchestration with systems like Kubernetes, necessitates their effective integration into undergraduate computer science curricula. However, the complexity of production-grade cloud-native systems present a steep learning curve. Traditional pedagogical approaches often involve either oversimplified toy projects built from scratch, which lack real-world relevance, or the direct use of complex cloud systems, which can obscure fundamental concepts. To overcome these barriers, we introduce miniK8s, a lightweight, Kubernetes-like platform designed for an undergraduate cloud computing course. miniK8s distinguishes itself by promoting the pedagogical vision of teaching students to build a substantial and realistic system by integrating existing, robust open-source components with a hand-written, simplified cornerstone component. With miniK8s, students can both learn the key concepts inside the cloud architecture (e.g., resource scaling) and build practical systems with open-source building blocks. We have successfully used miniK8s as a project in a cloud computing course for hundreds of undergraduate students, which can significantly enhance students' understanding of how real-world cloud-native systems work and their ability to build a complex system.
Dong Du 0003, Mingyu Wu 0001, Haibo Chen 0001, Binyu Zang
SIGCSE (1)2
2025 Towards Serialization/Deserialization-free State Transfer in Serverless Workflows
abstract
Serialization and deserialization dominate the state transfer time of serverless workflows, leading to substantial performance penalties when executing various serverless workflow applications. We identify the key reason for serialization and deserialization as a lack of ability to efficiently access the (remote) memory of another function. To this end, we propose RMMap , an OS primitive for remote memory map, which allows a serverless function to directly access the memory of another function, even if it is located remotely. RMMap is the first to completely eliminate serialization and deserialization overhead when transferring states between any pairs of functions in (unmodified) serverless workflows. To make remote memory map efficient and feasible, we co-design it with modern networking (RDMA), OS, language runtime, and serverless platform. Evaluations using real-world serverless workloads show that integrating RMMap with Knative reduces the serverless workflow execution time on Knative by up to 2.6× and improves resource utilizations by 86.3%.
Xingda Wei, Fangming Lu, Zhuobin Huang, Rong Chen 0001, Mingyu Wu 0001, Haibo Chen 0001
ACM Trans. Comput. Syst.5
2024 Jade: A High-throughput Concurrent Copying Garbage Collector
abstract
Garbage collection (GC) pauses are a notorious issue threatening the latency of applications. To mitigate this problem, state-of-the-art concurrent copying collectors allow GC threads to run simultaneously with application threads (mutators) in nearly all GC phases. However, the design of concurrent copying collectors does not always lead to low application latency. To this end, this work studies the behaviors of mainstream concurrent copying collectors in OpenJDK and mainly focuses on long application pauses under heavy workloads. By analyzing the design of those collectors, this work uncovers that lengthy pre-reclamation cycles (including GC phases before actual memory release), high GC frequency, and large metadata maintenance overhead are major factors for long pauses. Therefore, this work proposes Jade, a concurrent copying collector aiming to achieve both short pauses and high GC efficiency. Compared with existing collectors, Jade provides a group-wise collection mechanism to shorten pre-reclamation cycles while controlling GC frequency. It also embraces a generational heap layout and a single-phase algorithm to maximize young GC's throughput. The evaluation results on representative latency-critical applications show that Jade can reach sub-millisecond-level pauses even under heavy workloads and significantly improve applications' peak throughput compared with state-of-the-art concurrent collectors.
Mingyu Wu 0001, Yude Lin, Yifeng Jin, Zhe Li 0037, Hongtao Lyu, Denghui Dong, Haibo Chen 0001, Binyu Zang
EuroSys1
2024 Characterization and Reclamation of Frozen Garbage in Managed FaaS Workloads
abstract
FaaS (function-as-a-service) is becoming a popular workload in cloud environments due to its virtues such as auto-scaling and pay-as-you-go. High-level languages like JavaScript and Java are commonly used in FaaS for programmability, but their managed runtimes complicate memory management in the cloud. This paper first observes the issue of frozen garbage, which is caused by freezing cached function instances where their threads have been paused but the unused memory (e.g., garbage) is not reclaimed due to the semantic gap between FaaS and the managed runtime. This paper presents the first characterization of the negative effects induced by frozen garbage with various functions, which uncovers that it can occupy more than half of FaaS instances' memory resources on average. To this end, this paper proposes Desiccant, a freeze-aware memory manager for managed workloads in FaaS, which reclaims idle memory resources consumed by frozen garbage from managed runtime instances and thus notably improves memory efficiency. The evaluation on various FaaS workloads shows that Desiccant can reduce FaaS functions' peak memory consumption by up to 6.72×. Such saved memory consumption allows caching more FaaS instances to reduce the frequency of cold boots (creating instances before function execution) and p99 latency by up to 4.49× and 37.5%, respectively.
Ziming Zhao 0003, Mingyu Wu 0001, Haibo Chen 0001, Binyu Zang
EuroSys2
2024 Serialization/Deserialization-free State Transfer in Serverless Workflows
abstract
Serialization and deserialization play a dominant role in the state transfer time of serverless workflows, leading to substantial performance penalties during workflow execution. We identify the key reason as a lack of ability to efficiently access the (remote) memory of another function. We propose RMMap, an OS primitive for remote memory map. It allows a serverless function to directly access the memory of another function, even if it is located remotely. RMMap is the first to completely eliminates serialization and deserialization when transferring states between any pairs of functions in (unmodified) serverless workflows. To make remote memory map efficient and feasible, we co-design it with fast networking (RDMA), OS, language runtime, and serverless platform. Evaluations using real-world serverless workloads show that integrating RMMap with Knative reduces the serverless workflow execution time on Knative by up to 2.6 × and improves resource utilizations by 86.3%.
Fangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen 0001, Mingyu Wu 0001, Haibo Chen 0001
EuroSys5
2024 Toward an SGX-Friendly Java Runtime
abstract
Hardware enclaves assist in constructing a trusted execution environment (TEE) to store private code and data and thus become an appealing solution to enhance applications’ security. Nevertheless, state-of-the-art enclave implementations like Intel Software Guard Extensions (SGX) have severe performance issues and hinder the deployment of more complicated applications, especially those written in high-level languages like Java. To reduce the performance overhead, prior work has partitioned applications or rebuilt lightweight language runtimes, but they either require manual labor from developers or fail to provide full-fledged support for existing applications. This work instead providesSAJ, a runtime built upon a full-fledged Java virtual machine (JVM) and thus requires no modifications to applications.SAJfirst analyzes the performance of vanilla JVMs running in enclaves and finds that the memory management overhead and boot phase are culprits for performance slowdown. For memory management,SAJintroduces SGX-aware heap layout and garbage collector, which reduces both GC and application execution time. As for the boot phase,SAJintroduces an address-conscious launching mechanism to improve the boot performance. The evaluation under representative Java applications shows thatSAJcan reduce the overall GC pause time, application time, and boot time by 2.93$\boldsymbol{\times}$, 2.58$\boldsymbol{\times}$, and 2.73$\boldsymbol{\times}$on average, respectively.
Mingyu Wu 0001, Zhe Li 0037, Haibo Chen 0001, Binyu Zang, Sanhong Li, Haitao Song 0001
IEEE Trans. Computers1
2023 BeeHive: Sub-second Elasticity for Web Services with Semi-FaaS Execution
abstract
Function-as-a-service (FaaS), an emerging cloud computing paradigm, is expected to provide strong elasticity due to its promise to auto-scale fine-grained functions rapidly. Although appealing for applications with good parallelism and dynamic workload, this paper shows that it is non-trivial to adapt existing monolithic applications (like web services) to FaaS due to their complexity. To bridge the gap between complicated web services and FaaS, this paper proposes a runtime-based Semi-FaaS execution model, which dynamically extracts time-consuming code snippets (closures) from applications and offloads them to FaaS platforms for execution. It further proposes BeeHive, an offloading framework for Semi-FaaS, which relies on the managed runtime to provide a fallback-based execution model and addresses the performance issues in traditional offloading mechanisms for FaaS. Meanwhile, the runtime system of BeeHive selects offloading candidates in a user-transparent way and supports efficient object sharing, memory management, and failure recovery in a distributed environment. The evaluation using various web applications suggests that the Semi-FaaS execution supported by BeeHive can reach sub-second resource provisioning on commercialized FaaS platforms like AWS Lambda, which is up to two orders of magnitude better than other alternative scaling approaches in cloud computing.
Ziming Zhao 0003, Mingyu Wu 0001, Binyu Zang, Haibo Chen 0001
ASPLOS (2)2
2023 Flock: Towards Multitasking Virtual Machines for Function-as-a-Service
abstract
FaaS, or function as a service, promises unprecedented cost-efficiency and elasticity thanks to its on-demand and fine-grained execution nature. However, modern FaaS platforms mainly adopt virtual machines (VMs) or containers as a computing abstraction, which incurs costs like high startup latency, large memory footprint, and high communication overhead. Multi-tasking virtual machines (MVMs), which allow co-executing multiple functions in the same managed language runtime, are appealing for FaaS due to their lightweight nature. Unfortunately, existing MVMs are not designed for FaaS. The proposed abstraction of MVMs does not provide specialized support for fine-grained, latency-sensitive functions and their chain-like execution patterns. Meanwhile, the underlying runtime still contains many global modules and lacks essential support for function-level resource accounting and isolation. To this end, this work proposesFlock, a retrofitted MVM for FaaS execution, which provides FaaS-aware abstractions namedfuncletsand enhanced runtime support for isolation.Flockis implemented atop the HotSpot JVM of OpenJDK 8. Performance evaluation shows thatFlockresults in up to three orders of magnitude performance improvement over state-of-the-art FaaS platforms like OpenWhisk while providing sufficient isolation support for FaaS functions.
Ziming Zhao 0003, Mingyu Wu 0001, Xujie Cao, Haibo Chen 0001, Binyu Zang
IEEE Trans. Computers2
2022 Zero-Change Object Transmission for Distributed Big Data Analytics
Mingyu Wu 0001, Shuaiwei Wang, Haibo Chen 0001, Binyu Zang
USENIX ATC1
2022 Transparent and lightweight object placement for managed workloads atop hybrid memories
abstract
Managed workloads show strong demand for large memory capacity, which can be satisfied by a hybrid memory sub-system composed of traditional DRAM and the emerging non-volatile memory (NVM) technology. Nevertheless, NVM devices are limited by deficiencies like write endurance and asymmetric bandwidth, which threatens managed applications’ performance and reliability. Prior work has proposed different object placement mechanisms to mitigate problems introduced by NVM, but they require domain-specific knowledge on applications or significant change on managed runtime. By analyzing the performance of representative data-intensive workloads atop NVM, this paper finds that reducing write operations is key for performance and wear-leveling. To this end, this paper proposes GCMove, a transparent and efficient object placement mechanism for hybrid memories. GCMove embraces a lightweight write barrier for write detection and relies on garbage collections (GC) to copy objects into different devices according to their write-related behaviors. Compared with prior work, GCMove does not require significant changes in heap layout and thus can be easily integrated with mainstream copy-based garbage collection. The evaluation on various managed workloads shows that GCMove can eliminate 99.8% of NVM write operations on average and improve the performance by up to 19.81× compared with the NVM-only version.
Zhe Li 0037, Mingyu Wu 0001
VEE2
2021 Bridging the performance gap for copy-based garbage collectors atop non-volatile memory
abstract
Non-volatile memory (NVM) is expected to revolutionize the memory hierarchy with not only non-volatility but also large capacity and power efficiency. Memory-intensive applications, which are often written in managed languages like Java, would run atop NVM for better cost-efficiency. Unfortunately, such applications may suffer from performance slowdown due to the unmanaged performance gap between DRAM and NVM. This paper studies the performance of a series of Java applications atop NVM and uncovers that the copy-based garbage collection (GC), the mainstream GC algorithm, is an NVM-unfriendly component in JVM. GC becomes a severe performance bottleneck especially when memory resource is scarce. To this end, this paper analyzes the memory behavior of copy-based GC and uncovers that its inappropriate usage on NVM bandwidth is the main reason for its performance slowdown. This paper thus proposes two NVM-aware optimizations: write cache and header map, to effectively manage the limited NVM bandwidth. It further improves the GC performance with hardware instructions like non-temporal memory accesses and prefetching. We have implemented the optimizations on two mainstream copy-based garbage collectors in OpenJDK. Evaluation with various memory-intensive applications shows that our optimizations can improve the GC time, application execution time, application tail latency by up to 2.69×, 11.0%, and 5.09×, respectively.
Yanfei Yang, Mingyu Wu 0001, Haibo Chen 0001, Binyu Zang
EuroSys2
2020 A Survey on Serverless Computing and Its Implications for JointCloud Computing
abstract
Serverless computing is known as an appealing alternative cloud computing paradigm with its auto-scaling nature and pay-as-you-go charging model. Mainstream cloud vendors have proposed their own serverless platforms, while various kinds of applications have been refactored in a serverless manner for execution. However, the serverless computing model still entails refinement as it introduces performance and security issues. In this paper, we conduct a comprehensive survey on the serverless computing, mainly in three aspects: the type of applications suitable for serverless, the performance issues, and the security issues. We specifically elaborate previous efforts on resolving issues in serverless and shed light on the unresolved issues. We also discuss the opportunities and challenges in integrating serverless computing in the jointcloud infrastructure.
Mingyu Wu 0001, Zeyu Mi, Yubin Xia
JCC1
2020 Platinum: A CPU-Efficient Concurrent Garbage Collector for Tail-Reduction of Interactive Services
Mingyu Wu 0001, Ziming Zhao 0003, Yanfei Yang, Haibo Chen 0001, Binyu Zang, Haibing Guan, Sanhong Li, Chuansheng Lu, Tongbao Zhang
USENIX ATC1
2020 GCPersist: an efficient GC-assisted lazy persistency framework for resilient Java applications on NVM
abstract
The emergence of non-volatile memory (NVM) has stimulated broad interests in building efficient and persistent systems and programming models. However, most prior work is built atop an eager persistency model, which mandates applications to persist their data as soon as possible and thus causes considerable overhead. Besides, prior work mainly focuses on native languages and overlooks the interactions with the managed runtime system in a high-level language. Such issues limit the scope of applications on NVM, especially for resilient applications that already have reliable but inefficient recovery mechanisms. This paper proposes GCPersist, an easy-to-use NVM programming framework atop a lazy persistency model to defer the persistency of user data for better performance, with the assistance of the garbage collection (GC) module in the managed runtime. GCPersist further provides differentiated persistency modes to reduce the runtime overhead. We have implemented GCPersist on the HotSpot JVM of OpenJDK and the evaluation results on Intel Optane DC persistent memory devices show that GCPersist performs well with resilient applications (like Spark) by reducing the recovery time by up to 3.26X while introducing only 1--6% runtime overhead during normal execution.
Mingyu Wu 0001, Haibo Chen 0001, Binyu Zang, Haibing Guan
VEE1
2019 ScissorGC: scalable and efficient compaction for Java full garbage collection
abstract
Java runtime frees applications from manual memory management through automatic garbage collection (GC). This, however, is usually at the cost of stop-the-world pauses. State-of-the-art collectors leverage multiple generations, which will inevitably suffer from a full GC phase scanning and compacting the whole heap. This induces a pause tens of times longer than normal collections, which largely affects both throughput and latency of applications.
Mingyu Wu 0001, Binyu Zang, Haibo Chen 0001
VEE2
2018 Espresso: Brewing Java For More Non-Volatility with Non-volatile Memory
abstract
Fast, byte-addressable non-volatile memory (NVM) embraces both near-DRAM latency and disk-like persistence, which has generated considerable interests to revolutionize system software stack and programming models. However, it is less understood how NVM can be combined with managed runtime like Java virtual machine (JVM) to ease persistence management. This paper proposes Espresso, a holistic extension to Java and its runtime, to enable Java programmers to exploit NVM for persistence management with high performance. Espresso first provides a general persistent heap design called Persistent Java Heap (PJH) to manage persistent data as normal Java objects. The heap is then strengthened with a recoverable mechanism to provide crash consistency for heap metadata. Espresso further provides a new abstraction called Persistent Java Object (PJO) to provide an easy-to-use but safe persistence programming model for programmers to persist application data. Evaluation confirms that Espresso significantly outperforms state-of-art NVM support for Java (i.e., JPA and PCJ) while being compatible to data structures in existing Java programs.
Mingyu Wu 0001, Ziming Zhao 0003, Heting Li, Haibo Chen 0001, Binyu Zang, Haibing Guan
ASPLOS1
2017 POSTER: Recovering Performance for Vector-based Machine Learning on Managed Runtime
abstract
No abstract available.
Mingyu Wu 0001, Haibing Guan, Binyu Zang, Haibo Chen 0001
PPoPP1