VLDB 2026 Research / reviewers in the wild / expert
Mingzhe Zhang 0003
dblp:118/5481-3
· DBLP profile ↗
2ranked-venue papers
1as first author
0since 2021 · last 2019
0000-0002-4672-7884ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 77% Distributed systems · 23% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
non-volatile memory |
0.3 | 1 | 2018 | SIMPO: A Scalable In-Memory Persistent Object Framework Using NVRAM for Reliable Big Data Computing · ACM Trans. Archit. Code Optim. 2018 |
Memory systems › non-volatile memory › persistent memory
persistent object management |
0.3 | 1 | 2018 | SIMPO: A Scalable In-Memory Persistent Object Framework Using NVRAM for Reliable Big Data Computing · ACM Trans. Archit. Code Optim. 2018 |
Distributed systems › fault tolerance
checkpointing |
0.1 | 1 | 2018 | SIMPO: A Scalable In-Memory Persistent Object Framework Using NVRAM for Reliable Big Data Computing · ACM Trans. Archit. Code Optim. 2018 |
Distributed systems
fault tolerance |
0.1 | 1 | 2018 | SIMPO: A Scalable In-Memory Persistent Object Framework Using NVRAM for Reliable Big Data Computing · ACM Trans. Archit. Code Optim. 2018 |
Methods — techniques the papers use, named apart from their topics
write-combining · 0.3lazy evaluation · 0.3consolidated flushing · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | EC-Shuffle: Dynamic Erasure Coding Optimization for Efficient and Reliable Shuffle in SparkabstractFault-tolerance capabilities attract increasing attention from existing data processing frameworks, such as Apache Spark. To avoid replaying costly distributed computation, like shuffle, local checkpoint and remote replication are two popular approaches. They incur significant runtime overhead, such as extra storage cost or network traffic. Erasure coding is another emerging technology, which also enables data resilience. It is perceived as capable of replacing the checkpoint and replication mechanisms for its high storage efficiency. However, it suffers heavy network traffic due to distributing data partitions to different locations. In this paper, we propose EC-Shuffle with two encoding schemes and optimize the shuffle-based operations in Spark or MapReduce-like frameworks. Specifically, our encoding schemes concentrate on optimizing the data traffic during the execution of shuffle operations. They only transfer the parity chunks generated via erasure coding, instead of a whole copy of all data chunks. EC-Shuffle also provides a strategy, which can dynamically select the per-shuffle biased encoding scheme according to the number of senders and receivers in each shuffle. Our analyses indicate that this dynamic encoding selection can minimize the total size of parity chunks. The extensive experimental results using BigDataBench with hundreds of mappers and reducers shows this optimization can reduce up to 50% network traffic and achieve up to 38% performance improvement. Xin Yao 0008, Cho-Li Wang, Mingzhe Zhang 0003 |
CCGRID | 3 |
| 2018 | SIMPO: A Scalable In-Memory Persistent Object Framework Using NVRAM for Reliable Big Data ComputingabstractWhile CPU architectures are incorporating many more cores to meet ever-bigger workloads, advance in fault-tolerance support is indispensable for sustaining system performance under reliability constraints. Emerging non-volatile memory technologies are yielding fast, dense, and energy-efficient NVRAM that can dethrone SSD drives for persisting data. Research on using NVRAM to enable fast in-memory data persistence is ongoing. In this work, we design and implement a persistent object framework, dubbed scalable in-memory persistent object (SIMPO) , which exploits NVRAM, alongside DRAM, to support efficient object persistence in highly threaded big data applications. Based on operation logging, we propose a new programming model that classifies functions into instant and deferrable groups. SIMPO features a streamlined execution model, which allows lazy evaluation of deferrable functions and is well suited to big data computing workloads that would see improved data locality and concurrency. Our log recording and checkpointing scheme is effectively optimized towards NVRAM, mitigating its long write latency through write-combining and consolidated flushing techniques. Efficient persistent object management with features including safe references and memory leak prevention is also implemented and tailored to NVRAM. We evaluate a wide range of SIMPO-enabled applications with machine learning, high-performance computing, and database workloads on an emulated hybrid memory architecture and a real hybrid memory machine with NVDIMM. Compared with native applications without persistence, experimental results show that SIMPO incurs less than 5% runtime overhead on both platforms and even gains up to 2.5× speedup and 84% increase in throughput in highly threaded situations on the two platforms, respectively, thanks to the streamlined execution model. Mingzhe Zhang 0003, King Tin Lam, Xin Yao 0008, Cho-Li Wang |
ACM Trans. Archit. Code Optim. | 1 |