VLDB 2026 Research / reviewers in the wild / expert
Junxian Zhao
dblp:259/7298
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2023
0009-0000-4761-6847ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Let It Go: Relieving Garbage Collection Pain for Latency Critical Applications in GolangabstractGarbage Collection (GC) is a representative automatic memory manager widely deployed in popular programming languages, such as Java, C\#, and Golang (Go). Through GC, these languages provide programmers with flexibility and safety. However, GC leads to non-trivial overhead in compute and memory resources during application runtime. GC threads compete with non-GC threads (mutators) of an application, which particularly impacts latency-critical (LC) applications and causes long tail latency. Existing GC approaches do not efficiently address the interference, as GC is triggered passively without a global insight of the application; or they employ incremental GC to reduce the interference, while the incremental progress is not dynamically tailored during GC process according to runtime characteristics, which leads to significant performance degradation upon bursty requests. Junxian Zhao, Xiaobo Zhou 0002, Sang-Yoon Chang, Cheng-Zhong Xu 0001 |
HPDC | 1 |
| 2022 | Improving Concurrent GC for Latency Critical Services in Multi-tenant SystemsabstractFor resource utilization efficiency, latency critical (LC) services are commonly co-located with best-effort batch jobs in datacenter servers. Many LC services, such as Cassandra and HBase, run in Java Virtual Machine (JVM). We find that LC services often experience heavy-tailed latency due to performance interference of the concurrent garbage collection (GC) as well as multi-tenancy. The root cause is a semantic gap of resource allocation between JVM and the underlying Linux OS in multi-tenant systems. That is, the OS is unaware of the characteristics of different kinds of threads in JVM (i.e., GC threads and LC worker threads), which may lead to GC threads competing for CPUs; JVM is unaware of the resource utilization in the OS, which may trigger CPU-intensive GC operations when CPUs are busy. Furthermore, we find that co-located batch jobs can interfere with LC services due to Simultaneous Multi-Threading (SMT). Junxian Zhao, Aidi Pi, Xiaobo Zhou 0002, Sang-Yoon Chang, Cheng-Zhong Xu 0001 |
Middleware | 1 |
| 2021 | FlashByte: Improving Memory Efficiency with Lightweight Native StorageabstractIn-memory caching of intermediate data is effective in reducing re-computation and I/O cost in distributed data-analytics frameworks, but it also generates a large amount of data in Java heap which increases the overhead of garbage collection (GC). An alternative off-heap approach caches data in native storage by transmitting the data from heap to native storage so as to reduce GC overhead. However, it incurs severe serialization and de-serialization overheads. Serialization also generates non-trivial metadata of the cached data in native storage. We propose and develop FlashByte, a lightweight native storage that efficiently caches intermediate data. FlashByte improves memory efficiency by achieving low GC overhead, low data transmission overhead, and low memory consumption. Specifically, the cached data are divided into two parts: metadata stored in Java heap and raw data stored in native storage. The metadata is generated based on the profile of workloads. Its size is trivial because it only contains a concise format of raw data, which achieves low memory consumption in the heap as well as low GC overhead. Native storage only stores the raw data to reduce its memory consumption. According to the metadata, the raw data are efficiently transmitted between the heap and native storage without serialization and de-serialization. We implement FlashByte in Spark and conduct evaluation with benchmark workloads. Experimental results show that, compared with the in-heap approach of Vanilla Spark, FlashByte achieves up to 4x speedup of the job execution time, reduces GC time by up to 96%, and reduces the memory consumption in heap by up to 36%. Compared with the alternative off-heap approach, FlashByte achieves up to 2.3x speedup of the job execution time, reduces the data transmission time by up to 84%, and reduces the cache size in native storage by up to 34%. Junxian Zhao, Aidi Pi, Xiaobo Zhou 0002 |
CCGRID | 1 |
| 2021 | Memory at your service: fast memory allocation for latency-critical servicesabstractCo-location and memory sharing between latency-critical services, such as key-value store and web search, and best-effort batch jobs is an appealing approach to improving memory utilization in multi-tenant datacenter systems. However, we find that the very diverse goals of job co-location and the GNU/Linux system stack can lead to severe performance degradation of latency-critical services under memory pressure in a multi-tenant system. Aidi Pi, Junxian Zhao, Xiaobo Zhou 0002 |
Middleware | 2 |