Tianhai Zhao

dblp:17/5513 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
11since 2021 · last 2026
0009-0006-3192-3192ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 5 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Wiseswap: Elastic Datacenter Network-Aware Disaggregated Memory for Multi-Tenant Cloud
abstract
Disaggregated Memory Systems (DMS) hold substantial potential for cloud datacenters but face critical deployment barriers in multi-tenant RDMA environments. Existing DMS designs rely on idealized assumptions-overlooking interference from co-located RDMA applications, oversimplifying fabric topology considerations, and lacking elastic service-level objectives (SLOs) guarantees-resulting in performance degradation and resource inefficiency.
Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao
WWW3
2025 ServerlessRec: Fast Serverless Inference for Embedding-Based Recommender Systems with Disaggregated Memory
Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao
Euro-Par (1)3
2025 Container Workload Prediction Using Deep Domain Adaptation in Transfer Learning
Yunlan Wang, Tianhai Zhao, Jianhua Gu, Zhengxiong Hou, Chengwen Zhong
Euro-Par (1)3
2025 RapidNet: Software-Based Virtual RapidIO for Containerized Intra-Satellite Serving Network
Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao
ICA3PP (5)3
2025 ServerlessLSM: Fast RDMA-Codesigned Disaggregated Compaction for Elastic Serverless LSM-Tree Key-Value Store
abstract
The Log-Structured Merge-tree (LSM-tree) has become a cornerstone of modern key-value stores (KVSs) due to its efficiency in handling write-intensive workloads. However, traditional monolithic LSM-tree designs suffer from write stalls caused by resource contention between Memtable flushing and SSTable compaction, while existing distributed systems adopt coarse-grained elasticity that limits resource utilization and responsiveness. This paper introduces ServerlessLSM, a kernelspace RDMA-odesigned Serverless workflow architecture for LSM-trees. By decoupling Memtable flushing and compaction into independent Serverless functions, ServerlessLSM enables fine-grained elasticity and low-latency state transfers through distributed OS primitives (e.g., remote fork, remote memory mapping). Evaluations demonstrate that ServerlessLSM reduces cold-start latency by$\mathbf{9 8 \%}$and achieves$\mathbf{2. 4} \times$higher throughput compared to state-of-the-art solutions, while maintaining space amplification below 11 %, validating its efficiency and costeffectiveness in cloud environments.
Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao
ICWS3
2025 ServerlessPD: Fast RDMA-Codesigned Disaggregated Prefill-Decoding for Serverless Inference of Large Language Models
abstract
Large Language Model (LLM) inference suffers from inefficiencies in coupled prefill (P) and decoding (D) phases, leading to resource underutilization and scheduling bottlenecks. While disaggregated P-D architectures address this by isolating phases across asymmetric clusters, serverless deployments introduce critical challenges: cold-start latency during autoscaling and costly intermediate state transfers (e.g., KV cache) between distributed prefill and decoding instances. We present ServerlessPD, a system that co-designs remote fork with RDMA to enable near-instant autoscaling and zero-copy state transferring for serverless LLM inference. ServerlessPD introduces a RDMA-based OS kernel-integrated primitive that remotely forks active prefill instances into decoding instances across machines, bypassing cold starts by reusing pre-materialized GPU states, which grants child containers direct copy-on-write access to parent GPU memory. The system further employs GPU context interception to efficiently capture and replicate execution states, ensuring seamless state transfer. To optimize resource utilization, ServerlessPD integrates a dynamic launch-point algorithm that schedules fork operations based on real-time prefill-decoding dynamics, minimizing idle time and overlapping computation with state transfers. ServerlessPD demonstrates that RDMA-codeigned remote fork can unlock near-instant autoscaling and efficient state disaggregation for LLM serving.
Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao
ICWS3
2024 A Hierarchical Storage Mechanism for Hot and Cold Data Based on Temperature Model
Shicong Ma, Tianhai Zhao, Jianhua Gu, Yunlan Wang
DEXA (1)2
2024 LCKV: Learner-Cleaner Optimized Adaptive Key-Value Separated LSM-Tree Store
abstract
Persistent key-value store based on LSM-trees represents one of the most advanced designs. Recent research shows that key-value separation has become a popular optimization method for LSM-tree. However, this storage architecture still incurs significant overhead when dealing with some query- and update-intensive workloads. In this paper, we propose$\text{LC}\text{KV}$, a key-value separated LSM-tree storage system built using the$\underline{L}earner-\underline{C}leaner$optimization to increase throughput. Learner represents the construction of learned indexes to increase query throughput, responsible for building models for the hot-readcold-written keys stored in the LSM-tree and values stored in the cold-written$\mathrm{v}\text{alue logs}$(vLogs). Cleaner represents the garbage$\mathrm{c}\text{ollector}$(GC) aimedat increasing update throughput, responsible not only for garbage collection but also for maintaining the sorting of the cold-written vLog. Evaluations show that LCKV outperforms other state-of-the-art solutions.
Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao
ICCD3
2024 Stochastic Network Calculus Based Quality of Service Guarantee for Multi-class Traffic
abstract
With the rapid development of cloud computing technology, a variety of emerging network traffic types have higher quality of service(QoS) requirements. However, the traditional Internet’s best-effort service cannot meet the demands of cloud computing applications. Therefore, we propose a stochastic network calculus(SNC) based QoS guarantee mechanism consisting of two parts. The first part is MTACC, a multi-threshold adaptive admission control algorithm based on network calculus. MTACC classifies network traffic, calculates resource requirements, and introduces admission probabilities and demarcation parameters. As the network load reaches various thresholds, MTACC adjusts the demarcation parameter to modify the admission probabilities accordingly. The second part is TSRA, a two-stage resource allocation algorithm. In the first stage, basic resources are allocated to ensure the minimum resource requirements of the network traffic based on SNC. In the second stage, additional idle resources are allocated to improve the QoS of the network traffic based on the resource utility function. Finally, we simulate the QoS guarantee techniques using NS3. By comparing with Simple Sum and SCAC, we verify the effectiveness of our proposed QoS guarantee techniques. They provide robust performance guarantees for multi-class traffic such as delay-sensitive, bandwidth-sensitive, and packet loss-sensitive traffic.
Yunlan Wang, Tianhai Zhao, YongKuo Hu, Jianhua Gu, Zhengxiong Hou
IPCCC3
2024 Enhancing campus OS community engagement through the miniOS pilot class: A nine-year journey
Jianhua Gu, Mingxuan Liu 0007, Tianhai Zhao
Future Gener. Comput. Syst.3
2022 Prediction of job characteristics for intelligent resource allocation in HPC systems: a survey and future directions
Zhengxiong Hou, Hong Shen 0001, Xingshe Zhou 0001, Jianhua Gu, Yunlan Wang, Tianhai Zhao
Frontiers Comput. Sci.6
2018 Prediction Method of Blasting Vibration by Optimized GEP Based on Spark
abstract
In order to minimize the damage of engineering blasting vibration, it is very important to accurately predict the blasting peak velocity. The GEP algorithm is used to analyze the relationship between blasting vibration and related parameters. Aiming to improve the time performance of GEP in dealing with large-scale engineering blasting data, the gene structure of GEP is adjusted to head, body and tail. Then the GEP algorithm is paralleled on Spark cluster. The optimized parallel GEP can greatly improves the global search efficiency. The experimental results show that the method can obviously reduce the running time and can improve the prediction accuracy in most cases.
Yunlan Wang, Tianhai Zhao, Zhengxiong Hou, Hussain Khanzada Muzammil
COMPSAC (2)3
2015 Optimizing the fault-tolerance overheads of HPC systems using prediction and multiple proactive actions
Jianhua Gu, Yunlan Wang, Tianhai Zhao
J. Supercomput.4
2013 Research on Optimum Checkpoint Interval for Hybrid Fault Tolerance
Jianhua Gu, Yunlan Wang, Tianhai Zhao
APPT4
2012 A Hybrid Heuristic-Genetic Algorithm for Task Scheduling in Heterogeneous Multi-core System
Jianhua Gu, Yunlan Wang, Tianhai Zhao
ICA3PP (1)4
2012 mHLogGP: A Parallel Computation Model for CPU/GPU Heterogeneous Computing Cluster
Gangfeng Liu, Yunlan Wang, Tianhai Zhao, Jianhua Gu
NPC3
2010 ASAAS: Application Software as a Service for High Performance Cloud Computing
abstract
Currently, SAAS (Software as a Service) solutions are usually provided for business, such as salesforce.com. Few work focus on the application software for high performance scientific computing. However, in the high performance cloud computing environment, traditional application software is not intrinsically service oriented. And the limitation of traditional software licenses is a bottleneck problem for large scale of dynamic users. To enable on-demand services for applications, we propose a solution: Application Software as a Service (ASAAS). It provides a web services portal, an on-demand software license service for the users. Application software is wrapped as web services on the basis of underlying computational resources. With a pay-for-use mode, there is no limitation for the licenses any more. The instant service rate, average job response time, and cost are analyzed for an evaluation. A case of implementation and the evaluation show that ASAAS can bring a much better effect than traditional mechanism.
Zhengxiong Hou, Xingshe Zhou 0001, Jianhua Gu, Yunlan Wang, Tianhai Zhao
HPCC5
2007 A Study on Context-aware Privacy Protection for Personal Information
abstract
By using personal information in a pervasive computing environment, context-aware applications can provide appropriate services for people. This personal information is often involved in personal privacy. In order to protect personal privacy concerns about personal information, privacy role is proposed to control access personal information. We also construct an information system about the privacy decision of personal information disclosure based on people's interaction history. In the initial period of personal information disclosure, the privacy decision is made by people and the information system is constructed based on the decision data. Then privacy disclosure policies are extracted from this information system using rough set theory. According to deducing from the privacy disclosure policies and people's context information, the contextaware application is assigned to an adequate privacy role. It reduces the distraction of privacy decision for people. A case study further shows the proposed method is effective. Finally, it provides about the overload performance of privacy role analysis personaengine.
Qingsheng Zhang, Yong Qi 0001, Jizhong Zhao, Di Hou, Tianhai Zhao, Liang Liu 0010
ICCCN5