EDBT 2026 Demo / reviewers in the wild / expert
Jianhua Gu
dblp:22/4705
· DBLP profile ↗
27ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-0804-3593ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Wiseswap: Elastic Datacenter Network-Aware Disaggregated Memory for Multi-Tenant CloudabstractDisaggregated Memory Systems (DMS) hold substantial potential for cloud datacenters but face critical deployment barriers in multi-tenant RDMA environments. Existing DMS designs rely on idealized assumptions-overlooking interference from co-located RDMA applications, oversimplifying fabric topology considerations, and lacking elastic service-level objectives (SLOs) guarantees-resulting in performance degradation and resource inefficiency. Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao |
WWW | 2 |
| 2025 | ServerlessRec: Fast Serverless Inference for Embedding-Based Recommender Systems with Disaggregated Memory
Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao |
Euro-Par (1) | 2 |
| 2025 | Container Workload Prediction Using Deep Domain Adaptation in Transfer Learning
Yunlan Wang, Tianhai Zhao, Jianhua Gu, Zhengxiong Hou, Chengwen Zhong |
Euro-Par (1) | 5 |
| 2025 | RapidNet: Software-Based Virtual RapidIO for Containerized Intra-Satellite Serving Network
Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao |
ICA3PP (5) | 2 |
| 2025 | ServerlessLSM: Fast RDMA-Codesigned Disaggregated Compaction for Elastic Serverless LSM-Tree Key-Value StoreabstractThe Log-Structured Merge-tree (LSM-tree) has become a cornerstone of modern key-value stores (KVSs) due to its efficiency in handling write-intensive workloads. However, traditional monolithic LSM-tree designs suffer from write stalls caused by resource contention between Memtable flushing and SSTable compaction, while existing distributed systems adopt coarse-grained elasticity that limits resource utilization and responsiveness. This paper introduces ServerlessLSM, a kernelspace RDMA-odesigned Serverless workflow architecture for LSM-trees. By decoupling Memtable flushing and compaction into independent Serverless functions, ServerlessLSM enables fine-grained elasticity and low-latency state transfers through distributed OS primitives (e.g., remote fork, remote memory mapping). Evaluations demonstrate that ServerlessLSM reduces cold-start latency by$\mathbf{9 8 \%}$and achieves$\mathbf{2. 4} \times$higher throughput compared to state-of-the-art solutions, while maintaining space amplification below 11 %, validating its efficiency and costeffectiveness in cloud environments. Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao |
ICWS | 2 |
| 2025 | ServerlessPD: Fast RDMA-Codesigned Disaggregated Prefill-Decoding for Serverless Inference of Large Language ModelsabstractLarge Language Model (LLM) inference suffers from inefficiencies in coupled prefill (P) and decoding (D) phases, leading to resource underutilization and scheduling bottlenecks. While disaggregated P-D architectures address this by isolating phases across asymmetric clusters, serverless deployments introduce critical challenges: cold-start latency during autoscaling and costly intermediate state transfers (e.g., KV cache) between distributed prefill and decoding instances. We present ServerlessPD, a system that co-designs remote fork with RDMA to enable near-instant autoscaling and zero-copy state transferring for serverless LLM inference. ServerlessPD introduces a RDMA-based OS kernel-integrated primitive that remotely forks active prefill instances into decoding instances across machines, bypassing cold starts by reusing pre-materialized GPU states, which grants child containers direct copy-on-write access to parent GPU memory. The system further employs GPU context interception to efficiently capture and replicate execution states, ensuring seamless state transfer. To optimize resource utilization, ServerlessPD integrates a dynamic launch-point algorithm that schedules fork operations based on real-time prefill-decoding dynamics, minimizing idle time and overlapping computation with state transfers. ServerlessPD demonstrates that RDMA-codeigned remote fork can unlock near-instant autoscaling and efficient state disaggregation for LLM serving. Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao |
ICWS | 2 |
| 2024 | A Hierarchical Storage Mechanism for Hot and Cold Data Based on Temperature Model
Shicong Ma, Tianhai Zhao, Jianhua Gu, Yunlan Wang |
DEXA (1) | 3 |
| 2024 | LCKV: Learner-Cleaner Optimized Adaptive Key-Value Separated LSM-Tree StoreabstractPersistent key-value store based on LSM-trees represents one of the most advanced designs. Recent research shows that key-value separation has become a popular optimization method for LSM-tree. However, this storage architecture still incurs significant overhead when dealing with some query- and update-intensive workloads. In this paper, we propose$\text{LC}\text{KV}$, a key-value separated LSM-tree storage system built using the$\underline{L}earner-\underline{C}leaner$optimization to increase throughput. Learner represents the construction of learned indexes to increase query throughput, responsible for building models for the hot-readcold-written keys stored in the LSM-tree and values stored in the cold-written$\mathrm{v}\text{alue logs}$(vLogs). Cleaner represents the garbage$\mathrm{c}\text{ollector}$(GC) aimedat increasing update throughput, responsible not only for garbage collection but also for maintaining the sorting of the cold-written vLog. Evaluations show that LCKV outperforms other state-of-the-art solutions. Mingxuan Liu 0007, Jianhua Gu, Tianhai Zhao |
ICCD | 2 |
| 2024 | Stochastic Network Calculus Based Quality of Service Guarantee for Multi-class TrafficabstractWith the rapid development of cloud computing technology, a variety of emerging network traffic types have higher quality of service(QoS) requirements. However, the traditional Internet’s best-effort service cannot meet the demands of cloud computing applications. Therefore, we propose a stochastic network calculus(SNC) based QoS guarantee mechanism consisting of two parts. The first part is MTACC, a multi-threshold adaptive admission control algorithm based on network calculus. MTACC classifies network traffic, calculates resource requirements, and introduces admission probabilities and demarcation parameters. As the network load reaches various thresholds, MTACC adjusts the demarcation parameter to modify the admission probabilities accordingly. The second part is TSRA, a two-stage resource allocation algorithm. In the first stage, basic resources are allocated to ensure the minimum resource requirements of the network traffic based on SNC. In the second stage, additional idle resources are allocated to improve the QoS of the network traffic based on the resource utility function. Finally, we simulate the QoS guarantee techniques using NS3. By comparing with Simple Sum and SCAC, we verify the effectiveness of our proposed QoS guarantee techniques. They provide robust performance guarantees for multi-class traffic such as delay-sensitive, bandwidth-sensitive, and packet loss-sensitive traffic. Yunlan Wang, Tianhai Zhao, YongKuo Hu, Jianhua Gu, Zhengxiong Hou |
IPCCC | 5 |
| 2024 | Qualitative QoS-aware Scheduling of Moldable Parallel Jobs on HPC Clusters*abstractIn service oriented high-performance computing (HPC) clusters, end users have various Quality of Service (QoS) requirements. Most of the existing research work focuses on quantitative QoS requirements, such as deadlines, for rigid jobs. While, in many cases, it is more convenient for users to qualitatively state QoS requirements (such as performance-sensitive) at the submission of their jobs. Almost all kinds of QoS requirements will be greatly impacted by job scheduling, which determine the degree of job parallelism, execution time and waiting time, etc. Most modern parallel applications are moldable in the sense that they can choose a resource allocation before execution. Traditional sequential job scheduling mechanism with fixed resource allocation appears to be an obstacle to improve QoS for end users. To address this issue, we propose a novel qualitative QoS-aware sub-queue simultaneous scheduling method for moldable parallel jobs (with variable resource allocation) on HPC Clusters. We first define the qualitative QoS models for end users, then present our sub-queue simultaneous scheduling method, including job sequencing and resource allocation algorithms for a set of moldable parallel jobs on multi-core clusters. Our method can efficiently sequence queuing jobs and allocate appropriate resources for simultaneously running some performance-sensitive jobs in a sub-queue rather than running them one by one. Experimental results demonstrate the effectiveness of our method to improving QoS for end users. Zhengxiong Hou, Yubing Liu, Hong Shen 0001, Jianhua Gu |
ISPA | 5 |
| 2024 | Optimizing job scheduling by using broad learning to predict execution times on HPC clusters
Zhengxiong Hou, Hong Shen 0001, Qiying Feng, Zhiqi Lv, Xingshe Zhou 0001, Jianhua Gu |
CCF Trans. High Perform. Comput. | 7 |
| 2024 | Enhancing campus OS community engagement through the miniOS pilot class: A nine-year journey
Jianhua Gu, Mingxuan Liu 0007, Tianhai Zhao |
Future Gener. Comput. Syst. | 1 |
| 2022 | Prediction of job characteristics for intelligent resource allocation in HPC systems: a survey and future directions
Zhengxiong Hou, Hong Shen 0001, Xingshe Zhou 0001, Jianhua Gu, Yunlan Wang, Tianhai Zhao |
Frontiers Comput. Sci. | 4 |
| 2021 | Traffic Congestion Prediction: A Spatial-Temporal Context Embedding and Metric Learning ApproachabstractIn urban informatics, traffic congestion prediction is of great importance for travel route planning and traffic management, and has received extensive attention from academia and industry. However, most previous works fail to implement a citywide traffic congestion prediction on fine-grained road segment, and without comprehensively considering strong spatial-temporal correlations. To overcome these concerns, in this paper, we propose a spatial-temporal context embedding and metric learning approach (STE-ML) to predict the traffic congestion level. In particular, our STE-ML consists of a traffic spatial-temporal context embedding component, and a metric learning component. From local and global perspectives, the context embedding component can simultaneously integrate local spatial-temporal correlation features and global traffic statistics information, and compress into an unified and abstract embedding representation. Meanwhile, metric learning component benefits from learning a more suitable distance function tuned to specific task. The combination of these models together could enhance traffic congestion prediction performance. We conduct extensive experiments on real traffic data set to evaluate the performance of our proposed STE-ML approach, and make comparison with other existing techniques. The experimental results demonstrate that the proposed STE-ML outperforms the existing methods. Hongsheng Hao, Liang Wang 0017, Zenggang Xia, Zhiwen Yu 0001, Jianhua Gu, Ning Fu |
ICPADS | 5 |
| 2019 | Machine Learning Based Performance Analysis and Prediction of Jobs on a HPC ClusterabstractThere are a lot of middle-class or small-class high-performance computing clusters at universities and research institutes, etc. Large volumes of job logs have been accumulated after many years of operation. In this paper, on the basis of accumulated job logs on a high-performance computing cluster, we examine and analyze the job logs. Then, we study machine learning based performance analysis and prediction methods for parallel jobs. Various machine learning methods such as multivariate linear fitting, artificial neural network are used to build performance prediction models. We compare the errors of each model, and select the optimal prediction model for different users. The experimental results show that we can obtain reasonable prediction accuracy using the selected machine learning algorithms. Zhengxiong Hou, Shuxin Zhao, Yunlan Wang, Jianhua Gu, Xingshe Zhou 0001 |
PDCAT | 5 |
| 2019 | Evolving an optimal kernel extreme learning machine by using an enhanced grey wolf optimization strategy
Jianhua Gu, Jie Luo 0002, Qian Zhang 0049, Huiling Chen 0001, Zhifang Pan, Chengye Li |
Expert Syst. Appl. | 2 |
| 2019 | An integration approach of hybrid databases based on SQL in cloud computing environmentabstractSummary As the applications with big data in cloud computing environment grow, many existing systems expect to expand their service to support the dramatic increase of data, and modern software development for services computing and cloud computing software systems is no longer based on a single database but on existing multidatabases and this convergence needs new software architecture design. This paper proposes an integration approach to support hybrid database architecture, including MySQL, MongoDB, and Redis, to make it possible of allowing users to query data simultaneously from both relational SQL systems and NoSQL systems in a single SQL query. Two mechanisms are provided for constructing Redis's indexes and semantic transforming between SQL and MongoDB API to add the SQL feature for these NoSQL databases. With the proposed approach, hybrid database systems can be performed in a flexible manner, ie, access can be either relational database or NoSQL, depending on the size of data. The approach can effectively reduce development complexity and improve development efficiency of the software systems with multidatabases. This is the result of further research on the related topic, which fills the gap ignored by relevant scholars in this field to make a little contribution to the further development of NoSQL technology. Jianhua Gu |
Softw. Pract. Exp. | 2 |
| 2015 | Optimizing the Overheads for Uncoordinated Proactive Checkpointing
Jianhua Gu |
ICA3PP (4) | 2 |
| 2015 | Optimizing the fault-tolerance overheads of HPC systems using prediction and multiple proactive actions
Jianhua Gu, Yunlan Wang, Tianhai Zhao |
J. Supercomput. | 2 |
| 2014 | DDSF: A Data Deduplication System Framework for Cloud Environments
Jianhua Gu |
CLOSER | 1 |
| 2013 | Research on Optimum Checkpoint Interval for Hybrid Fault Tolerance
Jianhua Gu, Yunlan Wang, Tianhai Zhao |
APPT | 2 |
| 2012 | A Hybrid Heuristic-Genetic Algorithm for Task Scheduling in Heterogeneous Multi-core System
Jianhua Gu, Yunlan Wang, Tianhai Zhao |
ICA3PP (1) | 2 |
| 2012 | mHLogGP: A Parallel Computation Model for CPU/GPU Heterogeneous Computing Cluster
Gangfeng Liu, Yunlan Wang, Tianhai Zhao, Jianhua Gu |
NPC | 4 |
| 2010 | ASAAS: Application Software as a Service for High Performance Cloud ComputingabstractCurrently, SAAS (Software as a Service) solutions are usually provided for business, such as salesforce.com. Few work focus on the application software for high performance scientific computing. However, in the high performance cloud computing environment, traditional application software is not intrinsically service oriented. And the limitation of traditional software licenses is a bottleneck problem for large scale of dynamic users. To enable on-demand services for applications, we propose a solution: Application Software as a Service (ASAAS). It provides a web services portal, an on-demand software license service for the users. Application software is wrapped as web services on the basis of underlying computational resources. With a pay-for-use mode, there is no limitation for the licenses any more. The instant service rate, average job response time, and cost are analyzed for an evaluation. A case of implementation and the evaluation show that ASAAS can bring a much better effect than traditional mechanism. Zhengxiong Hou, Xingshe Zhou 0001, Jianhua Gu, Yunlan Wang, Tianhai Zhao |
HPCC | 3 |
| 2009 | A Runtime Reputation Based Grid Resource Selection Algorithm on the Open Science GridabstractThe scheduling and execution for grid application is an important problem in the grid environment. To get the high reliability and efficiency, we propose a runtime reputation based grid resource selection algorithm. According to the accumulated raw score, the runtime reputation degree for a grid resource is quantified as an evaluating score in the runtime of an application. Instead of being dependent on the historical experiences, it is dynamically adaptive to the runtime availability, load, and performance of the grid resources. The execution framework on the grid is based on Globus Toolkit and Swift system. In a real production grid, Open Science Grid (OSG), a typical grid application with large scale independent jobs was experimented, which was based on BLAST application. The experimental results for the performance of different policies are presented, with a benchmarking workload size of 10,000 jobs. The runtime reputation and behavior statistics for the grid resources are also presented. Zhengxiong Hou, Xingshe Zhou 0001, Michael Wilde, Jianhua Gu, Mihael Hategan |
ICPADS | 4 |
| 2006 | TV Program Recommendation for Multiple Viewers Based on user Profile Merging
Zhiwen Yu 0001, Xingshe Zhou 0001, Yanbin Hao, Jianhua Gu |
User Model. User Adapt. Interact. | 4 |
| 2005 | A flexible hybrid communication model based messaging middlewareabstractAs network-centric computing becomes more pervasive and applications become more distributed, the demand for flexible and efficient message delivery is increasing. Messaging middleware is a key technique of message delivery between distributed computing applications. The main communication models of current messaging middleware are publish/subscribe and point-to-point. The two communication models have their own characteristics. However, neither of them can achieve efficiency and flexibility simultaneously. In this paper, we proposed and implemented a hybrid communication model based messaging middleware (HCM3), which takes full advantages of two models. Performance result shows that HCM3 can deliver message in heterogeneous environment with the features of high efficiency, flexibility and reliability. Huifang Pan, Xingshe Zhou 0001, Zhiyi Yang, Jianhua Gu |
ISADS | 4 |