VLDB 2026 Research / reviewers in the wild / expert
Kan Hu
dblp:92/2627
· DBLP profile ↗
20ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-4775-7273ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 2Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 65% GPUs and heterogeneous computing · 33% Performance modeling and evaluation · 2% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU graph processing |
0.4 | 1 | 2020 | AsynGraph: Maximizing Data Parallelism for Efficient Iterative Graph Processing on GPUs · ACM Trans. Archit. Code Optim. 2020 |
Parallel and multicore computing › graph processing
iterative graph processing |
0.4 | 1 | 2020 | AsynGraph: Maximizing Data Parallelism for Efficient Iterative Graph Processing on GPUs · ACM Trans. Archit. Code Optim. 2020 |
Data mining › pattern mining
frequent pattern mining |
0.0 | 1 | 2000 | Towards Data Mining Benchmarking: A Testbed for Performance Study of Frequent Pattern Mining · SIGMOD Conference 2000 |
Data mining
pattern mining |
0.0 | 1 | 2000 | Towards Data Mining Benchmarking: A Testbed for Performance Study of Frequent Pattern Mining · SIGMOD Conference 2000 |
Performance modeling and evaluation
benchmarking |
0.0 | 1 | 2000 | Towards Data Mining Benchmarking: A Testbed for Performance Study of Frequent Pattern Mining · SIGMOD Conference 2000 |
Methods — techniques the papers use, named apart from their topics
structure-aware asynchronous processing · 0.9forward-backward intra-path processing · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TD3-Sched: Learning to Orchestrate Container-Based Cloud-Edge Resources via Distributed Reinforcement Learning
Shengye Song, Minxian Xu, Kan Hu, Wenxia Guo, Kejiang Ye |
PDCAT | 3 |
| 2025 | LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient DescentabstractABSTRACT Objective The microservices architecture has become a dominant paradigm in cloud computing due to its advantages in development, deployment, modularity, and scalability. Ensuring Quality of Service (QoS) through efficient Service Level Objective (SLO) resource allocation is a critical challenge. Current frameworks for microservice autoscaling based on SLOs often rely on heavy and complex models that are time‐consuming and resource‐intensive, making them unsuitable for rapidly changing environments and highly dynamic workloads. Methods This study proposes LSRAM (Lightweight SLO Resource Allocation Management), a novel framework designed to overcome the limitations of existing SLO‐based autoscaling methods. LSRAM operates in two stages: 1). Lightweight SLO Resource Allocation Model: Computes optimal SLO resource allocation for each microservice using a gradient descent method, ensuring rapid computation and minimal computational overhead. 2). SLO Resource Update Model: Adapts resource allocation dynamically in response to changes in the cluster environment, such as varying loads and application types, without requiring extensive retraining. Results LSRAM effectively addresses scenarios involving bursty traffic and fluctuating workloads. Compared to state‐of‐the‐art SLO allocation frameworks, LSRAM achieves the following: 1). Reduces resource usage by 17%. 2). Maintains QoS guarantees for users, even under dynamic conditions. 3). Demonstrates faster adaptability to changes in the system environment due to its lightweight design. Conclusion LSRAM offers a scalable, efficient, and adaptive solution for SLO‐based resource allocation in microservices architectures. By reducing resource usage while maintaining QoS, it provides a robust framework for managing dynamic and unpredictable workloads in cloud environments. Its lightweight design ensures practical applicability and superior performance compared to traditional, resource‐intensive methods. Kan Hu, Minxian Xu, Kejiang Ye, Cheng-Zhong Xu 0001 |
Softw. Pract. Exp. | 1 |
| 2024 | MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for MicroservicesabstractService Level Objectives (SLOs) aim to set threshold for service time in cloud services to ensure acceptable quality of service (QoS) and user satisfaction. Currently, many studies consider SLOs as a system resource to be allocated, ensuring QoS that to meet the SLOs. Existing microservice auto-scaling frameworks that rely on SLO resources often utilize complex and computationally intensive models, requiring significant time and resources to determine appropriate resource allocation. This paper aims to rapidly allocate SLO resources and minimize resource costs while ensuring application QoS meets the SLO requirements in a dynamically changing microservice environment. We propose MSARS, a framework that leverages meta-learning to quickly derive SLO resource allocation strategies and employs reinforcement learning for adaptive scaling of microservice resources. It features three innovative components: first, MSARS uses graph convolutional networks to predict the most suitable SLO resource allocation scheme for the current environment. Second, MSARS utilizes meta-learning to enable the graph neural network to quickly adapt to environmental changes ensuring adaptability in highly dynamic microservice environments. Third, MSARS generates auto-scaling policies for each microservice based on an improved Twin Delayed Deep Deterministic Policy Gradient (TD3) model. The adaptive auto-scaling policy integrates the SLO resource allocation strategy into the scheduling algorithm to satisfy SLOs. Finally, we compare MSARS with state-of-the-art resource auto-scaling algorithms that utilize neural networks and reinforcement learning, MSARS takes 40% less time to adapt to new environments, 38% reduction of SLO violations, and 8% less resources cost. Kan Hu, Linfeng Wen 0001, Minxian Xu, Kejiang Ye |
ISPA | 1 |
| 2021 | A Survey of Non-Volatile Main Memory Technologies: State-of-the-Arts, Practices, and Future Directions
Haikun Liu, Hai Jin 0001, Xiaofei Liao, Binsheng He, Kan Hu, Yu Zhang 0027 |
J. Comput. Sci. Technol. | 6 |
| 2021 | FDGLib: A Communication Library for Efficient Large-Scale Graph Processing in FPGA-Accelerated Data Centers
Qinggang Wang, Long Zheng 0003, Xiaofei Liao, Hai Jin 0001, Wenbin Jiang 0001, Kan Hu |
J. Comput. Sci. Technol. | 8 |
| 2020 | AsynGraph: Maximizing Data Parallelism for Efficient Iterative Graph Processing on GPUsabstractRecently, iterative graph algorithms are proposed to be handled by GPU-accelerated systems. However, in iterative graph processing, the parallelism of GPU is still underutilized by existing GPU-based solutions. In fact, because of the power-law property of the natural graphs, the paths between a small set of important vertices (e.g., high-degree vertices) play a more important role in iterative graph processing’s convergence speed. Based on this fact, for faster iterative graph processing on GPUs, this article develops a novel system, called AsynGraph , to maximize its data parallelism. It first proposes an efficient structure-aware asynchronous processing way . It enables the state propagations of most vertices to be effectively conducted on the GPUs in a concurrent way to get a higher GPU utilization ratio through efficiently handling the paths between the important vertices. Specifically, a graph sketch (consisting of the paths between the important vertices) is extracted from the original graph to serve as a fast bridge for most state propagations. Through efficiently processing this sketch more times within each round of graph processing, higher parallelism of GPU can be utilized to accelerate most state propagations. In addition, a forward-backward intra-path processing way is also adopted to asynchronously handle the vertices on each path, aiming to further boost propagations along paths and also ensure smaller data access cost. In comparison with existing GPU-based systems, i.e., Gunrock, Groute, Tigr, and DiGraph, AsynGraph can speed up iterative graph processing by 3.06–11.52, 2.47–5.40, 2.23–9.65, and 1.41–4.05 times, respectively. Yu Zhang 0027, Xiaofei Liao, Lin Gu 0002, Hai Jin 0001, Kan Hu, Haikun Liu, Bingsheng He |
ACM Trans. Archit. Code Optim. | 5 |
| 2018 | Disk Failure Prediction in Data Centers via Online LearningabstractDisk failure has become a major concern with the rapid expansion of storage systems in data centers. Based on SMART (Self-Monitoring, Analysis and Reporting Technology) attributes, many researchers derive disk failure prediction models using machine learning techniques. Despite the significant developments, the majority of works rely on offline training and thereby hinder their adaption to the continuous update of forthcoming data, suffering from the 'model aging' problem. We are therefore motivated to uncover the root cause -- the dynamic SMART distribution for 'model aging', aiming to resolve the performance degradation as to pave a comprehensive study in practice. Jiang Xiao 0001, Song Wu 0001, Yusheng Yi, Hai Jin 0001, Kan Hu |
ICPP | 6 |
| 2017 | Fairness-aware dynamic rate control and flow scheduling for network function virtualizationabstractBy softwarizing traditional dedicated hardware based functions to virtualized network functions (VNFs) that can run on standard commodity servers, network function virtualization (NFV) technology promises high efficiency, flexibility and scalability. To NFV service providers, one primary concern is to maximize network throughput and reduce service time. To reach this goal, two main challenges should be tackled: 1) how to schedule the unpredictable and burst network flows; 2) how to fairly allocate resources between various flows with different resource requirements. In this paper, we are motivated to investigate a throughput maximization problem with joint consideration of fairness between multiple flows using a discrete time queuing model. By taking advantages of Lyapunov optimization techniques, we propose a low-complexity online distributed algorithm that can achieve arbitrary optimal utility with different fairness levels by tuning the fairness bias. The high efficiency of our proposal is validated by both theoretical analysis and extensive simulation studies. Sheng Tao, Lin Gu 0002, Deze Zeng, Hai Jin 0001, Kan Hu |
IWQoS | 5 |
| 2013 | Scheduling overcommitted VM: Behavior monitoring and dynamic switching-frequency scaling
Huacai Chen, Hai Jin 0001, Kan Hu |
Future Gener. Comput. Syst. | 3 |
| 2013 | Design and implementation of a trusted monitoring framework for cloud platforms
Deqing Zou, Wenrong Zhang, Weizhong Qiang, Guofu Xiang, Laurence T. Yang, Hai Jin 0001, Kan Hu |
Future Gener. Comput. Syst. | 7 |
| 2011 | Optimization of Sparse Matrix-Vector Multiplication with Variant CSR on GPUsabstractSparse Matrix-Vector multiplication (SpMV) is one of the most significant yet challenging issues in computational science area. It is a memory-bound application whose performance mostly depends on the input matrix and the underlying architecture. Many researchers have paid more attentions on exploring a variety of optimization techniques to SpMV. One of the most promising respects is how to adapt the storage format to satisfy the underlying architecture. Alterative storage formats can largely lessen memory pressure, however, the computational resources are often underutilized. Therefore, a new storage format, which is called Compressed Sparse Row with Segmented Interleave Combination (SIC), is proposed. Stemming from Compressed Sparse Row format (CSR), SIC format employs an interleave combination pattern that combines certain amount of CSR rows to form a new SIC row. In order to further improve performance, segmented processing is also brought in. According to the empirical data, we also develop an automatic SIC-based SpMV suitable for all the matrices. Experimental results show that our approach outperforms the NVIDIA CSR vector kernel, achieving up to 12.6 × speedup. It also demonstrates a comparable performance with the Hybrid format, even with the highest 2.89 × speedup. Xiaowen Feng, Hai Jin 0001, Kan Hu, Jingxiang Zeng, Zhiyuan Shao |
ICPADS | 4 |
| 2010 | Adaptive audio-aware scheduling in Xen virtual environmentabstractWith the development of client virtualization technology, it has become an important tendency to apply soft real-time applications in virtual environment. Currently, most schedulers in VMM (i.e., virtual machine monitor) take the fairly sharing of processor resources and load balancing as a main concern, while show less regard to application diversity and I/O responsiveness. This would be unable to meet the requirement for latency-sensitive tasks, such as audio application. Audio stream may suffer from severe input buffer overrun or output buffer underrun, especially in the case that there is no real-time guarantee on virtualized clients under heavy load. In this paper, we introduce experiments to illustrate that current scheduler in Xen does a poor job in guaranteeing fluent audio playing, and then formulate the fluent playing conditions a scheduler should satisfy. A scheduling strategy with soft real-time support is proposed to improve the responsiveness of latency-sensitive guests. To implement our proposition, we extend the Credit scheduler by using flexible time slice and real-time priority. Our solution is audio-aware and capable of adjusting the real-time priority of guest domains adaptively, achieving a better experience for end-users. The experimental results show that audio glitches can be completely eliminated via our extended scheduler even when the system load is very high. Huacai Chen, Hai Jin 0001, Kan Hu, Minhao Yuan |
AICCSA | 3 |
| 2010 | Dynamic Switching-Frequency Scaling: Scheduling Overcommitted Domains in Xen VMMabstractVirtualization enables multiple guest operating systems run on a single physical platform. These virtual machines may host any types of application, including concurrent HPC programs. Traditionally, VMM schedulers have focused on fairly sharing the processor resources among domains, rarely consider VCPUs’ behaviors. However, this can result in poor application performance to overcommitted domains if there are concurrent programs hosted in them. In this paper we review the properties of both Xen’s Credit and SEDF schedulers, and show how these schedulers may seriously impact the performance of the communication-intensive and I/O-intensive concurrent applications in overcommitted domains. We discuss the origination of the problem theoretically, and confirm the derived conclusion on benchmarks. A novel approach, that dynamically scales the context switching-frequency by selecting variable time slices according to VCPUs` behaviors, is then proposed to improve the Credit scheduler more adaptive for concurrent applications. The experimental results show that this extended Credit scheduler can improve the performance of communication-intensive and I/O-intensive concurrent applications in overcommitted domains to the same magnitude as in undercommitted domains. Huacai Chen, Hai Jin 0001, Kan Hu |
ICPP | 3 |
| 2010 | XenHVMAcct: Accurate CPU Time Accounting for Hardware-Assisted Virtual MachineabstractCPU time accounting is a basis of performance measurement and process scheduling in operating system. Accounting operations are traditionally completed in timer interrupt handler since timer interrupt is periodically delivered to OS. However, when virtualization introduced, the CPU time is shared by multiple virtual CPUs (i.e., VCPU for short) and the virtual timer interrupt is paused for those ones be scheduled out. This makes the time accounting be inaccurate, and we should consider new method for VM to provide a stable and reliable data source, especially for the hardware-assisted virtual machines (i.e., HVM for short) which are not aware of VMM. The key point of accurate CPU time accounting is to distinguish the time allocated to “this VCPU” and “other VCPUs”. Para-virtualization (i.e., PV for short) achieves this goal by modifying the timer handling routines. For HVM, we propose an accurate accounting method (named XenHVMAcct) within Xen virtual platform. XenHVMAcct is designed by using the mechanisms of virtual interrupt and loadable kernel module, without direct modifications to guest OS. Experimental results show that our accounting method is as accurate as the PV solution. Huacai Chen, Hai Jin 0001, Kan Hu |
PDCAT | 3 |
| 2010 | FTDS: Adjusting Virtual Computing Resources in Threshing CasesabstractIn a virtual execution environment, dynamic computing resource adjustment technique, configuring the computing resource of virtual machines automatically according to the actual loads generated by applications, is often adopted in virtual machine monitor to improve the resource utilization rate. Traditionally, the simple Additive Increase Subtractive Decrease (AISD) scheme is used as an adjusting rule. However, in some special situations, for example, compiling kernel in a virtual machine, the configuration of virtual machines may change abruptly because of the violent vibration of workload during a short interval, and the threshing can inevitably result in additional overhead under AISD rule. In this paper, we extend the Proportional-Integral-Derivative (PID) algorithm and present a feedback control model for configuring virtual computing resources, and propose an innovative adjusting scheme called Forecasting and Time Delayed Subtraction (FTDS) to reduce the overhead caused by threshing. The FTDS uses both statistic history of resource requests and current utilization of computing resource to predict whether threshing happens and to determine how many and when to adjust the amount of virtual CPUs. Experimental results show that FTDS can effectively reduce the jitter occurred in adjusting and make the performance penalty for threshing decreased from 9% to 0.3% compared with AISD, while maintaining that in non-threshing cases the same as AISD. Hai Jin 0001, Kan Hu, Zhiyuan Shao |
PDP | 3 |
| 2001 | Towards the building of a dense-region-based OLAP system
David Wai-Lok Cheung, Ben Kao, Kan Hu, Sau Dan Lee |
Data Knowl. Eng. | 4 |
| 2001 | An Adaptive Algorithm for Mining Association Rules on Shared-Memory Parallel Machines
David Wai-Lok Cheung, Kan Hu, Shaowei Xia |
Distributed Parallel Databases | 2 |
| 2000 | Towards Data Mining Benchmarking: A Testbed for Performance Study of Frequent Pattern MiningabstractPerformance benchmarking has played an important role in the research and development in relational DBMS, object-relational DBMS, data warehouse systems, etc. We believe that benchmarking data mining algorithms is a long overdue task, and it will play an important role in the research and development of data mining systems as well. Jian Pei 0001, Runying Mao, Kan Hu |
SIGMOD Conference | 3 |
| 1999 | DROLAP - A Dense-Region Based Approach to On-Line Analytical Processing
David Wai-Lok Cheung, Ben Kao, Kan Hu, Sau Dan Lee |
DEXA | 4 |
| 1998 | Asynchronous Parallel Algorithm for Mining Association Rules on a Shared-Memory MultiPprocessorsabstractMining association rules from large databases is an important problem in data mining.There is a need to develop parallel algorithm for this problem because it is a very costly computation process.However, all proposed parallel algorithms for mining association rules follow the conventional level-wise approach.On a sharedmemory multi-processors, they will impose a synchronization in every iteration which degrades greatly their performance.The deficiency comes from the contention on the shared I/O channel when all processors are accessing their database partitions in the shared storage synchronously.An asynchronous algorithm APM has been proposed for mining association rules on shared-memory multiprocessors.All participating processors in APM generate candidates and count their supports independently without synchronization.Furthermore, it can finish the computation with less I/O than required in the level-wise approach.The algorithm has been implemented on a Sun Enterprise 4000 multi-processors with 12 nodes.The experiments show that APM has super performance than other proposed synchronous algorithms. David Wai-Lok Cheung, Kan Hu, Shaowei Xia |
SPAA | 2 |