VLDB 2026 Research / reviewers in the wild / expert
Hao Fu 0021
dblp:64/3069-21
· DBLP profile ↗
14ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-4349-8748ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MixLoRA: An Efficient Multi-Tenant Framework for Concurrently Serving Diverse LoRA Models in Large Language ModelsabstractThe rapid advancement of large language models (LLMs) has driven the widespread adoption of Low-Rank Adaptation (LoRA) for efficient fine-tuning. However, existing multi-LoRA inference systems, such as Punica, face significant challenges in GPU utilization when handling concurrent requests with diverse ranks, leading to suboptimal throughput and increased latency. To address these limitations, we propose MixLoRA, a novel CUDA stream-based framework that enhances GPU efficiency by enabling parallel execution of mixed-rank LoRA requests. MixLoRA introduces a multi-worker scheduling architecture that selects the optimal number of workers for concurrent execution based on the model size, inference request characteristics, and device information. Each worker manages requests of a fixed rank, eliminating the need for rank-aligned batching and maximizing resource utilization through independent CUDA streams. Experimental results demonstrate that MixLoRA achieves up to 1.2× higher throughput compared to state-of-the-art solutions, offering a scalable and efficient approach for multi-tenant LLM inference services. Ronghuai Chen, Ce Yu, Hao Fu 0021, Xiaoteng Hu, Bin Yang 0043 |
ICPP | 3 |
| 2024 | Accuracy-Efficiency Optimization for Multi-Stage Small Object Detection in Surveillance Video with Collaborative Frame SamplingabstractIn video analytics, accuracy and efficiency are two important metrics and there tend to be a tradeoff between each other. In this paper, we consider accuracy-efficiency optimization for small object detection in surveillance video, which is important and has been widely used in many scenarios such as license plate detection in the traffic domain. Given that small objects tend to being attached to big objects, multi-stage object detection is supposed to be an effective approach to achieve high accuracy for small objects by detecting big objects first and then small objects within the ROIs (Region of interests) of big objects. However, existing studies considered the accuracy-efficiency optimization for small object detection only within the single-stage scenario by changing the frame resolution or sampling rate configuration of video data, which are not suitable for multi-stage detection given that its accuracy-efficiency result is determined by the results of all stages jointly. In this paper, we propose an Adaptive and Collaborative frame Sampling approach named ACS for accuracy-efficiency optimization in the multi-stage small object detection. To improve the efficiency significantly while guaranteeing a given accuracy threshold, ACS dynamically adjusts the sampling rates of all stages collaboratively and periodically using the Karush-Kuhn-Tucker (KKT) condition based on the Lagrangian multiplier method. Additionally, we introduce a tuning knob to allow users to flexibly balance accuracy and efficiency, while ensuring a given accuracy threshold λ. Extensive experiments demonstrate the effectiveness of our approach in improving detection efficiency while guaranteeing diverse accuracy requirements. Chunhong Du, Shanjiang Tang, Song Meng, Jiekai Gou, Ce Yu, Yusen Li, Hao Fu 0021 |
CLUSTER | 7 |
| 2024 | A large-scale heterogeneous computing framework for non-uniform sampling two-dimensional convolution applications
Ce Yu, Jian Xiao 0001, Hao Fu 0021, Bo Kang |
CCF Trans. High Perform. Comput. | 5 |
| 2024 | Fairness-Efficiency Scheduling for Pay-as-You-Go Shared Caching Systems With Long-Term Fairness GuaranteesabstractPay-as-you-go caching systems are now widely used as storage services in cloud computing. However, users’ data caching requirements not only change over time, but also they are affected by workload characteristics, making it difficult to always ensure high efficient use of cache resources. Cache resource sharing is an effective way to improve the efficiency of cache usage. To incentivize users to share caches, it is essential to ensure long-term fairness among multiple users. However, traditional resource allocation strategies canonly guarantee memory less fairness among users,which is not applicable to long-term cache sharing systems. In this paper, we propose a fair allocation policy named FairCache for Pay-as-you-go cache resources. First, FairCache can satisfy four desirable properties of resource allocation: sharing incentive, pay-as-you-usefairness, strategy proofness, and pare to efficiency. Second, FairCache is an efficiency-fairness resource allocation policy based on the efficiency knob θ. The strategy keeps sensitive to the constantly changing cache demands of multiple users within the system by adjusting the efficiency knob θ, thus ensuring long-term multi-user fairness while maximizing the efficiency of cache usage. In addition, FairCache also has an anti-cheating mechanism to avoid possible free-rider problems when multiple users cache access files. Finally, this paper implements the FairCache policy in Alluxio. The experimental results show that FairCache is a lightweight scheduler and that it can maximize the efficiency usage of cache resources while ensuring the long-term for multiple users in the pay-as-you-go Cache systems fairness. Shanjiang Tang, Zhongyu Zhou, Jiekai Gou, Ce Yu, Yusen Li, Hao Fu 0021, Chao Sun 0008, Jian Xiao 0001 |
IEEE Trans. Serv. Comput. | 6 |
| 2022 | EasyNUSC: An Efficient Heterogeneous Computing Framework for Non-uniform Sampling Two-Dimensional Convolution Applications
Ce Yu, Jian Xiao 0001, Hao Fu 0021, Shanjiang Tang, Bo Kang |
ICA3PP | 5 |
| 2022 | Long-Term Fairness Scheduler for Pay-as-You-Use Cache Sharing Systems
Zhongyu Zhou, Shanjiang Tang, Hao Fu 0021, Wanqing Chang, Ce Yu, Chao Sun 0008, Yusen Li, Jian Xiao 0001 |
ICA3PP | 3 |
| 2022 | A method for efficient radio astronomical data gridding on multi-core vector processor
Ce Yu, Jian Xiao 0001, Shanjiang Tang, Hao Fu 0021, Bo Kang, Chenzhou Cui |
Parallel Comput. | 6 |
| 2021 | DVQShare: An Analytics System for DNN-based Video QueriesabstractApplying deep neural networks (DNNs) to video analytics tasks has drawn attention from both academic and industry communities. However, due to the high computational complexity of DNN models and the explosion of video data, it is challenging to process massive concurrent video queries efficiently and effectively. In this paper, we propose a video analytics system named DVQShare to process DNN-based video queries in a batch mode. The key idea is sharing, including time sharing, spatial sharing, and logical sharing. In principle, sharing across queries can help us reduce the overall amount of frames to be analyzed, which can help us improve the overall performance and reduce the monetary cost. Two modules are designed to process video queries by exploiting the above three sharing opportunities. First, an analysis module is integrated to guide the generation of query processing plans. Within this module, temporal sharing is considered to reuse historical results produced by other queries to remove pending frames that have been analyzed, and spatial sharing is adopted to avoid redundant processing over overlapping video clips. Additionally, we utilize logical sharing to further improve system's overall performance by considering the logical relationship between queries. Second, a query processing engine is devised to execute the query pipeline generated by the analysis module and return the final results. In experiments, we implement a prototype of the DVQShare system based on MXNet, and results show that it can achieve up to 2X performance speedup. Hao Fu 0021, Shanjiang Tang, Ce Yu, Yusen Li, Yanjie Liu |
CCGRID | 1 |
| 2021 | HGP4CNN: an efficient parallelization framework for training convolutional neural networks on modern GPUs
Hao Fu 0021, Shanjiang Tang, Bingsheng He, Ce Yu |
J. Supercomput. | 1 |
| 2020 | Accelerating Exact Constrained Shortest Paths on GPUsabstractThe recently emerging applications such as software-defined networks and autonomous vehicles require efficient and exact solutions for constrained shortest paths (CSP), which finds the shortest path in a graph while satisfying some user-defined constraints. Compared with the common shortest path problems without constraints, CSP queries have a significantly larger number of subproblems. The most widely used labeling algorithm becomes prohibitively slow and impractical. Other existing approaches tend to find approximate solutions and build costly indices on graphs for fast query processing, which are not suitable for emerging applications with the requirement of exact solutions. A natural question is whether and how we can efficiently find the exact solution for CSP. In this paper, we propose Vine , a framework that parallelizes the labeling algorithm to efficiently find the exact CSP solution using GPUs. The major challenge addressed in Vine is how to deal with a large number of subproblems that are mostly unpromising but require a significant amount of memory and computational resources. Our solution is twofold. First, we develop a two-level pruning approach to eliminate the subproblems by making good use of the GPU's hierarchical memory. Second, we propose an adaptive parallelism control model based on the observations that the degree of parallelism (DOP) is the key to performance optimization with the given amount of computational resources. Extensive experiments show that Vine achieves 18× speedup on average over the widely adopted CPU-based solution running on 40 CPU threads. Vine also has over 5× speedup compared with a GPU approach that statically controls the DOP. Compared to the state-of-the-art approximate solution with preprocessed indices, Vine provides exact results with competitive or even better performance. Shengliang Lu, Bingsheng He, Yuchen Li 0001, Hao Fu 0021 |
Proc. VLDB Endow. | 4 |
| 2018 | GLP4NN: A Convergence-invariant and Network-agnostic Light-weight Parallelization Framework for Deep Neural Networks on Modern GPUsabstractIn this paper, we propose a network-agnostic and convergence-invariant light-weight parallelization framework, namely GLP4NN, to accelerate the training of Deep Neural Networks (DNNs) by taking advantage of emerging GPU features, especially concurrent kernel execution. To determine the number of concurrent kernels on the fly, we design an analytical model in the kernel analyzer module and integrate a compact asynchronous resource tracker in the resource tracker module for collecting runtime configurations of kernels with low memory and time overheads. We further develop a runtime scheduler module and a pool-based stream manager for handling GPU work queues in GLP4NN to avoid consuming too many CPU threads or processes while dispatching workloads to GPU devices. In our experiments, we integrate GLP4NN into Caffe to accelerate the batch-based training of four well-known networks on NVIDIA GPUs. Experimental results show GLP4NN is able to achieve a speedup of up to 4X over the original implementation as well as keep the convergence property of networks. Hao Fu 0021, Shanjiang Tang, Bingsheng He, Ce Yu |
ICPP | 1 |
| 2017 | KD-Tree and HEALPix-Based Distributed Cone Search Indexing System for Multi-Band Astronomical Catalogs
Ce Yu, Jian Xiao 0001, Xiaoteng Hu, Hao Fu 0021, Kun Li 0001, Yanyan Huang |
ICA3PP | 5 |
| 2015 | A Multilevel Fault-Tolerance Technique for the DAG Data Driven ModelabstractFault tolerance of hardware failure is a challenging work for parallel programming in massively parallel processing environment. However, traditional rollback-recovery techniques, which an be classified into checkpoint-based and log-based, would introduce extra overhead for recording an overall snapshot of an application. For a specialized programming model, a private recovery technique is valuable and can achieve a better performance.In this paper, a multilevel fault-tolerance technique designed for the DAG data driven model is proposed. It utilized the checkpoint-based fault tolerance technique for system recovery, and timeout to detect and revoery from performance faults. It consists of two kinds of checkpoints: the DAG pattern checkpoint and the intermediate result checkpoint. The DAG pattern checkpoint is designed for tracing the current processing progress of the DAG model, while the intermediate results checkpoint is used to record outputs of compute nodes. Moreover, we also implement this technique in the EasyHPS runtime system. Experimental results show that the check pointing overhead is as low as 2.6%. Hao Fu 0021, Ce Yu |
CCGRID | 1 |
| 2015 | A List Scheduling Algorithm for DAG-Based Parallel Computing Models
Hao Fu 0021, Ce Yu |
ICA3PP (2) | 1 |