Bao Li 0002

dblp:51/3716-2 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Similarity-Aware Function Pre-Loading for Serverless Inference
abstract
The ubiquity of cold starts in serverless architectures poses a critical barrier to low-latency inference. While existing prewarming methods leverage idle container memory to pre-load functions, they often neglect resource contention among functions within the shared container, frequently resulting in severe request blocking. To address these challenges, this paper proposes SFP, a Similarity-Aware Function Pre-Loading strategy which optimizes function distribution by selecting target containers for pre-loading functions. SFP deploys functions with low invocation similarity within the same container while distributing those with high invocation similarity across distinct containers. Function invocation similarity, a metric proposed in this study, is derived from the Jaccard similarity coefficient and quantifies the temporal overlap between disparate function invocations. Experimental results based on real-world workloads demonstrate that, compared to state-of-the-art methods, the proposed strategy improves inference request throughput by up to 230% and achieves memory savings ranging from 6.7% to 34.8%.
Jichang Dong, Bao Li 0002, Yusong Tan
CF2
2025 Heterogeneity-Aware Two-Tier GPU Resource Scheduling for Machine Learning Tasks
abstract
The widespread use of machine learning (ML) tasks has led to a rapid expansion of GPU clusters. Heterogeneous GPUs exhibit different performance characteristics across ML tasks, which leads to challenges for resource scheduling. As cluster scale expands, computational overhead for resource allocation grows while overall utilization remains low. Existing schedulers also lack adaptability and flexibility to different workloads and users requirements. In this work, we design a two-tier allocation framework, Gsched. We introduce an allocation metric, classify-put, to measure the relationship between tasks and GPUs. Based on classifyput, we propose a grouping mechanism, which breaks down ML tasks and GPU resource into smaller and parallel groups. In each group, we can use different scheduling policies and optimize parameters of ML tasks. Gsched can maintain high resource allocation efficiency while reducing computational overhead. Experimental results have shown that our framework achieves the best performance in most cases, reducing the average task completion time by up to 50% and tail latency by up to 15% compared with other advanced schedulers. The scheduling computational overhead can be reduced by several orders of magnitude.
Xilong Gu, Bao Li 0002, Chunbo Jia, Chenlin Huang
HPCC4
2025 Promoting Resource Utilization in HPC via Scheduling
abstract
ABSTRACT The increasing complexity of supercomputing workloads poses challenges to efficient resource management, especially in balancing computational and I/O demands. Shared burst buffers, as high‐speed intermediate storage, offer a promising avenue to mitigate I/O bottlenecks. However, naive job scheduling strategies often neglect the potential of burst buffers, relying on heuristic methods with limited adaptability. To address this, we propose an innovative burst‐buffer‐aware scheduling framework that integrates burst buffer capacity into job scheduling as a key resource. Through multi‐objective optimization, the framework intelligently balances trade‐offs among job waiting time, slowdown, and completion time, surpassing the rigidity of conventional approaches. Leveraging real‐world workload traces, the framework dynamically adapts to varying windows and job characteristics, combining computation and burst buffer demands to optimize scheduling decisions. Experimental results reveal that the proposed framework enhances scheduling efficiency and system adaptability, establishing a smarter and more effective approach to supercomputing job scheduling. This work underscores the importance of burst‐buffer‐aware strategies in advancing high‐performance computing, offering novel insights into intelligent resource management.
Gang Xian, Yusong Tan, Jie Yu 0006, Wenxiang Yang, Bao Li 0002
Concurr. Comput. Pract. Exp.5
2025 SNCD: A fast and scalable distributed near-miss code clone detector for big code based on partial index
Rulin Xie, Yi Ren 0008, Jianbo Guan, Bao Li 0002, Jun Ma 0015, Yusong Tan
Future Gener. Comput. Syst.6
2024 Dimac: Dynamic Integrity Measurement Architecture for Containers with ARM TrustZone
abstract
How to implement dynamic trusted measurement for multi-tenant container platforms is an important issue in establishing trust in cloud-native services. Existing TPM-based schemes lack support for scalability, dynamism, and isolation in measurements, making integrity measurement implementation based on trusted execution environment(TEE) an excellent solution to this problem. However, challenges such as identifying the objects and timing for dynamic trustworthy measurement, bridging semantic gaps across domains, and providing multi-tenant privacy protection need to be addressed. This paper focuses on the trust invariants of container processes and proposes capturing page faults to perform cross-domain memory direct measurement on these invariants as early as possible, minimizing the potential for TOC-TOU issues. Based on this, we propose a fine-grained remote attestation method for container task execution flow information while considering privacy protection. We have implemented a prototype system, and experimental results show a performance loss of only about 8%.
Liantao Song, Bao Li 0002
ICWS4
2022 SCORE: A Resource-Efficient Microservice Orchestration Model Based on Spectral Clustering in Edge Computing
Yusong Tan, Bao Li 0002
ICSOC4
2022 Evaluation Ranking is More Important for NAS
abstract
Search space, searching method, and candidate evaluation scheme are critical to the success of Neural Architecture Search (NAS), especially the evaluation strategy. An effective and efficient neural architecture performance evaluator could successfully save computing costs and search time while guiding the NAS process to the optimal solution as fast as possible. Most existing NAS algorithms attempt to compute the absolute accuracy of the candidate architecture, which is almost impossible to achieve and meaningless to the final performance improvement. In this paper, we propose ERNAS, a novel neural architecture performance evaluation approach that optimizes the ranking of the candidate architecture performance, rather than the absolute accuracy itself. With the help of ERNAS, many existing NAS methods could achieve better performance without further evaluation. The experimental results demonstrate that ERNAS can be trained effectively enough with extremely limited training data (423 neural architectures randomly sampled form NAS-Bench-101, which is only 0.1% of the entire search space). The accuracy of the neural architecture search result produced by ERNAS is greater than that of the SOTA methods.
Yusen Zhang 0007, Bao Li 0002, Yusong Tan, Songlei Jian
IJCNN2
2022 ProxyDWRR: A Dynamic Load Balancing Approach for Heterogeneous-CPU Kubernetes Clusters
abstract
Edge computing is booming as a promising paradigm to push the service and computation resources from the cloud to the edge of network. As the de-facto standard for container orchestration, Kubernetes is more and more widely used not only in cloud computing but also in edge computing. However, Kubernetes is designed for homogenous cloud data centers, and it does not take into account heterogeneous scenarios, which is ubiquitous is the edge. This will lead to load imbalance among containers with its default rough load balancing mechanism. To deal with this problem, we firstly propose a Dynamically Weighted Random Routing (DWRR) algorithm based on the default random algorithm in Kubernetes. Besides, we design and implement ProxyDWRR, a load balancing plugin for the Kubernetes cluster with heterogeneous CPU. It is fully compatible with the existing load balancing mechanism in Kubernetes. We validated our solution based on a cloud-native microservices application. The experimental results show that ProxyDWRR can effectively balance the load between containers in clusters with heterogeneous CPU. In our experiments, DWRR can improve the CPU utilization of the containers by about 25% and the throughput of the application by about 22.6% compared to the default load balancing algorithms, which enables the cluster to evacuate bursty load more effectively.
Qingkun Wang, Yi Ren 0008, Saqing Yang, Jianbo Guan, Bao Li 0002, Yusong Tan
JCC5
2021 FastDCF: A Partial Index Based Distributed and Scalable Near-Miss Code Clone Detection Approach for Very Large Code Repositories
Yi Ren 0008, Jianbo Guan, Bao Li 0002, Jun Ma 0015, Yusong Tan
PDCAT4
2013 iFlatLFS: Performance optimization for accessing massive small files
abstract
The processing of massive small files is a challenge in the design of distributed file systems. Currently, the combined-block-storage approach is prevalent. However, the approach employs traditional file systems like ExtFS and may cause inefficiency for random access to small files. This paper focuses on optimizing the performance of data servers in accessing massive small files. We present a Flat Lightweight File System (iFlatLFS) to manage small files, which is based on a simple metadata scheme and a flat storage architecture. iFlatLFS aims to substitute the traditional file system on data servers that are mainly used to store small files, and it can greatly simplify the original data access procedure. The new metadata proposed in this paper occupies only a fraction of the original metadata size based on traditional file systems. We have implemented iFlatLFS in CentOS 5.5 and integrated it into an open source Distributed File System (DFS), called Taobao FileSystem (TFS), which is developed by a top B2C service provider, Alibaba, in China and is managing over 28.6 billion small photos. We have conducted extensive experiments to verify the performance of iFlatLFS. The results show that when the file size ranges from 1KB to 64KB, iFlatLFS is faster than Ext4 by 48% and 54% on average for random read and write in the DFS environment, respectively. Moreover, after iFlatLFS is integrated into TFS, iFlatLFS-based TFS is faster than the existing Ext4-based TFS by 45% and 49% on average for random read access and hybrid access (the mix of read and write accesses), respectively.
Songling Fu, Chenlin Huang, Ligang He, Nadeem Chaudhary, Xiangke Liao, Shazhou Yang, Bao Li 0002
HiPC8
2010 Non-rigid Registration in 3D Implicit Vector Space
abstract
We present an implicit approach for pair-wise non-rigid registration of moving and deforming objects. Shapes of interest are implicitly embedded in the 3D implicit vector space. In this implicit embedding space, registration is performed using a global-to-local framework. Firstly, a non-linear optimization functional defined on the vector distance function is used to find the global alignment between shapes. Secondly, an incremental cubic B-spline free form deformation is used to recover the non-rigid transformation parameters. Local non-rigid registration is posed in terms of minimising an energy functional, for which we give a closed-form linear system and solve it using an improved iterative Gauss-Seidel method. Our approach can consistently produce smooth and continuous registration fields, and correctly establish dense one-to-one correspondences. It can naturally deal with both open partial and closed shapes, and imperfect models with gaps and noise, through its use of the implicit vector representation. Experimental results on several datasets demonstrate the robustness of the proposed method.
Zhi-Quan Cheng, Gang Dang, Ralph R. Martin, Jun Li 0042, Honghua Li, Yin Chen 0003, Bao Li 0002, Kai Xu 0004, Shiyao Jin
Shape Modeling International9
2010 Robust normal estimation for point clouds with sharp features
Bao Li 0002, Ruwen Schnabel, Reinhard Klein, Zhi-Quan Cheng, Gang Dang, Shiyao Jin
Comput. Graph.1
2009 An Adaptive Octree Textures Painting Algorithm
abstract
Traditional texturing using a set of two dimensional image maps is an established and widespread practice. However, it is difficult to parameterize a model in texture space, particularly with representations such as implicit surfaces, subdivision surfaces, and very dense or detailed polygonal meshes. Based on an adaptive octree textures definition, this paper proposes a direct reverse-projecting pixel-level painting approach which has less storage requirements to general octree textures maps. In addition, it depends on texture lookup in the GPU, which particularly lookup faster than the non-GPU program.
Gang Dang, Zhi-Quan Cheng, Kai Xu 0004, Bao Li 0002
SMC6
2009 Variational Surface Approximation and Model Selection
abstract
Abstract We consider the problem of approximating an arbitrary generic surface with a given set of simple surface primitives. In contrast to previous approaches based on variational surface approximation, which are primarily concerned with finding an optimal partitioning of the input geometry, we propose to integrate a model selection step into the algorithm in order to also optimize the type of primitive for each proxy. Our method is a joint global optimization of both the partitioning of the input surface as well as the types and number of used shape proxies. Thus, our method performs an automatic trade‐off between representation complexity and approximation error without relying on a user supplied predetermined number of shape proxies. This way concise surface representations are found that better exploit the full approximative power of the employed primitive types.
Bao Li 0002, Ruwen Schnabel, Shiyao Jin, Reinhard Klein
Comput. Graph. Forum1
2008 An Error-Resilient Arithmetic Coding Algorithm for Compressed Meshes
abstract
The effort of on-the-fly accessing 3D contents over the Internet has been done in recent years. And 3D streaming has been investigated to represent 3D models in the compact format and progressively transmit them on the limited-bandwidth and lossy channel. In the paper, an error resilient 3D mesh coding algorithm is presented, which employs an extended multiple quantization (EMQ) arithmetic coder method, inspired by the error-resilient JPEG 2000 image coding standards. With periodic arithmetic coder restarting and termination markers, the error resilient EMQ coder divides bit stream into little independent parts and enables basic transmission error containment. Furthermore, the EMQ coder has the intrinsic capacity of handling noise by using the maximum a posteriori (MAP) decoder. Experiments show that the method improves the mesh transmission quality in a simulated network environment.
Zhi-Quan Cheng, Bao Li 0002, Kai Xu 0004, Gang Dang, Shiyao Jin
CW2
2008 Meaningful Mesh Segmentation Guided by the 3D Short-Cut Rule
Zhi-Quan Cheng, Bao Li 0002, Gang Dang, Shiyao Jin
GMP2