VLDB 2026 Research / reviewers in the wild / expert
Binbin Feng
dblp:309/7506
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-9134-3484ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WIET: Harmonizing Group-aware Model Weighting and Worker Allocation for Ensemble Temporal Prediction MaaSabstractEnsemble Temporal Prediction Model-as-a-Service (ETP-MaaS) has become crucial in fields like financial modeling and cloud monitoring. Existing solutions fail to co-optimally address a two-fold challenge of dynamic collaboration and heterogeneity, treating models as independent entities and employing simplistic worker allocation rules. However, at the model level, data volatility means that optimal performance requires identifying and weighting constantly shifting subgroups of base models, not just individual ones; at the system level, these model groups must be efficiently mapped to a pool of heterogeneous and dynamically available workers. To this end, we introduce WIET, an efficient ETP-MaaS system that co-optimizes model weighting and worker allocation. For adaptive weighting, WIET identifies evolving group behaviors among base models and propose a novel group temporal locality-enhanced weighting method. Additionally, WIET develops an efficient, multi-dimensional worker allocation method powered by hybrid heuristic optimization, effectively reducing bottlenecks and resource waste. Experiments show WIET consistently outperforms state-of-the-art methods in terms of accuracy, latency, and resource usage across various workloads and tasks. Binbin Feng, Shikun He, Yingxin Wang, Pengwei Wang 0001, Zhijun Ding |
AAAI | 1 |
| 2026 | MGroup: Multi-Instance Workload Prediction Approach Based on Group Behavior PerceptionabstractWorkload prediction is a key step in artificial intelligence for IT operations (AIOps) on cloud platforms, enabling proactive application management for performance assurance, cost reduction, and energy optimization. With the popularity of microservice architectures, user requests are now handled collaboratively by multiple service instances, so the workload variation is no longer the individual behavior of each instance but the group behavior of multiple instances. However, existing approaches typically analyze each instance independently and fail to explicitly model group-level workload evolution, leading to suboptimal predictions. To address this issue, we propose MGroup, a workload group behavior-aware multi-instance workload prediction method. First, we define the concept of highly-coupled multi-instance workload group behavior and its evolution, shifting the analytical focus from individuals to groups; second, we propose an adaptive method for identifying and characterizing the workload group behavior based on both static and dynamic correlations, shifting from similarity-based to correlation-based representation; third, we propose a multi-instance parallel prediction neural network that jointly captures local and global workload evolution, shifting from implicit modeling to explicit modeling. Based on this approach, we design a workload prediction system tailored to cloud-native applications. Finally, experimental results on public datasets show that MGroup reduces MAE by 14.62%-21.60% and RMSE by 21.97%-29.27% compared to existing state-of-the-art methods, which provides an effective solution for realizing workload prediction for cloud-native applications. Binbin Feng, Zhijun Ding, Changjun Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | Multi-agent-Driven Dual-Layer Serverless Adaptive Ensemble Inference Method
Yingxin Wang, Binbin Feng, Zhijun Ding |
ICSOC (1) | 2 |
| 2025 | DesFaaS: Cross-Layer Joint Dynamic Deployment System for Serverless Stateful FunctionsabstractThe on-demand resource model of serverless computing has driven its growing popularity. However, stateful applications require external mechanisms to manage their state, often through dynamic migration for resource allocation. This migration causes function locations to change dynamically, necessitating the scheduling of state data accordingly. Consequently, both stateful functions and their associated states must adapt efficiently to environmental changes-a co-adjustment relationship that is often overlooked or oversimplified in existing systems. Therefore, we design and build DesFaaS, a cross-layer joint dynamic deployment system, to provide an open source solution supporting real Kubernetes clusters for the runtime management of serverless stateful applications with variable locations. Specifically, first, we propose a cross-layer joint dynamic deployment framework, introducing three-layer hybrid migration media and designing a three-layer state management architecture. The designed automated management processes and interfaces support efficient function migration and state management under different load environments. Second, we propose an adaptive co-scheduling method for multi-function state data to optimize the access latency and resource utilization of state data in the system and support in-situ computing by collecting and analyzing the cost. Finally, we developed a prototype system, DesFaaS, and conducted real and simulated experiments, which demonstrated that DesFaaS has 19.36%, 39.06%, and 5.36% improvement in service performance, cost efficiency, and energy consumption compared with the state-of-the-art system. Yuquan Jing, Binbin Feng, Zhijun Ding |
IEEE Trans. Cloud Comput. | 2 |
| 2024 | LGDCloudSim: A Resource Management Simulation System for Large-Scale Geographically Distributed Cloud Data Center ScenariosabstractCurrent IaaS providers have deployed data centers worldwide, with resources continually increasing. Meanwhile, there is a rising trend in the concurrency of user requests and the diversity of user request types. To achieve better resource allocation, various complex scheduling architectures have been proposed. However, due to the challenges associated with real-world experiments, simulation systems are needed to build exper-imental environments for related research. As existing systems do not perform well enough, we construct LGDCloudSim. It is designed with full consideration of the characteristics of the large-scale geographically distributed cloud data center scenarios. To support large-scale simulations, we propose state management optimization and operation process optimization methods. Exper-iments show that LGDCloudSim can simulate up to 5 X 108hosts and 107request concurrency. It also supports diverse scheduling architectures and different request types. Yuehao Xu, Binbin Feng, Zhijun Ding |
CLOUD | 3 |
| 2024 | Context-Aware Runtime Type Prediction for Heterogeneous Microservices
Yibing Lin, Binbin Feng, Zhijun Ding |
Euro-Par (1) | 2 |
| 2024 | Less: Large-scale Workload Forecasting Model Based on Multiple Sequence CompressionabstractAs application migration to the cloud becomes the mainstream way of application deployment, application runtime management presents a significant need for large-scale workload prediction technology. Existing large-scale workload approaches commonly construct training sets based on all the training samples generated from the original data to generate prediction models. However, due to the similar behavior between different container instances of the microservice application, training and modeling in this way results in a huge number of redundant samples, which produces a significant redundant training overhead. Therefore, this paper proposes Less, a largescale workload forecasting model based on multiple sequence compression. First, based on the grouping results of similar containers, a container workload feature recognition algorithm is proposed to determine the common and individual features of container workloads in each prediction period, so as to guide the compression of workload sequences within each group; second, a fitness function that takes into account the common features, individual features, and the number of sequences are designed, and the optimal compressed sequences are solved by the Whale Optimization Algorithm to efficiently reduce the number of redundant training workload sequences, and then the Bidirectional Gated Recurrent Unit model is built and trained based on the compressed sequences, which effectively reduces the model complexity and overhead while ensuring the accuracy. Finally, we validate the comprehensive advantages of Less in terms of accuracy and overhead based on public datasets and verify the effectiveness of each subpart of our model through ablation experiments. Zeyuan Ding, Binbin Feng, Wangyang Yu 0001, Bolan Zhang |
ICWS | 2 |
| 2024 | Proactive Elastic Scheduling for Serverless Ensemble Inference ServicesabstractRecently, AI inference services have adopted ensemble architectures, which are widely recognized and used for their advanced performance. However, the existing ensemble inference services are mainly created and managed using the platform-as-a-service model with the static ensemble service architecture and rigid persistent resource allocation. This makes the highly heterogeneous inference requests rely on a fixed combination of basic learners and manual homogeneous resource management, resulting in insufficient precision, waste of resources, and high management costs. Serverless computing, represented by Functions-as-a-Service (FaaS), realizes transparent on-demand resource allocation to developers, which is suitable for ensemble inference services. Therefore, we propose a serverless proactive elastic scheduling solution PESEI for ensemble inference services. First, a two-level hierarchical dynamic ensemble service framework with joint model precision and overhead sensing is proposed to ensure inference precision while improving cost efficiency; second, a proactive elastic resource allocation algorithm with dynamic sensing of workload patterns is proposed to optimize the quality of service and cost efficiency for heterogeneous base learners; based on this, a prototype system is designed and developed to support the autonomous management of ensemble inference services, which implements the benign combination and adaptation of dynamic service architecture and elastic resource management. Real cluster experiments on public datasets demonstrate the effectiveness and robustness of PESEI, providing a new solution for building serverless ensemble inference services. Shikun He, Binbin Feng, Zhijun Ding |
ICWS | 2 |
| 2024 | Adaptive Selecting Algorithm for Runtime Types of MicroservicesabstractIn recent years, serverless computing has become increasingly popular in the domain of microservices. Compared to serverful computing, serverless computing significantly reduces developers’ expenses due to its resource elasticity and on-demand allocation features. However, serverless computing suffers from long cold start time and high function invocation latency, leading to suboptimal service performance. Due to the dynamic workload, microservices exhibit varying demands for different runtime types over time, which is overlooked by existing approaches. Therefore, we propose the Adaptive Selecting Algorithm for Runtime Types of microservices, which optimizes resource usage in cloud service providers (CSPs) and ensures efficient execution of developers’ applications. Specifically, the algorithm dynamically switches microservices to the optimal runtime type by analyzing the workload patterns, resource requirements, and execution efficiency of microservices. We conducted experiments on real clusters, demonstrating that algorithm enhances the service quality of microservices while improving their cost efficiency. Binbin Feng, Zhijun Ding |
ICWS | 2 |
| 2024 | Heterogeneity-Aware Proactive Elastic Resource Allocation for Serverless ApplicationsabstractServerless computing is a popular cloud computing model that offers on-demand resource allocation and pay-as-you-go application execution. However, there are still challenges in allocating resources for workflow applications: inaccurate and inefficient resource estimation, high-latency inter-function communication, and long server readiness time. Therefore, we propose the heterogeneity-aware Proactive serverLess wOrkflow Elastic Allocation method (PLOEA) to address these issues and optimize infrastructure costs for cloud service providers (CSPs) while meeting the diverse needs of developers. Specifically, we propose a resource configuration estimation method for heterogeneous workflow applications that builds an ensemble multi-task expert classifier to analyze individual and common resource usage patterns, ensuring estimation accuracy and efficiency. Further, we propose a group allocation strategy for multiple applications that optimizes the spatiotemporal distribution of instances by considering the allocation urgency, communication affinity between functions, and the multi-core architecture of servers. Furthermore, we present a proactive server elastic scaling method that senses workload features, including workload level, trend, and magnitude changes, and combines them with CSP's attention differences to guide the server scaling size. Finally, experiments based on public datasets prove that PLOEA provides better service quality and cost efficiency than existing methods. Binbin Feng, Zhijun Ding, Changjun Jiang 0002 |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | Microservice Extraction Based on a Comprehensive Evaluation of Logical Independence and PerformanceabstractMonolithic architectures are becoming increasingly difficult to cope with complex applications, and microservice architectures, which offer flexibility and logical independence in development and maintenance, are the new choice for companies and developers. Migrating a legacy monolithic architecture application to a microservice architecture rather than building it from scratch is considered an easy way to use it. To ensure that the migrated microservice applications can take advantage of their benefits, we need to propose a reasonable and effective microservice extraction method. Considering the single responsibility principle in the microservice design principle, most existing microservice extraction methods only pursue the high logical independence of the extraction results and pay little attention to whether the extraction results have good performance. Applications need to perform well, and studies have shown that poor microservice extraction schemes can negatively impact the performance of the migrated application. As a result, when extracting, we should also consider the performance of the results. A few studies consider the performance of extraction results, but only in terms of a few factors affecting performance, such as network overhead, rather than considering all factors affecting performance comprehensively, which leads to an inaccurate evaluation of performance. Therefore, oriented toward the most widely used managed languages today, we propose an effective Microservice Extraction method based on a Comprehensive Evaluation of logical independence and performance (MECE). Firstly, we propose a workflow-based approach to evaluate the performance of microservice extraction results by considering multiple influencing factors, focusing on the management cost ignored in existing studies, and designing an effective management cost evaluation model. After that, we propose a meta-heuristic search-based algorithm to obtain feasible microservice extraction results. In experiments based on actual deployments, the extraction results of the MECE method obtained a performance improvement of up to 46.15% without significant loss of logical independence compared to existing methods, which verifies the effectiveness of the method. Zhijun Ding, Yuehao Xu, Binbin Feng, Changjun Jiang 0002 |
IEEE Trans. Software Eng. | 3 |
| 2023 | LogoNet: A Fine-Grained Network for Instance-Level Logo Sketch Retrieval
Binbin Feng, Jun Li 0033 |
ICIG (4) | 1 |
| 2023 | Work-in-Progress-A Large-Scale Workload Forecasting Model for ContainersabstractAs application migration to the cloud becomes the mainstream way of application deployment, application runtime management presents a significant need for large-scale workload prediction technology. However, existing large-scale workload forecasting models focus more on improving model accuracy and ignore the models’ storage, training time, and testing time, which leads to colossal overhead. Therefore, this paper proposes a large-scale workload forecasting model for containers. First, based on the workload value features and waveform features, a feature-enhanced workload similarity calculation algorithm is proposed to determine the grouping of containers with similar workload patterns in real time by analyzing the historical similarity and recent similarity of workloads among different containers; second, we employ Transformer as the base model to design position encoding and attention mask based on the real-time workload similarity relationship and achieve forecasting model parallelized training based on the multi-head self-attention mechanism, which balances the workload prediction accuracy and model overhead. Finally, we will validate the comprehensive advantages of our model in terms of accuracy and overhead based on public datasets and verify the effectiveness of each subpart of our model through ablation experiments. Zeyuan Ding, Binbin Feng, Wangyang Yu 0001 |
ICWS | 2 |
| 2023 | GROUP: An End-to-end Multi-step-ahead Workload Prediction Approach Focusing on Workload Group BehaviorabstractAccurately forecasting workloads can enable web service providers to achieve proactive runtime management for applications and ensure service quality and cost efficiency. For cloud-native applications, multiple containers collaborate to handle user requests, making each container’s workload changes influenced by workload group behavior. However, existing approaches mainly analyze the individual changes of each container and do not explicitly model the workload group evolution of containers, resulting in sub-optimal results. Therefore, we propose a workload prediction method, GROUP, which implements the shifts of workload prediction focus from individual to group, workload group behavior representation from data similarity to data correlation, and workload group behavior evolution from implicit modeling to explicit modeling. First, we model the workload group behavior and its evolution from multiple perspectives. Second, we propose a container correlation calculation algorithm that considers static and dynamic container information to represent the workload group behavior. Third, we propose an end-to-end multi-step-ahead prediction method that explicitly portrays the complex relationship between the evolution of workload group behavior and the workload changes of each container. Lastly, enough experiments on public datasets show the advantages of GROUP, which provides an effective solution to achieve workload prediction for cloud-native applications. Binbin Feng, Zhijun Ding |
WWW | 1 |
| 2023 | FAST: A Forecasting Model With Adaptive Sliding Window and Time Locality Integration for Dynamic Cloud WorkloadsabstractThe workload predictor has attracted attention as a key component of the proactive service operation management framework. However, the request and resource workloads of cloud applications are highly dynamic. Existing approaches decompose the original workload into trend, seasonal, and random components, establish models accordingly, and then combine all outputs to generate results. Indeed, the random component usually has significant heteroscedasticity and noise, having little or even a negative effect on model accuracy improvement. In our model, trend and seasonal components are seen as macro workload changes, and the micro workload changes are obtained by an adaptive sliding window algorithm. Therefore, we propose an ensembling model named FAST for Forecasting workloads with Adaptive Sliding window and Time locality integration. Notably, we propose an adaptive sliding window algorithm that considers trend correlation, time correlation, and random fluctuations of workload for online regression to achieve higher accuracy with lower overhead; and for the error-based integration strategy, we propose a time locality concept for local-predictor behavior and develop a multi-class regression algorithm for model integration. Finally, we conduct experiments on Google cluster trace datasets which show FAST has better accuracy than all other state-of-the-art models for dynamic workloads. Binbin Feng, Zhijun Ding, Changjun Jiang 0002 |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | Q-percentile Bandwidth Billing Based Geo-Scheduling AlgorithmabstractCurrent IaaS providers deploy cheaper computing resources in newly built data centers and provide cross-regional network services to improve the interoperability of computing resources in different regions. Third-party service providers can use part of their budget to purchase cross-regional communication resources to use cheaper resources in remote areas to reduce the cost of processing massive task requests. The Q-percentile charging model is widely used in cross-regional communication resources billing, but there is little task scheduling research on that billing method. Therefore, this paper studies a geo-distributed task scheduling scenario using the Q-percentile charging model. We design a geo-scheduling algorithm specifically for Q-percentile charging model to allocate resources in the two dimensions of computing resources and communication resources. Furthermore, referring to three existing communication resource allocation strategies, we design three bandwidth allocation algorithms considering the Q-percentile charging characteristics to provide suitable solutions for different scenarios. We conducted experiments based on public well-known datasets such as LIGO workflow. Results show that, compared with the baseline, the scheduling algorithm proposed in this paper can reduce the task scheduling cost between geo-distributed data centers by 10%-20% based on various task loads and show differences in the applicability of different communication resource allocation strategies. Yaoyin You, Binbin Feng, Zhijun Ding |
CLOUD | 2 |
| 2022 | COIN: A Container Workload Prediction Model Focusing on Common and Individual Changes in WorkloadsabstractRecently, containers have become the primary deployment form for cloud applications. Predicting container workload accurately is critical to ensure the quality of service (QoS) and cost-efficiency of the applications and meet service level agreements (SLAs) with users. However, facing multiple challenges, including model unavailability due to insufficient data, model maladaptation due to dynamic workload changes, and model non-generalization due to changeable workload patterns in container workload prediction, existing methods have not yet provided a united and effective solution. To this end, we propose a novel integrated forecasting model named COIN that combines COmmon and INdividual changes in container workloads to ensure the availability, adaptivity, and generality of the prediction model based on transfer learning and online learning. Besides, we present a container similarity calculation algorithm for real cloud scenarios, which combines the static and dynamic information of containers and comprehensively depicts the similarity between containers. Through experiments based on two public datasets, the COIN model achieves a higher accuracy than existing state-of-the-art solutions, demonstrating the effectiveness and robustness of our proposed model, which provides a new solution to container workload prediction. Zhijun Ding, Binbin Feng, Changjun Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |