Haogang Tong

dblp:356/3049 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0005-7179-9053ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 BASE: Burst-Adaptive Autoscaling via Stacked Ensembles for SLO Assurance and Cost Efficiency
abstract
Autoscaling is a technology that automatically scales resources for applications without human intervention to ensure runtime Quality of Service (QoS) while reducing costs. However, user-facing cloud applications serve dynamic workloads that often exhibit variability and contain bursts, posing challenges to autoscaling in maintaining QoS within Service-Level Objectives (SLOs). Conservative strategies risk over-provisioning, while aggressive ones may cause SLO violations, making it more challenging to design effective autoscaling. This paper introduces BASE, a burst-adaptive autoscaling framework that leverages a stacked ensemble of machine learning models to mitigate SLO violations and reduce costs for containerized services and applications operating under time-varying workloads. BASE incorporates a novel prediction-based burst detection mechanism that distinguishes between predictable workload spikes and actual uncertain bursts. When bursts are detected, BASE appropriately overestimates them and allocates resources accordingly to address the rapid growth in resource demand. On the other hand, BASE employs reinforcement learning to rectify potential inaccuracies in resource estimation, enabling more precise resource allocation during non-burst periods. Experiments across ten real-world workloads demonstrate BASE's effectiveness, achieving a significant reduction in SLO violations with lower resource costs compared to other prominent methods.
Chunyang Meng, Haogang Tong, Tianyang Wu, Maolin Pan, Yang Yu 0027, Yi Jiang 0012
IEEE Trans. Serv. Comput.2
2026 SynScale: Spatiotemporal Collaborative Autoscaling for Microservices in Edge-Clouds
abstract
Edge-cloud environments constitute a heterogeneous computing paradigm that integrates resource-constrained edge servers with high-performance cloud servers. While microservices have revolutionized large-scale applications development by enhancing scalability and flexibility, achieving efficient microservice autoscaling—the dynamic adjustment of instances to maintain quality of service (QoS) and meet service-level agreements (SLA) targets—remains challenging in such environments. Most existing autoscaling techniques rely on metric forecasting and centralized or per-node control, causing them to overlook temporal evolution and server relationships and resulting in unstable and weakly coordinated scaling. To address these challenges, we present SynScale, a distributed collaborative autoscaling framework that strengthens both temporal sensitivity and structural coordination. SynScale introduces a spatiotemporal representation module that couples a temporal attention network with a multi graph convolutional model. The temporal component distills behavioral trends from recent observations to ensure coherent responses to time-varying dynamics, while the spatial component performs relational reasoning over explicitly modeled inter server correlations to yield structure-aware representations that support effective cross-node coordination. These spatiotemporal embeddings are then used within a multi-agent reinforcement learning paradigm, enabling distributed agents to generate context-aware scaling decisions that align local adaptability with system-wide efficiency. Experimental evaluations against state of-the-art autoscaling techniques show that SynScale reduces average response time by 59.32%, SLA violations by 85.71%, and P95 latency by 73.2%, while also lowering resource cost. It further improves scaling stability—reflected by lower instance time and longer instance lifetimes—and maintains lightweight, scale-stable runtime overhead, ensuring practical deployability in heterogeneous edge-cloud environments.
Haogang Tong, Chunyang Meng, Maolin Pan, Yang Yu 0027
IEEE Trans. Serv. Comput.1
2024 FuncScaler: Cold-Start-Aware Holistic Autoscaling for Serverless Resource Management
abstract
Serverless computing, an emerging paradigm, bolsters operational efficiency and cost savings by enabling the dynamic execution of fine-grained functions through a Function as a Service (FaaS). It incorporates autoscaling, thereby eliminating the need for infrastructure management. However, the conventional autoscaling strategies employed by cloud service providers frequently result in cold starts and inefficiencies. The complexity of variable workloads and the intricate interdependencies between functions amplify this challenge. For the above, this paper introduces FuncScaler, a deep learning- enhanced, cold-start-aware, holistic autoscaling approach for FaaS. FuncScaler employs a Gated Recurrent GCN to capture spatiotemporal relationships among functions, aggregating neighbour information and controlling information flow. Additionally, the framework integrates a cold-start-aware Jackson Queuing Network (JQN), which utilizes predicted workloads to initiate pre-warming and holistically autoscales. This empowers FuncScaler to reveal deeper function relationships, precisely forecast workloads, and concurrently reconfigure interconnected resources, effectively mitigating cold starts and hotspots. Our study shows FuncScaler excels beyond five comparative methods in ensuring Quality of Service (QoS) and enhancing resource utilization, additionally reducing cold start latency by 42.3% compared to the second-best approach.
Haogang Tong, Chunyang Meng, Maolin Pan, Yang Yu 0027
ICWS2
2023 DeepScaler: Holistic Autoscaling for Microservices Based on Spatiotemporal GNN with Adaptive Graph Learning
abstract
Autoscaling functions provide the foundation for achieving elasticity in the modern cloud computing paradigm. It enables dynamic provisioning or de-provisioning resources for cloud software services and applications without human intervention to adapt to workload fluctuations. However, autoscaling microservice is challenging due to various factors. In particular, complex, time-varying service dependencies are difficult to quantify accurately and can lead to cascading effects when allocating resources. This paper presents DeepScaler, a deep learning-based holistic autoscaling approach for microservices that focus on coping with service dependencies to optimize service-level agreements (SLA) assurance and cost efficiency. DeepScaler employs (i) an expectation-maximization-based learning method to adaptively generate affinity matrices revealing service dependencies and (ii) an attention-based graph convolutional network to extract spatio-temporal features of microservices by aggregating neighbors' information of graph-structural data. Thus DeepScaler can capture more potential service dependencies and accurately estimate the resource requirements of all services under dynamic workloads. It allows DeepScaler to reconfigure the resources of the interacting services simultaneously in one resource provisioning operation, avoiding the cascading effect caused by service dependencies. Experimental results demonstrate that our method implements a more effective autoscaling mechanism for microservice that not only allocates resources accurately but also adapts to dependencies changes, significantly reducing SLA violations by an average of 41% at lower costs.
Chunyang Meng, Haogang Tong, Maolin Pan, Yang Yu 0027
ASE3