EDBT 2026 Demo / reviewers in the wild / expert
Shiyi Zhu
dblp:321/4194
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 83% Energy-efficient computing · 17% | |
| Artificial intelligence
2 papers |
Deep learning architectures and training · 48% Reinforcement learning · 32% Language models and text generation · 21% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
autoscaling |
2.0 | 3 | 2024 | DeepScaling: Autoscaling Microservices With Stable CPU Utilization for Large Scale Production Cloud Systems · IEEE/ACM Trans. Netw. 2024 Full Scaling Automation for Sustainable Development of Green Data Centers · IJCAI 2023 A Meta Reinforcement Learning Approach for Predictive Autoscaling in the Cloud · KDD 2022 |
Cloud and datacenter computing
cluster resource management and scheduling |
1.4 | 2 | 2024 | DeepScaling: Autoscaling Microservices With Stable CPU Utilization for Large Scale Production Cloud Systems · IEEE/ACM Trans. Netw. 2024 Full Scaling Automation for Sustainable Development of Green Data Centers · IJCAI 2023 |
Cloud and datacenter computing
workload prediction |
0.8 | 2 | 2023 | Full Scaling Automation for Sustainable Development of Green Data Centers · IJCAI 2023 A Meta Reinforcement Learning Approach for Predictive Autoscaling in the Cloud · KDD 2022 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.8 | 1 | 2024 | CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending · ACL (1) 2024 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context window extension |
0.8 | 1 | 2024 | CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending · ACL (1) 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending · ACL (1) 2024 |
Cloud and datacenter computing › autoscaling
microservice autoscaling |
0.8 | 1 | 2024 | DeepScaling: Autoscaling Microservices With Stable CPU Utilization for Large Scale Production Cloud Systems · IEEE/ACM Trans. Netw. 2024 |
Energy-efficient computing
datacenter energy efficiency |
0.7 | 1 | 2023 | Full Scaling Automation for Sustainable Development of Green Data Centers · IJCAI 2023 |
Energy-efficient computing
power management |
0.7 | 1 | 2023 | Full Scaling Automation for Sustainable Development of Green Data Centers · IJCAI 2023 |
Machine learning › Reinforcement learning
meta-reinforcement learning |
0.6 | 1 | 2022 | A Meta Reinforcement Learning Approach for Predictive Autoscaling in the Cloud · KDD 2022 |
Machine learning › Reinforcement learning › meta-reinforcement learning
model-based meta reinforcement learning |
0.6 | 1 | 2022 | A Meta Reinforcement Learning Approach for Predictive Autoscaling in the Cloud · KDD 2022 |
Cloud and datacenter computing › autoscaling
predictive autoscaling |
0.6 | 1 | 2022 | A Meta Reinforcement Learning Approach for Predictive Autoscaling in the Cloud · KDD 2022 |
Cloud and datacenter computing
resource management |
0.6 | 1 | 2022 | A Meta Reinforcement Learning Approach for Predictive Autoscaling in the Cloud · KDD 2022 |
Machine learning › Deep learning architectures and training
positional encoding |
0.2 | 1 | 2024 | CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending · ACL (1) 2024 |
Methods — techniques the papers use, named apart from their topics
deep periodic workload prediction · 1.1spatio-temporal graph neural network · 0.8rotary position embedding · 0.8deep q-network · 0.8deep neural network · 0.8collinear constrained attention · 0.8deep representation learning · 0.7neural processes · 0.6neural process · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window ExtendingabstractSelf-attention and position embedding are two crucial modules in transformer-based Large Language Models (LLMs).However, the potential relationship between them is far from well studied, especially for long context window extending.In fact, anomalous behaviors that hinder long context extrapolation exist between Rotary Position Embedding (RoPE) and vanilla self-attention.Incorrect initial angles between Q and K can cause misestimation in modeling rotary position embedding of the closest tokens.To address this issue, we propose Collinear Constrained Attention mechanism, namely CoCA.Specifically, we enforce a collinear constraint between Q and K to seamlessly integrate RoPE and self-attention.While only adding minimal computational and spatial complexity, this integration significantly enhances long context window extrapolation ability.We provide an optimized implementation, making it a drop-in replacement for any existing transformer-based models.Extensive experiments demonstrate that CoCA excels in extending context windows.A CoCAbased GPT model, trained with a context length of 512, can extend the context window up to 32K (60×) without any fine-tuning.Additionally, incorporating CoCA into LLaMA-7B achieves extrapolation up to 32K within a training length of only 2K.Our code is publicly available at: https://github.com/codefuse- ai/Collinear-Constrained-Attention Shiyi Zhu, Wei Jiang 0041, Siqiao Xue, Yifan Wu 0002 |
ACL (1) | 1 |
| 2024 | DeepScaling: Autoscaling Microservices With Stable CPU Utilization for Large Scale Production Cloud SystemsabstractCloud service providers often provision excessive resources to meet the desired Service Level Objectives (SLOs), by setting lower CPU utilization targets. This can result in a waste of resources and a noticeable increase in power consumption in large-scale cloud deployments. To address this issue, this paper presents DeepScaling, an innovative solution for minimizing resource cost while ensuring SLO requirements are met in a dynamic, large-scale production microservice-based system. We propose DeepScaling, which introduces three innovative components to adaptively refine the target CPU utilization of servers in the data center, and we maintain it at a stable value to meet SLO constraints while using minimum amount of system resources. First, DeepScaling forecasts workloads for each service using a Spatio-temporal Graph Neural Network. Secondly, it estimates CPU utilization with a Deep Neural Network, considering factors such as periodic tasks and traffic. Finally, it uses a modified Deep Q-Network (DQN) to generate an autoscaling policy that controls service resources to maximize service stability while meeting SLOs. Evaluation of DeepScaling in Ant Group’s large-scale cloud environment shows that it outperforms state-of-the-art autoscaling approaches in terms of maintaining stable performance and resource savings. The deployment of DeepScaling in the real-world environment of 1900+ microservices saves the provisioning of over 100,000 CPU cores per day, on average. Shiyi Zhu, Wei Jiang 0041, K. K. Ramakrishnan, Meng Yan 0001, Xiaohong Zhang 0002, Alex X. Liu |
IEEE/ACM Trans. Netw. | 2 |
| 2023 | Continual Learning in Predictive AutoscalingabstractPredictive Autoscaling is used to forecast the workloads of servers and prepare the resources in advance to ensure service level objectives (SLOs) in dynamic cloud environments. However, in practice, its prediction task often suffers from performance degradation under abnormal traffics caused by external events (such as sales promotional activities and applications' re-configurations), for which a common solution is to re-train the model with data of a long historical period, but at the expense of high computational and storage costs. To better address this problem, we propose a replay-based continual learning method, i.e., Density-based Memory Selection and Hint-based Network Learning Model (DMSHM), using only a small part of the historical log to achieve accurate predictions. First, we discover the phenomenon of sample overlap when applying replay-based continual learning in prediction tasks. In order to surmount this challenge and effectively integrate new sample distribution, we propose a density-based sample selection strategy that utilizes kernel density estimation to calculate sample density as a reference to compute sample weight, and employs weight sampling to construct a new memory set. Then we implement hint-based network learning based on hint representation to optimize the parameters. Finally, we conduct experiments on public and industrial datasets to demonstrate that our proposed method outperforms state-of-the-art continual learning methods in terms of memory capacity and prediction accuracy. Furthermore, we demonstrate remarkable practicability of DMSHM in real industrial applications. Hongyan Hao, Zhixuan Chu, Shiyi Zhu, Gangwei Jiang, Yan Wang 0002, Caigao Jiang, James Y. Zhang, Siqiao Xue, Jun Zhou 0011 |
CIKM | 3 |
| 2023 | Full Scaling Automation for Sustainable Development of Green Data CentersabstractThe rapid rise in cloud computing has resulted in an alarming increase in data centers' carbon emissions, which now accounts for >3% of global greenhouse gas emissions, necessitating immediate steps to combat their mounting strain on the global climate. An important focus of this effort is to improve resource utilization in order to save electricity usage. Our proposed Full Scaling Automation (FSA) mechanism is an effective method of dynamically adapting resources to accommodate changing workloads in large-scale cloud computing clusters, enabling the clusters in data centers to maintain their desired CPU utilization target and thus improve energy efficiency. FSA harnesses the power of deep representation learning to accurately predict the future workload of each service and automatically stabilize the corresponding target CPU usage level, unlike the previous autoscaling methods, such as Autopilot or FIRM, that need to adjust computing resources with statistical models and expert knowledge. Our approach achieves significant performance improvement compared to the existing work in real-world datasets. We also deployed FSA on large-scale cloud computing clusters in industrial data centers, and according to the certification of the China Environmental United Certification Center (CEC), a reduction of 947 tons of carbon dioxide, equivalent to a saving of 1538,000 kWh of electricity, was achieved during the Double 11 shopping festival of 2022, marking a critical step for our company’s strategic goal towards carbon neutrality by 2030. Shiyu Wang 0001, Yinbo Sun, Xiaoming Shi 0001, Shiyi Zhu, Lintao Ma, James Zhang, Yangfei Zheng, Liu Jian |
IJCAI | 4 |
| 2022 | DeepScaling: microservices autoscaling for stable CPU utilization in large scale cloud systemsabstractCloud service providers conservatively provision excessive resources to ensure service level objectives (SLOs) are met. They often set lower CPU utilization targets to ensure service quality is not degraded, even when the workload varies significantly. Not only does this potentially waste resources, but it can also consume excessive power in large-scale cloud deployments. This paper aims to minimize resource costs while ensuring SLO requirements are met in a dynamically varying, large-scale production microservice environment. We propose DeepScaling, which introduces three innovative components to adaptively refine the target CPU utilization to a level that is maintained at a stable value to meet SLO constraints while using minimum resources. First, DeepScaling forecasts the workload for each service using a Spatio-temporal Graph Neural Network. Second, DeepScaling estimates the CPU utilization by mapping the workload intensity to an estimated CPU utilization with a Deep Neural Network, while taking into account multiple factors in the cloud environment (e.g., periodic tasks and traffic). Third, DeepScaling generates an autoscaling policy for each service based on an improved Deep Q Network (DQN). The adaptive autoscaling policy updates the target CPU utilization to be a maximum, stable value, while ensuring SLOs is not violated. We compare DeepScaling with state-of-the-art autoscaling approaches in the large-scale production cloud environment of the Ant Group. It shows that DeepScaling outperforms other approaches both in terms of maintaining stable service performance, and saving resources, by a significant margin. The deployment of DeepScaling in Ant Group's real production environment with 135 microservices saves the provisioning of over 30,000 CPU cores per day, on average. Shiyi Zhu, Wei Jiang 0041, K. K. Ramakrishnan, Yangfei Zheng, Meng Yan 0001, Xiaohong Zhang 0002, Alex X. Liu |
SoCC | 2 |
| 2022 | A Meta Reinforcement Learning Approach for Predictive Autoscaling in the CloudabstractPredictive autoscaling (autoscaling with workload forecasting) is an important mechanism that supports autonomous adjustment of computing resources in accordance with fluctuating workload demands in the Cloud. In recent works, Reinforcement Learning (RL) has been introduced as a promising approach to learn the resource management policies to guide the scaling actions under the dynamic and uncertain cloud environment. However, RL methods face the following challenges in steering predictive autoscaling, such as lack of accuracy in decision-making, inefficient sampling and significant variability in workload patterns that may cause policies to fail at test time. To this end, we propose an end-to-end predictive meta model-based RL algorithm, aiming to optimally allocate resource to maintain a stable CPU utilization level, which incorporates a specially-designed deep periodic workload prediction model as the input and embeds the Neural Process [11, 16] to guide the learning of the optimal scaling actions over numerous application services in the Cloud. Our algorithm not only ensures the predictability and accuracy of the scaling strategy, but also enables the scaling decisions to adapt to the changing workloads with high sample efficiency. Our method has achieved significant performance improvement compared to the existing algorithms and has been deployed online at Alipay, supporting the autoscaling of applications for the world-leading payment platform. Siqiao Xue, Chao Qu, Xiaoming Shi 0001, Cong Liao, Shiyi Zhu, Xiaoyu Tan, Lintao Ma, Shiyu Wang 0001, Yun Hu 0001, Lei Lei 0001, Yangfei Zheng, James Zhang |
KDD | 5 |