Yuanning Gao

dblp:142/3841 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
8since 2021 · last 2023
0000-0002-9072-1084ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Computer networks · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2023 DBAugur: An Adversarial-based Trend Forecasting System for Diversified Workloads
abstract
Trend forecasting is vital to optimize the workload performance. It becomes even more urgent with an increasing number of applications and database configurations. However, DBAs mainly target at historical workloads and may give suboptimal configuration advice when the workload trends have changed. Although there are some studies on trend forecasting, they have several limitations. First, they mainly predict the changes of query numbers, which do not combine other critical factors (e.g., disk utilization) and cannot fully reflect the future workload trends. Besides, there are numerous queries in the workloads and exact clustering algorithms like K-means cannot effectively merge similar queries which contain noises like time shifts. Second, basic machine learning models like RNN may have relatively low prediction accuracy on complex workloads (e.g., no cycles but random bursts). Third, real-world workloads may have diverse patterns, while previous models cannot efficiently and reliably predict for all the different workload patterns.To address these challenges, we propose a trend forecasting system (DBAugur) that utilizes adversarial neural networks to predict the trends of different workloads. First, DBAugur collects the important features (e.g., queries, resource metrics) to characterize workloads, and reduces the number of involved queries by separately merging similar queries based on the SQL semantics and trend patterns. Second, DBAugur utilizes Generative Adversarial Networks (GANs) to capture the latent patterns, correlations between different metrics, and occasional bursts within the complicated and time-varying workloads. Moreover, we further propose a time-sensitive ensemble algorithm that takes advantage of various machine learning models (e.g., generative models, convolutional models, feed-forward models) to accommodate the various workload patterns. The experimental results show that DBAugur outperformed state-of-the-art methods on various real-world workloads.
Yuanning Gao, Xiuqi Huang, Xuanhe Zhou, Xiaofeng Gao 0001, Guoliang Li 0001, Guihai Chen
ICDE1
2023 An Adaptive Metadata Management Scheme Based on Deep Reinforcement Learning for Large-Scale Distributed File Systems
abstract
A major challenge confronting today’s distributed metadata management schemes is how to meet the dynamic requirements of various applications through effectively mapping and migrating metadata nodes to different metadata servers (MDS’s). Most of the existing works dynamically reallocate nodes to different servers adopting history-based coarse-grained methods, failing to make a timely and efficient update on the distribution of nodes. In this paper, we present the first deep reinforcement learning-leveraged distributed metadata management scheme, AdaM, to address the aforementioned dilemma. AdaM is an adaptive fine-grained metadata management scheme that trains an actor-critic network to migrate “hot” metadata nodes to different MDS’s based on its observations of the current “states” (i.e., access pattern, the structure of namespace tree and current distribution of nodes on MDS’s). Adaptive to varying access patterns, AdaM can automatically migrate hot metadata nodes among servers to keep load balancing while maintaining metadata locality. Besides, we propose a self-adaptive metadata cache policy, which dynamically combines the two strategies of managing caches on the server side and the client side to gain better query performance. Last but not least, we design a distributed metadata processing 2PC Protocol called MST-based 2PC to ensure data consistency. Experiments on a real-world dataset demonstrate the superiority of our proposed method over other schemes.
Xiuqi Huang, Yuanning Gao, Xinyi Zhou 0006, Xiaofeng Gao 0001, Guihai Chen
IEEE/ACM Trans. Netw.2
2022 MetisRL: A Reinforcement Learning Approach for Dynamic Routing in Data Center Networks
Yuanning Gao, Xiaofeng Gao 0001, Guihai Chen
DASFAA (2)1
2022 LightPro: Lightweight Probabilistic Workload Prediction Framework for Database-as-a-Service
abstract
Nowadays, Database-as-a-Service (DBaaS) has become more and more popular among users as it can largely reduce the complexity of managing databases and applications. Considering the increasing complexity of different applications, system management and automation such as self-provisioning and performance tuning can be challenging. To better achieve autonomous optimization of the system, the ability to predict future workload patterns is of great essence. In this paper, we propose a novel lightweight probabilistic workload forecasting framework (LIGHTPRO) that is easy to train and robust, to help the system predict future workload patterns, leveraging multi-head attention mechanism and convolution operations. Experiments on real-world query traces demonstrate the superiority of LIGHTPRO in reducing training time and capturing both long-term and short-term temporal patterns of the workload compared with other baselines.
Xiuqi Huang, Shiyi Cao, Yuanning Gao, Xiaofeng Gao 0001, Guihai Chen
ICWS3
2022 An End-to-End Learning-Based Metadata Management Approach for Distributed File Systems
abstract
Current distributed file systems are designed to support PB-scale even EB-scale data storage. Metadata service, which manages file attribute information and the global namespace tree, is crucial to system performance. Distributed metadata management, using multiple metadata servers (MDS's) to store metadata, provides effective approaches to alleviate the workload of a single server. However, maintaining good metadata locality and keeping load balancing among MDS's at the same time is a nontrivial problem. To better take advantage of the current distribution of the metadata, in this article, we present the first machine learning based model called DeepHash, which leverages the neural network to learn a locality preserving hashing (LPH) mapping scheme. DeepHash first converts the metadata nodes to feature vectors by the network embedding technology. Due to the absence of training labels, i.e., the hash values of metadata nodes, we design a pair loss function with distinctive characters to train DeepHash, and introduce the sampling strategy to improve the training efficiency. Besides, we propose an efficient algorithm to dynamically balance the workload and adopt the cache model to improve query efficiency. The experiments on the Amazon EC2 platform demonstrate that the DeepHash can preserve the metadata locality meanwhile maintaining a high load balancing, which denotes the effectiveness and efficiency of DeepHash compared with traditional and state-of-the-art schemes.
Yuanning Gao, Xiaofeng Gao 0001, Ruisi Zhang, Guihai Chen
IEEE Trans. Computers1
2022 An Embedded GRASP-VNS based Two-Layer Framework for Tour Recommendation
abstract
The advance of tour recommendation allows people to get well-fit route plans, which contain a sequence of Points of Interest (POIs) based on tourists’ constraints and preferences. Large-scale POIs, called Super-POIs in this article, often contain multiple scenic spots and entrances. Tourists have to specify a suitable tour route inside Super-POI to obtain good tour experience. However, most of existing tour recommendation algorithms ignore the detailed information inside Super-POIs. By taking super-POIs into account, we proposeEmbedded Tour(eTOUR), a two-layer framework considering route design of POIs (Outer Model) and scenic routes inside the Super-POIs (Inner Model) respectively. To combine two models, an Embedded GRASP-VNS Algorithm is introduced based on an embedding strategy. For Outer Model, we apply Greedy Randomized Adaptive Search Procedure (GRASP) for route construction and Variable Neighborhood Search (VNS) for local improvement. Super-POI is treated as a “meta node” in outer route construction. For Inner Model, the optimal route inside Super-POI obtained by DFS-based Tree Search with Pruning is revised dynamically to adapt to the outer route. Furthermore, we discuss a special case in the Super-POI where a key graph is defined and treated as the “must go” route. We modify the solution of Chinese Postman Problem in this case to reduce the time complexity. Finally, experiments based on two real datasets demonstrate the effectiveness of our proposal.
Yuanning Gao, Xiaofeng Gao 0001, Xianyue Li, Bin Yao 0002, Guihai Chen
IEEE Trans. Serv. Comput.1
2021 An Attention-Based Bi-GRU for Route Planning and Order Dispatch of Bus-Booking Platform
Yucen Gao, Yuanning Gao, Xiaofeng Gao 0001, Xiang Li 0006, Guihai Chen
DASFAA (1)2
2021 NETR-Tree: An Efficient Framework for Social-Based Time-Aware Spatial Keyword Query
abstract
The development of global positioning system stimulates the popularity of location-based social network (LBSN) services. With a large volume of data containing locations, texts, check-in information, and social relationships, spatial keyword queries in LBSNs have become increasingly complex. In this paper, we identify and solve the Social-based Time-aware Spatial Keyword Query (STSKQ) that returns the top-k objects by considering geo-spatial score, keywords similarity, visiting time score, and social relationship effect. To tackle STSKQ, we propose a two-layer hybrid index structure called Network Embedding Time-aware R-tree (NETR-Tree). In the user layer, we exploit the network embedding strategy to measure the relationship effect in users' relationship network. In the location layer, we build a Time-aware R-tree (TR-tree) considered spatial objects' spatiotemporal check-in information, and present a corresponding query processing algorithm. Finally, extensive experiments on two different real-life LBSNs demonstrate the effectiveness and efficiency of our methods, compared with existing state-of-the-art methods.
Xiuqi Huang, Yuanning Gao, Xiaofeng Gao 0001, Guihai Chen
ICWS2
2020 Adaptive Recollected RNN for Workload Forecasting in Database-as-a-Service
Chenzhengyi Liu, Weibo Mao, Yuanning Gao, Xiaofeng Gao 0001, Shifu Li, Guihai Chen
ICSOC3
2019 AdaM: An Adaptive Fine-Grained Scheme for Distributed Metadata Management
abstract
Distributed metadata management, administrating the distribution of metadata nodes on different metadata servers (MDS's), can substantially improve overall performance of large-scale distributed storage systems if well designed. A major difficulty confronting many metadata management schemes is the trade-off between two conflicting aspects: system load balance and metadata locality preservation. It becomes even more challenging as file access pattern inevitably varies with time. However, existing works dynamically reallocate nodes to different servers adopting history-based coarse-grained methods, failing to make timely and efficient update on distribution of nodes. In this paper, we propose an adaptive fine-grained metadata management scheme, AdaM, leveraging Deep Reinforcement Learning, to address the trade-off dilemma against time-varying access pattern. At each time step, AdaM collects environmental "states" including access pattern, the structure of namespace tree and current distribution of nodes on MDS's. Then an actor-critic network is trained to reallocate hot metadata nodes to different servers according to the observed "states". Adaptive to varying access pattern, AdaM can automatically migrate hot metadata nodes among servers to keep load balancing while maintaining metadata locality. We test AdaM on real-world data traces. Experimental results demonstrate the superiority of our proposed method over other schemes.
Shiyi Cao, Yuanning Gao, Xiaofeng Gao 0001, Guihai Chen
ICPP2
2019 DeepHash: An End-to-End Learning Approach for Metadata Management in Distributed File Systems
abstract
In distributed file systems, distributed metadata management can be considered as a mapping problem, i.e., how to effectively map the metadata namespace tree to multiple metadata servers (MDS's). In general, all traditional distributed metadata management schemes simply presume a rigid mapping function, thus failing to adaptively meet the requirements of different applications. To better take advantage of the current distribution of the metadata, in this exploratory paper, we present the first machine learning based model called DeepHash, which leverages the deep neural network to learn a locality preserving hashing (LPH) mapping. To help learn a good position relationship of metadata nodes in the namespace tree, we first present a metadata representation strategy. Due to the absence of training labels, i.e., the hash values of metadata nodes, we design two kinds of loss functions with distinctive characters to train DeepHash respectively, including a pair loss and a triplet loss, and introduce some sampling strategies for these two approaches. We conduct extensive experiments on Amazon EC2 platform to compare the performance of DeepHash with traditional and state-of-the-art schemes. The results demonstrate that DeepHash can preserve the metadata locality well while maintaining a high load balancing, which denotes the effectiveness and efficiency of DeepHash.
Yuanning Gao, Xiaofeng Gao 0001, Guihai Chen
ICPP1
2019 An efficient and scalable multi-dimensional indexing scheme for modular data centers
Yuanning Gao, Xiaofeng Gao 0001, Yichen Zhu 0002, Guihai Chen
Data Knowl. Eng.1
2019 U2-Tree: A Universal Two-Layer Distributed Indexing Scheme for Cloud Storage System
abstract
The indices in cloud storage systems manage the stored data and support diverse queries efficiently. Secondary index, the index built on the attributes other than the primary key, facilitates a variety of queries for different purposes. An efficient design of secondary indices is called two-layer indexing scheme. It divides indices in the system into the global index layer and the local index layer. However, previous works on two-layer indexing are mainly on a P2P overlay network. In this paper, we propose U2-Tree, a universal two-layer distributed indexing scheme built on data center networks with tree-like topologies. To construct the U2-Tree, we first build local index according to data features and, then, assign potential indexing range of the global index for each host based on the distribution rule of local data. After that, we use several false positives control techniques, including gap elimination and Bloom filter, to publish meta-data about local index to global index host. In the final step, the global index collects published information and uses tree data structures to organize them. In our design, we take advantage of the topological properties of tree-like topologies, introduce and compare detailed optimization techniques in the construction of two-layer indexing scheme. Furthermore, we discuss the index updating, index tuning, and the fault tolerance of U2-Tree. Finally, we validate the effectiveness and efficiency of U2-Tree by giving a series of theoretical analyses and conducting numerical experiments on Amazon EC2 platform.
Xiaofeng Gao 0001, Yuanning Gao, Yichen Zhu 0002, Guihai Chen
IEEE/ACM Trans. Netw.2
2019 An Efficient Ring-Based Metadata Management Policy for Large-Scale Distributed File Systems
abstract
The growing size of modern file system is expected to reach EB-scale. Therefore, an efficient and scalable metadata service is critical to system performance. Distributed metadata management schemes, which use multiple metadata servers (MDS's) to store metadata, provide a highly effective approach to alleviate the workload of a single server. However, it is difficult to maintain good metadata locality and load balancing among MDS's at the same time. In this paper, we propose a novel hashing scheme called AngleCut to partition metadata namespace tree and serve large-scale distributed storage systems. AngleCut first uses a locality preserving hashing (LPH) function to project the namespace tree into linear keyspace, i.e., multiple Chord-like rings. Then we design a history-based allocation strategy to adjust the workload of MDS's dynamically. Besides, we propose a two-layer metadata cache mechanism, including server-side cache and client-side cache to provide the two stage access acceleration. Last but not least, we introduce a distributed metadata processing 2PC Protocol Based on Message Queue (2PC-MQ) to ensure data consistency. In general, our scheme preserves good metadata locality as well as maintains a high load balancing between MDS's. The theoretical proof and extensive experiments on Amazon EC2 demonstrate the superiority of AngleCut over previous literature.
Yuanning Gao, Xiaofeng Gao 0001, Xiaochun Yang 0001, Guihai Chen
IEEE Trans. Parallel Distributed Syst.1
2018 eTOUR: A Two-Layer Framework for Tour Recommendation with Super-POIs
Chunwei Wang, Yuanning Gao, Xiaofeng Gao 0001, Bin Yao 0002, Guihai Chen
ICSOC2