Jinlong Cai

dblp:131/9951 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2024 Online Index Recommendation for Slow Queries
abstract
Database autonomy service (DAS) is a platform that provides assistance to database maintainers or administrators in managing a large number of database instances in major internet companies. An important task of DAS is to find missing indexes to improve the performance of slow queries reported from its managed online database instances. Traditional database systems provide the “what-if” function or the hypothetical index technique. Index metadata is modified to simulate the benefits of indexes for queries without creating physical index files. Decades of research have led to plenty of ideas for index recommendation through the use of the “what-if” function and different search strategies. However, the popular open-source database system MySQL, used by most internet companies, has not provided the “what-if” function. In Meituan, tens of thousands of MySQL instances have been deployed across many business lines. Consequently, the DAS platform has accumulated lots of index creation samples. In this paper, we introduce index learner (IdxL), designed to learn index creation knowledge from these informative index data. IdxL resolves the problem of index recommendation by formulating it into an end-to-end supervised learning problem. Given a slow query, IdxL uses learned index creation knowledge to directly predict the missing indexes. Experimental results demonstrate: (1) IdxL is superior to the state-of-the-art index recommendation methods, especially when the error in cost estimation was propagated to the search in candidate index space, and (2) in particular, IdxL achieves up to 97% performance gain over a state-of-the-art method relying on the optimizer's cost estimation in the Meituan-specific index recommendation scenario. Finally, we present the applied results of IdxL in the Meituan DAS platform, demonstrating its ability to transfer index creation knowledge from certain databases to others.
Gan Peng, Kaikai Ye, Jinlong Cai, Yufeng Shen, Weiyuan Xu
ICDE5
2024 DB-MAGS: Multi-Anomaly Data Generation System for Transactional Databases
abstract
Existing database performance anomaly datasets have the problems of comprehensiveness in anomaly types, coarse-grained root causes, and unrealistic simulation for reproducing concurrent anomalies. To address these issues, we propose a data generation system tailored for Multi-Anomaly Reproduction in Databases (DB-MAGS). DB-MAGS guarantees unified, authentic, and comprehensive data generation, while also providing fine-grained root causes. In the case of only a single anomaly occurred in the database, we categorize the factors affecting database performance anomalies, select five major categories of anomalies, and further subdivide each category into eighteen minor categories. This finer granularity of anomaly classification facilitates more specific and targeted anomaly remediation. For multiple anomalies simultaneously occurred in a database system, we categorize the relationships between anomalies into causal and concurrent, and enumerate different combinations of multiple anomalies, making the simulation of multiple anomaly scenarios more comprehensive and enhancing the diversity of generated data.
Yiqi Shen, Miaodong Shen, Weiyuan Xu, Li Kai, Jinlong Cai
Proc. VLDB Endow.7
2023 A Data-Driven Index Recommendation System for Slow Queries
abstract
The Database Autonomy Service (DAS) is a platform designed to assist database administrators in managing a large number of database instances within major internet companies. One of the key tasks in DAS is to find missing indexes to improve the slow query execution. In Meituan, a vast array of business lines deploy tens of thousands of MySQL database instances. Consequently, a great number of human-generated index cases are accumulated in the DAS platform. This motivates us to build a data-driven index recommendation system, referred to as idxLearner, which can learn index creation knowledge from human-generated index cases. In this demonstration, users can interact with idxLearner by choosing source databases to construct the training data, training the recommendation model, inputting slow queries for various target databases, and observing the recommended indexes and their evaluation results.
Gan Peng, Peng Cai 0001, Kaikai Ye, Jinlong Cai, Yufeng Shen
CIKM5
2013 Adaptive scale based entropy-like estimator for robust fitting
abstract
In this paper, we propose a novel robust estimator, called ASEE (Adaptive Scale based Entropy-like Estimator) which minimizes the entropy of inliers. This estimator is based on IKOSE (Iterative Kth Ordered Scale Estimator) and LEL (Least Entropy-Like Estimator). Unlike LEL, ASEE only considers inliers' entropy while excluding outliers, which makes it very robust in parametric model estimation. Compared with other robust estimators, ASEE is simple and computationally efficient. From the experiments on both synthetic and real-image data, ASEE is more robust than several state-of-the-art robust estimators, especially in handling extreme outliers.
Jinlong Cai, Hanzi Wang
ICASSP1
2013 AMSAC: An adaptive robust estimator for model fitting
abstract
In this paper, we firstly propose a novel robust scale estimator called AIKOSE. It can estimate the scale of inlier noises by adaptively selecting the optimal value of K in the IKOSE scale estimator. Moreover, based on AIKOSE, we propose a novel robust estimator called AMSAC, which can fit a model without requiring a manually tuned threshold. In the experiments, we demonstrate the performance of AMSAC on line fitting and homography estimation by using both synthetic data and real images. Experimental results show that AM-SAC is more robust than other competing robust estimators.
Hanzi Wang, Jinlong Cai, Jianyu Tang
ICIP2
2013 Gaussian Function Assisted Neural Networks Decoding Algorithm for Turbo Product Codes
Xingcheng Liu, Jinlong Cai
ISNN (2)2