Rong Kang

dblp:71/3106 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorComputer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 VIDEX: A Disaggregated and Extensible Virtual Index for the Cloud and AI Era
abstract
Virtual indexes play a crucial role in database query optimization. However, with the rapid advancement of cloud computing and AI-driven models for database optimization, traditional virtual index approaches face significant challenges. Cloud-native environments often prohibit direct conducting query optimization process on production databases due to stability requirements and data privacy concerns. Moreover, while AI models show promising progress, their integration with database systems poses challenges in system complexity, inference acceleration, and model hot updates. In this paper, we present VIDEX, a three-layer disaggregated architecture that decouples database instances, the virtual index optimizer, and algorithm services, providing standardized interfaces for AI model integration. Users can configure VIDEX by either collecting production statistics or loading from a prepared file, enabling high-accuracy what-if analysis using virtual indexes that yield query plans identical to production instances. Additionally, users can freely integrate new AI-driven algorithms into VIDEX. VIDEX has been deployed at ByteDance, serving thousands of MySQL instances daily and over millions of SQL queries for index optimization tasks.
Rong Kang, Tieying Zhang, Xianghong Xu 0001, Linhui Xu, Zhimin Liang, Lei Zhang 0213, Jianjun Chen 0001
Proc. VLDB Endow.1
2025 Data-Agnostic Cardinality Learning from Imperfect Workloads
abstract
Cardinality estimation (CardEst) is a critical aspect of query optimization. Traditionally, it leverages statistics built directly over the data. However, organizational policies (e.g., regulatory compliance) may restrict global data access. Fortunately, query-driven cardinality estimation can learn CardEst models using query workloads. However, existing query-driven models often require access to data or summaries for best performance, and they assume perfect training workloads with complete and balanced join templates (or join graphs). Such assumptions rarely hold in real-world scenarios, in which join templates are incomplete and imbalanced. We present GRASP, a data-agnostic cardinality learning system designed to work under these real-world constraints. GRASP's compositional design generalizes to unseen join templates and is robust to join template imbalance. It also introduces a new pertable CardEst model that handles value distribution shifts for range predicates, and a novel learned count sketch model that captures join correlations across base relations. Across three database instances, we demonstrate that GRASP consistently outperforms existing query-driven models on imperfect workloads, both in terms of estimation accuracy and query latency. Remarkably, GRASP achieves performance comparable to, or even surpassing, traditional approaches built over the underlying data on the complex CEB-IMDb-full benchmark — despite operating without any data access and using only 10% of all possible join templates.
Peizhi Wu, Rong Kang, Tieying Zhang, Jianjun Chen 0001, Ryan Marcus, Zachary G. Ives
Proc. VLDB Endow.2
2024 AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse Estimators
abstract
Estimating the Number of Distinct Values (NDV) is fundamental for numerous data management tasks, especially within database applications. However, most existing works primarily focus on introducing new statistical or learned estimators, while identifying the most suitable estimator for a given scenario remains largely unexplored. Therefore, we propose AdaNDV, a learned method designed to adaptively select and fuse existing estimators to address this issue. Specifically, (1) we propose to use learned models to distinguish between overestimated and underestimated estimators and then select appropriate estimators from each category. This strategy provides a complementary perspective by integrating overestimations and underestimations for error correction, thereby improving the accuracy of NDV estimation. (2) To further integrate the estimation results, we introduce a novel fusion approach that employs a learned model to predict the weights of the selected estimators and then applies a weighted sum to merge them. By combining these strategies, the proposed AdaNDV fundamentally distinguishes itself from previous works that directly estimate NDV. Moreover, extensive experiments conducted on real-world datasets, with the number of individual columns being several orders of magnitude larger than in previous studies, demonstrate the superior performance of our method.
Xianghong Xu 0001, Tieying Zhang, Xiao He 0008, Haoyang Li 0015, Rong Kang, Wang Shuai, Linhui Xu, Zhimin Liang, Shangyu Luo, Lei Zhang 0213, Jianjun Chen 0001
Proc. VLDB Endow.5
2022 G-VIDO: A Vehicle Dynamics and Intermittent GNSS-Aided Visual-Inertial State Estimator for Autonomous Driving
abstract
This paper proposes G-VIDO, a vehicle dynamics, and intermittent Global Navigation Satellite System (GNSS)-aided visual-inertial state estimator, to address the state estimation problem of autonomous vehicle localization (i.e., position and orientation estimation in the global coordinate system) under various GNSS states. A dynamics pre-integration theory is proposed on the basis of a two-degree-of-freedom (DOF) vehicle dynamics model, and dynamics constraints are built in the optimization back-end, considering the unobservable problem of the monocular visual-inertial system under degenerate motions. The proposed highly nonlinear system can be robustly initialized by loosely aligning the monocular structure from motion (SfM) results, pre-integrated IMU measurements, and vehicle motion information. GNSS is used for reference frame transformation and constraint construction in the sliding window. The cumulative error can be corrected with the aid of GNSS, and the vehicle’s position in the global coordinate system can be determined. A GNSS anomaly detection algorithm is proposed to improve the system robustness under intermittent GNSS. Experiments have shown that G-VIDO can provide real-time, robust, and seamless localization in multiple GNSS states, with an RMSE of less than 30 cm (with GNSS). Moreover, we proved that the initialization and local odometry modules in G-VIDO outperform several state-of-the-art VIO systems and our preliminary work VINS-Vehicle.
Lu Xiong 0001, Rong Kang, Junqiao Zhao, Peizhi Zhang, Ran Ju, Chen Ye 0002, Tiantian Feng
IEEE Trans. Intell. Transp. Syst.2
2020 VINS-PL-Vehicle: Points and Lines-based Monocular VINS Combined with Vehicle Kinematics for Indoor Garage
abstract
In this paper, we propose VINS-PL-Vehicle, a points and lines-based monocular visual-inertial navigation system(VINS) combined with vehicle kinematics for indoor garage. The indoor garage contains some texture-less regions. Thus, it is difficult to ensure the robustness of VINS only using point features. Therefore, in addition to point features, we also add line features on the pillars, parking slots, and top structures in the garage. Besides, we utilize the accurately known velocity and steering wheel angle from the vehicle chassis to form the kinematic constraints between image frames, which can not only solve the problem that the scale is not observable due to insufficient excitation of accelerometer when the vehicle is running at a constant speed, but also utilize its drift-free characteristics to improve the accuracy of system initialization and optimization. The vehicle tests show that our algorithm can achieve significantly higher positioning accuracy than VINS-Mono in the indoor garage.
Peizhi Zhang, Lu Xiong 0001, Zhuoping Yu, Rong Kang, Dequan Zeng
IV4
2020 Apache IoTDB: Time-series database for Internet of Things
abstract
The amount of time-series data that is generated has exploded due to the growing popularity of Internet of Things (IoT) devices and applications. These applications require efficient management of the time-series data on both the edge and cloud side that support high throughput ingestion, low latency query and advanced time series analysis. In this demonstration, we present Apache IoTDB managing time-series data to enable new classes of IoT applications. IoTDB has both edge and cloud versions, provides an optimized columnar file format for efficient time-series data storage, and time-series database with high ingestion rate, low latency queries and data analysis support. It is specially optimized for time-series oriented operations like aggregations query, down-sampling and sub-sequence similarity search. An edge-to-cloud time-series data management application is chosen to demonstrate how IoTDB handles time-series data in real-time and supports advanced analytics by integrating with Hadoop and Spark. An end-to-end IoT data management solution is shown by integrating IoTDB with PLC4x, Calcite, and Grafana.
Chen Wang 0018, Xiangdong Huang 0001, Jialin Qiao, Lei Rui, Rong Kang, Julian Feinauer, Kevin Mcgrail, Peng Wang 0027, Diaohan Luo, Jianmin Wang 0001, Jia-Guang Sun 0001
Proc. VLDB Endow.7
2019 Maximum-Margin Hamming Hashing
abstract
Deep hashing enables computation and memory efficient image search through end-to-end learning of feature representations and binary codes. While linear scan over binary hash codes is more efficient than over the high-dimensional representations, its linear-time complexity is still unacceptable for very large databases. Hamming space retrieval enables constant-time search through hash lookups, where for each query, there is a Hamming ball centered at the query and the data points within the ball are returned as relevant. Since inside the Hamming ball implies retrievable while outside irretrievable, it is crucial to explicitly characterize the Hamming ball. The main idea of this work is to directly embody the Hamming radius into the loss functions, leading to Maximum-Margin Hamming Hashing (MMHH), a new model specifically optimized for Hamming space retrieval. We introduce a max-margin t-distribution loss, where the t-distribution concentrates more similar data points to be within the Hamming ball, and the margin characterizes the Hamming radius such that less penalization is applied to similar data points within the Hamming ball. The loss function also introduces robustness to data noise, where the similarity supervision may be inaccurate in practical problems. The model is trained end-to-end using a new semi-batch optimization algorithm tailored to extremely imbalanced data. Our method yields state-of-the-art results on four datasets and shows superior performance on noisy data.
Rong Kang, Yue Cao 0001, Mingsheng Long, Jianmin Wang 0001, Philip S. Yu
ICCV1
2007 Adaptive Radial Basis Function Detector for Beamforming
abstract
We consider nonlinear detection in rank-deficient multiple-antenna assisted beamforming systems. By exploiting the inherent symmetry of the underlying optimal Bayesian detection solution, a symmetric radial basis function (RBF) detector is proposed and two adaptive algorithms are developed for training the proposed RBF detector. The first adaptive algorithm, referred to as the nonlinear least bit error, is a stochastic approximation to the Parzen window estimation of the detector output's probability density function while the second algorithm is based on a clustering. The proposed adaptive solutions are capable of providing a signal to noise ratio gain in excess of 8 dB against the theoretical linear minimum bit error rate benchmarker, when supporting four users with the aid of two receive antennas or five users employing three antenna elements.
Sheng Chen 0001, Khaled Labib, Rong Kang, Lajos Hanzo
ICC3