VLDB 2026 Research / reviewers in the wild / expert
Xingyan Li
dblp:59/1357
· DBLP profile ↗
10ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 3 since 2021Systems, architecture and hardware · 5 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EAST: An Interpretable Knob Estimation System for Cloud DatabaseabstractDatabase vendors have made significant redesigns and developments to the relational database for providing cloud-hosted and cloud-native database services. Thus, the original knob-tuning experiences of DBAs are no longer applicable to the cloud database era. An interpretable estimation service is urgently needed to provide explicit guidance for database knob tuning. Unfortunately, less attention has been paid to estimating the performance of the knob configuration. To fill this gap, we propose EAST, a knob estimation system to provide interpretable & transferable knob estimation services for cloud databases. Firstly, we design an interpretable knob-embedding-based estimator to achieve the trusted and white-box knob estimation for researchers, practitioners, and even artificial intelligence knob tuners. Secondly, we design a two-stage transfer estimation approach by stacking ensemble learning to utilize historical experiences, improving time efficiency. Thirdly, EAST provides a user-friendly interface to support direct knob estimation and transfer knob estimation services. We have deployed our EAST11https://gitee.com/opengauss/openGauss-DBMind/tree/incubator/dbmind/components/knob_estimator to the DBMind component of OpenGauss and demonstrated the effectiveness of our system under the open-source benchmark TPCC. Hongzhi Wang 0001, Jian Geng, Zixuan Wang 0032, Xingyan Li |
ICDE | 5 |
| 2025 | Graph Invariant Learning with Feature and Structure Denoising for Out-of-Distribution DataabstractOut-of-distribution (OOD) data presents significant challenges to traditional Graph Neural Networks (GNNs). The latest research common to creates new out-of-distribution samples through invariant and variant data combinations to enhance the model’s generalization ability. However, these researches mainly focus on distinguishing the variant(non-critical) features from the invariant (critical) features in the new sample training while paying less attention to reinforcing the invariant features. This insufficient emphasis on invariant features can adversely affect the model’s accuracy in out-of-distribution scenarios. To this end, we proposes a graph invariant learning method both creating the out-of-distribution train environment and focusing on learning graph invariant features. Specifically, we use feature level and graph level patterns comixup strategy, which can simulate various complex mixed environments and capture invariant patterns from mixed graph data. In details, first, we obtain representations of variant subgraphs and invariant sub-graphs through a representation network, and then perturb the invariant subgraphs by mixing up the variant subgraphs in the representation space to simulate complex environments in reality. Afterward, we integrate invariant subgraphs and variable subgraphs at the graph level, and utilize the discrepancy in mutual information between the graph-level and representation-level to train the model, thereby encouraging the model to capture invariant patterns underlying the graphs. Extensive experiments on eight real-world datasets demonstrate that our method significantly out-performs state-of-the-art under various distribution shifts. Xingyan Li, Qingyao Xu |
INDIN | 1 |
| 2024 | Comparative Evaluation of Causal Discovery and Inference Approaches on Arctic Sea Ice Time Series DataabstractSea ice extent plays a crucial role in the Arctic system, and thus the study on causal relationships between sea ice extent and other climate variables comes to our interest to better understand the system. To find the causal relationship we applied various state-of-the-art causal discovery techniques from the time-independent and time-dependent domains. Then we employed several causal inference models to quantify the causal effects of different causal relationships in the Arctic system. The NSIDC Sea Ice Concentration observation data and the ERA-5 global reanalysis data were used in our study. Our analysis shows that the GES and VarLiNGAM from causal discovery methods and the conditional instrumental variable (CIV) causal inference model perform better on the Arctic Sea Ice time series dataset. Omar Faruque, Xingyan Li, Md. Azim Khan, Homayra Alam, Jianwu Wang 0001 |
IEEE Big Data | 2 |
| 2024 | MT-HCCAR: Multi-task Deep Learning with Hierarchical Classification and Attention-Based Regression for Cloud Property Retrieval
Xingyan Li, Andrew M. Sayer, Ian T. Carroll, Xin Huang 0005, Jianwu Wang 0001 |
ECML/PKDD (10) | 1 |
| 2023 | Reproducible and Portable Big Data Analytics in the CloudabstractCloud computing has become a major approach to help reproduce computational experiments. Yet there are still two main difficulties in reproducing batch based Big Data analytics (including descriptive and predictive analytics) in the cloud. The first is how to automate end-to-end scalable execution of analytics including distributed environment provisioning, analytics pipeline description, parallel execution, and resource termination. The second is that an application developed for one cloud is difficult to be reproduced in another cloud, a.k.a. vendor lock-in problem. To tackle these problems, we leverage serverless computing and containerization techniques for automated scalable execution and reproducibility, and utilize the adapter design pattern to enable application portability and reproducibility across different clouds. We propose and develop an open-source toolkit that supports 1) fully automated end-to-end execution and reproduction via a single command, 2) automated data and configuration storage for each execution, 3) flexible client modes based on user preferences, 4) execution history query, and 5) simple reproduction of existing executions in the same environment or a different environment. We did extensive experiments on both AWS and Azure using four Big Data analytics applications that run on virtual CPU/GPU clusters. The experiments show our toolkit can achieve good execution performance, scalability, and efficient reproducibility for cloud-based Big Data analytics. Xin Wang 0122, Pei Guo, Xingyan Li, Aryya Gangopadhyay, Carl E. Busart, Jade Freeman, Jianwu Wang 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2022 | Enhanced Deep Learning Super-Resolution for Bathymetry DataabstractSpatial resolution is critical for observing and monitoring environmental phenomena. Acquiring high-resolution bathymetry data directly from satellites is not always feasible due to limitations on equipment, so spatial data scientists and researchers turn to single image super-resolution (SISR) methods that utilize deep learning techniques as an alternative method to increase pixel density. While super resolution residual networks (e.g., SR-ResNet) are promising for this purpose, several challenges still need to be addressed: (1) Earth data such as bathymetry is expensive to obtain and relatively limited in its data record amount; (2) certain domain knowledge needs to be complied with during model training; (3) certain areas of interest require more accurate measurements than other areas. To address these challenges, following the transfer learning principle, we study how to leverage an existing pre-trained super-resolution deep learning model, namely SR-ResNet, for high-resolution bathymetry data generation. We further enhance the SR-ResNet model to add corresponding loss functions based on domain knowledge. To let the model perform better for certain spatial areas, we add additional loss functions to increase the penalty of the areas of interest. Our experiments show our approaches achieve higher accuracy than most baseline models when evaluating using metrics including MSE, PSNR, and SSIM. Xingyan Li, Zachary Williams, Xin Huang 0005, Mark Carroll, Jianwu Wang 0001 |
BDCAT | 1 |
| 2022 | Benchmarking Probabilistic Machine Learning Models for Arctic Sea Ice ForecastingabstractThe Arctic is a region with unique climate features, motivating new AI methodologies to study it. Unfortunately, Arctic sea ice has seen a continuous decline since 1979. This not only poses a significant threat to Arctic wildlife and surrounding coastal communities but is also adversely affecting the global climate patterns. To study the potential of AI in tackling climate change, we analyze the performance of four probabilistic machine learning methods in forecasting sea-ice extent for lead times of up to 6 months, further comparing them with traditional machine learning methods. Our comparative analysis shows that Gaussian Process Regression is a good fit to predict sea-ice extent for longer lead times with lowest RMSE score. Sahara Ali, Seraj Al Mahmud Mostafa, Xingyan Li, Sara Khanjani, Jianwu Wang 0001, James R. Foulds, Vandana Pursnani Janeja |
IGARSS | 3 |
| 2009 | Distributed sensor analysis for fault detection in tightly-coupled multi-robot team tasksabstractThis paper presents a distributed version of our previous work, called SAFDetection, which is a sensor analysis-based fault detection approach that is used to monitor tightly-coupled multi-robot team tasks.While the centralized version of SAFDetection was shown to be successful, a shortcoming of the approach is that it does not scale well to large team sizes. The distributed SAFDetection approach addresses this problem by adapting and distributing the approach across team members. Distributed SAFDetection has the same theoretic foundation as centralized SAFDetection, which maps selected robot sensor data to a robot state by using a clustering algorithm, and builds state transition diagrams to describe the normal behavior of the robot system. However, rather than processing multiple robots' sensor data centralized on a server, distributed SAFDetection performs feature selection and clustering on individual robots to build the normal behavior model of an individual robot and the entire robot team. Fault detection is also accomplished in a distributed manner. We have implemented this distributed approach on a physical robot team and in simulation. This paper presents the results of these experiments, showing that distributed SAFDetection is an efficient approach to detect both local and interactive faults in tightly-coupled multi-robot team tasks. Compared to the centralized version, this approach provides more scalability and reliability. Xingyan Li, Lynne E. Parker |
ICRA | 1 |
| 2009 | Expressive facial speech synthesis on a robotic platformabstractThis paper presents our expressive facial speech synthesis system Eface, for a social or service robot. Eface aims at enabling a robot to deliver information clearly with empathetic speech and an expressive virtual face. The empathetic speech is built on the Festival speech synthesis system and provides robots the capability to speak with different voices and emotions. Two versions of a virtual face have been implemented to display the robot's expressions. One with just over 100 polygons has a lower hardware requirement but looks less natural. The other has over 1000 polygons; it looks realistic, but costs more CPU resource and requires better video hardware. The whole system is incorporated into the popular open source robot interface Player, which makes client programs easy to write and debug. Also, it is convenient to use the same system with different robot platforms. We have implemented this system on a physical robot and tested it with a robotic nurse assistant scenario. Xingyan Li, Bruce A. MacDonald, Catherine I. Watson |
IROS | 1 |
| 2007 | Sensor Analysis for Fault Detection in Tightly-Coupled Multi-Robot Team TasksabstractThis paper presents a sensor analysis based fault detection approach (which we call SAFDetection) that is used to monitor tightly-coupled multi-robot team tasks. Our approach aims at detecting both physical and logic faults of a robot system with little prior knowledge on the system. We do not need the motion model or a priori knowledge of the possible fault types of the monitored system. Our approach treats the monitored robot system as a black box, with only sensor data available. Thus, we believe the approach is general, and can be used in a wide variety of robot systems performing many different kinds of tasks. Our approach combines data clustering techniques with the generation of a probabilistic state diagram to model the normal operation of the multi-robot system. We have implemented this approach on a physical robot team. This paper presents the results of these experiments, which show that sensor data analyzed from a training phase of normal operation can be used to generate a model of normal robot team operation. This model can then be used to detect many types of abnormal behavior of the system, based purely on monitoring the sensor data of the system. Xingyan Li, Lynne E. Parker |
ICRA | 1 |