VLDB 2026 Research / reviewers in the wild / expert
Shengkun Xie
dblp:79/4553
· DBLP profile ↗
22ranked-venue papers
17as first author
11since 2021 · last 2026
0000-0002-9533-2096ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 12 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatially Constrained Clustering for Analyzing Spatio-Temporal Dynamics of Automobile Insurance Risk
Shengkun Xie, Samy El Mouttalibi, Jin Zhang 0026, Clare Chua-Chow |
DATA (1) | 1 |
| 2026 | ASTAC: An adaptive spatio-temporal autocorrelation clustering for complex traffic flow pattern discovery
Jiyan Wang, Xiaojun Ban, Shengkun Xie |
Expert Syst. Appl. | 3 |
| 2025 | A Comparative Study of Non-Linear Modelling Capabilities of T-S Fuzzy Models and Back-Propagation Neural NetworksabstractIn the era of data-driven decision-making, selecting appropriate nonlinear modeling techniques is critical for building robust and interpretable predictive systems. While both Takagi-Sugeno (T-S) fuzzy models and Back-Propagation (BP) neural networks are well-established universal approximators in the data science domain, their fundamentally different structural characteristics lead to varied performance across application scenarios. This paper presents a systematic comparative study of these two modeling approaches through a series of simulation experiments designed to reflect key tasks in predictive analytics, including static function approximation, dynamic system modeling and forecasting, and real-time state tracking of time-varying systems. By evaluating performance across multiple dimensions modeling accuracy, robustness to noise, and adaptability to temporal dynamics, this work provides actionable insights into the practical strengths and limitations of each model type. The results show that T-S fuzzy models offer superior accuracy in clean, stable environments, while BP neural networks demonstrate strong resilience to noise and generalization ability in uncertain conditions. Additionally, T-S fuzzy models exhibit higher adaptability in real-time, dynamic contexts, making them valuable for time-sensitive predictive applications. This study contributes to the broader data science community by offering a structured framework for model selection based on application-specific demands, helping practitioners and researchers alike to optimize predictive modeling strategies for complex, nonlinear systems. Linxiang Li, Kainan Liu, Xiaojun Ban, Shengkun Xie |
DSAA | 4 |
| 2024 | Spatial and Spatio-Temporal Modelling of Auto Insurance Claim Frequencies During Pre-and Post-COVID-19 Pandemic
Jin Zhang 0026, Shengkun Xie, Anna T. Lawniczak, Clare Chua-Chow |
DATA | 2 |
| 2023 | Exploring Functional Patterns of Driving Records by Interacting with Major Classes and Territory Using Generalized Additive Models
Shengkun Xie, Anna T. Lawniczak, Clare Chua-Chow |
DATA | 1 |
| 2023 | Modelling auto insurance Size-of-Loss distributions using Exponentiated Weibull distribution and de-grouping methodsabstractRate regulation plays an important role in the financial service of the auto insurance industry. Modelling the Size of Loss distributions, particularly the large loss distribution at the aggregate level, is a key component of rate regulation to produce a benchmark for the industry. However, traditionally actuarial practice is constructing the group-based Size of Loss with irregular intervals. The frequency values associated with each group are counted to form a frequency distribution. However, this approach is limited by lacking a parametric distribution that can capture the underlying randomness of the Size of Loss. In this work, we propose a de-grouping method based on the simulation of the uniform random variate, with and without constraint on knowing the average incurred loss per claim count. Using the de-grouping method, we simulate the individual observation based on the limited information available from the grouped Size of Loss. We have successfully identified the Exponentiated Weibull distribution as a good candidate as it outperforms other heavy-tailed distributions being considered. Our findings provide critical guidance and direction for the practical use of the proposed de-grouping method in auto insurance rate regulation practice and other fields in business and economics. Shengkun Xie |
Expert Syst. Appl. | 1 |
| 2023 | TOPSIS-based comprehensive measure of variable importance in predictive modelling
Shengkun Xie, Jin Zhang 0026 |
Expert Syst. Appl. | 1 |
| 2022 | Fuzzy Clustering and Non-negative Sparse Matrix Approximation on Estimating Territory Risk RelativitiesabstractThis work proposes using a fuzzy clustering approach to design territories for auto insurance rate-making and regulation in rate filing review. Our approach moves the focus from the current hard clustering method to a soft approach to improve the evaluation of territory risk for the rate-making purpose. Furthermore, we use non-negative sparse principal component analysis to smooth the estimate of risk relativities of basic rating units, which are Forward Sortation Areas (FSA) used in Canada. This additional step helps achieve the double sparsity of the fuzzy membership matrix by removing the small membership values and minor principal components to achieve a smoothing effect on the risk relativity estimate of FSA. Combining fuzzy clustering with non-negative sparse principal component analysis is particularly novel as it enables better decision-making of auto insurance rate regulation, where high-level statistics and major data patterns are desired. Shengkun Xie, Chong Gan |
FUZZ-IEEE | 1 |
| 2022 | A Novel Variable Selection Approach Based on Multi-criteria Decision Analysis
Shengkun Xie, Jin Zhang 0026 |
IPMU (2) | 1 |
| 2022 | Feature extraction of auto insurance size of loss data using functional principal component analysis
Shengkun Xie |
Expert Syst. Appl. | 1 |
| 2021 | Estimating Territory Risk Relativity for Auto Insurance Rate Regulation using Generalized Linear Mixed Models
Shengkun Xie, Chong Gan, Clare Chua-Chow |
DATA | 1 |
| 2020 | Improving Statistical Reporting Data Explainability via Principal Component Analysis
Shengkun Xie, Clare Chua-Chow |
DATA | 1 |
| 2019 | Feature Extraction of Epileptic EEG in Spectral Domain via Functional Data Analysis
Shengkun Xie, Anna T. Lawniczak |
ICPRAM | 1 |
| 2017 | Spatially Constrained Clustering to Define Geographical Rating Territories
Shengkun Xie, Anna T. Lawniczak |
ICPRAM | 1 |
| 2016 | Effects of model parameter interactions on Naïve creatures' success of learning to cross a highwayabstractWe investigate a learning strategy for a swarm of autonomous robots. We identify the robots with cognitive agents and describe a model of naïve creatures learning to cross a highway. The creatures use a type of “observational social learning”, in which each creature learns from observing the outcomes of the other creatures that have crossed the highway. The learning outcomes are influenced by many of the simulation model's independent variables/parameters, which are the creatures' emotional states of fear and/or desire, their crossing point locations, the creatures ability or not to change a crossing point, the highway traffic conditions characterized by cars density, the duration of the creatures' learning process within the same environment (i.e., under the same highway traffic conditions characterized by cars density) and the transfer of the knowledge base acquired in one learning environment to another one (i.e., from an environment with one car traffic density to another one). We study how these factors, in particular their interactions, affect the model performance measured by the number of successful, killed and queued creatures and the creatures' ability to learn. Anna T. Lawniczak, Leslie Ly, Fei Yu 0007, Shengkun Xie |
CEC | 4 |
| 2015 | Learning Dictionary Via Wavelet Sparse Principal Component Analysis
Shengkun Xie, Anna T. Lawniczak, Sridhar Krishnan 0001 |
ICPRAM (1) | 1 |
| 2013 | Noise Effects on Spatial Pattern Data Classification Using Wavelet Kernel PCA - A Monte Carlo Simulation Study
Shengkun Xie, Anna T. Lawniczak, Sridhar Krishnan 0001 |
ISNN (1) | 1 |
| 2012 | Sparse principal component extraction and classification of long-term biomedical signalsabstractThis article focuses on finding a solution of sparse representation for signal classification in long-term observational studies. An approach that involves sparse principal component analysis (SPCA) is proposed. This method first uses a non-overlapping moving window for signal segmentation and makes use of SPCA to select a limited number of signal segments for constructing sparse principal components. A set of supervised predictive models based on sparse principal components of training signal segments is then constructed for signal approximation. Within this approach, their model residuals are estimated and used for signal classification. A nearly perfect classification accuracy is obtained for both the synthetic data and EEG signals that we considered. This highly positive result suggests that the proposed method may be useful for automatic event detection in long-term observational signals. Shengkun Xie, Sridhar Krishnan 0001, Anna T. Lawniczak |
CBMS | 1 |
| 2012 | Analysis of communication network surveillance using functional ANOVA model with unequal variancesabstractAnalysis of simulation result plays an important role in helping decision-making when simulation modeling is used as an approach to investigate complex systems. Analysis of variance (ANOVA) is a popular statistical technique for analyzing simulation results when various designs of simulation experiments are implemented and investigated. However, conventional ANOVA technique is based on the assumption of homogeneous variances of output data, which may not be realistic for data coming from the real-world complex systems. The existence of heterogeneous structures of data coming from complex systems requires a more reliable analysis method, in order to analyze such data. In this paper, a functional ANOVA model with unequal variances is proposed for meeting this goal. The proposed model aims to better capture the heterogeneous data structure that is caused by various experimental simulation setups. The applicability of the proposed method is illustrated by using simulated communication network traffic data. Our work contributes to the development of new approaches for analysis of simulation and real-world complex systems data with heterogeneous structures. The proposed method can be useful for defense problems, e.g. for analysis of communication network surveillance or for the purpose of improving the operational efficiency. Shengkun Xie, Sridhar Krishnan 0001, Anna T. Lawniczak |
CISDA | 1 |
| 2012 | Time-Frequency Analysis via Ramanujan SumsabstractResearch in signal processing shows that a variety of transforms have been introduced to map the data from the original space into the feature space, in order to efficiently analyze a signal. These techniques differ in their basis functions, that is used for projecting the signal into a higher dimensional space. One of the widely used schemes for quasi-stationary and non-stationary signals is the time-frequency (TF) transforms, characterized by specific kernel functions. This work introduces a novel class of Ramanujan Fourier Transform (RFT) based TF transform functions, constituted by Ramanujan sums (RS) basis. The proposed special class of transforms offer high immunity to noise interference, since the computation is carried out only on co-resonant components, during analysis of signals. Further, we also provide a 2-D formulation of the RFT function. Experimental validation using synthetic examples, indicates that this technique shows potential for obtaining relatively sparse TF-equivalent representation and can be optimized for characterization of certain real-life signals. Lakshmi Sugavaneswaran, Shengkun Xie, Karthikeyan Umapathy, Sridhar Krishnan 0001 |
IEEE Signal Process. Lett. | 2 |
| 2011 | Detection of stationary network load increase using univariate network aggregate traffic data by dynamic PCAabstractNetwork operators are now facing bandwidth outages as well as a growing pressure to ensure good Quality of Service (QoS). An important practical issue for network service providers is to pay close attention to the load changes of network traffic, in particular, the stationary increase of load from a normal demand. Many network monitoring applications and performance analysis tools are based on the study of an aggregate measure of network traffic, e.g. number of packets in transit (NPT), which is a long-term univariate time series. To classify this type of network traffic data and detect any increase of network source load, we propose a dynamic principal component analysis (PCA) method, first to extract data features and then to detect a stationary load increase of network traffic. The proposed detection schemes are based on either the major or the minor principal components of network traffic data. To demonstrate the applications of the proposed feature extraction method and the detection schemes, we applied them to network traffic data simulated from the packet switching network (PSN) model. Additionally, we propose a combined detection scheme that uses both the major and the minor principal components. The proposed detection schemes, based on dynamic PCA, show enhanced performance in detecting an increase of network load for the simulated network traffic data. These results offer a new feature extraction method based on dynamic PCA that creates additional feature variables for event detection in a univariate time series. Shengkun Xie, Anna T. Lawniczak |
CISDA | 1 |
| 2009 | A Case Study of ICA with Multi-scale PCA of Simulated Traffic Data
Shengkun Xie, Pietro Liò, Anna T. Lawniczak |
ICANN (2) | 1 |