VLDB 2026 Research / reviewers in the wild / expert
You-Gan Wang
dblp:70/8020
· DBLP profile ↗
17ranked-venue papers
1as first author
14since 2021 · last 2025
0000-0003-0901-4671ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An adaptive regression algorithm with a clustering process for multi-modal data predictionabstractAbstract Prediction problems regularly exist in practical application problems, where real application systems are usually complex. For multi-model data in complex systems, effectively identifying the patterns of each data set is significant for subsequent prediction. This paper focuses on the prediction problem of multi-model data and proposes an adaptive regression algorithm (CFM-MSVR) that combines a clustering feedback mechanism with improved support vector regression. The clustering feedback mechanism (CFM) clusters samples based on their residuals in each forecasting model, enabling the discovery of original data generation models. Meanwhile, it can intelligently estimate the number of clusters based on the number of samples in each cluster, reducing computational cost and dependence on empirical settings. In the regression stage, the multi-model support vector regression (MSVR) leverages the non-dominated sorting genetic algorithm II (NSGA-II) to optimise the parameters of the support vector regression, thereby improving the generalisation of each sub-model. The proposed method is evaluated on a simulated dataset, four real-world datasets, and the 2012 Global Energy Forecasting Competition dataset. Results show that CFM-MSVR achieves a MAPE% of 1.52 on the energy prediction task, demonstrating its strong performance in complex forecasting scenarios. Shangrui Zhao, Weiqi Yu, Yulu Wu, Jinran Wu, Xi-An Li 0004, You-Gan Wang |
Discov. Comput. | 6 |
| 2025 | LF-HGRILF: A law-fact heterogeneous graph representation and iterative learning framework for legal judgment prediction
Yinying Kong, You-Gan Wang, Haodong Deng, Zhanhao Xiao |
Knowl. Based Syst. | 2 |
| 2025 | Multi-Granularity Autoformer for long-term deterministic and probabilistic power load forecastingabstractLong-term power load forecasting is critical for power system planning but is constrained by intricate temporal patterns. Transformer-based models emphasize modeling long- and short-term dependencies yet encounter limitations from complexity and parameter overhead. This paper introduces a novel Multi-Granularity Autoformer (MG-Autoformer) for long-term load forecasting. The model leverages a Multi-Granularity Auto-Correlation Attention Mechanism (MG-ACAM) to effectively capture fine-grained and coarse-grained temporal dependencies, enabling accurate modeling of short-term fluctuations and long-term trends. To enhance efficiency, a shared query-key (Q-K) mechanism is utilized to identify key temporal patterns across multiple resolutions and reduce model complexity. To address uncertainty in power load forecasting, the model incorporates a quantile loss function, enabling probabilistic predictions while quantifying uncertainty. Extensive experiments on benchmark datasets from Portugal, Australia, America, and ISO New England demonstrate the superior performance of the proposed MG-Autoformer in long-term power load point and probabilistic forecasting tasks. Yang Yang 0052, Yuchao Gao, Jinran Wu, Shangce Gao, You-Gan Wang |
Neural Networks | 6 |
| 2023 | Accelerated Genetic Algorithm with Population Control for Energy-Aware Virtual Machine Placement in Data Centers
Yu-Chu Tian, Maolin Tang, You-Gan Wang, Jiong Jin, Weizhe Zhang |
ICONIP (2) | 4 |
| 2023 | Inferring circadian gene regulatory relationships from gene expression data with a hybrid frameworkabstractBACKGROUND: The central biological clock governs numerous facets of mammalian physiology, including sleep, metabolism, and immune system regulation. Understanding gene regulatory relationships is crucial for unravelling the mechanisms that underlie various cellular biological processes. While it is possible to infer circadian gene regulatory relationships from time-series gene expression data, relying solely on correlation-based inference may not provide sufficient information about causation. Moreover, gene expression data often have high dimensions but a limited number of observations, posing challenges in their analysis. METHODS: In this paper, we introduce a new hybrid framework, referred to as Circadian Gene Regulatory Framework (CGRF), to infer circadian gene regulatory relationships from gene expression data of rats. The framework addresses the challenges of high-dimensional data by combining the fuzzy C-means clustering algorithm with dynamic time warping distance. Through this approach, we efficiently identify the clusters of genes related to the target gene. To determine the significance of genes within a specific cluster, we employ the Wilcoxon signed-rank test. Subsequently, we use a dynamic vector autoregressive method to analyze the selected significant gene expression profiles and reveal directed causal regulatory relationships based on partial correlation. CONCLUSION: The proposed CGRF framework offers a comprehensive and efficient solution for understanding circadian gene regulation. Circadian gene regulatory relationships are inferred from the gene expression data of rats based on the Aanat target gene. The results show that genes Pde10a, Atp7b, Prok2, Per1, Rhobtb3 and Dclk1 stand out, which have been known to be essential for the regulation of circadian activity. The potential relationships between genes Tspan15, Eprs, Eml5 and Fsbp with a circadian rhythm need further experimental research. Shuwen Hu, Yi Jing, You-Gan Wang, Jing Gao 0006, Yu-Chu Tian |
BMC Bioinform. | 4 |
| 2023 | QQLMPA: A quasi-opposition learning and Q-learning based marine predators algorithm
Shangrui Zhao, Yulu Wu, Shuang Tan, Jinran Wu, Zhesen Cui, You-Gan Wang |
Expert Syst. Appl. | 6 |
| 2023 | Mixture extreme learning machine algorithm for robust regressionabstractThe extreme learning machine (ELM) is a well-known approach for training single hidden layer feedforward neural networks (SLFNs) in machine learning. However, ELM is most effective when used for regression on datasets with simple Gaussian distributed error because it often employs a squared loss in its objective function. In contrast, real-world data is often collected from unpredictable and diverse contexts, which may contain complex noise that cannot be characterized by a single distribution. To address this challenge, we propose a robust mixture ELM algorithm, called Mixture-ELM, that enhances modeling capability and resilience to both Gaussian and non-Gaussian noise. The Mixture-ELM algorithm uses an adjusted objective function that blends Gaussian and Laplacian distributions to approximate any continuous distribution and match the noise. The Gaussian mixture accurately models the residual distribution, while the inclusion of the Laplacian distribution addresses the limitations of the Gaussian distribution in identifying outliers. We derive a solution to the novel objective function using the expectation maximization (EM) and iteratively reweighted least squares (IRLS) algorithms. We evaluate the effectiveness of the algorithm through numerical simulation and experiments on benchmark datasets, thereby demonstrating its superiority over other state-of-the-art machine learning methods in terms of robustness and generalization. Shangrui Zhao, Xuan-Ang Chen, Jinran Wu, You-Gan Wang |
Knowl. Based Syst. | 4 |
| 2023 | Accelerated computation of the genetic algorithm for energy-efficient virtual machine placement in data centersabstractAbstract Energy efficiency is a critical issue in the management and operation of cloud data centers, which form the backbone of cloud computing. Virtual machine (VM) placement has a significant impact on energy-efficiency improvement for virtualized data centers. Among various methods to solve the VM-placement problem, the genetic algorithm (GA) has been well accepted for the quality of its solution. However, GA is also computationally demanding, particularly in the computation of its fitness function. This limits its application in large-scale systems or specific scenarios where a fast VM-placement solution of good quality is required. Our analysis in this paper reveals that the execution time of the standard GA is mostly consumed in the computation of its fitness function. Therefore, this paper designs a data structure extended from a previous study to reduce the complexity of the fitness computation from quadratic to linear one with respect to the input size of the VM-placement problem. Incorporating with this data structure, an alternative fitness function is proposed to reduce the number of instructions significantly, further improving the execution-time performance of GA. Experimental studies show that our approach achieves 11 times acceleration of GA computation for energy-efficient VM placement in large-scale data centers with about 1500 physical machines in size. Yu-Chu Tian, You-Gan Wang, Weizhe Zhang |
Neural Comput. Appl. | 3 |
| 2023 | A new algorithm for support vector regression with automatic selection of hyperparameters
You-Gan Wang, Jinran Wu, Zhi-Hua Hu, Geoffrey J. McLachlan |
Pattern Recognit. | 1 |
| 2023 | Robust Adaptive Rescaled Lncosh Neural Network Regression Toward Time-Series ForecastingabstractIn time series forecasting with outliers and random noise, parameter estimation in a neural network via minimizing the$l_{2}$loss is unreliable. Therefore, an adaptive rescaled lncosh loss function is proposed in this article to handle time series modeling with outliers and random noise. It overcomes the limitation of the single distribution of traditional loss functions and can switch among$l_{1}$,$l_{2}$, and the Huber losses. A tuning parameter in the loss function is estimated by using a “working” likelihood approach according to estimated residuals. From the proposed loss function, a robust adaptive rescaled lncosh neural network (RARLNN) regression model is developed for highly accurate predictions. In the training phase of the model, an iterative learning procedure is presented to estimate the tuning parameter and train the neural network in iterations. A new prediction interval construction method is also developed based on quantile theory. The proposed RARLNN model is applied to two groups of wind speed forecasting tasks. The results show that the proposed RARLNN model is more conducive to enhancing forecasting accuracy and stability from the perspectives of noise distribution and outliers. Yang Yang 0052, Jinran Wu, Yu-Chu Tian, Dong Yue 0001, You-Gan Wang |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2022 | An opposition learning and spiral modelling based arithmetic optimization algorithm for global continuous optimization problems
Yang Yang 0052, Yuchao Gao, Shuang Tan, Shangrui Zhao, Jinran Wu, Shangce Gao, Tengfei Zhang 0001, Yu-Chu Tian, You-Gan Wang |
Eng. Appl. Artif. Intell. | 9 |
| 2022 | Robust penalized extreme learning machine regression with applications in wind speed forecasting
Yang Yang 0052, Yuchao Gao, Jinran Wu, You-Gan Wang, Liya Fu |
Neural Comput. Appl. | 5 |
| 2021 | mUSP: a high-accuracy map of the in situ crosstalk of ubiquitylation and SUMOylation proteome predicted via the feature enhancement approachabstractReversible post-translational modification (PTM) orchestrates various biological processes by changing the properties of proteins. Since many proteins are multiply modified by PTMs, identification of PTM crosstalk site has emerged to be an intriguing topic and attracted much attention. In this study, we systematically deciphered the in situ crosstalk of ubiquitylation and SUMOylation that co-occurs on the same lysine residue. We first collected 3363 ubiquitylation-SUMOylation (UBS) crosstalk site on 1302 proteins and then investigated the prime sequence motifs, the local evolutionary degree and the distribution of structural annotations at the residue and sequence levels between the UBS crosstalk and the single modification sites. Given the properties of UBS crosstalk sites, we thus developed the mUSP classifier to predict UBS crosstalk site by integrating different types of features with two-step feature optimization by recursive feature elimination approach. By using various cross-validations, the mUSP model achieved an average area under the curve (AUC) value of 0.8416, indicating its promising accuracy and robustness. By comparison, the mUSP has significantly better performance with the improvement of 38.41 and 51.48% AUC values compared to the cross-results by the previous single predictor. The mUSP was implemented as a web server available at http://bioinfo.ncu.edu.cn/mUSP/index.html to facilitate the query of our high-accuracy UBS crosstalk results for experimental design and validation. Hao-Dong Xu, Ru-Ping Liang, You-Gan Wang, Jian-Ding Qiu |
Briefings Bioinform. | 3 |
| 2021 | A temporal LASSO regression model for the emergency forecasting of the suspended sediment concentrations in coastal oceans: Accuracy and interpretability
Shaotong Zhang, Jinran Wu, Yonggang Jia, You-Gan Wang, Qibin Duan |
Eng. Appl. Artif. Intell. | 4 |
| 2020 | An improved firefly algorithm for global continuous optimization problems
Jinran Wu, You-Gan Wang, Kevin Burrage, Yu-Chu Tian, Brodie Lawson |
Expert Syst. Appl. | 2 |
| 2020 | Exact algorithms for energy-efficient virtual machine placement in data centers
Chen Wei 0003, Zhi-Hua Hu, You-Gan Wang |
Future Gener. Comput. Syst. | 3 |
| 2019 | Significance tests for analyzing gene expression data with small sample sizesabstractMOTIVATION: Under two biologically different conditions, we are often interested in identifying differentially expressed genes. It is usually the case that the assumption of equal variances on the two groups is violated for many genes where a large number of them are required to be filtered or ranked. In these cases, exact tests are unavailable and the Welch's approximate test is most reliable one. The Welch's test involves two layers of approximations: approximating the distribution of the statistic by a t-distribution, which in turn depends on approximate degrees of freedom. This study attempts to improve upon Welch's approximate test by avoiding one layer of approximation. RESULTS: We introduce a new distribution that generalizes the t-distribution and propose a Monte Carlo based test that uses only one layer of approximation for statistical inferences. Experimental results based on extensive simulation studies show that the Monte Carol based tests enhance the statistical power and performs better than Welch's t-approximation, especially when the equal variance assumption is not met and the sample size of the sample with a larger variance is smaller. We analyzed two gene-expression datasets, namely the childhood acute lymphoblastic leukemia gene-expression dataset with 22 283 genes and Golden Spike dataset produced by a controlled experiment with 13 966 genes. The new test identified additional genes of interest in both datasets. Some of these genes have been proven to play important roles in medical literature. AVAILABILITY AND IMPLEMENTATION: R scripts and the R package mcBFtest is available in CRAN and to reproduce all reported results are available at the GitHub repository, https://github.com/iullah1980/MCTcodes. SUPPLEMENTARY INFORMATION: Supplementary data is available at Bioinformatics online. Insha Ullah, Sudhir Paul, Zhenjie Hong, You-Gan Wang |
Bioinform. | 4 |