Wanghu Chen

dblp:80/3115 · DBLP profile ↗
← Back
14ranked-venue papers in the field
12as first author
10since 2021 · last 2024
0000-0002-9233-7609ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 13 (12 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2024 Clustering of Students' Behavior Using the GPT-2 Module Based on the Macroscopic Attention Model
abstract
A reasonable group classification of students is beneficial for the process management of university students on campus. To this end, a Macroscopic Attention (MA) model is developed to characterize individual students. Existing clustering methods find it difficult to take into account the temporal characteristics of the data and the relationship between different dimensions. Consequently a time series clustering method, GPT-K, is proposed to optimize the clustering effect by capturing the complex temporal dependencies and interactions between the MA quality features in the time series data through the GPT-2 module. Experiments indicate that the method has a great advantage over other clustering methods like AE-K and CNN-K in terms of Silhouette Coefficient, Davies-Bouldin Index and Calinski-Harabaz Index. We analysis the results of student group classification, it is found that the probability of students in one group to receive scholarships is significantly higher than that of students in another group. This means that there is a significant gap in academic performance between the two groups. Based on the research in this paper, early intervention and guidance can be provided to students in a timely manner.
Wanghu Chen, Siqi Zeng 0002, Jing Li 0131
IEEE Big Data1
2024 PV-DETR: A Multimodal Fault Detection Model of PV Arrays based on Parallel Block Attention
abstract
Detection of faults in photovoltaic arrays can reduce power generation losses and extend the equipment’s lifespan. Traditional operation and maintenance of photovoltaic power stations primarily rely on electrical characteristics or infrared images. However, data from a single modality are susceptible to environmental interference, affecting detection accuracy. To address these issues, we propose a model called PV-DETR for fault detection in photovoltaic arrays under complex environmental conditions. This model is an extension of RT-DETRv2, which leverages the Transformer architecture for feature extraction and decoding. The model employs a PResNet50 module instead of the original ResNet50, along with haar wavelet downsampling and a parallel block attention mechanism. The PResNet50 module can reduce dimensionality while minimizing information loss. Haar wavelet downsampling retains the original global information and compresses feature maps effectively, and the parallel block attention mechanism significantly enhances the detection of small infrared targets. Experimental results show that the final PV-DETR model achieves an average accuracy of 89% and an average recall of 85% in fault detection using multimodal data, outperforming existing models, including the original RT-DETRv2.
Wanghu Chen, Yihua Luo, Long Li 0019, Jing Li 0131
IEEE Big Data1
2024 GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian Optimization
abstract
Modern database management systems (DBMS) expose hundreds of configurable knobs to control system behaviours. Determining the appropriate values for these knobs to improve DBMS performance is a long-standing problem in the database community. As there is an increasing number of knobs to tune and each knob could be in continuous or categorical values, manual tuning becomes impractical. Recently, automatic tuning systems using machine learning methods have shown great potentials. However, existing approaches still incur significant tuning costs or only yield sub-optimal performance. This is because they either ignore the extensive domain knowledge available (e.g., DBMS manuals and forum discussions) and only rely on the runtime feedback of benchmark evaluations to guide the optimization, or they utilize the domain knowledge in a limited way. Hence, we propose GPTuner, a manual-reading database tuning system that leverages domain knowledge extensively and automatically to optimize search space and enhance the runtime feedback-based optimization process. Firstly, we develop a Large Language Model (LLM)-based pipeline to collect and refine heterogeneous knowledge, and propose a prompt ensemble algorithm to unify a structured view of the refined knowledge. Secondly, using the structured knowledge, we (1) design a workload-aware and training-free knob selection strategy, (2) develop a search space optimization technique considering the value range of each knob, and (3) propose a Coarse-to-Fine Bayesian Optimization Framework to explore the optimized space. Finally, we evaluate GPTuner under different benchmarks (TPC-C and TPC-H), metrics (throughput and latency) as well as DBMS (PostgreSQL and MySQL). Compared to the state-of-the-art approaches, GPTuner identifies better configurations in 16x less time on average. Moreover, GPTuner achieves up to 30% performance improvement (higher throughput or lower latency) over the best-performing alternative.
Jiale Lao, Yibo Wang 0015, Zhiyuan Cheng 0010, Wanghu Chen, MingJie Tang, Jianguo Wang 0001
Proc. VLDB Endow.7
2023 Enhanced Segmentation of PV Arrays in Infrared Images using an Improved SegFormer Approach
abstract
An infrared image segmentation model for photovoltaic arrays is proposed based on the improved Segformer. The inception-enhanced attention mechanism and multi-scale spatial feature extraction is leveraged to address problems, such as segmentation holes and environmental misclassification. Furthermore, the encoder of the proposed model is improved using the Feature Pyramid Network and bilinear interpolation operation to enhance the completeness of edge details. Experiments on infrared images gathered from a real-world power station shows that the model achieves improvements of 0.48 in mIoU, 0.3 in mAcc, and 1.38 in mDice, compared to existing models besides the original Segformer.
Wanghu Chen, Shengfang Jin, Yihua Luo, Jing Li 0131
IEEE Big Data1
2022 Anomaly detection of high-dimensional data based on Ensemble GANs with Dropout
abstract
An unsupervised anomaly detection approach DGANs is proposed based on ensemble GANs with Dropout. The comparisons with representative approaches on 10 public datasets show it has advantages in accuracy, recall and F1 scores. DGANs can address the overfitting problem in ensemble GANs training on high-dimensional datasets.
Wanghu Chen, Jilong Yao, Meilin Zhou, Jing Li 0131, Mengyang Shen
BDCAT1
2022 Pavement Condition Detection Method Based on Time-Frequency Features and Capsule Neural Network
abstract
Pavement condition detection is beneficial to road maintenance and driving experience. Acceleration sensors of smart phones can provide an economical and ubiquitous way to gather pavement condition data. A pavement condition detection method is proposed based on acceleration sensor data of smart phones, which incorporates time-frequency features into capsule networks. The method well addresses the problems of low accuracy caused by the length change of time series, and the high dimensionality of sensor data. Experiments show that the method proposed outperforms the representative methods in accuracy, precision and F1 scores. Especially, it has an improvement in F1 score up to 31.62% compared with the benchmark methods.
Wanghu Chen, Pengbo Lv, Jing Li 0131
IEEE Big Data1
2022 Psychological Attention-based Analytics of Multivariate Campus Behaviors of University Students
abstract
Psychological Attention (PA) is introduced to characterize multivariate campus behaviors of university students, and a computation model for PA qualities is proposed driven by online behavioral big data. The PA-based behavior clustering is applied into the analytics of multivariate behaviors of university students to reveal the impacts of PA qualities on academic performances. Experiments show PA-based behavior clustering has great advantages in the Silhouette Coefficient (SC), Davies-Bouldin Index (DBI), and Calinski-Harabasz index (CHI) over the clustering based on the traditional features. It means that PA qualities can well distinguish the potential patterns in multivariate behaviors, since the students in one cluster have a higher possibility of 38.10% to get the top-level scholarship than those in another one. Experimental results also represents that the Stability and Distribution of the PA in multivariate behaviors have active impacts on students’ academic performances. The studies in the paper can be applied into the prediction of students’ academic performances, and the personalized in-advance guidances from them.
Wanghu Chen, Jilong Yao, Jing Li 0131, Chunyu Pang
IEEE Big Data1
2022 Outlier Detection Based on Stacked Autoencoder and Gaussian Mixture Model
abstract
The outlier detection of high-dimensional data is still of challenge. The performance of existing unsupervised approaches will be affected with the increase of outliers in a dataset. The stacked autoencoder and GMM are introduced to the detection of outliers, and an approach termed SAGMM is proposed. The stacked autoencoder can reduce the reconstruction error of observations, and the GMM determines the outliers based on the mixture distributions of observations obtained in model training. Experiments on public datasets show that the proposed approach SAGMM outperforms the similar approaches in precision, and has a good balance between the precision and recall rate, since it improves the F1 scores compared with them.
Jing Li 0131, Pengbo Lv, Wanghu Chen
IEEE Big Data4
2021 Spatio-temporal Clustering based on HHT and Its Applications in Thermal Boiler Controlling
abstract
The heating surface temperature controlling of thermal boilers are critical to safe production, energy saving and emission reduction. With the background of temperature prediction of heating surfaces in thermal boilers, the paper proposes a novel time-series clustering approach at first. Considering time series as arbitrary signals, features extracted from their Marginal Spectrums based on Hilbert Huang Transform is introduced to the clustering. From the proposed time-series clustering approach, a Spatio-temporal clustering approach to enabling local heating surface partitioning is derived. The temperature of the heating surfaces partitioned is then predicted using an LSTM model depending on multiple time-series related to the work conditions of a boiler. The proposed time-series clustering is compared with the traditional approaches on public datasets, and shows great advantages, and the temperature prediction of local heating surfaces of thermal boilers in practice also verify that the proposed approaches are effective.
Wanghu Chen, Jing Li 0131, Chenhan Zhai, Pengbo Lv, Shengfang Jin
IEEE BigData1
2021 Anomaly detection of high-dimensional sparse data based on Ensemble Generative Adversarial Networks
abstract
Anomaly detection has drawn public attentions in past decades. However, in a high-dimensional sparse data space, anomaly detection still faces big challenges. In this paper, the Generative Adversarial Network (GAN) combined with Ensemble Learning is introduced to anomaly detection in high-dimensional sparse data. On one hand, the generator of GAN can produce noise data to avoid the data space to be too sparse based on the potential data distribution patterns. On the other hand, the exchanges of the pairing of generators and discriminators can enable the model proposed to learn complex distribution of the data, which may be composed of some various distributions, and to avoid the training process to drop into over-fitting to some extent. Experiments on public datasets show that the proposed approach can improve AUC by 7% compared with traditional GAN based approaches, and by 7.5% to 21.8% compared with other representative anomaly detection approaches.
Wanghu Chen, Meilin Zhou, Chenhan Zhai, Mengyang Shen, Pengbo Lv, Ali Arshad
IEEE BigData1
2020 Autoencoder-based outlier detection for sparse, high dimensional data
abstract
Outlier detection is essential in many data mining tasks. For high-dimensional data, its outlier detection often faces two challenges caused by sparse spatial distribution of data and big difficulties to get enough class labels. Therefore, it is valuable to explore a simpler and more effective approach to unsupervised outlier detection. In this paper, focusing on high-dimensional sparse data, an unsupervised outlier detection approach based on autoencoders and Robust PCA is proposed. Because Robust PAC has greater advantages in feature extraction of high-dimensional data and autoencoder has powerful capabilities in the reconstruction of normal data, the proposed approach can effectively address the two problems concerned above. The proposed approach is compared with some representative approaches, including ABOD, KNN, LOF, SOS and SOD, on eight well-known public datasets. The experiments show that compared with them, the proposed approach has advantages in both precision and recall rate, and can more accurately distinguish between normal data and outliers.
Wanghu Chen, Jing Li 0131, Ali Arshad
IEEE BigData1
2019 Benchmarking Discretisation Level of Continuous Attributes: Theoretical and Experimental Approaches
abstract
The discretisation of an attribute refers to partitioning its continuous numerical values into intervals, each of which is associated a categorical label. The amount of such different categorical labels is called as target discretisation level of the continuous attribute. For data mining algorithms that can only work on discrete data, the discretisation will be necessary. At the same time, the discretisation can also make the original data more concise and interpretable. However, it is challenging to balance the target discretisation level and the information loss during the discretisation process. In this paper, we propose to use entropy of a continuous attribute as a benchmark to determine its target discretisation level for the first time. An entropy based naive unsupervised discretisation approach is also proposed and shows big advantages in terms of both data reduction and accuracy, which is evaluated by performing classifiers on the dataset whose continuous attributes are discretised based on the proposed approach. Our experiments on 28 datasets and 9 popular classifiers show that the accuracy of a discretisation approach will be largely affected when the target discretisation level of each continuous attribute is lower than the entropy benchmark. Meanwhile increasing the target discretisation level from the benchmark does not always improve the accuracy of the discretizer. These discoveries can provide valuable guidance to explore or optimise the approaches to the discretisation of continuous attributes.
Wanghu Chen, Jing Li 0131, Bo Yang 0043, Jianwu Wang 0001
IEEE BigData1
2018 Blockchain Based Provenance Sharing of Scientific Workflows
abstract
In a research community, the provenance sharing of scientific workflows can enhance distributed research cooperation, experiment reproducibility verification and experiment repeatedly doing. Considering that scientists in such a community are often in a loose relation and distributed geographically, traditional centralized provenance sharing architectures have shown their disadvantages in poor trustworthiness, reliabilities and efficiency. Additionally, they are also difficult to protect the rights and interests of data providers. All these have been largely hindering the willings of distributed scientists to share their workflow provenance. Considering the big advantages of blockchain in decentralization, trustworthiness and high reliability, an approach to sharing scientific workflow provenance based on blockchain in a research community is proposed. To make the approach more practical, provenance is handled on-chain and original data is delivered off-chain. A kind of block structure to support efficient provenance storing and retrieving is designed, and an algorithm for scientists to search workflow segments from provenance as well as an algorithm for experiments backtracking are provided to enhance the experiment result sharing, save computing resource and time cost by avoiding repeated experiments as far as possible. Analyses show that the approach is efficient and effective.
Wanghu Chen, Xiaoyan Liang, Jing Li 0131, Hongwu Qin, Yuxiang Mu, Jianwu Wang 0001
IEEE BigData1
2017 Enhancing the MapReduce training of BP neural networks based on local weight matrix evolution
abstract
Training Back-Propagation Neural Networks (BPNNs) on big datasets faces two challenges, the hight time cost and the possibility of getting trapped into local optimum. MapReduce has been introduced to improve the efficiency of BPNN training on big datasets in recent years. After each turn of BPNN training on each split of the dataset concurrently, lots of local BPNNs that are only convergent on the specific split will be produced, and a global BPNN candidate convergent on the whole dataset needs to be generated from them. This process is full of challenges because it has a high impact on the training efficiency as well as the training accuracy. The paper introduces the evolution of the local BPNNs into the MapReduce training of BPNN, and proposes a novel approach. Profiting from the advantages of EAs in global optimum searching, the approach can reduce the iterations to get the global convergent BPNN candidate and avoid the training process to get trapped into local optimum. Experiments show the approach can improve the training efficiency and accuracy remarkably. The approach has also been applied into a real-world big data application and verified it can work well on big and high dimension datasets.
Wanghu Chen, Xintian Li, Jing Li 0131, Jianwu Wang 0001
IEEE BigData1