Bin Xia 0003

dblp:80/846-3 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PawsPro: A proactive Paw-shaped framework for predictive modeling and optimal decision in hard drive failure prevention
Bin Xia 0003, Junmin Cao, Guozi Sun
Future Gener. Comput. Syst.1
2024 SSR-TA: Sequence-to-Sequence-based expert recurrent recommendation for ticket automation
Chenhan Cao, Xiaoyu Fang, Bingqing Luo, Bin Xia 0003
Neural Comput. Appl.4
2024 SiaDFP: A Disk Failure Prediction Framework Based on Siamese Neural Network in Large-Scale Data Center
abstract
With the rapid development of cloud services, service providers increasingly rely on a dependable storage system equipped with large-capacity disks to ensure data availability. The primary source of unreliability in such storage systems attributes to disk failures. In recent years, some proactive methods base on machine learning models have emerged, aiming to predict impending disk failures by leveraging the SMART attributes of disks. These methods enable service providers to timely back up storage data. While the methods prove more effective and efficient in disk failure prediction, they still face challenges, such as inadequate mining of abnormal information and imbalanced classification. In this paper, we mainly analyzed the change of data distribution in hard disks. From the data analysis, we observed that the distribution change in the failed disk is obvious during the period before the disk damage, while that in the healthy disk is insignificant during running time. Motivated by the observation, we propose a novel framework namedSiaDFP, based on Siamese neural network, designed to predict impending disk failures by capturing the distribution changes in failed disks. Additionally, we observed that the failed disks exhibit some change points as an abnormal feature by analyzing the disk data trend. To fully mining abnormal information inhere in failed disks, we propose CP-MAP mechanism and 2D-Attention mechanism. Furthermore, we present a subsampling approach named Region Balanced Sampling to address the challenge of imbalanced classification. Experiments on the real-world dataset Backblaze and Baidu demonstrate that the performance ofSiaDFPis outstanding in the task of disk failure prediction.
Xiaoyu Fang, Wenbai Guan, Chenhan Cao, Bin Xia 0003
IEEE Trans. Serv. Comput.5
2023 Cost-sensitive Tensor-based Dual-stage Attention LSTM with Feature Selection for Data Center Server Power Forecasting
abstract
Power forecasting has a guiding effect on power-aware scheduling strategies to reduce unnecessary power consumption in data centers. Many metrics related to power consumption can be collected in physical servers, such as the status of CPU, memory, and other components. However, most existing methods empirically exploit a small number of metrics to forecast power consumption. To this end, this article uses feature selection based on causality to explore the metrics that strongly influence the power consumption of different tasks. Moreover, we propose a tensor-based dual-stage attention LSTM to forecast the non-linear and non-periodic power consumption. In the proposed model, a multi-way delay embedding transform is utilized to convert the time series into tensors along the temporal direction. The LSTM combines with the tensor technique and the attention mechanism to capture the temporal pattern effectively. In addition, we adopt the cost-sensitive loss function to optimize the specific power forecasting problem in data centers. The experimental results demonstrate that our method can achieve up to 1.4% to 4.3% forecasting accuracy improvement compared with the state-of-the-art models.
Ziyu Shen, Binghui Liu, Zheng Liu 0001, Bin Xia 0003, Yun Li 0009
ACM Trans. Intell. Syst. Technol.5
2022 Data characteristics aware prediction model for power consumption of data center servers
abstract
Summary Due to the rapid increase in the number and scale of data centers, the information and communication technology (ICT) equipment in data centers consumes an enormous amount of power. A power prediction model is therefore essential for decision‐making optimization and power management of ICT equipment. However, it is difficult to predict the power consumption of data centers accurately due to the complex power patterns and nonlinear interdependencies among components. Existing methods either rely on standard formulas, or simply treat it as time series, both leading to poor power prediction accuracy. To overcome those limitations, in this article, we present a systematic power prediction framework called characteristic aware attention‐augmented deep learning‐based prediction method. In particular, we first analyze the different power consumption series to illustrate their different temporal characteristics. Second, we perform different data processing for the corresponding characteristics of power series samples. Third, we propose an accurate and efficient neural network model to predict future power consumption with the pretreated data. The experimental results show that the proposed model is able to achieve superior prediction accuracy.
Ziyu Shen, Bin Xia 0003, Zheng Liu 0001, Yun Li 0009
Concurr. Comput. Pract. Exp.4
2022 (AD)2: Adversarial domain adaptation to defense with adversarial perturbation removal
Keji Han, Bin Xia 0003, Yun Li 0009
Pattern Recognit.2
2022 TRAC: Traceable and Revocable Access Control Scheme for mHealth in 5G-Enabled IIoT
abstract
Mobile healthcare (mHealth) enables people to collect and share their personal health records (PHRs) and gain rapid medical treatment via mobile 5G-enabled Industrial Internet of Things (IIoT) devices, which also brings the challenge of keeping the PHRs confidentiality and preventing unauthorized access. By the emerging ciphertext-policy attribute-based encryption (CP-ABE), the PHR owner can encrypt his/her PHR data under self-defined access policies. However, existing CP-ABE schemes are suffering from either heavy computation cost and storage overhead or traitor tracing and direct revocation. In this article, we propose an efficient, traceable, and revocable access control scheme named TRAC for mHealth in 5G-enabled IIoT. In TRAC, the ciphertext is composed of the attribute-relevant ciphertext encrypted under anand-gate access structure and the identity-relevant ciphertext associated with some potential receivers. The malicious user who leaks his/her privilege to unauthorized entities will be precisely tracked and added in the revocation list, by which the cloud server can update the identity-relevant ciphertext by itself. The length of final ciphertext and the time of bilinear pairing operations used in decryption are constant. The security analysis and performance evaluation indicate the security, efficiency, and practicality of TRAC.
Qi Li 0011, Bin Xia 0003, Haiping Huang, Yinghui Zhang 0002, Tao Zhang 0029
IEEE Trans. Ind. Informatics2
2021 An Effective Classification-Based Framework for Predicting Cloud Capacity Demand in Cloud Services
abstract
The rapid development of pay-as-you-go cloud services motivates the increasing number of cloud resource demands. However, the volatile demands bring new challenges for current techniques to minimize the cost of cloud capacity planning and VM provisioning while satisfying the customer demands. The service vendors will incur enormous revenue loss within the long-term inappropriate planning, especially when the demands fluctuate abruptly and frequently. In this paper, we cast the cloud capacity planning as a classification problem and propose an integrated framework, which effectively predicts the abrupt changing demands, to reduce the cost of cloud resource provisioning. In this framework, we first apply Piecewise Linear Representation to segment the time series of cloud resource demands for labeling the changing trend of each period. Second, Weighted SVM is leveraged to fit the statistical information and the label of each period and predict the changing trend of the following period. Finally, an incremental learning strategy is utilized to ensure the low cost of updating the model using the upcoming requests. We evaluate our framework on the IBM Smart Cloud Enterprise (SCE) trace data and the experimental results show the effectiveness of our proposed framework.
Bin Xia 0003, Tao Li 0001, Qifeng Zhou, Qianmu Li, Hong Zhang 0021
IEEE Trans. Serv. Comput.1
2020 Estimating Power Consumption of Containers and Virtual Machines in Data Centers
abstract
Virtualization technologies provide solutions of cloud computing. Virtual resource scheduling is a crucial task in data centers, and the power consumption of virtual resources is a critical foundation of virtualization scheduling. Containers are the smallest unit of virtual resource scheduling and migration. Although many effective models for estimating power consumption of virtual machines (VM) have been proposed, few power estimation models of containers have been put forth. In this paper, we offer a fast-training piecewise regression model based on decision tree to build a VM power estimation model and estimate the containers' power by treating the container as a group of processes on the VM. In our model, we characterize the nonlinear relationship between power and features and realize the effective estimation of the containers on the VM. We evaluate the proposed model on 13 workloads in PARSEC and compare it with several models. The experimental results prove the effectiveness of our proposed model on most workloads. Moreover, the estimated power of the containers is in line with expectations.
Ziyu Shen, Bin Xia 0003, Zheng Liu 0001, Yun Li 0009
CLUSTER3
2020 DPAST-RNN: A Dual-Phase Attention-Based Recurrent Neural Network Using Spatiotemporal LSTMs for Time Series Prediction
Shajia Shan, Ziyu Shen, Bin Xia 0003, Zheng Liu 0001, Yun Li 0009
ICONIP (3)3
2020 Adversarial Named Entity Recognition with POS label embedding
abstract
Named Entity Recognition (NER) is dedicated to recognizing different types of named entity. Previous works have shown that part-of-speech, as an important feature, provides complementary syntactical information to NER systems. However, these studies suffer from two limitations: (i) the previous models do not consider the noise from part-of-speech; (ii) the previous models need to re-extract features from token representations. In this paper, we propose a novel approach that can alleviate the above issues as well as make full use of part-of-speech features via attention mechanism and adversarial training. We evaluate our model on three NER datasets, and the experimental results demonstrate that our model achieves a state-of-the-art F1-score of Twitter dataset while matching a state-of-the-art performance on the CoNLL-2003 and Weibo datasets.
Yu Wang 0072, Bin Xia 0003, Yun Li 0009, Ziye Zhu
IJCNN3
2019 ADPR: An Attention-based Deep Learning Point-of-Interest Recommendation Framework
abstract
With the development of location-based social networks (LBSNs), Point-of-Interest (POI) recommendation has attracted lots of attention. Most of the existing studies focus on recommending POIs to users based on their recent check-ins. However, the recent check-ins may contain some daily check-ins that users are not really interested in. If a model treats the recent check-ins equally, it is non-trivial to capture the actual preference of users. To address the issue of mining the actual preferences of users in the POI recommendation, we propose an attention-based deep learning POI recommendation framework (ADPR), which consists of a latent representation method and an attention-based deep convolutional neural network. To learn the embedding of users and POIs, we propose a latent representation method, which incorporates the geographical influence and the categories of POIs to capture the relationships between POIs better. Further, we propose an attention-based deep convolutional neural network, which employs the attention mechanism to filter the important information in the recent check-ins, to recommend POIs to users based on the latent representations of users and the recent check-ins. We conduct experiments on a real-world LBSN dataset to evaluate our framework, and the experimental results show the effectiveness of our framework.
Junjie Yin, Yun Li 0009, Zheng Liu 0001, Jian Xu 0009, Bin Xia 0003, Qianmu Li
IJCNN5
2019 SC-NER: A Sequence-to-Sequence Model with Sentence Classification for Named Entity Recognition
Yu Wang 0072, Yun Li 0009, Ziye Zhu, Bin Xia 0003, Zheng Liu 0001
PAKDD (1)4
2019 Joint prediction of time series data in inventory management
Qifeng Zhou, Ruyuan Han, Tao Li 0001, Bin Xia 0003
Knowl. Inf. Syst.4
2019 WE-Rec: A fairness-aware reciprocal recommendation based on Walrasian equilibrium
Bin Xia 0003, Junjie Yin, Jian Xu 0009, Yun Li 0009
Knowl. Based Syst.1
2018 Topic Modeling for Noisy Short Texts with Multiple Relations
abstract
Understanding contents in social networks by inferring high-quality latent topics from short texts is a significant task in social analysis, which is challenging because social network contents are usually extremely short, noisy and full of informal vocabularies.Due to the lack of sufficient word co-occurrence instances, well-known topic modeling methods such as LDA and LSA cannot uncover high-quality topic structures.Existing research works seek to pool short texts from social networks into pseudo documents or utilize the explicit relations among these short texts such as hashtags in tweets to make classic topic modeling methods work.In this paper, we explore this problem by proposing a topic model for noisy short texts with multiple relations called MRTM (Multiple Relational Topic Modeling).MRTM exploits both explicit and implicit relations by introducing a document-attribute distribution and a two-step random sampling strategy.Extensive experiments, compared with stateof-the-art topic modeling approaches, demonstrate that MRTM can alleviate the word co-occurrence sparsity and uncover highquality latent topics from noisy short texts.
Chi-Yu Liu, Zheng Liu 0001, Tao Li 0001, Bin Xia 0003
SEKE4
2018 Noise-tolerance matrix completion for location recommendation
Bin Xia 0003, Tao Li 0001, Qianmu Li, Hong Zhang 0021
Data Min. Knowl. Discov.1
2018 Multiple Relational Topic Modeling for Noisy Short Texts
abstract
Understanding contents in social networks by inferring high-quality latent topics from short texts is a significant task in social analysis, which is challenging because social network contents are usually extremely short, noisy and full of informal vocabularies. Due to the lack of sufficient word co-occurrence instances, well-known topic modeling methods such as LDA and LSA cannot uncover high-quality topic structures. Existing research works seek to pool short texts from social networks into pseudo documents or utilize the explicit relations among these short texts such as hashtags in tweets to make classic topic modeling methods work. In this paper, we explore this problem by proposing a topic model for noisy short texts with multiple relations called MRTM (Multiple Relational Topic Modeling). MRTM exploits both explicit and implicit relations by introducing a document-attribute distribution and a two-step random sampling strategy. Extensive experiments, compared with the state-of-the-art topic modeling approaches, demonstrate that MRTM can alleviate the word co-occurrence sparsity and uncover high-quality latent topics from noisy short texts.
Zheng Liu 0001, Chi-Yu Liu, Bin Xia 0003, Tao Li 0001
Int. J. Softw. Eng. Knowl. Eng.3
2017 FLAP: An End-to-End Event Log Analysis Platform for System Management
abstract
Many systems, such as distributed operating systems, complex networks, and high throughput web-based applications, are continuously generating large volume of event logs. These logs contain useful information to help system administrators to understand the system running status and to pinpoint the system failures. Generally, due to the scale and complexity of modern systems, the generated logs are beyond the analytic power of human beings. Therefore, it is imperative to develop a comprehensive log analysis system to support effective system management. Although a number of log mining techniques have been proposed to address specific log analysis use cases, few research and industrial efforts have been paid on providing integrated systems with an end-to-end solution to facilitate the log analysis routines.
Tao Li 0001, Yexi Jiang, Chunqiu Zeng, Bin Xia 0003, Zheng Liu 0001, Wubai Zhou, Wentao Wang 0006, Dewei Bao
KDD4
2017 VRer: Context-Based Venue Recommendation using embedded space ranking SVM in location-based social network
Bin Xia 0003, Zhen Ni, Tao Li 0001, Qianmu Li, Qifeng Zhou
Expert Syst. Appl.1
2017 FIU-Miner (a fast, integrated, and user-friendly system for data mining) and its applications
Tao Li 0001, Chunqiu Zeng, Wubai Zhou, Zheng Liu 0001, Qifeng Zhou, Bin Xia 0003, Qing Wang 0016, Wentao Wang 0006
Knowl. Inf. Syst.8
2015 A Classification-Based Demand Trend Prediction Model in Cloud Computing
Qifeng Zhou, Bin Xia 0003, Yexi Jiang, Qianmu Li, Tao Li 0001
WISE (2)2