Jian Xu 0009

dblp:73/1149-9 · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0003-4216-5500ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Computer networks · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PosParser: A Heuristic Online Log Parsing Method Based on Part-of-Speech Tagging
abstract
Log parsing, the process of transforming raw logs into structured data, is a key step in the complex computer system's intelligent operation and maintenance and therefore has received extensive attention. Among all log parsing methods, heuristic log parsing methods are lightweight and can work in a streaming mode to well meet the real-time parsing requirements. However, the existing log representations used in the heuristic log parsing methods are not powerful in distinguishing log messages, which leads to low parsing accuracy and weak generality. Inspired by trigger word extraction of the event detection task in natural language processing (NLP), this paper proposes an online log parser, named PosParser, which employs the part-of-speech (PoS) tagging to extract a function token sequence (FTS) as the log message representation, and then identify event templates of log messages through the FTS. Experimental results on sixteen logs from real systems demonstrate that the FTS is powerful in distinguishing log messages from different event templates, and PosParser not only performs better in terms of parsing accuracy than state-of-the-art methods but is also comparable to them in efficiency.
Jinzhao Jiang, Jian Xu 0009
IEEE Trans. Big Data3
2024 LogTransformer: Transforming IT System Logs Into Events Using Tree-Based Approach
abstract
As an important outcome of complex IT systems in operation, logs provide valuable information for system operation and maintenance. Log event (or template) extraction plays a vital role in log analysis, as its accuracy significantly impacts follow-up tasks such as log anomaly detection and event pattern discovery. Despite achieving high accuracy on specific system logs, existing log event extraction approaches still struggle with low accuracy and instability when handling logs from heterogeneous systems or logs with variable-length parameters. To address these issues, this paper proposes LogTransformer, an online event extraction approach based on a tree structure. A tree-based log content parsing approach is proposed to perform log event extraction by comparing the similarity between a log tree representing an incoming log message and an event tree representing a specific log template. Extensive experiments are conducted on sixteen benchmark log datasets to evaluate the effectiveness, robustness, and efficiency of the proposed approach. The experimental results demonstrate that an average accuracy exceeds 90%, surpassing the state-of-the-art online log parser, Drain.
Jian Xu 0009
IEEE Trans. Netw. Serv. Manag.2
2024 Task allocation for unmanned aerial vehicles in mobile crowdsensing
Sunyue Xu, Jing Zhang 0015, Shunmei Meng, Jian Xu 0009
Wirel. Networks4
2023 Graph transformer-based self-adaptive malicious relation filtering for fraudulent comments detection in social network
Liangjun Li, Jian Xu 0009
Knowl. Based Syst.2
2023 MLog: Mogrifier LSTM-Based Log Anomaly Detection Approach Using Semantic Representation
abstract
Streaming logs provide valuable information for complex systems in diagnosing system faults or conducting security analysis. Although the log sequence anomaly detection has drawn more and more attention and achieved a satisfactory performance, it remains an extremely difficult task because of several intrinsic challenges including new event occurrences in a continuously evolving environment, making full use of rich dependency hidden in sequential events from the global and local view. To meet these challenges, in this article, we propose MLog, a hybrid deep neural network for detecting anomalies in log sequences. Specifically, MLog leverages the transformer encoder and a novel event Inverse Document Frequency (IDF) weighted mechanism to obtain a semantic vector for an individual log template. Log sequences represented by sequential template semantic vectors are then fed into a deep neural network combing the Mogrifier Long Short Term Memory (LSTM) with Convolutional Neural Network (CNN) to capture global and local sequential patterns simultaneously. We implement MLog and evaluate it by conducting extensive experiments on two well-known benchmark datasets, HDFS and BGL, from the aspects of detection accuracy and robustness. The results show that MLog outperforms the state-of-the-art approaches and is robust to the evolving logs. To encourage reproducibility, we make the implementation of MLog available.
Jian Xu 0009
IEEE Trans. Serv. Comput.3
2022 Prediction of disease-associated nsSNPs by integrating multi-scale ResNet models with deep feature fusion
abstract
More than 6000 human diseases have been recorded to be caused by non-synonymous single nucleotide polymorphisms (nsSNPs). Rapid and accurate prediction of pathogenic nsSNPs can improve our understanding of the principle and design of new drugs, which remains an unresolved challenge. In the present work, a new computational approach, termed MSRes-MutP, is proposed based on ResNet blocks with multi-scale kernel size to predict disease-associated nsSNPs. By feeding the serial concatenation of the extracted four types of features, the performance of MSRes-MutP does not obviously improve. To address this, a second model FFMSRes-MutP is developed, which utilizes deep feature fusion strategy and multi-scale 2D-ResNet and 1D-ResNet blocks to extract relevant two-dimensional features and physicochemical properties. FFMSRes-MutP with the concatenated features achieves a better performance than that with individual features. The performance of FFMSRes-MutP is benchmarked on five different datasets. It achieves the Matthew's correlation coefficient (MCC) of 0.593 and 0.618 on the PredictSNP and MMP datasets, which are 0.101 and 0.210 higher than that of the existing best method PredictSNP1. When tested on the HumDiv and HumVar datasets, it achieves MCC of 0.9605 and 0.9507, and area under curve (AUC) of 0.9796 and 0.9748, which are 0.1747 and 0.2669, 0.0853 and 0.1335, respectively, higher than the existing best methods PolyPhen-2 and FATHMM (weighted). In addition, on blind test using a third-party dataset, FFMSRes-MutP performs as the second-best predictor (with MCC and AUC of 0.5215 and 0.7633, respectively), when compared with the other four predictors. Extensive benchmarking experiments demonstrate that FFMSRes-MutP achieves effective feature fusion and can be explored as a useful approach for predicting disease-associated nsSNPs. The webserver is freely available at http://csbio.njust.edu.cn/bioinf/ffmsresmutp/ for academic use.
Fang Ge, Ying Zhang 0053, Jian Xu 0009, Muhammad Arif 0012, Jiangning Song, Dongjun Yu
Briefings Bioinform.3
2022 MAResNet: predicting transcription factor binding sites by combining multi-scale bottom-up and top-down attention and residual network
abstract
Accurate identification of transcription factor binding sites is of great significance in understanding gene expression, biological development and drug design. Although a variety of methods based on deep-learning models and large-scale data have been developed to predict transcription factor binding sites in DNA sequences, there is room for further improvement in prediction performance. In addition, effective interpretation of deep-learning models is greatly desirable. Here we present MAResNet, a new deep-learning method, for predicting transcription factor binding sites on 690 ChIP-seq datasets. More specifically, MAResNet combines the bottom-up and top-down attention mechanisms and a state-of-the-art feed-forward network (ResNet), which is constructed by stacking attention modules that generate attention-aware features. In particular, the multi-scale attention mechanism is utilized at the first stage to extract rich and representative sequence features. We further discuss the attention-aware features learned from different attention modules in accordance with the changes as the layers go deeper. The features learned by MAResNet are also visualized through the TMAP tool to illustrate that the method can extract the unique characteristics of transcription factor binding sites. The performance of MAResNet is extensively tested on 690 test subsets with an average AUC of 0.927, which is higher than that of the current state-of-the-art methods. Overall, this study provides a new and useful framework for the prediction of transcription factor binding sites by combining the funnel attention modules with the residual network.
Long-Chen Shen, Yiheng Zhu 0001, Jian Xu 0009, Jiangning Song, Dongjun Yu
Briefings Bioinform.4
2022 Adaptive processing rate based container provisioning for meshed Micro-services in Kubernetes Clouds
Zhicheng Cai, Yamin Lei, Jian Xu 0009, Rajkumar Buyya
CCF Trans. High Perform. Comput.4
2022 S3Feature: A static sensitive subgraph-based feature for android malware detection
Fan Ou, Jian Xu 0009
Comput. Secur.2
2022 LibRoad: Rapid, Online, and Accurate Detection of TPLs on Android
abstract
Third-party library (TPL) detection plays a very important role in Android malware analysis. The focus of recent works has been shifted to the signature-based approach. However, previous methods have several limitations such as high time complexity and low precision, especially with the presence of similar TPLs and various versions of a TPL. To solve these issues, we propose a rapid, online, and accurate TPL detection approach, named LibRoad, which also follows the line of the signature-based research. To reduce the time cost, our approach integrates an application preprocessing component and a pairwise package matching component. The former divides an application into primary modules and non-primary modules to enable us to focus on analyzing packages in non-primary modules that are the most possibly imported from a TPL. The latter adopts a combination of the package name based matching policy for non-obfuscated packages and the signature-based matching policy for obfuscated packages, where the package name based matching policy has a lower time complexity than the signature based one. Further, to improve performance, our approach integrates a perfectly matched package and TPL determination component, which adopts the package filter mechanism, online TPL detection, and local TPL discovery to identify TPLs with low false positive and false negative. We conduct several groups of experiments on real-world applications and two ground truth bases. Experimental results show that compared to state-of-the-art approaches, LibRoad can achieve a high recall of 99.86 percent and a low false positive rate of 11.48 percent without the loss of efficiency.
Jian Xu 0009, Qianting Yuan
IEEE Trans. Mob. Comput.1
2021 Leveraging the attention mechanism to improve the identification of DNA N6-methyladenine sites
abstract
DNA N6-methyladenine is an important type of DNA modification that plays important roles in multiple biological processes. Despite the recent progress in developing DNA 6mA site prediction methods, several challenges remain to be addressed. For example, although the hand-crafted features are interpretable, they contain redundant information that may bias the model training and have a negative impact on the trained model. Furthermore, although deep learning (DL)-based models can perform feature extraction and classification automatically, they lack the interpretability of the crucial features learned by those models. As such, considerable research efforts have been focused on achieving the trade-off between the interpretability and straightforwardness of DL neural networks. In this study, we develop two new DL-based models for improving the prediction of N6-methyladenine sites, termed LA6mA and AL6mA, which use bidirectional long short-term memory to respectively capture the long-range information and self-attention mechanism to extract the key position information from DNA sequences. The performance of the two proposed methods is benchmarked and evaluated on the two model organisms Arabidopsis thaliana and Drosophila melanogaster. On the two benchmark datasets, LA6mA achieves an area under the receiver operating characteristic curve (AUROC) value of 0.962 and 0.966, whereas AL6mA achieves an AUROC value of 0.945 and 0.941, respectively. Moreover, an in-depth analysis of the attention matrix is conducted to interpret the important information, which is hidden in the sequence and relevant for 6mA site prediction. The two novel pipelines developed for DNA 6mA site prediction in this work will facilitate a better understanding of the underlying principle of DL-based DNA methylation site prediction and its future applications.
Ying Zhang 0053, Yan Liu 0038, Jian Xu 0009, Xiaoyu Wang 0016, Xinxin Peng, Jiangning Song, Dongjun Yu
Briefings Bioinform.3
2021 PScL-HDeep: image-based prediction of protein subcellular location in human tissue using ensemble learning of handcrafted and deep learned features with two-layer feature selection
abstract
Protein subcellular localization plays a crucial role in characterizing the function of proteins and understanding various cellular processes. Therefore, accurate identification of protein subcellular location is an important yet challenging task. Numerous computational methods have been proposed to predict the subcellular location of proteins. However, most existing methods have limited capability in terms of the overall accuracy, time consumption and generalization power. To address these problems, in this study, we developed a novel computational approach based on human protein atlas (HPA) data, referred to as PScL-HDeep, for accurate and efficient image-based prediction of protein subcellular location in human tissues. We extracted different handcrafted and deep learned (by employing pretrained deep learning model) features from different viewpoints of the image. The step-wise discriminant analysis (SDA) algorithm was applied to generate the optimal feature set from each original raw feature set. To further obtain a more informative feature subset, support vector machine-based recursive feature elimination with correlation bias reduction (SVM-RFE + CBR) feature selection algorithm was applied to the integrated feature set. Finally, the classification models, namely support vector machine with radial basis function (SVM-RBF) and support vector machine with linear kernel (SVM-LNR), were learned on the final selected feature set. To evaluate the performance of the proposed method, a new gold standard benchmark training dataset was constructed from the HPA databank. PScL-HDeep achieved the maximum performance on 10-fold cross validation test on this dataset and showed a better efficacy over existing predictors. Furthermore, we also illustrated the generalization ability of the proposed method by conducting a stringent independent validation test.
Matee Ullah, Fazal Hadi, Jian Xu 0009, Jiangning Song, Dongjun Yu
Briefings Bioinform.4
2020 A multi-view similarity measure framework for trouble ticket mining
Jian Xu 0009, Jiapeng Mu, Gaorong Chen
Data Knowl. Eng.1
2019 DUE Distribution and Pairing in D2D Communication
abstract
The D2D (Device-to-Device) communication has been very popular as it is a promising and low-cost solution to reduce the burden on the cellular network. However, there are rare concerns about the distribution and pairing of DUEs(D2D user equipments), which have a significant impact on QoS (Quality of Service) of D2D communication. In this paper, we propose a novel algorithm based on the coalitional game to optimally adjust the distribution of DUEs. The proposed algorithm aims to form the optimal coalition structure, which achieves a balance between the throughput and power consumption of each coalition, obtaining the enhanced QoS of D2D. We show that our algorithm is superior to the benchmark models in terms of the throughput and energy efficiency of the DUE coalition. To further improve the QoS, we also propose a method to predict and maximize the pairing probability of DUEs. The proposed prediction method adopts the Logistic Regression to model the global pairing probability according to the communication parameters of DUEs. Experimental results show that the proposed prediction method is significantly superior to the benchmark methods in terms of prediction accuracy. In addition, the pairing probability maximization algorithm proposed also significantly improves the pairing probability.
Weifeng Lu, Xiaoqiang Ren, Jia Xu 0003, Siguang Chen, Jian Xu 0009
ICCCN6
2019 ADPR: An Attention-based Deep Learning Point-of-Interest Recommendation Framework
abstract
With the development of location-based social networks (LBSNs), Point-of-Interest (POI) recommendation has attracted lots of attention. Most of the existing studies focus on recommending POIs to users based on their recent check-ins. However, the recent check-ins may contain some daily check-ins that users are not really interested in. If a model treats the recent check-ins equally, it is non-trivial to capture the actual preference of users. To address the issue of mining the actual preferences of users in the POI recommendation, we propose an attention-based deep learning POI recommendation framework (ADPR), which consists of a latent representation method and an attention-based deep convolutional neural network. To learn the embedding of users and POIs, we propose a latent representation method, which incorporates the geographical influence and the categories of POIs to capture the relationships between POIs better. Further, we propose an attention-based deep convolutional neural network, which employs the attention mechanism to filter the important information in the recent check-ins, to recommend POIs to users based on the latent representations of users and the recent check-ins. We conduct experiments on a real-world LBSN dataset to evaluate our framework, and the experimental results show the effectiveness of our framework.
Junjie Yin, Yun Li 0009, Zheng Liu 0001, Jian Xu 0009, Bin Xia 0003, Qianmu Li
IJCNN4
2019 WE-Rec: A fairness-aware reciprocal recommendation based on Walrasian equilibrium
Bin Xia 0003, Junjie Yin, Jian Xu 0009, Yun Li 0009
Knowl. Based Syst.3
2018 Expert recommendation for trouble ticket routing
Jian Xu 0009, Rouying He
Data Knowl. Eng.1
2018 Signature based trouble ticket classification
Jian Xu 0009, Wubai Zhou, Rouying He, Tao Li 0001
Future Gener. Comput. Syst.1
2018 Trouble Ticket Routing Models and Their Applications
abstract
A trouble ticket is an important information carrier in system maintenance, which records problem symptoms, the resolving process, and resolutions. A critical challenge for the ticket management system is how to quickly deal with trouble tickets and fix problems. Thousands of tickets, bouncing among multiple expert groups before being fixed, will consume limited system maintenance resources and may also violate the service level agreement. Thus, trouble tickets should be routed to the right expert group as quickly as possible in order to reduce the processing delay. In this paper, to address the challenge in ticket routing, we exploit three different routing models by mining the combination of problem descriptions and resolution sequences from the historical resolved tickets, and develop the corresponding routing recommendation algorithms to determine the next expert group to solve the problem. To evaluate the performance of routing recommendation algorithms, we conduct extensive experiments on a real ticket data set. The experimental results show that the proposed models and algorithm can effectively shorten the mean number of steps to resolve with a high ratio of the number of successfully resolved tickets to the total number of tickets, especially for the long routing sequences generated from manual assignments. These models and algorithms have the potential of being used in a ticket routing recommendation engine to greatly reduce human intervention in the routing process.
Jian Xu 0009, Rouying He, Wubai Zhou, Tao Li 0001
IEEE Trans. Netw. Serv. Manag.1
2017 STAR: A System for Ticket Analysis and Resolution
abstract
In large scale and complex IT service environments, a problematic incident is logged as a ticket and contains the ticket summary (system status and problem description). The system administrators log the step-wise resolution description when such tickets are resolved. The repeating service events are most likely resolved by inferring similar historical tickets. With the availability of reasonably large ticket datasets, we can have an automated system to recommend the best matching resolution for a given ticket summary. In this paper, we first identify the challenges in real-world ticket analysis and develop an integrated framework to efficiently handle those challenges. The framework first quantifies the quality of ticket resolutions using a regression model built on carefully designed features. The tickets, along with their quality scores obtained from the resolution quality quantification, are then used to train a deep neural network ranking model that outputs the matching scores of ticket summary and resolution pairs. This ranking model allows us to leverage the resolution quality in historical tickets when recommending resolutions for an incoming incident ticket. In addition, the feature vectors derived from the deep neural ranking model can be effectively used in other ticket analysis tasks, such as ticket classification and clustering. The proposed framework is extensively evaluated with a large real-world dataset.
Wubai Zhou, Ramesh Baral, Qing Wang 0016, Chunqiu Zeng, Tao Li 0001, Jian Xu 0009, Zheng Liu 0001, Larisa Shwartz, Genady Grabarnik
KDD7
2016 System situation ticket identification using SVMs ensemble
Jian Xu 0009, Tao Li 0001
Expert Syst. Appl.1
2016 Resource allocation based on quantum particle swarm optimization and RBF neural network for overlay cognitive OFDM System
Lei Xu 0015, Fang Qian, Qianmu Li, Yuwang Yang, Jian Xu 0009
Neurocomputing6
2016 Pattern discovery via constraint programming
Jian Xu 0009, Chunqiu Zeng, Tao Li 0001
Knowl. Based Syst.1
2015 Node anomaly detection for homogeneous distributed environments
Jian Xu 0009, Yexi Jiang, Chunqiu Zeng, Tao Li 0001
Expert Syst. Appl.1
2014 Social network user influence sense-making and dynamics prediction
Wei Peng 0001, Tao Li 0001, Tong Sun 0001, Qianmu Li, Jian Xu 0009
Expert Syst. Appl.6
2012 A new evidential trust model for open distributed systems
Jian Xu 0009, Hong Zhang 0021
Expert Syst. Appl.2
2010 Agent service matchmaking algorithm for autonomic element with semantic and QoS constraints
Manwu Xu, Hong Zhang 0021, Jian Xu 0009
Knowl. Based Syst.4
2005 A neural-wavelet based methodology for software aging forecasting
abstract
A number of recent studies have reported the phenomenon of "software aging", characterized by progressive performance degradation and a sudden crash/hang of a software system due to exhaustion of operating system resources, fragmentation and accumulation of errors. To counteract this phenomenon, this paper proposed a novel four-stage method for software aging forecast in operation. The prior data of software performance parameters are treated as time series. The forecast method combines wavelet multiresolution decomposition and neural networks. First, we apply a smoothing unit based on the wavelet multiresolution analysis to reduce the influence of noise. Second, the special performance data is decomposed into different scales by nondecimated Haar wavelet decomposition. Third, each scale is predicted by a separate neural network. Lastly, the next sample of the original time series is predicted by another neural network. The proposed method is tested using the performance parameters data collected from a realistic software system to evaluate the forecasting performance.
Jian Xu 0009, Jing You
SMC1
2005 Modeling and availability analysis of nested software rejuvenation policy
abstract
Software rejuvenation is a proactive technique to counteract software aging. A new nested software rejuvenation policy is put forward in this paper. A finite-state automaton is used to model the working process of software with this policy. Comparing to the conventional periodic software rejuvenation policy, the nested policy takes into account the application-level and system-level rejuvenation simultaneously, specially giving emphasis on nesting. This paper solves the rejuvenation intervals based on the minimum downtime and compares the maximum system availability of the nested software rejuvenation policy with the conventional periodic software rejuvenation policy's. The numerical results demonstrate that the new policy consumes less downtime, enhances software availability and reliability.
Jing You, Jian Xu 0009, Xue-long Zhao, Feng-Yu Liu
SMC2