Maozhen Li 0001

dblp:l/MaozhenLi · DBLP profile ↗
← Back
95ranked-venue papers
16as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 22 since 2021Systems, architecture and hardware · 34 · 10 first-author · 2 since 2021Computer networks · 11 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 RGCNet: Riemannian graph convolutional networks for end-to-end smart contract vulnerability detection
abstract
Frequent security issues with smart contract vulnerabilities have become a pressing challenge in the industry. Conventional program analysis methods lack flexibility and extensibility, leading to high false positive rates. Deep learning approaches are emerging as a new trend to address this issue. Compared to other neural networks, graph convolutional networks can better capture the structural and logical information of smart contracts. However, existing methods do not fully consider the scale-free characteristics of smart contracts and fail to leverage their complex hierarchical structures and semantic information. Therefore, we develop an end-to-end vulnerability detection framework using Riemannian Graph Convolutional Networks (RGCNet). We first construct smart contract graphs that are rich in semantic and structural information. Next, we learn features of the smart contract graph in the Riemannian manifold, thereby better reflecting its actual topology. Simultaneously, the word embedding network extracts semantic features, forming an end-to-end network where modules promote one another. Extensive experiments are conducted on three vulnerabilities using real-world smart contracts. The results show that the proposed approach exhibits superior performance over state-of-the-art methodologies in terms of accuracy, precision, and recall.
Yaoxin Chen, Haiming Zhu, Qicong Wang, Maozhen Li 0001
Neurocomputing6
2026 Lightweight AI-driven traffic forecasting and shaping for 6G LEO satellite networks
Mingji Dong, Maozhen Li 0001, Zuqing Zhu
Neurocomputing5
2026 Fine-tuning CLIP with mixture of experts and cross-modal alignment via contrastive learning for multimodal sentiment analysis
Fengjun Zhou, Xueqiang Gao, Zhiquan Feng, Maozhen Li 0001
Neurocomputing6
2026 Auxiliary domain joint adaptation and selection for cross-domain few-shot object detection
Nianyin Zeng, Zerui Cheng, Da Teng, Peishu Wu, Maozhen Li 0001
Neurocomputing6
2026 Vendor-Independent Design Space Exploration and Resource Optimization Framework for 3-D Networks-on-Chip Using Hypergraph-Genetic Algorithm Integration
abstract
This paper presents a novel methodology for design space exploration and resource optimisation of three-dimensional Networks-on-Chip (3D NoC) architectures using hypergraph modelling and genetic algorithms. The proposed approach combines mathematical rigour with evolutionary search capabilities to efficiently explore the vast design space of 3D NoC configurations, providing a vendor-independent solution for NoC architects. The key contribution is the development of Performance-Cost-Ratio (PCR) functions that enable quantitative evaluation of different topologies and routing algorithms, extended to include power and thermal considerations with dynamic adaptation mechanisms for runtime traffic variations. Validation through four compute-intensive use cases demonstrates significant improvements, with optimised architectures achieving up to 33% reduction in latency, 40% increase in throughput, and 30% reduction in power consumption compared to baseline implementations. Validation against published silicon implementations shows 92-96% correlation accuracy, confirming the framework’s practical applicability for developing efficient and scalable NoC solutions as processor designs advance towards kilo-core scales and beyond.
Ahmed Al-Alousi, Maozhen Li 0001, Hongying Meng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Enhanced air pollution spatiotemporal forecast model using frequency domain convolution and attention mechanism
Haiwei Yang, Ru Yang 0001, Ling Ding 0003, Shiqiang Du, Maozhen Li 0001, Bo Zhang 0004
Eng. Appl. Artif. Intell.5
2025 Reinforcement learning-based secure training for adversarial defense in graph neural networks
abstract
The security of Graph Neural Networks (GNNs) is crucial for ensuring the reliability and protection of the systems they are integrated within real-world applications. However, current approaches lack the ability to prevent GNNs from learning high-risk information, including edges, nodes, convolutions, etc. In this paper, we propose a secure GNN learning framework called Reinforcement Learning-based Secure Training Algorithm . We first introduce a model conversion technique that transforms the training process of GNNs into a verifiable Markov Decision Process model. To maintain the security of model we employ Deep Q-Learning algorithm to prevent high-risk information messages. Additionally, to verify whether the strategy derived from Deep Q-Learning algorithm meets safety requirements, we design a model transformation algorithm that converts MDPs into probabilistic verification models, thereby ensuring our method’s security through formal verification tools. The effectiveness and feasibility of our proposed method are demonstrated by achieving a 6.4% improvement in average accuracy on open-source datasets under adversarial attack graphs.
Dongdong An, Yi Yang 0001, Hongda Qi, Maozhen Li 0001
Neurocomputing7
2025 Piecewise convolutional neural network relation extraction with self-attention mechanism
Bo Zhang 0004, Kehao Liu, Ru Yang 0001, Maozhen Li 0001
Pattern Recognit.5
2024 SAC-based UAV mobile edge computing for energy minimization and secure data transmission
Xu Zhao 0005, Yichuan Wu, Maozhen Li 0001
Ad Hoc Networks5
2024 Normalizing flow based uncertainty estimation for deep regression analysis
abstract
Uncertainty estimation is a critical component of building safe and reliable machine learning models. Accurate estimation of uncertainties is essential for identifying and mitigating potential risks and ensuring that machine learning systems operate reliably in real-world scenarios. Various approaches, such as ensemble and Bayesian neural networks have been developed by sampling probability predictions from submodels, which is computatinally expensive. At present, these techniques are incapable of precisely delineating the boundary separating in-distribution (ID) and out-of-distribution (OOD) data. To fill up this research gap, this paper presents a normalizing flow based framework to directly predict parameters of prior distributions over the probability with a neural network, the proposed model is able to effectively differentiate between ID and OOD data in regression problems. The posterior distributions learned by the model precisely represent uncertainties for OOD data based solely on ID data, without the need for OOD data during training. This approach has shown promising results in a number of applications, including image depth estimation and image adversarial attacks.
Baobing Zhang, Wanxin Sui, Zhengwen Huang, Maozhen Li 0001, Man Qi
Neurocomputing4
2024 Federated deep reinforcement learning for task offloading and resource allocation in mobile edge computing-assisted vehicular networks
Xu Zhao 0005, Yichuan Wu, Maozhen Li 0001
J. Netw. Comput. Appl.5
2024 Modeling Group Opinion Evolution on Online Social Networks: A Gravitational Field Perspective
abstract
The research on group behavior is effective for establishing a good network environment since people in social networks tend to form groups spontaneously. Most studies on group behavior on online social networks assume that all individuals are reduced to one cluster, ignoring the existence of potential clusters and their importance in group opinion dynamics. This article introduces a novel group-gravitational field (GGF) model to investigate the opinion evolution based on group behavior by the following aspects: 1) the GGF model reduces a cluster in the social network into a charge and the whole network into a gravitational field; 2) the GGF model calculates the initial influence of a cluster according to the topology information and further constructs a gravity matrix of the network based on the Coulomb law; and 3) opinion-leader clusters exert the internal field force on common opinion clusters inside the gravitational field. The GGF model simulates the evolution of opinions among clusters in a network and studies the law of group behavior according to the influence between clusters based on Coulomb’s law. Experiments on real social networks verify that the GGF model enhances the speed of opinion evolution significantly. The simulation experiments indicate that the existence of clusters promotes the rapid convergence of opinions, a gathering of followers influences information dissemination in social networks, and the GGF model fits the reality better. This article provides a new approach to network supervision and control.
Meizi Li, Xinyi Zhang 0006, Maozhen Li 0001, Yunwen Chen, Yanhong Bai, Bo Zhang 0004, Ru Yang 0001
IEEE Trans. Comput. Soc. Syst.3
2024 Cross-Block Sparse Class Token Contrast for Weakly Supervised Semantic Segmentation
abstract
Most existing Vision Transformer-based frameworks for weakly supervised semantic segmentation utilize class activation maps to generate pseudo masks. Although it mitigates the class-agnostic issue, this approach still suffers from misclassification and noise in segmentation results. To overcome these limitations, we propose an attention-based framework named Cross-block Sparse Class Token Contrast (CB-SCTC), which incorporates Dynamic Sparse Attention module (DSA) and Cross-block Class Token Contrast scheme (CB-CTC). Specifically, the proposed Cross-block Class Token Contrast scheme forces diversity between the final class tokens by learning from the lower similarity of the class tokens in the relatively shallower blocks. Moreover, the Dynamic Sparse Attention module is designed to post-process the output from the softmax function in the attention mechanism to reduce noise. Extensive experiments prove the proposed framework is a valid alternative to class activation maps. Our framework demonstrates competitive mIoU scores on the PASCAL VOC 2012(val:75.5%, test:75.2%) and MS COCO 2014 dataset(val:46.9%). Our code is available athttps://github.com/Jingfeng-Tang/CB-SCTC.
Keyang Cheng, Jingfeng Tang, Hongjian Gu, Maozhen Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Learning Transactional Behavioral Representations for Credit Card Fraud Detection
abstract
Credit card fraud detection is a challenging task since fraudulent actions are hidden in massive legitimate behaviors. This work aims to learn a new representation for each transaction record based on the historical transactions of users in order to capture fraudulent patterns accurately and, thus, automatically detect a fraudulent transaction. We propose a novel model by improving long short-term memory with a time-aware gate that can capture the behavioral changes caused by consecutive transactions of users. A current-historical attention module is designed to build up connections between current and historical transactional behaviors, which enables the model to capture behavioral periodicity. An interaction module is designed to learn comprehensive and rational behavioral representations. To validate the effectiveness of the learned behavioral representations, experiments are conducted on a large real-world transaction dataset provided to us by a financial company in China, as well as a public dataset. Experimental results and the visualization of the learned representations illustrate that our method delivers a clear distinction between legitimate behaviors and fraudulent ones, and achieves better fraud detection performance compared with the state-of-the-art methods.
Yu Xie 0019, Guanjun Liu, ChunGang Yan, Changjun Jiang 0002, MengChu Zhou, Maozhen Li 0001
IEEE Trans. Neural Networks Learn. Syst.6
2023 Task offloading strategy and scheduling optimization for internet of vehicles based on deep reinforcement learning
Xu Zhao 0005, Mingzhen Liu, Maozhen Li 0001
Ad Hoc Networks3
2023 Corrigendum to 'Research on Lightweight Anomaly Detection of Multimedia Traffic in Edge Computing' Computer & Security, 111(2021) 102463
Xu Zhao 0005, Guangqiu Huang, Maozhen Li 0001
Comput. Secur.5
2023 Air pollutant diffusion trend prediction based on deep learning for targeted season - North China as an example
Bo Zhang 0004, Zhihao Wang 0005, Yunjie Lu, Maozhen Li 0001, Ru Yang 0001, Jianguo Pan, Zuliang Kou
Expert Syst. Appl.4
2023 Cold-start item recommendation for representation learning based on heterogeneous information networks with fusion side information
Meizi Li, Weiqiao Que, Ziyao Geng, Maozhen Li 0001, Zuliang Kou, Jisheng Chen, Bo Zhang 0004
Future Gener. Comput. Syst.4
2023 Sonar image garbage detection via global despeckling and dynamic attention graph optimization
Keyang Cheng, Liuyang Yan, Yi Ding 0001, Maozhen Li 0001, Humaira abdul Ghafoor
Neurocomputing5
2023 A spatial correlation prediction model of urban PM2.5 concentration based on deconvolution and LSTM
abstract
Precise prediction of air pollutants can effectively reducre the occurrence of heavy pollution incidents. With the current surge of massive data, deep learning appears to be a promising technique to achieve dynamic prediction of air pollutant concentration from both the spatial and temporal dimensions. This paper presents Dev-LSTM, a prediction model building on deconvolution and LSTM. The novelty of Dev-LSTM lies in its capability to fully extract the spatial feature correlation of air pollutant concentration data, preventing the excessive loss of information caused by traditional convolution. At the same time, the feature associations in the time dimension are mined to produce accurate prediction results. Experimental results show that Dev-LSTM outperforms traditional prediction models on a variety of indicators.
Bo Zhang 0004, Ruihan Yong, Guojian Zou, Ru Yang 0001, Jianguo Pan, Maozhen Li 0001
Neurocomputing7
2023 Implicit Negative Link Prediction With a Network Topology Perspective
abstract
Sign prediction in signed social networks is a new research direction in the field of social relation mining, which reveals underlying links between users. Traditional sign prediction research focuses on the prediction of positive signs and neglects the mining of potential implicit links, and there is little research on negative sign prediction. To address these problems, we propose a two-stage model that uses implicit link detection and link sign prediction. First, we use the preference attachment closeness degree (PACD) to predict possible implicit links by adding a measure of relationship closeness to the traditional link prediction algorithm (PA). Next, we propose a negative link sign prediction (Ne-LP) method to predict relation types through multidimensional negative sign-related features, including those of nodes, user similarity, and structural balance, and merge them by a logistic regression model. Finally, we evaluate PACD and Ne-LP through extensive experiments on three real-world social network datasets, whose results demonstrate that the method can effectively mine implicit relations and accurately predict negative links.
Bo Zhang 0004, Wenqing Liu, Ru Yang 0001, Maozhen Li 0001
IEEE Trans. Comput. Soc. Syst.5
2023 Logical Topology Inference via CPGCN Joint Optimizing With Pedestrian Re-Id
abstract
With the rise of artificial intelligence, deep learning has become the main research method of pedestrian recognition re-identification (re-id). However, most of the existing researches usually just determine the retrieval order based on the geographical location of cameras, which ignore the spatio-temporal logic characteristics of pedestrian flow. Furthermore, most of these methods rely on common object detection to detect and match pedestrians directly, which will separate the logical connection between videos from different cameras. In this research, a novel pedestrian re-identification model assisted by logical topological inference is proposed, which includes: 1) a joint optimization mechanism of pedestrian re-identification and multicamera logical topology inference, which makes the multicamera logical topology provides the retrieval order and the confidence for re-identification. And meanwhile, the results of pedestrian re-identification as a feedback modify logical topological inference; 2) a dynamic spatio-temporal information driving logical topology inference method via conditional probability graph convolution network (CPGCN) with random forest-based transition activation mechanism (RF-TAM) is proposed, which focuses on the pedestrian's walking direction at different moments; and 3) a pedestrian group cluster graph convolution network (GC-GCN) is designed to measure the correlation between embedded pedestrian features. Some experimental analyses and real scene experiments on datasets CUHK-SYSU, PRW, SLP, and UJS-reID indicate that the designed model can achieve a better logical topology inference with an accuracy of 87.3% and achieve the top-1 accuracy of 77.4% and the mAP accuracy of 74.3% for pedestrian re-identification.
Keyang Cheng, Qing Liu 0015, Rabia Tahir, Liangmin Wang 0001, Maozhen Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 CEModule: A Computation Efficient Module for Lightweight Convolutional Neural Networks
abstract
Lightweight convolutional neural networks (CNNs) rely heavily on the design of lightweight convolutional modules (LCMs). For an LCM, lightweight design based on repetitive feature maps (LoR) is currently one of the most effective approaches. An LoR mainly involves an extraction of feature maps from convolutional layers (CE) and feature map regeneration through cheap operations (RO). However, existing LoR approaches carry out lightweight improvements only from the aspect of RO but ignore the problems of poor generalization, low stability, and high computation workload incurred in the CE part. To alleviate these problems, this article introduces the concept of key features from a CNN model interpretation perspective. Subsequently, it presents a novel LCM, namely CEModule, focusing on the CE part. CEModule increases the number of key features to maintain a high level of accuracy in classification. In the meantime, CEModule employs a group convolution strategy to reduce floating-point operations (FLOPs) incurred in the training process. Finally, this article brings forth a dynamic adaptation algorithm ( α -DAM) to enhance the generalization of CEModule-enabled lightweight CNN models, including the developed CENet in dealing with datasets of different scales. Compared with the state-of-the-art results, CEModule reduces FLOPs by up to 54% on CIFAR-10 while maintaining a similar level of accuracy in classification. On ImageNet, CENet increases accuracy by 1.2% following the same FLOPs and training strategies.
Maozhen Li 0001, Changjun Jiang 0002, Guanjun Liu
IEEE Trans. Neural Networks Learn. Syst.2
2023 An Alternating-Direction-Method of Multipliers-Incorporated Approach to Symmetric Non-Negative Latent Factor Analysis
abstract
Large-scale undirected weighted networks are frequently encountered in big-data-related applications concerning interactions among a large unique set of entities. Such a network can be described by a Symmetric, High-Dimensional, and Incomplete (SHDI) matrix whose symmetry and incompleteness should be addressed with care. However, existing models fail in either correctly representing its symmetry or efficiently handling its incomplete data. For addressing this critical issue, this study proposes an Alternating-Direction-Method of Multipliers (ADMM)-based Symmetric Non-negative Latent Factor Analysis (ASNL) model. It adopts fourfold ideas: 1) implementing the data density-oriented modeling for efficiently representing an SHDI matrix's incomplete and imbalanced data; 2) separating the non-negative constraints from the decision parameters to avoid truncations during the training process; 3) incorporating the ADMM principle into its learning scheme for fast model convergence; and 4) parallelizing the training process with load balance considerations for high efficiency. Empirical studies on four SHDI matrices demonstrate that ASNL significantly outperforms several state-of-the-art models in both prediction accuracy for missing data of an SHDI and computational efficiency. It is a promising model for handling large-scale undirected networks raised in real applications.
Xin Luo 0001, Yurong Zhong, Zidong Wang 0001, Maozhen Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Chinese named-entity recognition via self-attention mechanism and position-aware influence propagation embedding
Bo Zhang 0004, Kehao Liu, Maozhen Li 0001, Jianguo Pan
Data Knowl. Eng.4
2022 RCL-Learning: ResNet and convolutional long short-term memory-based spatiotemporal air pollutant concentration prediction model
Bo Zhang 0004, Guojian Zou, Dongming Qin, Hongwei Mao, Maozhen Li 0001
Expert Syst. Appl.6
2022 Task offloading of cooperative intrusion detection system based on Deep Q Network in mobile edge computing
Xu Zhao 0005, Guangqiu Huang, Maozhen Li 0001
Expert Syst. Appl.5
2022 Generating self-attention activation maps for visual interpretations of convolutional neural networks
Maozhen Li 0001
Neurocomputing2
2022 SKG-Learning: a deep learning model for sentiment knowledge graph construction in social networks
Bo Zhang 0004, Maozhen Li 0001, Meizi Li
Neural Comput. Appl.4
2021 Research on lightweight anomaly detection of multimedia traffic in edge computing
Xu Zhao 0005, Guangqiu Huang, Maozhen Li 0001
Comput. Secur.5
2021 Explaining the black-box model: A survey of local interpretation methods for deep neural networks
Siguang Li, ChunGang Yan, Maozhen Li 0001, Changjun Jiang 0002
Neurocomputing4
2021 Low load DIDS task scheduling based on Q-learning in edge computing environment
Xu Zhao 0005, Guangqiu Huang, Maozhen Li 0001, Quanli Gao
J. Netw. Comput. Appl.4
2021 Hybridization between Neural Computing and Nature-Inspired Algorithms for a Sentence Similarity Model Based on the Attention Mechanism
abstract
Sentence similarity analysis has been applied in many fields, such as machine translation, the question answering system, and voice customer service. As a basic task of natural language processing, sentence similarity analysis plays an important role in many fields. The task of sentence similarity analysis is to establish a sentence similarity scoring model through multi-features. In previous work, researchers proposed a variety of models to deal with the calculation of sentence similarity. But these models do not consider the association information of sentence pairs, but only input sentence pairs into the model. In this article, we propose a sentence feature extraction model based on multi-feature attention. In addition, with the development of deep learning and the application of nature-inspired algorithms, researchers have proposed various hybrid algorithms that combine nature-inspired algorithms with neural networks. The hybrid algorithms not only solve the problem of decision-making based on multiple features but also improve the performance of the model. In the model, we use the attention mechanism to extract sentence features and assign weight. Then, the convolutional neural network is used to reduce the dimension of the matrix. In the training process, we integrate the firefly algorithm in the neural networks. The experimental results show that the accuracy of our model is 74.21%.
Peiying Zhang 0001, Xingzhe Huang, Maozhen Li 0001, Yu Xue 0003
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2020 An artificial bee colony-based kernel ridge regression for automobile insurance fraud identification
Chun Yan, Wei Liu 0051, Maozhen Li 0001, Jindong Chen
Neurocomputing4
2020 An analysis of generative adversarial networks and variants for image synthesis on MNIST dataset
Keyang Cheng, Rabia Tahir, Lubamba Kasangu Eric, Maozhen Li 0001
Multim. Tools Appl.4
2020 Hierarchical attributes learning for pedestrian re-identification via parallel stochastic gradient descent combined with momentum correction and adaptive learning rate
Keyang Cheng, Yongzhao Zhan 0001, Maozhen Li 0001, Kenli Li 0001
Neural Comput. Appl.4
2020 Improved TrAdaBoost and its Application to Transaction Fraud Detection
abstract
AdaBoost is a boosting-based machine learning method under the assumption that the data in training and testing sets have the same distribution and input feature space. It increases the weights of those instances that are wrongly classified in a training process. However, the assumption does not hold in many real-world data sets. Therefore, AdaBoost is extended to transfer AdaBoost (TrAdaBoost) that can effectively transfer knowledge from one domain to another. TrAdaBoost decreases the weights of those instances that belong to the source domain but are wrongly classified in a training process. It is more suitable for the case that data are of different distribution. Can it be improved for some special transfer scenarios, e.g., the data distribution changes slightly over time? We find that the distribution of credit card transaction data can change with the changes in the transaction behaviors of users, but the changes are slow most of the time. These changes are yet important for detecting transaction fraud since they result in a so-called concept drift problem. In order to make TrAdaBoost more suitable for the abovementioned case, we, thus, propose an improved TrAdaBoost (ITrAdaBoost) in this article. It updates (i.e., increases or decreases) the weight of a wrongly classified instance in a source domain according to the distribution distance from the instance to a target domain, and the calculation of distance is based on the theory of reproducing kernel Hilbert space. We do a series of experiments over five data sets, and the results illustrate the advantage of ITrAdaBoost.
Lutao Zheng, Guanjun Liu, ChunGang Yan, Changjun Jiang 0002, MengChu Zhou, Maozhen Li 0001
IEEE Trans. Comput. Soc. Syst.6
2020 Distributed Set-Membership Filtering for Multirate Systems Under the Round-Robin Scheduling Over Sensor Networks
abstract
In this paper, the distributed set-membership filtering problem is dealt with for a class of time-varying multirate systems in sensor networks with the communication protocol. For relieving the communication burden, the round-Robin (RR) protocol is exploited to orchestrate the transmission order, under which each sensor node only broadcasts partial information to both the corresponding local filter and its neighboring nodes. In order to meet the practical transmission requirements as well as reduce communication cost, the multirate strategy is proposed to govern the sampling/update rate of the plant, the sensors, and the filters. By means of the lifting technique, the augmented filtering error system is established with a unified sampling rate. The main purpose of the addressed filtering problem is to design a set of distributed filters such that, in the simultaneous presence of the RR transmission protocol, the multirate mechanism, and the bounded noises, there exists a certain ellipsoid that includes all possible error states at each time instant. Then, the desired distributed filter gains are obtained by minimizing such an ellipsoid in the sense of the minimum trace of the weighted matrix. The proposed resource-efficient filtering algorithm is of a recursive form, thereby facilitating the online implementation. A numerical simulation example is given to demonstrate the effectiveness of the proposed protocol-based distributed filter design method.
Shuai Liu 0007, Zidong Wang 0001, Guoliang Wei, Maozhen Li 0001
IEEE Trans. Cybern.4
2019 MSML: A Novel Multilevel Semi-Supervised Machine Learning Framework for Intrusion Detection System
abstract
Intrusion detection technology has received increasing attention in recent years. Many researchers have proposed various intrusion detection systems using machine learning (ML) methods. However, there are two noteworthy factors affecting the robustness of the model. One is the severe imbalance of network traffic in different categories and the other is the nonidentical distribution between training set and test set in feature space. This paper presents a multilevel intrusion detection model framework named multilevel semi-supervised ML (MSML) to address these issues. The MSML framework includes four modules: 1) pure cluster extraction; 2) pattern discovery; 3) fine-grained classification (FC); and 4) model updating. In the pure cluster module, we introduce an concept of “pure cluster” and propose a hierarchical semi-supervised k-means algorithm with an aim to find out all the pure clusters. In the pattern discovery module, we define the “unknown pattern” and apply cluster-based method aiming to find those unknown patterns. Then a test sample is sentenced to labeled known pattern or unlabeled unknown pattern. The FC module can achieves FC for those unknown pattern samples. The model updating module provides a mechanism for retraining. KDDCUP99 dataset is applied to evaluate MSML. Experimental results show that MSML is superior to other existing intrusion detection models in terms of overall accuracy, F1-score, and unknown pattern recognition capability.
Haipeng Yao, Danyang Fu, Peiying Zhang 0001, Maozhen Li 0001, Yunjie Liu 0001
IEEE Internet Things J.4
2019 Virtual network embedding based on modified genetic algorithm
Peiying Zhang 0001, Haipeng Yao, Maozhen Li 0001, Yunjie Liu 0001
Peer-to-Peer Netw. Appl.3
2018 Data-driven pedestrian re-identification based on hierarchical semantic representation
abstract
Summary Limited number of labeled data of surveillance video causes the training of supervised model for pedestrian re‐identification to be a difficult task. Besides, applications of pedestrian re‐identification in pedestrian retrieving and criminal tracking are limited because of the lack of semantic representation. In this paper, a data‐driven pedestrian re‐identification model based on hierarchical semantic representation is proposed, extracting essential features with unsupervised deep learning model and enhancing the semantic representation of features with hierarchical mid‐level ‘attributes’. Firstly, CNNs, well‐trained with the training process of CAEs, is used to extract features of horizontal blocks segmented from unlabeled pedestrian images. Then, these features are input into corresponding attribute classifiers to judge whether the pedestrian has the attributes. Lastly, with a table of ‘attributes‐classes mapping relations’, final result can be calculated. Under the premise of improving the accuracy of attribute classifier, our qualitative results show its clear advantages over the CHUK02, VIPeR, and i‐LIDS data set. Our proposed method is proved to effectively solve the problem of dependency on labeled data and lack of semantic expression, and it also significantly outperforms the state‐of‐the‐art in terms of accuracy and semanteme.
Keyang Cheng, Fangjie Xu, Man Qi, Maozhen Li 0001
Concurr. Comput. Pract. Exp.5
2018 Compressive tracking combined with sample weights and adaptive learning factor
abstract
Summary The compressive tracking algorithm introduces a compressive sensing theory into the target tracking field and produces good real‐time performance. However, the original compressive tracking algorithm ignores the fact that individual samples make different contributions to the target and that the learning factor is an empirical value that remains constant when the template is updated. Therefore, adverse factors (such as noise) and errors can infiltrate into the parametric model during the updating of the model when the object is obscured or receives interference from external factors, which will lead to tracking drift. In view of these problems, the weights of samples are given according to the distance between the sample and the target when training the Naive Bayesian classifier; hence, the stability of the tracking is improved. While the introduction of the Bhattacharyya coefficient is utilized to adjust the learning factor, this can help parameters to self‐adapt effectively. Experimental results show that the improved tracking algorithm has a better adaption to the target appearance variations, illumination changes, occlusion, and so on, and has better robustness than the original algorithm.
Yong Jin 0004, Hong-ying Li, Maozhen Li 0001
Concurr. Comput. Pract. Exp.5
2018 High performance deep learning techniques for big data analytics
abstract
High performance deep learning techniques for big data analyticsThe past few years have witnessed the momentum of big data, which continuously receives a growing effort from both the academia and industry.The challenge with big data is how to extract meaningful information and knowledge from it.Recently, deep learning 1 as an advanced machine learning technique has been widely taken up by the research community due to its multi-layered structure and effectiveness in extracting low-level features.Therefore, it is critical to explore advanced and high performance deep learning techniques for big data analytics especially for heterogeneous big data analytics including the process of data acquisition, feature extraction and representation, time series data analysis, knowledge representation, and semantic modeling.This special issue aimed to solicit high quality research articles and reviews reflecting the advances in deep learning for big data analytics of a high volume, velocity, variety, and veracity.Potential topics include parallel deep neural networks for data analytics of high volumes, high performance deep neural networks for data stream analytics, semantic modeling in big data analytics, optimized architectural designs of deep neural networks, parameter tuning in deep neural networks, and distributed deep neural networks for big data analytics.We received a large number of submissions for this special issue and conducted a rigor review process.The papers to be included in this special bring a wide scope of topics related to deep learning.A portion of the papers included in this special issue focus on traditional machine learning.Xu et al 2 present a hybrid interpretable model for ORCID
Maozhen Li 0001
Concurr. Comput. Pract. Exp.1
2018 Semantic enhanced deep learning for image classification
Siguang Li, Maozhen Li 0001
Concurr. Comput. Pract. Exp.2
2018 Semantic enhanced deep learning for image classification
abstract
Summary Extracting training data semantics mainly depends on manual annotation samples prepared in advance, and semantic mapping largely relies on the prior conditions in image classification. A majority of the data samples in reality are unlabeled data, which necessitates a more efficient way to extract image semantics in classification. Deep learning, as an advanced machine learning technique, has recently received significant attention from both the academia and industry. Deep learning has the potential to extract semantic features from images in an automated way to fill up the semantic gap. This paper presents 2 novel deep learning models (ie, stacked denoising auto‐encoder and convolution deep Boltzmann machine). A stacked denoising auto‐encoder builds on a stacked auto‐encoder but uses a denoising auto‐encoder to improve the learning quality. Convolution deep Boltzmann machine combines the convolutional neural network model with the deep Boltzmann machine. These 2 models are evaluated on CIFAR‐10 and STL‐10 datasets, respectively. Experimental results show that the proposed 2 models outperform the original models in terms of both precision and recall.
Siguang Li, Maozhen Li 0001, Changjun Jiang 0001
Concurr. Comput. Pract. Exp.2
2018 Parallelizing Hartley transform with Hadoop for fast detection of glass defects
abstract
Summary Glass defect detection methods based on grating projection can effectively detect various glass defects. The Fourier transform in general can be used as an online processing method for detecting glass defects based on fringe images. Processing fringe images with Fourier transform needs a large amount of computation as Fourier transform is a complex computation method. In order to reduce the amount of computation, an improved fringe image processing method based on the Hartley transform is proposed in this paper. To further speed up the computation process, the Hartley transform is parallelized with Hadoop, which is a major computing technology in support of data intensive applications. Experimental results show that the parallel Hartley transform significantly reduces computation complexity in detection of glass defects.
Maozhen Li 0001, Yong Jin 0004, Zhaoba Wang, Guodong Guo
Concurr. Comput. Pract. Exp.1
2018 Emotion detection from EEG recordings based on supervised and unsupervised dimension reduction
abstract
Summary In recent years, researchers have been trying to detect human emotions from recorded brain signals such as electroencephalogram (EEG) signals. However, due to the high levels of noise from the EEG recordings, a single feature alone cannot achieve good performance. A combination of distinct features is the key for automatic emotion detection. In this paper, we present a hybrid dimension feature reduction scheme using a total of 14 different features extracted from EEG recordings. The scheme combines these distinct features in the feature space using both supervised and unsupervised feature selection processes. Maximum Relevance Minimum Redundancy (mRMR) is applied to re‐order the combined features into max‐relevance with the labels and min‐redundancy of each feature. The generated features are further reduced with principal component analysis (PCA) for extracting the principal components. Experimental results show that the proposed work outperforms the state‐of‐art methods using the same settings in the publicly available DEAP data set.
Hongying Meng, Maozhen Li 0001, Fan Zhang 0101, Asoke K. Nandi
Concurr. Comput. Pract. Exp.3
2018 MapReduce-based parallel GEP algorithm for efficient function mining in big data applications
abstract
Summary Gene expression programming (GEP) algorithm is one of the most effective function mining algorithms in enabling the mathematical equation fitting for the input dataset. However, GEP algorithm encounters low efficiency issue in big data processing due to large overhead in its evolution when it handles the large‐scale data. In order to solve the issue, this paper presents two parallelized GEP algorithms using MapReduce. Based on data separation, the first algorithm aims at speeding up the large‐scale classification. However, it is lack of ability to output the mined equation explicitly. Therefore, based on the further improvements of the first algorithm, the second parallelized GEP algorithm aims at mining the equation efficiently and also outputs the equation explicitly and directly. The experimental results show that both algorithms are effective for processing large volume of data.
Yang Liu 0108, Chenxiao Ma, Lixiong Xu, Maozhen Li 0001
Concurr. Comput. Pract. Exp.5
2018 Reduced alignment based on Petri nets
abstract
Summary Alignment is the state‐of‐the‐art technique in conformance checking and becoming more important for the analysis of business processes. To improve the efficiency of alignment, a new alignment approach is presented based on Petri net models and traces. It takes artificial logs and models as an example to illustrate the procedure of the new alignment approach. The approach can generate an optimal alignment tree including all of the optimal alignments between the given trace and the Petri net model based on standard likelihood cost function. This paper gives the approach a specific and rigorous characterization. The approach is implemented on ProM as a plugin and has been evaluated using complex logs and models as a case study.
Yinhua Tian, Yuyue Du, Maozhen Li 0001, Qiang Hu 0002
Concurr. Comput. Pract. Exp.3
2018 Transfer learning-based online multiperson tracking with Gaussian process regression
abstract
Summary Most existing tracking‐by‐detection approaches are affected by abrupt pedestrian pose changes, lighting conditions, scale changes, and real‐time processing, which leads to issues such as detection errors and drifts. To deal with these issues, we present a novel multi‐person tracking framework by introducing a new Gaussian Process Regression based observation model, which learns in a semi‐supervised manner. The background information is taken into consideration to build the discriminative tracker, training samples are re‐weighted appropriately to ease the impact of the potential sample misalignment and noisy during model updating. Unlabeled samples from the current frame provide rich information, which is used for enhancing the tracking inference. Experimental results show that the proposed approach outperforms a number of state‐of‐the‐art methods on some benchmark datasets.
Baobing Zhang, Siguang Li, Zhengwen Huang, Babak H. Rahi, Qicong Wang, Maozhen Li 0001
Concurr. Comput. Pract. Exp.6
2018 Robust convolution kernel quantity determination based on corner radiation area adaptation
Yuyue Du, Siguang Li, Maozhen Li 0001
Neurocomputing4
2018 A novel reinforcement learning algorithm for virtual network embedding
Haipeng Yao, Maozhen Li 0001, Peiying Zhang 0001
Neurocomputing3
2018 CLOTHO: A Large-Scale Internet of Things-Based Crowd Evacuation Planning System for Disaster Management
abstract
In recent years, different kinds of natural hazards or man-made disasters happened that were diversified and difficult to control with heavy casualties. In this paper, we focus on the rapid and systematic evacuation of large-scale densities of people after disasters to reduce loss in an effective manner. The optimal evacuation planning is a key challenge and becomes a hotspot of research and development. We design our system based on an Internet of Things (IoT) scenario that utilizes a mobile cloud computing platform in order to develop the crowd lives oriented track and help optimization system (CLOTHO). CLOTHO is an evacuation planning system for large-scale densities of people in disasters. It includes the mobile terminal (IoT side) for data collection and the cloud backend system for storage and analytics. We build our solution upon a typical IoT/fog disaster management scenario and we propose an IoT application based on an evacuation planning algorithm that uses the artificial potential field (APF), which is the core of CLOTHO. APF is conceptualized as an IoT service, and can determine the direction of evacuation automatically according to the gradient direction of the potential field, suitable for rapid evacuation of large population. Based on APF, we propose an evacuation planning algorithm names as APF with relationship attraction (APF-RA). APF-RA guides the evacuees with relationship to move to the same shelter as much as possible, to calm evacuees and realize a more humanitarian evacuation. The experimental results show that CLOTHO (using APF and APF-RA) can effectively improve convergence rate, shorten the evacuation route length and evacuation time, and make the remaining capacity of the surrounding shelters well balanced.
Xiaolong Xu 0002, Lei Zhang 0001, Stelios Sotiriadis, Eleana Asimakopoulou, Maozhen Li 0001, Nik Bessis
IEEE Internet Things J.5
2018 NetworkAI: An Intelligent Network Architecture for Self-Learning Control Strategies in Software Defined Networks
abstract
The past few years have witnessed a wide deployment of software defined networks facilitating a separation of the control plane from the forwarding plane. However, the work on the control plane largely relies on a manual process in configuring forwarding strategies. To address this issue, this paper presents NetworkAI, an intelligent architecture for self-learning control strategies in software defined networking networks. NetworkAI employs deep reinforcement learning and incorporates network monitoring technologies, such as the in-band network telemetry to dynamically generate control policies and produces a near optimal decision. Simulation results demonstrated the effectiveness of NetworkAI.
Haipeng Yao, Tianle Mai, Xiaobin Xu 0004, Peiying Zhang 0001, Maozhen Li 0001, Yunjie Liu 0001
IEEE Internet Things J.5
2018 Chinese Open Relation Extraction and Knowledge Base Establishment
abstract
Named entity relation extraction is an important subject in the field of information extraction. Although many English extractors have achieved reasonable performance, an effective system for Chinese relation extraction remains undeveloped due to the lack of Chinese annotation corpora and the specificity of Chinese linguistics. Here, we summarize three kinds of unique but common phenomena in Chinese linguistics. In this article, we investigate unsupervised linguistics-based Chinese open relation extraction (ORE), which can automatically discover arbitrary relations without any manually labeled datasets, and research the establishment of a large-scale corpus. By mapping the entity relations into dependency-trees and considering the unique Chinese linguistic characteristics, we propose a novel unsupervised Chinese ORE model based on Dependency Semantic Normal Forms (DSNFs). This model imposes no restrictions on the relative positions among entities and relationships and achieves a high yield by extracting relations mediated by verbs or nouns and processing the parallel clauses. Empirical results from our model demonstrate the effectiveness of this method, which obtains stable performance on four heterogeneous datasets and achieves better precision and recall in comparison with several Chinese ORE systems. Furthermore, a large-scale knowledge base of entity and relation, called COER, is established and published by applying our method to web text, which conquers the trouble of lack of Chinese corpora.
Shengbin Jia, Shijia E, Maozhen Li 0001, Yang Xiang 0006
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2018 Schema Theory-Based Data Engineering in Gene Expression Programming for Big Data Analytics
abstract
Gene expression programming (GEP) is a data driven evolutionary technique that well suits for correlation mining. Parallel GEPs are proposed to speed up the evolution process using a cluster of computers or a computer with multiple CPU cores. However, the generation structure of chromosomes and the size of input data are two issues that tend to be neglected when speeding up GEP in evolution. To fill the research gap, this paper proposes three guiding principles to elaborate the computation nature of GEP in evolution based on an analysis of GEP schema theory. As a result, a novel data engineered GEP is developed which follows closely the generation structure of chromosomes in parallelization and considers the input data size in segmentation. Experimental results on two data sets with complementary features show that the data engineered GEP speeds up the evolution process significantly without loss of accuracy in data correlation mining. Based on the experimental tests, a computation model of the data engineered GEP is further developed to demonstrate its high scalability in dealing with potential big data using a large number of CPU cores.
Zhengwen Huang, Maozhen Li 0001, Christos Chousidis, Ali Mousavi 0001, Changjun Jiang 0002
IEEE Trans. Evol. Comput.2
2017 gSched: a resource aware Hadoop scheduler for heterogeneous cloud computing environments
abstract
Summary MapReduce has become a major programming model for data‐intensive applications in cloud computing environments. Hadoop, an open source implementation of MapReduce, has been adopted by an increasingly wide user community. However, Hadoop suffers from task scheduling performance degradation in heterogeneous contexts because of its homogeneous design focus. This paper presents gSched, a resource‐aware Hadoop scheduler that takes into account both the heterogeneity of computing resources and provisioning charges in task allocation in cloud computing environments. gSched is initially evaluated in an experimental Hadoop cluster and demonstrates enhanced performance compared with the default Hadoop scheduler. Further evaluations are conducted on the Amazon EC2 cloud that demonstrates the effectiveness of gSched in task allocation in heterogeneous cloud computing environments. Copyright © 2016 John Wiley & Sons, Ltd.
Godwin Caruana, Maozhen Li 0001, Man Qi, Mukhtaj Khan, Omer F. Rana
Concurr. Comput. Pract. Exp.2
2017 Optimizing hadoop parameter settings with gene expression programming guided PSO
abstract
Summary Hadoop MapReduce has become a major computing technology in support of big data analytics. The Hadoop framework has over 190 configuration parameters, and some of them can have a significant effect on the performance of a Hadoop job. Manually tuning the optimum or near optimum values of these parameters is a challenging task and also a time consuming process. This paper optimizes the performance of Hadoop by automatically tuning its configuration parameter settings. The proposed work first employs gene expression programming technique to build an objective function based on historical job running records, which represents a correlation among the Hadoop configuration parameters. It then employs particle swarm optimization technique, which makes use of the objective function to search for optimal or near optimal parameter settings. Experimental results show that the proposed work enhances the performance of Hadoop significantly compared with the default settings. Moreover, it outperforms both rule‐of‐thumb settings and the Starfish model in Hadoop performance optimization. © 2016 The Authors.Concurrency and Computation: Practice and ExperiencePublished by John Wiley & Sons Ltd.
Mukhtaj Khan, Zhengwen Huang, Maozhen Li 0001, Gareth A. Taylor, Mushtaq Khan
Concurr. Comput. Pract. Exp.3
2017 Internet of People
abstract
This Editorial introduces the articles to be included in the Special Issue on Internet of People. Internet of People (IoP) refers to digital connectivity of people through the Internet infrastructure forming a network of collective intelligence and stimulating interactive communication among people. The purpose of the special issue is to collate a selection of representative articles that were primarily presented at the IEEE International Conference on Internet of People (IoP 2015) on August 10 to 14, 2015, in Beijing, China. The special issue was also made open to public submissions for a wide inclusion. The scope of the special issue is broad and represents a multidisciplinary nature of IoP. It covers topics from the network enabling technologies at the physical layer to services at the application layer. It is reported that around 40% of the world population has an Internet connection today, and the number of Internet users reached 3 billion in 2014. The sheer size of the Internet population has led to big data challenges of volume, velocity, variety, and veracity of the digital data generated. For this purpose, this special issue also includes a few articles on high performance computing techniques that can speed up the computation process in analyzing big volumes of data. Mobile and wireless networks facilitate the communication among people. In the work of Huang et al. 1 they present a multiple-user partner selection algorithm to improve the overall secrecy rate of the network. Zhong et al. 2 focus on key management for multiple groups of the IoP using multicasting. A novel area-based multiple group key management scheme is proposed to facilitate the movement of mobile users in wireless communication networks with minimized communication overhead. On the basis of a self-organizing feature map neural network model, Yao et al. 3 present wireless local area network interference self-optimization method to quickly locate the fault access point and optimize the network performance to smoothen the communication process of people. Neural networks are also used in the work of Jin et al. 4 for online recognition of glass defects. Crowdsourcing enables IoP by soliciting contributions from a large group of people connected by the Internet. Wang et al. 5 conduct a survey on mobile crowdsourcing from the aspects of real-time and location-sensitive crowdsourced tasks. Related research challenges and possible solutions are discussed. Shao et al. 6 use crowdsensing in vehicular networks to predict the traffic condition in intelligent transportation systems. This solves the problem of traditional approaches being inefficiency ineffectiveness in data uploading and usage. Guo et al. 7 extract information from external-related attributes and improve latent Dirichlet allocation to build a topic mining model to facilitate services in telephone call centers. Knowledge diffusion as a component of crowdsourcing also plays an increasing role in modern communication networks and social networks. Zhang et al. 8 study the topological structure and research a new knowledge diffusion model taking into account both learning and forgetting attributes. The results from this research reveal that the social networks with a high degree of heterogeneity well suit for knowledge diffusion. In a mobile environment in IoP, network latency would have a significant impact on the communication of mobile services. For this purpose, Ding et al. 9 target at service composition so that mobile services can be optimized and provisioned to users with low communication latency. In the work of Gopalakrishna et al. 10 they assess relevance in cyber-physical systems. For this purpose, a new metric called relevance score is proposed for evaluation of a number of machine learning techniques. Big data has received a momentum from both academia and industry since the US government announced the big data initiative in 2012. MapReduce11-13 has become a major computing model in support of big data applications especially in dealing with data of a huge volume. Hadoop,14 which is an open source implementation of MapReduce, has been widely used in developing MapReduce applications. However, Hadoop does not have a sophisticated scheme in job scheduling. For this purpose, Liu et al. 15 present a dynamic load balancing algorithm on the basis of sliding windows with an aim to target at heterogeneous Hadoop cluster systems. Hadoop has over 190 parameters for user to configure. It has become a challenging issue for user to optimize the performance of Hadoop through a manual process of tuning these parameters. For this purpose, Khan et al. 16 use gene expression programming to dynamically mine the correlation of Hadoop parameters and further use particle swarm optimization technique to optimize the parameter settings to enhance the performance of Hadoop. Compared with the default Hadoop settings, this work can achieve 50% faster in Hadoop performance. On the basis of the MapReduce model, Cheng et al. 17 develop a MapReduce style parallel and distributed deep convolutional neural network for person re-identification in surveillance videos. Qi18 presents a parallel motion detection algorithm using a cluster of inexpensive computing nodes. Zhao et al. 19 parallelize the work on anomalous subgraph detection using Spark,20 an in-memory fast processing technology that can be deployed on MapReduce. Finally, Ren et al. 21 look at job scheduling problems in virtualized cloud environments from the aspect of power consumption. We hope that the perspectives presented in this special issue would be of a great interest to the readers. We also expect the readers to contribute to this exciting and fast growing research area. We would like to thank Professor Geoffrey Fox, the editor-in-chief of Concurrency and Computation: Practice and Experience for his timely advice on this special issue. A big thanks also goes to Rizza Mostar-Salterio, the production editor of Wiley for her great support in publication of the special issue.
Maozhen Li 0001
Concurr. Comput. Pract. Exp.1
2017 Sparse representations based distributed attribute learning for person re-identification
Keyang Cheng, Kaifa Hui, Yongzhao Zhan 0001, Maozhen Li 0001
Multim. Tools Appl.4
2016 Explore the Brain Response to Naturalistic and Continuous Music Using EEG Phase Characteristics
Jie Li 0016, Hongfei Ji, Rong Gu 0003, Lusong Hou, Qiang Wu 0009, Rongrong Lu, Maozhen Li 0001
ICIC (1)8
2016 A MapReduce-based parallel K-means clustering for large-scale CIM data verification
abstract
Summary The Common Information Model (CIM) has been heavily used in electric power grids for data exchange among a number of auxiliary systems such as communication systems, monitoring systems, and marketing systems. With a rapid deployment of digitalized devices in electric power networks, the volume of data continuously grows, which makes verification of CIM data a challenging issue. This paper presents a parallelK‐meansclustering algorithm for large‐scale CIM data verification. The parallelK‐meansbuilds on the MapReduce computing model which has been widely taken up by the community in dealing with data‐intensive applications. A genetic algorithm‐based load‐balancing scheme is designed to balance the workloads among the heterogeneous computing nodes for a further improvement in computation efficiency. The performance of the parallelK‐meansis initially evaluated in a small‐scale in‐house MapReduce cluster and subsequently evaluated in a commercial cloud computing platform. Finally, the parallelK‐meansis evaluated in large‐scale simulated MapReduce environments. Both the experimental and simulation results show that the parallelK‐meansreduces the CIM data‐verification time significantly compared with the sequentialK‐meansclustering, while generating a high level of precision in data verification. Copyright © 2015 John Wiley & Sons, Ltd.
Chuang Deng, Yang Liu 0010, Lixiong Xu, Junyong Liu, Siguang Li, Maozhen Li 0001
Concurr. Comput. Pract. Exp.7
2016 Preface
Kenli Li 0001, Yong Liu 0012, Maozhen Li 0001
Int. J. Pattern Recognit. Artif. Intell.3
2016 A Resource Aware MapReduce Based Parallel SVM for Large Scale Image Classifications
Wenming Guo, Nasullah Khalid Alham, Yang Liu 0010, Maozhen Li 0001, Man Qi
Neural Process. Lett.4
2016 Hadoop Performance Modeling for Job Estimation and Resource Provisioning
abstract
MapReduce has become a major computing model for data intensive applications. Hadoop, an open source implementation of MapReduce, has been adopted by an increasingly growing user community. Cloud computing service providers such as Amazon EC2 Cloud offer the opportunities for Hadoop users to lease a certain amount of resources and pay for their use. However, a key challenge is that cloud service providers do not have a resource provisioning mechanism to satisfy user jobs with deadline requirements. Currently, it is solely the user's responsibility to estimate the required amount of resources for running a job in the cloud. This paper presents a Hadoop job performance model that accurately estimates job completion time and further provisions the required amount of resources for a job to be completed within a deadline. The proposed model builds on historical job execution records and employs Locally Weighted Linear Regression (LWLR) technique to estimate the execution time of a job. Furthermore, it employs Lagrange Multipliers technique for resource provisioning to satisfy jobs with deadline requirements. The proposed model is initially evaluated on an in-house Hadoop cluster and subsequently evaluated in the Amazon EC2 Cloud. Experimental results show that the accuracy of the proposed model in job execution estimation is in the range of 94.97 and 95.51 percent, and jobs are completed within the required deadlines following on the resource provisioning scheme of the proposed model.
Mukhtaj Khan, Yong Jin 0004, Maozhen Li 0001, Yang Xiang 0006, Changjun Jiang 0002
IEEE Trans. Parallel Distributed Syst.3
2014 Discovery of Rare Sequential Topic Patterns in Document Stream
abstract
Plain text documents created and distributed on the Internet are ever changing in various forms. Mining topics of these documents has significant applications in many domains. Most of the literature is devoted to topic modeling, while sequential patterns of topics in document streams are ignored. Moreover, traditional sequential pattern mining algorithms mainly focused on frequent patterns for deterministic data sets, and thus not suitable for document streams with topic uncertainty and rare patterns. In this paper, we formulate and handle the mining problem of rare Sequential Topic Patterns (STPs) for Internet document streams, which are rare on the whole but relatively often for specific users, so also interesting. Since this type of rare STPs reflects users’ specific behaviors, our work can be applied in many fields, such as personalized context-aware recommendation and real-time monitoring on abnormal user behaviors on the Internet. We propose a novel approach to discovering user-related rare STPs based on the temporal and probabilistic information of concerned topics. After extracting topics from documents by LDA and sorting the document stream into sessions for different users during different time periods, the proposed algorithms discover rare STPs by (1) mining STP candidates for each user through an efficient algorithm based on pattern-growth, and (2) generating user-related rare STPs by pattern rarity analysis. Experiments on both synthetic and real data sets show that our approach can discover interesting rare STPs very effectively and efficiently.
Zhongyi Hu 0004, Hongan Wang, Jiaqi Zhu 0001, Maozhen Li 0001, Ying Qiao 0001, Changzhi Deng
SDM4
2014 Parallelizing multiclass support vector machines for scalable image annotation
Nasullah Khalid Alham, Maozhen Li 0001, Yang Liu 0010
Neural Comput. Appl.2
2013 HSim: A MapReduce simulator in enabling Cloud Computing
Yang Liu 0010, Maozhen Li 0001, Nasullah Khalid Alham, Suhel Hammoud
Future Gener. Comput. Syst.2
2013 An ontology enhanced parallel SVM for scalable spam filter training
Godwin Caruana, Maozhen Li 0001, Yang Liu 0010
Neurocomputing2
2012 Enhancing list scheduling heuristics for dependent job scheduling in grid computing environments
Geoffrey Falzon, Maozhen Li 0001
J. Supercomput.2
2012 Enhancing genetic algorithms for dependent job scheduling in grid computing environments
Geoffrey Falzon, Maozhen Li 0001
J. Supercomput.2
2012 Dealing With Uncertain Entities in Ontology Alignment Using Rough Sets
abstract
Ontology alignment facilitates exchange of knowledge among heterogeneous data sources. Many approaches to ontology alignment use multiple similarity measures to map entities between ontologies. However, it remains a key challenge in dealing with uncertain entities for which the employed ontology alignment measures produce conflicting results on similarity of the mapped entities. This paper presents OARS, a rough-set based approach to ontology alignment which achieves a high degree of accuracy in situations where uncertainty arises because of the conflicting results generated by different similarity measures. OARS employs a combinational approach and considers both lexical and structural similarity measures. OARS is extensively evaluated with the benchmark ontologies of the ontology alignment evaluation initiative (OAEI) 2010, and performs best in the aspect of recall in comparison with a number of alignment systems while generating a comparable performance in precision.
Sadaqat Jan, Maozhen Li 0001, Hamed S. Al-Raweshidy, Ali Mousavi 0001, Man Qi
IEEE Trans. Syst. Man Cybern. Part C2
2011 FARM: file annotation and retrieval on mobile devices
Sadaqat Jan, Maozhen Li 0001, Hamed S. Al-Raweshidy
Pers. Ubiquitous Comput.2
2010 Optimizing peer selection in BitTorrent networks with genetic algorithms
Tiejun Wu, Maozhen Li 0001, Man Qi
Future Gener. Comput. Syst.2
2009 Distributed Indexing for Resource Discovery in P2P Networks
abstract
P2P networks facilitate people belonging to a community to share resources of interest. However, discovering resources in a large scale P2P network poses a number of challenges. Although distributed hash table (DHT) structured P2P networks have shown enhanced scalability in routing messages, they only support key based exact matches. This paper presents DIndex, a distributed indexing component that can be used in P2P networks in support of range queries. DIndex introduces the concept of search dimensions for partitioning a search space, and it organizes peer nodes in a three-layered structure. Experimental results show that, for aP2P network with N number of peers, the average number of hops per message is less than log(N).
Marco Hentschel, Maozhen Li 0001, Mahesh Ponraj, Man Qi
CCGRID2
2009 Facilitating resource discovery in grid environments with peer-to-peer structured tuple spaces
Maozhen Li 0001, Man Qi
Peer-to-Peer Netw. Appl.1
2009 A grouped P2P network for scalable grid information services
Vijay Sahota, Maozhen Li 0001, Mark A. Baker, Nick Antonopoulos
Peer-to-Peer Netw. Appl.2
2008 Automatically wrapping legacy software into services: A grid case study
Maozhen Li 0001, Bin Yu 0005, Man Qi, Nick Antonopoulos
Peer-to-Peer Netw. Appl.1
2008 Grid Service Discovery with Rough Sets
abstract
The computational grid is rapidly evolving into a service-oriented computing infrastructure that facilitates resource sharing and large-scale problem solving over the Internet. Service discovery becomes an issue of vital importance in utilizing grid facilities. This paper presents ROSSE, a Rough sets-based search engine for grid service discovery. Building on the Rough sets theory, ROSSE is novel in its capability to deal with the uncertainty of properties when matching services. In this way, ROSSE can discover the services that are most relevant to a service query from a functional point of view. Since functionally matched services may have distinct nonfunctional properties related to the quality of service (QoS), ROSSE introduces a QoS model to further filter matched services with their QoS values to maximize user satisfaction in service discovery. ROSSE is evaluated from the aspects of accuracy and efficiency in discovery of computing services.
Maozhen Li 0001, Bin Yu 0005, Omer F. Rana, Zidong Wang 0001
IEEE Trans. Knowl. Data Eng.1
2006 Service Matchmaking with Rough Sets
abstract
With the wide adoption of open grid services architecture (OGSA) and Web services resource framework (WSRF), the grid is emerging as a service-oriented computing infrastructure for engineers and scientists to solve data and computationally intensive problems. It is envisioned that computing resources in a future grid environment will be exposed as services. Service discovery becomes an issue of vital importance for a wider uptake of the grid. This paper presents RSSM, a rough sets based service matchmaking algorithm for service discovery with an aim to tolerate uncertainty in identifying service properties. The evaluation results show that the RSSM algorithm is more effective in service discovery compared with other mechanisms such as UDDI and OWLS.
Maozhen Li 0001, Bin Yu 0005, Chang Huang, Yong-Hua Song
CCGRID1
2006 Analysis of Interoperability Issues Between EGEE and VEGA Grid Infrastructures
Bartosz Kryza, Lukasz Skital, Jacek Kitowski, Maozhen Li 0001, Takebumi Itagaki
HPCC4
2006 PGGA: A predictable and grouped genetic algorithm for job scheduling
Maozhen Li 0001, Bin Yu 0005, Man Qi
Future Gener. Comput. Syst.1
2006 Stability analysis for stochastic Cohen-Grossberg neural networks with mixed time delays
abstract
In this letter, the global asymptotic stability analysis problem is considered for a class of stochastic Cohen-Grossberg neural networks with mixed time delays, which consist of both the discrete and distributed time delays. Based on an Lyapunov-Krasovskii functional and the stochastic stability analysis theory, a linear matrix inequality (LMI) approach is developed to derive several sufficient conditions guaranteeing the global asymptotic convergence of the equilibrium point in the mean square. It is shown that the addressed stochastic Cohen-Grossberg neural networks with mixed delays are globally asymptotically stable in the mean square if two LMIs are feasible, where the feasibility of LMIs can be readily checked by the Matlab LMI toolbox. It is also pointed out that the main results comprise some existing results as special cases. A numerical example is given to demonstrate the usefulness of the proposed global stability criteria.
Zidong Wang 0001, Yurong Liu, Maozhen Li 0001, Xiaohui Liu 0001
IEEE Trans. Neural Networks3
2004 SGrid: a service-oriented model for the Semantic Grid
Maozhen Li 0001, P. van Santen, David W. Walker, Omer F. Rana, Mark A. Baker
Future Gener. Comput. Syst.1
2004 Migrating legacy codes to distributed computing environments: a CORBA approach
Maozhen Li 0001, David W. Walker, Omer F. Rana, Coral Walker
Inf. Softw. Technol.1
2004 Leveraging legacy codes to distributed problem-solving environments: a Web services approach
abstract
Abstract This paper presents WSOWG, a Web‐services‐oriented wrapper generator for automatically wrapping non‐networked legacy codes as Web services for reuse in distributed problem‐solving environments. Using WSOWG, a finite element based computational fluid dynamics (CFD) legacy code has been wrapped as a Web service. A problem‐solving environment for simulating incompressible Navier–Stokes flows has also been implemented. A user makes use of the CFD service through a Web page without knowing the exact implementation of the service. In this way, a user's computing environment can be extended to a heterogeneous distributed computing environment. Performance evaluation shows that the overhead to invoke the CFD Web service generated by WSOWG using Simple Object Access Protocol (SOAP) and CORBA Internet Inter‐ORB Protocol (IIOP) is reasonable compared with that of invoking another CFD Web service manually wrapped from the CFD legacy code using SOAP only. Copyright © 2004 John Wiley & Sons, Ltd.
Maozhen Li 0001, Man Qi
Softw. Pract. Exp.1
2003 PortalLab: A Web Services Toolkit for Building Semantic Grid Portals
abstract
Grid is computer-based infrastructure that provides dependable, consistent, pervasive access to distributed resources. Built on top of a Grid, a Semantic Grid is a service-oriented infrastructure that provides a range of computation, information and knowledge services. A purpose of a Grid portal is to provide easy and seamless access to Grid heterogeneous resources and services through a Web-based user interface. This paper presents PortalLab, a Web Services oriented toolkit for designing, integrating and building Semantic Grid portals. Portals built from PortalLab are composed from a collection of reusable Web Services oriented portlets that are themselves semantic Grid services. Each portlet has a WSDL interface and a semantic registry defined in a domain ontology repository. The use of software agents assists end users in formulating domain problems, searching possible solutions (solvers) and submitting user tasks to the Grid. Multiple agents work in a peer-to-peer environment to allow users to access federated Grid services across different domains to improve fault tolerance and quality of service in user job submission and execution on the Grid. Since portlets are context independent, a PortalLab portal provides the ability to interoperate with different Grid systems at a portal level.
Maozhen Li 0001, P. van Santen, David W. Walker, Omer F. Rana, Mark A. Baker
CCGRID1
2003 MAPBOT: a Web based map information retrieval system
Maozhen Li 0001, Man Qi
Inf. Softw. Technol.1
2003 Engineering high-performance legacy codes as CORBA components for problem-solving environments
Maozhen Li 0001, David W. Walker, Omer F. Rana, Coral Walker, P. T. Williams, R. C. Ward
J. Parallel Distributed Comput.1
2001 Wrapping MPI-based legacy codes as Java/CORBA components
Maozhen Li 0001, Omer F. Rana, David W. Walker
Future Gener. Comput. Syst.1
2000 Implementing Problem Solving Environments for Computational Science (Research Note)
Omer F. Rana, Maozhen Li 0001, Matthew S. Shields, David W. Walker, David Golby
Euro-Par2
2000 PaDDMAS: Parallel and Distributed Data Mining Application Suite
abstract
Discovering complex associations, anomalies and patterns in distributed data sets is gaining popularity in a range of scientific, medical and business applications. Various algorithms are employed to perform data analysis within a domain, and range from statistical to machine learning and AI based techniques. Several issues need to be addressed however to scale such approaches to large data sets, particularly when these are applied to data distributed at various sites. As new analysis techniques are identified, the core tool set must enable easy integration of such analytical components. Similarly, results from an analysis engines must be sharable, to enable storage, visualisation or further analysis of results. We describe the architecture of PaDDMAS, a component based system for developing distributed data mining applications. PaDDMAS provides a tool set for combining pre-developed or custom components using a dataflow approach, with components performing analysis, data extraction or data management and translation. Each component is wrapped as a Java/CORBA object, and has an interface defined in XML. Components can be serial or parallel objects, and may be binary or contain a more complex internal structure. We demonstrate a prototype using a neural network analysis algorithm.
Omer F. Rana, David W. Walker, Maozhen Li 0001, Steven J. Lynden, Mike Ward
IPDPS3
2000 A Wrapper Generator for Wrapping High Performance Legacy Codes as Java/CORBA Components
abstract
This paper describes a Wrapper Generator for wrapping high performance legacy codes as Java/CORBA components for use in a distributed component-based problem- solving environment. Using the Wrapper Generator we ave automatically wrapped an MPI-based legacycode as a single CORBA object, and implemented a problem- solving environment for molecular dynamic simulations. Performance comparisons between runs of the CORBA object and the original legacy code on a cluster of workstations and on a parallel computer are also presented.
Maozhen Li 0001, Omer F. Rana, Matthew S. Shields, David W. Walker
SC1
2000 A Java/CORBA-based visual program composition environment for PSEs
abstract
A problem solving environment (PSE) is a complete, integrated computing environment for composing, compiling and running applications in a specific problem area or domain. A visual programming composition environment (VPCE) is described, which serves as a user interface for a PSE, and uses Java and CORBA to provide a framework of tools to enable the construction of scientific applications from components. The VPCE consists of a component repository, from which the user can select off-the-shelf or in-house components, a graphical composition area on which components can be combined, various tools that facilitate the configuration of components, the integration of legacy codes into components and the design and building of new components. The VPCE produces output using dataflow techniques in the form of a task graph, annotated with a performance model plus constraints for each component, expressed in XML. In addition, the VPCE supports a domain specific expert system based on JESS (Ernest Friedman-Hill, JESS: The Java Expert System Shell. See web site at: http://herzberg.ca.sandia.gov/jess/, 1999) to guide the user in component selection and to perform integrity checking. Copyright © 2000 John Wiley & Sons, Ltd.
Matthew S. Shields, Omer F. Rana, David W. Walker, Maozhen Li 0001, David Golby
Concurr. Pract. Exp.4
2000 The software architecture of a distributed problem-solving environment
abstract
This paper describes the functionality and software architecture of a generic problem-solving environment (PSE) for collaborative computational science and engineering. A PSE is designed to provide transparent access to heterogeneous distributed computing resources, and is intended to enhance research productivity by making it easier to construct, run, and analyze the results of computer simulations. Although implementation details are not discussed in depth, the role of software technologies such as CORBA, Java, and XML is outlined. An XML-based component model is presented. The main features of a Visual Component Composition Environment for software development, and an Intelligent Resource Management System for scheduling components, are described. Some prototype implementations of PSE applications are also presented. Copyright © 2000 John Wiley & Sons, Ltd.
David W. Walker, Maozhen Li 0001, Omer F. Rana, Matthew S. Shields, Coral Walker
Concurr. Pract. Exp.2