VLDB 2026 Research / reviewers in the wild / expert
Zili Zhang 0001
dblp:17/1185-1
· DBLP profile ↗
79ranked-venue papers
10as first author
13since 2021 · last 2025
0000-0002-8721-9333ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 26 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing decomposition-based hybrid models for forecasting multivariate and multi-source time series by federated transfer learningabstractTime series forecasting is a complex task that demands both accuracy and efficiency. Hybrid models have shown promising forecasting performance. These models integrate decomposition algorithms (e.g., the multivariate variational mode decomposition , MVMD) with individual models such as the classical long short-term memory network (LSTM). However, these models encounter challenges such as high time consumption in the model training process and data leakage issues when dealing with multivariate time series from different data owners (i.e., multi-source). To address these challenges, a new framework for hybrid models aimed at multivariate and multi-source time series (MMTS) forecasting, referred to as MVMD-x-FTL (e.g., MVMD-LSTM-FTL), has been proposed. In this framework, MVMD refers to the federated MVMD (Fed-MVMD) algorithm we proposed with the federated transfer learning (FTL) technique, while the x represents any selected individual model. The framework is evaluated on 8 real-world datasets from various domains such as network traffic, solar radiation, energy consumption, etc. Compared to the classical framework MVMD-x (i.e., MVMD-LSTM), the proposed MVMD-LSTM-FTL framework reduces the LSTM model fitting time by an average of 56.17% and mean absolute error (MAE), root mean square error (RMSE), normalized RMSE (NRMSE) for forecasting accuracy by up to 1.79%. And the privacy budget ɛ in Fed-MVMD has no significant impact on the final forecasting results. These experimental results demonstrate that our proposed framework effectively addresses the concerns of time consumption and privacy leakage for multivariate and multi-source time series forecasting tasks. Yonghou He, Zili Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2023 | Inertial projection neural network for nonconvex sparse signal recovery with prior information
Xiaohu Luo, Zili Zhang 0001 |
Frontiers Comput. Sci. | 2 |
| 2022 | Temporal Neighborhood Change Centrality for Important Node Identification in Temporal Networks
Langzhou He, Yi Wang 0045, Zili Zhang 0001 |
ICONIP (1) | 5 |
| 2022 | MEOD: A Robust Multi-stage Ensemble Model Based on Rank Aggregation and Stacking for Outlier Detection
Zhengchao Jiang, Fan Zhang 0094, Zili Zhang 0001 |
KSEM (3) | 5 |
| 2022 | Sparse Dense Transformer Network for Video Action Recognition
Xiaochun Qu, Jinye Ran, Zili Zhang 0001 |
KSEM (2) | 6 |
| 2022 | Identifying Multiple Influential Nodes for Complex Networks Based on Multi-agent Deep Reinforcement Learning
Shengzhou Kong, Langzhou He, Guilian Zhang, Zili Zhang 0001 |
PRICAI (3) | 5 |
| 2022 | Learning Spatial Fusion and Matching for Visual Object Tracking
Zili Zhang 0001 |
PRICAI (3) | 2 |
| 2022 | Source-Free Implicit Semantic Augmentation for Domain Adaptation
Zili Zhang 0001 |
PRICAI (2) | 2 |
| 2022 | CMAL: Cost-Effective Multi-Label Active Learning by Querying SubexamplesabstractMulti-label active learning (MAL) aims to learn an accurate multi-label classifier by selecting which examples (or example-label pairs) will be annotated and reducing query effort. MAL is a more complicated and expensive process than single-label active learning, due to one example can be associated with a set of non-exclusive labels and the annotator has to scrutinize the whole example and label space to provide correct annotations. Instead of scrutinizing the whole example for annotation, we may just examine some of its subexamples with respect to a label for annotation. In this way, we can not only save the annotation cost but also speedup the annotation process. Given this observation, we introduce CMAL, a two-stage Cost-effective MAL strategy (CMAL) by querying subexamples. CMAL first selects the most informative example-label pairs by leveraging uncertainty, label correlation and label space sparsity. Specifically, the uncertainty of a label to an example can be reduced if its correlated labels already annotated to the example, and its uncertainty can be reduced also if more examples annotated to this label. Next, CMAL greedily queries the most probable positive subexample-label pairs of the selected example-label pair. In addition, we propose rCMAL to account for the representative of examples to more reliably select example-label pairs in the first stage. Extensive experiments on multi-label datasets from diverse domains show that our proposed CMAL and rCMAL can better save the query cost than state-of-the-art MAL methods. The contribution of leveraging label correlation, label sparsity, and representative for saving cost is also confirmed. Guoxian Yu, Xia Chen 0004, Carlotta Domeniconi, Jun Wang 0035, Zhao Li 0007, Zili Zhang 0001, Xiangliang Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2021 | Progressive Text-to-Face Synthesis with Generative Adversarial NetworkabstractText-to-Face synthesis has considerable challenges and potentials in the field of public safety. Compared with the Text-to-Image synthesis models, the text descriptions of facial features are more complex and diverse. For the text embedding, most of the previous Text-to-Face synthesis models only deal with a single sentence containing several features of face images, and the generated images are vague and lack of details. In this paper, a novel Progressive Text-to-Face synthesis with Generative Adversarial Network (PFGAN) is proposed to generate natural face images from text descriptions. Firstly, a new text encoding method Convolution-Deconvolution Word Embedding LSTM (CDWE-BLSTM) is leveraged as the text encoder, which tackles more complex sentences and improves the accuracy of text encoding. Secondly, the PFGAN is composed of multiple generators and discriminators arranged in a tree-like structure. Furthermore, face images at multiple scales are progressively generated from different branches of the tree, corresponding to the same descriptions. images at multiple scales corresponding to the same scene are generated from different branches of the tree. By comparing with three existing Text-to-Face synthesis methods, extensive experiments demonstrate that the proposed PFGAN is very competitive in the IS (Inception Scores), FID (Frechet Inception Distance) and resolution of the generated face images. Xing Qiao, Yanghong Han, Zili Zhang 0001 |
FG | 4 |
| 2021 | A Novel Sigma-Lognormal Parameter Extractor for Online Signatures
Jianhuan Huang, Zili Zhang 0001 |
ICDAR (3) | 2 |
| 2021 | Correcting Large Knowledge Bases Using Guided Inductive Logic Learning Rules
Zili Zhang 0001 |
PRICAI (1) | 2 |
| 2021 | Discovering Multiple Co-Clusterings With Matrix FactorizationabstractClustering is a fundamental data exploration task which aims at discovering the hidden grouping structure in the data. The traditional clustering methods typically compute a single partition. However, there often exist different and equally meaningful clusterings in complex data. To solve this issue, multiple clustering approaches have emerged with the goal of exploring alternative clusterings from different perspectives. Existing solutions to this problem mainly focus on one-way clustering, that is, they cluster either the samples or the features. However, for many practical tasks, it is meaningful and desirable to explore alternative two-way clusterings (or co-clusterings), which capture not only the sample cluster structure but also the feature cluster structure. To tackle this interesting and unresolved task, we introduce an approach, called multiple co-clusterings (MultiCCs), to generate multiple alternative co-clusterings at the same time. MultiCC takes advantage of matrix tri-factorization to seek the co-clustering indicator matrices for samples and features and defines the row and column redundancy quantification terms to enforce diversity among co-clusterings based on these indicator matrices. After that, it integrates matrix tri-factorization and two nonredundancy terms into a unified objective function and gives an alternative optimization procedure to optimize the objective function. Extensive experimental results demonstrate that MultiCC performs significantly better than the existing multiple clustering methods. In addition, MultiCC can find out interesting co-clusters, which cannot be made by those comparing methods. Jun Wang 0035, Guoxian Yu, Carlotta Domeniconi, Zhiwen Yu 0002, Zili Zhang 0001 |
IEEE Trans. Cybern. | 6 |
| 2020 | Improved Performance of GANs via Integrating Gradient Penalty with Spectral Normalization
Hongwei Tan, Linyong Zhou, Zili Zhang 0001 |
KSEM (2) | 4 |
| 2020 | A Predictive-Reactive Approach with Genetic Programming and Cooperative Coevolution for the Uncertain Capacitated Arc Routing ProblemabstractThe uncertain capacitated arc routing problem is of great significance for its wide applications in the real world. In the uncertain capacitated arc routing problem, variables such as task demands and travel costs are realised in real time. This may cause the predefined solution to become ineffective and/or infeasible. There are two main challenges in solving this problem. One is to obtain a high-quality and robust baseline task sequence, and the other is to design an effective recourse policy to adjust the baseline task sequence when it becomes infeasible and/or ineffective during the execution. Existing studies typically only tackle one challenge (the other being addressed using a naive strategy). No existing work optimises the baseline task sequence and recourse policy simultaneously. To fill this gap, we propose a novel proactive-reactive approach, which represents a solution as a baseline task sequence and a recourse policy. The two components are optimised under a cooperative coevolution framework, in which the baseline task sequence is evolved by an estimation of distribution algorithm, and the recourse policy is evolved by genetic programming. The experimental results show that the proposed algorithm, called Solution-Policy Coevolver, significantly outperforms the state-of-the-art algorithms to the uncertain capacitated arc routing problem for the ugdb and uval benchmark instances. Through further analysis, we discovered that route failure is not always detrimental. Instead, in certain cases (e.g., when the vehicle is on the way back to the depot) allowing route failure can lead to better solutions. Yuxin Liu 0003, Yi Mei 0001, Mengjie Zhang 0001, Zili Zhang 0001 |
Evol. Comput. | 4 |
| 2019 | Multi-View Multi-Instance Multi-Label Learning Based on Collaborative Matrix FactorizationabstractMulti-view Multi-instance Multi-label Learning (M3L) deals with complex objects encompassing diverse instances, represented with different feature views, and annotated with multiple labels. Existing M3L solutions only partially explore the inter or intra relations between objects (or bags), instances, and labels, which can convey important contextual information for M3L. As such, they may have a compromised performance.\ In this paper, we propose a collaborative matrix factorization based solution called M3Lcmf. M3Lcmf first uses a heterogeneous network composed of nodes of bags, instances, and labels, to encode different types of relations via multiple relational data matrices. To preserve the intrinsic structure of the data matrices, M3Lcmf collaboratively factorizes them into low-rank matrices, explores the latent relationships between bags, instances, and labels, and selectively merges the data matrices. An aggregation scheme is further introduced to aggregate the instance-level labels into bag-level and to guide the factorization. An empirical study on benchmark datasets show that M3Lcmf outperforms other related competitive solutions both in the instance-level and bag-level prediction. Yuying Xing, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zili Zhang 0001, Maozu Guo 0001 |
AAAI | 5 |
| 2019 | Small-Scale Data Classification Based on Deep Forest
Meiyang Zhang, Zili Zhang 0001 |
KSEM (1) | 2 |
| 2019 | Sentinel Nodes Identification for Infectious Disease Surveillance on Temporal Social NetworksabstractActive surveillance, which aims at detecting and controlling infectious diseases at an early stage, is essential to prevent the spread of infections, protect people’s health, and promote social good. One difficult problem in active surveillance is how to intelligently sample a small group of nodes as sentinels from a large number of individuals for detecting the outbreaks of infectious diseases as early as possible. To sample sentinels, the existing methods depending on the global information about a social network are infeasible for mapping out social connections is time-consuming and inaccurate. Instead, some existing studies utilize local information about individuals’ connected neighbors to heuristically select sentinels. However, few of them take into account the temporal structure of social connections, which is believed to have a direct effect on the spread of infectious diseases. In this paper, we propose two temporal-network surveillance strategies for selecting sentinels based on the friendship paradox theory, a sociological theory describing a phenomenon in social networks that most people have fewer friends than their friends have. By simulating our strategies with three existing strategies based on the susceptible-infected (SI) model, the results show that our proposed 1stAN and 2ndRN strategies can detect the outbreak of infectious diseases earlier than the other strategies on the synthetic temporal network and two real-world temporal social networks, respectively. Jiachen Geng, Yuanxi Li 0003, Zili Zhang 0001 |
WI | 3 |
| 2019 | Noise-Resistant Statistical Traffic ClassificationabstractNetwork traffic classification plays a significant role in cyber security applications and management scenarios. Conventional statistical classification techniques rely on the assumption that clean labelled samples are available for building classification models. However, in the big data era, mislabelled training data commonly exist due to the introduction of new applications and lack of knowledge. Existing statistical traffic classification techniques do not address the problem of mislabelled training data, so their performance become poor in the presence of mislabelled training data. To meet this challenge, in this paper, we propose a new scheme, Noise-resistant Statistical Traffic Classification (NSTC), which incorporates the techniques of noise elimination and reliability estimation into traffic classification. NSTC estimates the reliability of the remaining training data before it builds a robust traffic classifier. Through a number of traffic classification experiments on two real-world traffic data sets, the results show that the new NSTC scheme can effectively address the problem of mislabelled training data. Compared with the state of the art methods, NSTC can significantly improve the classification performance in the context of big unclean data. Binfeng Wang, Jun Zhang 0010, Zili Zhang 0001, Lei Pan 0002, Yang Xiang 0001, Dawen Xia |
IEEE Trans. Big Data | 3 |
| 2018 | Cost Effective Multi-label Active Learning via Querying SubexamplesabstractMulti-label active learning addresses the scarce labeled example problem by querying the most valuable unlabeled examples, or example-label pairs, to achieve a better performance with limited query cost. Current multi-label active learning methods require the scrutiny of the whole example in order to obtain its annotation. In contrast, one can find positive evidence with respect to a label by examining specific patterns (i.e., subexample), rather than the whole example, thus making the annotation process more efficient. Based on this observation, we propose a novel two-stage cost effective multi-label active learning framework, called CMAL. In the first stage, a novel example-label pair selection strategy is introduced. Our strategy leverages label correlation and label space sparsity of multi-label examples to select the most uncertain example-label pairs. Specifically, the unknown relevant label of an example can be inferred from the correlated labels that are already assigned to the example, thus reducing the uncertainty of the unknown label. In addition, the larger the number of relevant examples of a particular label, the smaller the uncertainty of the label is. In the second stage, CMAL queries the most plausible positive subexample-label pairs of the selected example-label pairs. Comprehensive experiments on multi-label datasets collected from different domains demonstrate the effectiveness of our proposed approach on cost effective queries. We also show that leveraging label correlation and label sparsity contribute to saving costs. Xia Chen 0004, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zhao Li 0007, Zili Zhang 0001 |
ICDM | 6 |
| 2018 | Multiple Co-clusteringsabstractThe goal of multiple clusterings is to discover multiple independent ways of organizing a dataset into clusters. Current approaches to this problem just focus on one-way clustering. In many real-world applications, though, it's meaningful and desirable to explore alternative two-way clustering (or co-clusterings), where both samples and features are clustered. To tackle this challenge and unexplored problem, in this paper we introduce an approach, called Multiple Co-Clusterings (MultiCC), to discover non-redundant alternative co-clusterings. MultiCC makes use of matrix tri-factorization to optimize the sample-wise and feature-wise co-clustering indicator matrices, and introduces two non-redundancy terms to enforce diversity among co-clusterings. We then combine the objective of matrix tri-factorization and two non-redundancy terms into a unified objective function and introduce an iterative solution to optimize the function. Experimental results show that MultiCC outperforms existing multiple clustering methods, and it can find interesting co-clusters which cannot be discovered by current solutions. Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zhiwen Yu 0002, Zili Zhang 0001 |
ICDM | 6 |
| 2018 | Feature-Induced Partial Multi-label LearningabstractCurrent efforts on multi-label learning generally assume that the given labels of training instances are noise-free. However, obtaining noise-free labels is quite difficult and often impractical, and the presence of noisy labels may compromise the performance of multi-label learning. Partial multi-label learning (PML) addresses the scenario in which each instance is annotated with a set of candidate labels, of which only a subset corresponds to the ground-truth. The PML problem is more challenging than partial-label learning, since the latter assumes that only one label is valid and may ignore the correlation among candidate labels. To tackle the PML challenge, we introduce a feature induced PML approach called fPML, which simultaneously estimates noisy labels and trains multi-label classifiers. In particular, fPML simultaneously factorizes the observed instance-label association matrix and the instance-feature matrix into low-rank matrices to achieve coherent low-rank matrices from the label and the feature spaces, and a low-rank label correlation matrix as well. The low-rank approximation of the instance-label association matrix is leveraged to estimate the association confidence. To predict the labels of unlabeled instances, fPML learns a matrix that maps the instances to labels based on the estimated association confidence. An empirical study on public multi-label datasets with injected noisy labels, and on archived proteomic datasets, shows that fPML can more accurately identify noisy labels than related solutions, and consequently can achieve better performance on predicting labels of instances than competitive methods. Guoxian Yu, Xia Chen 0004, Carlotta Domeniconi, Jun Wang 0035, Zhao Li 0007, Zili Zhang 0001, Xindong Wu 0001 |
ICDM | 6 |
| 2018 | Incomplete Multi-View Weak-Label LearningabstractLearning from multi-view multi-label data has wide applications. There are two main challenges of this learning task: incomplete views and missing (weak) labels. The former assumes that views may not include all data objects. The weak label setting implies that only a subset of relevant labels are provided for training objects while other labels are missing. Both incomplete views and weak labels can lead to significant performance degradation. In this paper, we propose a novel model (iMVWL) to jointly address the two challenges. iMVWL simultaneously learns a shared subspace from incomplete views with weak labels, the local label structure and the predictor in this subspace, which can not only capture cross-view relationships but also weak-label information of training samples. We further develop an alternative solution to optimize our model, this solution can avoid suboptimal results and reinforce their reciprocal effects, and thus further improve the performance. Extensive experimental results on several real-world datasets validate the effectiveness of our model against other competitive algorithms. Qiaoyu Tan, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zili Zhang 0001 |
IJCAI | 5 |
| 2018 | Multi-Label Co-TrainingabstractMulti-label learning aims at assigning a set of appropriate labels to multi-label samples. Although it has been successfully applied in various domains in recent years, most multi-label learning methods require sufficient labeled training samples, because of the large number of possible label sets. Co-training, as an important branch of semi-supervised learning, can leverage unlabeled samples, along with scarce labeled ones, and can potentially help with the large labeled data requirement. However, it is a difficult challenge to combine multi-label learning with co-training. Two distinct issues are associated with the challenge: (i) how to solve the widely-witnessed class-imbalance problem in multi-label learning; and (ii) how to select samples with confidence, and communicate their predicted labels among classifiers for model refinement. To address these issues, we introduce an approach called Multi-Label Co-Training (MLCT). MLCT leverages information concerning the co-occurrence of pairwise labels to address the class-imbalance challenge; it introduces a predictive reliability measure to select samples, and applies label-wise filtering to confidently communicate labels of selected samples among co-training classifiers. MLCT performs favorably against related competitive multi-label learning methods on benchmark datasets and it is also robust to the input parameters. Yuying Xing, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zili Zhang 0001 |
IJCAI | 5 |
| 2018 | Traffic Flow Fluctuation Analysis Based on Beijing Taxi GPS Data
Jingyi Guo, Xianghua Li, Zili Zhang 0001 |
KSEM (2) | 3 |
| 2018 | Matrix Factorization for Identifying Noisy Labels of Multi-label Instances
Xia Chen 0004, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zili Zhang 0001 |
PRICAI | 5 |
| 2018 | Multi-view Weak-label Learning based on Matrix CompletionabstractWeak-label learning is an important branch of multi-label learning; it deals with samples annotated with incomplete (weak) labels. Previous work on weak-label learning mainly considers data represented by a single view. An intuitive way to leverage multiple features obtained from different views is to concatenate the features into a single vector. However, this process is not only prone to over-fitting and often results in very high time-complexity, but also ignores the potentially useful complementary information spread across the different views. In this paper, we propose an approach based on Matrix Completion for multi-view Weak-label Learning (McWL). Matrix completion (MC) has sound theoretical properties and is robust to missing values in both feature and label spaces. Our method enforces the optimization of multiple view integration and of MC-based classification within a unified objective function. Specifically, a kernel target alignment technique and the loss function of an MC-based classifier are used to jointly and iteratively adjust the weights assigned to individual views, and to optimize the classifier. McWL can selectively integrate views and is able to assign small weights to views of low quality. Extensive experiments on a broad range of datasets validate the effectiveness of our approach against competitive algorithms. Qiaoyu Tan, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zili Zhang 0001 |
SDM | 5 |
| 2018 | Block-sparse signal recovery via ℓ 2 / ℓ 1 - 2 minimisation methodabstractMotivated by the recently emerged method for sparse signal recovery, in this study, the authors make an ongoing effect to extend this methodology to the setting of block sparsity, which directly leads to the proposed method for block‐sparse signal recovery. Some theoretical results are induced to guarantee the validity of proposed method. In particular, the obtained recovery condition rigorously includes the one induced by Yin et al ., and the obtained error estimate can be used to model both the (block‐) sparse and non‐sparse signals, which is more comprehensive than that induced by Yin et al . which applies only to the sparse signals. The authors also derive an alternating direction method of multipliers (ADMM)‐based algorithm to tackle the induced optimisation problem. Some experimental results that are based on the synthetic block‐sparse signals and the real‐world foetal electrocardiogram signals further demonstrate the better performance of the method when it is compared with the state‐of‐the‐art group‐lasso method and method for 0 < q < 1. Wendong Wang 0001, Jianjun Wang 0003, Zili Zhang 0001 |
IET Signal Process. | 3 |
| 2018 | Design, verification and robotic application of a novel recurrent neural network for computing dynamic Sylvester equation
Lin Xiao 0002, Zhijun Zhang 0003, Zili Zhang 0001, Weibing Li, Shuai Li 0002 |
Neural Networks | 3 |
| 2018 | Network Community Detection Based on the Physarum-Inspired Computational FrameworkabstractCommunity detection is a crucial and essential problem in the structure analytics of complex networks, which can help us understand and predict the characteristics and functions of complex networks. Many methods, ranging from the optimization-based algorithms to the heuristic-based algorithms, have been proposed for solving such a problem. Due to the inherent complexity of identifying network structure, how to design an effective algorithm with a higher accuracy and a lower computational cost still remains an open problem. Inspired by the computational capability and positive feedback mechanism in the wake of foraging process of Physarum, a kind of slime, a general Physarum-based computational framework for community detection is proposed in this paper. Based on the proposed framework, the inter-community edges can be identified from the intra-community edges in a network and the positive feedback of solving process in an algorithm can be further enhanced, which are used to improve the efficiency of original optimization-based and heuristic-based community detection algorithms, respectively. Some typical algorithms (e.g., genetic algorithm, ant colony optimization algorithm, and Markov clustering algorithm) and real-world datasets have been used to estimate the efficiency of our proposed computational framework. Experiments show that the algorithms optimized by Physarum-inspired computational framework perform better than the original ones, in terms of accuracy and computational cost. Chao Gao 0001, Mingxin Liang, Xianghua Li, Zili Zhang 0001, Zhen Wang 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | Special issue: Mobile Web Data Analytics (part I)abstractIt is essential to constantly collect data with various mobile applications from diverse sources, such as smartphones and ubiquitous sensors.However, how do you conduct the analysis on such a mass of mobile data or mobile web data aiming to solve issues in different areas of applications, including human behavior recognition, medication, recommendation and transportation?Nowadays, research in mobile and social computing environments is now turning to novel concepts to address the challenge of data processing and analyzing.The special issue Mobile web data analytics addresses issues of data management in mobile and social computing environments with a special focus on data processing and applications.The goal of the special issue is to build a forum for researchers from academy and industry to investigate challenging and innovative research issues on the subject, which combines data analytics within mobile and social environment and to explore creative concepts, theories, innovative technologies and intelligent solutions.We intend this special issue to act as an initial place where people from different areas can find a forum to discuss issues of data management and processing in new and emerging mobile computing environments.We accepted 11 papers that provide deep research results to report the advance of mobile web data analytics and applications.These papers are grouped into Zili Zhang 0001, Li Liu 0001, Li Li 0006, Xiangliang Zhang 0001 |
Web Intell. | 1 |
| 2018 | Special issue: Mobile web data analytics (part II)abstractIt is essential to constantly collect data with various mobile applications from diverse sources, such as smartphones and ubiquitous sensors.However, how do you conduct the analysis on such a mass of mobile data or mobile web data aiming to solve issues in different areas of applications, including human behavior recognition, medication, recommendation and transportation?Nowadays, research in mobile and social computing environments is turning to novel concepts to address the challenge of data processing and analyzing.This special issue Mobile web data analytics addresses issues of data management in mobile and social computing environments with a special focus on data processing and applications.The goal of this special issue is to build a forum for researchers from academia and industry to investigate challenging and innovative research issues on the subject, which combines data analytics within mobile and social environment and to explore creative concepts, theories, innovative technologies and intelligent solutions.We intend this special issue to act as an initial place where people from different areas can find a forum to discuss issues of data management and processing in new and emerging mobile computing environments.We accepted 11 papers that provide deep research results to report the advance of mobile web data analytics and applications.These papers are grouped into Zili Zhang 0001, Li Liu 0001, Li Li 0006, Xiangliang Zhang 0001 |
Web Intell. | 1 |
| 2017 | Automated heuristic design using genetic programming hyper-heuristic for uncertain capacitated arc routing problemabstractUncertain Capacitated Arc Routing Problem (UCARP) is a variant of the well-known CARP. It considers a variety of stochastic factors to reflect the reality where the exact information such as the actual task demand and accessibilities of edges are unknown in advance. Existing works focus on obtaining a robust solution beforehand. However, it is also important to design effective heuristics to adjust the solution in real time. In this paper, we develop a new Genetic Programming-based Hyper-Heuristic (GPHH) for automated heuristic design for UCARP. A novel effective meta-algorithm is designed carefully to address the failures caused by the environment change. In addition, it employs domain knowledge to filter some infeasible candidate tasks for the heuristic function. The experimental results show that the proposed GPHH significantly outperforms the existing GPHH methods and manually designed heuristics. Moreover, we find that eliminating the infeasible and distant tasks in advance can reduce much noise and improve the efficacy of the evolved heuristics. In addition, it is found that simply adding a slack factor to the expected task demand may not improve the performance of the GPHH. Yuxin Liu 0003, Yi Mei 0001, Mengjie Zhang 0001, Zili Zhang 0001 |
GECCO | 4 |
| 2017 | An Enhanced Markov Clustering Algorithm Based on Physarum
Mingxin Liang, Chao Gao 0001, Xianghua Li, Zili Zhang 0001 |
PAKDD (1) | 4 |
| 2017 | A Physarum-Inspired Ant Colony Optimization for Community Mining
Mingxin Liang, Chao Gao 0001, Xianghua Li, Zili Zhang 0001 |
PAKDD (1) | 4 |
| 2017 | Emphasizing Essential Words for Sentiment Classification Based on Recurrent Neural Networks
Fei Hu 0004, Li Li 0006, Zili Zhang 0001, Xiaofei Xu 0002 |
J. Comput. Sci. Technol. | 3 |
| 2017 | A new genetic algorithm based on modified Physarum network model for bandwidth-delay constrained least-cost multicast routing
Mingxin Liang, Chao Gao 0001, Zili Zhang 0001 |
Nat. Comput. | 3 |
| 2017 | A new multi-agent system to simulate the foraging behaviors of Physarum
Yuxin Liu 0003, Chao Gao 0001, Zili Zhang 0001, Mingxin Liang, Yuxiao Lu |
Nat. Comput. | 3 |
| 2017 | Robust Signal Recovery With Highly Coherent Measurement MatricesabstractBy embedding an ℓp-norm noise constraint for p ≥ 2 into the recently emerged ℓ1-2method, in this letter, we study theoretically and numerically an ℓ1-2/ℓpmethod for recovery of general noisy signals from highly coherent measurement matrices. In particular, the obtained theoretical results not only improve the condition deduced in [1] for Gaussian noisy signal recovery but also provide a new theoretical guarantee for generally nonGaussian noisy signal recovery. What is more, to better boost the recovery performance, a partial sum ℓ1-2/ℓpmethod is also proposed latter. This improved method, together with the previous ℓ1-2/ℓpmethod, becomes more competitive when compared with some of the state-of-the-art methods in recovering noisy signals from highly coherent measurement matrices. Wendong Wang 0001, Jianjun Wang 0003, Zili Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2017 | Solving NP-Hard Problems with Physarum-Based Ant Colony SystemabstractNP-hard problems exist in many real world applications. Ant colony optimization (ACO) algorithms can provide approximate solutions for those NP-hard problems, but the performance of ACO algorithms is significantly reduced due to premature convergence and weak robustness, etc. With these observations in mind, this paper proposes a Physarum-based pheromone matrix optimization strategy in ant colony system (ACS) for solving NP-hard problems such as traveling salesman problem (TSP) and 0/1 knapsack problem (0/1 KP). In the Physarum-inspired mathematical model, one of the unique characteristics is that critical tubes can be reserved in the process of network evolution. The optimized updating strategy employs the unique feature and accelerates the positive feedback process in ACS, which contributes to the quick convergence of the optimal solution. Some experiments were conducted using both benchmark and real datasets. The experimental results show that the optimized ACS outperforms other meta-heuristic algorithms in accuracy and robustness for solving TSPs. Meanwhile, the convergence rate and robustness for solving 0/1 KPs are better than those of classical ACS. Yuxin Liu 0003, Chao Gao 0001, Zili Zhang 0001, Yuxiao Lu, Mingxin Liang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | A bio-inspired method for locating the diffusion source with limited observersabstractLocating the source of diffusion is a challenging problem in complex networks and has great practical significance for restraining rumors propagation and controlling epidemics spreading. An efficient locating method should have a higher locating accuracy with the minimum required information. Although existing locating methods based on observers consider the time delays of edges, they compute the time delays based on the shortest path, which may differ from the actual diffusion process. Moreover, the higher locating accuracy of traditional method with observers has a great dependence on the assumption that the propagation delays along edges follow a definite distribution such as the Gaussian distribution. In order to solve these shortcomings, this paper proposes a Physarum-inspired method to locate the diffusion source that is independence of the distribution of propagation delays. Our method quantifies the nutrient transportation process in the adaptive network evolved by Physarum, which is used to simulate the information or epidemic diffusion routes in a social network. Simulation results on various benchmark networks show that our method has a better performance in terms of error distance than that of Gaussian method without assuming the definite distribution of time delays. Together with the advantage that our method does not require the sender information of observers compared with existing methods, our method allows for a wider range of applications in the real-world networks. Yuxin Liu 0003, Chao Gao 0001, Xinyan She, Zili Zhang 0001 |
CEC | 4 |
| 2016 | An Evidential Spam-Filtering FrameworkabstractSpam, also known as unsolicited bulk e-mail (UBE), has recently become a serious threat that negatively impacts the usability of legitimate mails. In this article, an evidential spam-filtering framework is proposed. As a useful tool to handle uncertainty, the Dempster–Shafer theory of evidence (D–S theory) is integrated into the proposed approach. Five representative features from an e-mail header are analyzed. With a machine-learning algorithm, e-mail headers with known classifications are used to train the framework. When using the framework for a given e-mail header, its representative features are quantified. Although in classical probability theory, possibilities are forcedly assigned even when information is not adequate, in our approach, for every word in an e-mail subject, basic probability assignments (BPA) are assigned in a more flexible way, thus providing a more reasonable result. Finally, BPAs are combined and transformed into pignistic probabilities for decision-making. Empirical trials on real-world datasets show the efficiency of the proposed framework. Xiaoyan Su, Yong Hu 0002, Zili Zhang 0001, Yong Deng 0001 |
Cybern. Syst. | 4 |
| 2016 | A distributed spatial-temporal weighted model on MapReduce for short-term traffic flow forecasting
Dawen Xia, Binfeng Wang, Huaqing Li 0001, Yantao Li 0001, Zili Zhang 0001 |
Neurocomputing | 5 |
| 2015 | Robust Traffic Classification with Mislabelled Training SamplesabstractTraffic classification plays the significant role in the network security and management. However, accurate classification is challenging if the training data is contaminated with unclean traffic. Recent researches often assume clean training data, and hence performance reduced on real-time network traffic. To meet this challenge, in this paper, we propose a robust method, Unclean Traffic Classification (UTC), which incorporates noise elimination and suspected noise reweighting. Firstly, UTC eliminates strong noisy training data identified by a consensus filtering with multiple classifiers. Furthermore, UTC estimates the relevance of remaining training data and learns a robust traffic classifier. Through a number of experiments on a real-world traffic dataset, we show that the new method outperforms existing state-of-the-art traffic classification methods, under the extremely difficult circumstance with unclean training data. Binfeng Wang, Jun Zhang 0010, Zili Zhang 0001, Wei Luo 0001, Dawen Xia |
ICPADS | 3 |
| 2015 | A new model to imitate the foraging behavior of Physarum polycephalum on a nutrient-poor substrate
Zili Zhang 0001, Yong Deng 0001 |
Neurocomputing | 2 |
| 2015 | Semi-supervised classification based on subspace sparse representation
Guoxian Yu, Guoji Zhang, Zili Zhang 0001, Zhiwen Yu 0002, Lin Deng 0001 |
Knowl. Inf. Syst. | 3 |
| 2015 | Predicting Protein Function Using Multiple KernelsabstractHigh-throughput experimental techniques provide a wide variety of heterogeneous proteomic data sources. To exploit the information spread across multiple sources for protein function prediction, these data sources are transformed into kernels and then integrated into a composite kernel. Several methods first optimize the weights on these kernels to produce a composite kernel, and then train a classifier on the composite kernel. As such, these approaches result in an optimal composite kernel, but not necessarily in an optimal classifier. On the other hand, some approaches optimize the loss of binary classifiers and learn weights for the different kernels iteratively. For multi-class or multi-label data, these methods have to solve the problem of optimizing weights on these kernels for each of the labels, which are computationally expensive and ignore the correlation among labels. In this paper, we propose a method called Predicting Protein Function using Multiple Kernels (ProMK). ProMK iteratively optimizes the phases of learning optimal weights and reduces the empirical loss of multi-label classifier for each of the labels simultaneously. ProMK can integrate kernels selectively and downgrade the weights on noisy kernels. We investigate the performance of ProMK on several publicly available protein function prediction benchmarks and synthetic datasets. We show that the proposed approach performs better than previously proposed protein function prediction approaches that integrate multiple data sources and multi-label multiple kernel learning methods. The codes of our proposed method are available at https://sites.google.com/site/guoxian85/promk. Guoxian Yu, Huzefa Rangwala, Carlotta Domeniconi, Guoji Zhang, Zili Zhang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2014 | A C-DBSCAN Algorithm for Determining Bus-Stop Locations Based on Taxi GPS Data
Chao Gao 0001, Binfeng Wang, Zili Zhang 0001 |
ADMA | 6 |
| 2014 | A Semantic-Based EMRs Integration Framework for Diagnosis Decision-Making
Huili Jiang, Zili Zhang 0001 |
KSEM | 2 |
| 2014 | Dividing Traffic Sub-areas Based on a Parallel K-Means Algorithm
Binfeng Wang, Chao Gao 0001, Dawen Xia, Zhuobo Rong, Zili Zhang 0001 |
KSEM | 7 |
| 2014 | Guest Editors' Introduction
Zili Zhang 0001, Zhi Jin 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2014 | Coverage enhancement by using the mobility of mobile sensor nodes
Can Fang, Peng Zhang 0005, Zili Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2014 | Sample Subset Optimization Techniques for Imbalanced and Ensemble Learning Problems in Bioinformatics ApplicationsabstractData sampling is a widely used technique in a broad range of machine learning problems. Traditional sampling approaches generally rely on random resampling from a given dataset. However, these approaches do not take into consideration additional information, such as sample quality and usefulness. We recently proposed a data sampling technique, called sample subset optimization (SSO). The SSO technique relies on a cross-validation procedure for identifying and selecting the most useful samples as subsets. In this paper, we describe the application of SSO techniques to imbalanced and ensemble learning problems, respectively. For imbalanced learning, the SSO technique is employed as an under-sampling technique for identifying a subset of highly discriminative samples in the majority class. In ensemble learning, the SSO technique is utilized as a generic ensemble technique where multiple optimized subsets of samples from each class are selected for building an ensemble classifier. We demonstrate the utilities and advantages of the proposed techniques on a variety of bioinformatics applications where class imbalance, small sample size, and noisy data are prevalent. Pengyi Yang, Paul D. Yoo, Juanita I. Fernando, Bing Bing Zhou, Zili Zhang 0001, Albert Y. Zomaya |
IEEE Trans. Cybern. | 5 |
| 2013 | The Spontaneous Behavior in Extreme Events: A Clustering-Based Quantitative Analysis
Ning Shi, Chao Gao 0001, Zili Zhang 0001, Lu Zhong, Jiajin Huang |
ADMA (1) | 3 |
| 2013 | Protein Function Prediction by Integrating Multiple Kernels
Guoxian Yu, Huzefa Rangwala, Carlotta Domeniconi, Guoji Zhang, Zili Zhang 0001 |
IJCAI | 5 |
| 2013 | Learning to Map Chinese Sentences to Logical Forms
Zhihua Liao, Zili Zhang 0001 |
KSEM | 2 |
| 2013 | A Semantic Technology Supported Precision Agriculture System: A Case Study for Citrus Fertilizing
Ye Yuan 0013, Zili Zhang 0001 |
KSEM | 3 |
| 2012 | A Generic Classifier-Ensemble Approach for Biomedical Named Entity Recognition
Zhihua Liao, Zili Zhang 0001 |
PAKDD (1) | 2 |
| 2011 | Sample Subset Optimization for Classifying Imbalanced Biological Data
Pengyi Yang, Zili Zhang 0001, Bing Bing Zhou, Albert Y. Zomaya |
PAKDD (2) | 2 |
| 2011 | Missing Value Estimation for Mixed-Attribute Data SetsabstractMissing data imputation is a key issue in learning from incomplete data. Various techniques have been developed with great successes on dealing with missing values in data sets with homogeneous attributes (their independent attributes are all either continuous or discrete). This paper studies a new setting of missing data imputation, i.e., imputing missing data in data sets with heterogeneous attributes (their independent attributes are of different types), referred to as imputing mixed-attribute data sets. Although many real applications are in this setting, there is no estimator designed for imputing mixed-attribute data sets. This paper first proposes two consistent estimators for discrete and continuous missing target values, respectively. And then, a mixture-kernel-based iterative estimator is advocated to impute mixed-attribute data sets. The proposed method is evaluated with extensive experiments compared with some typical algorithms, and the result demonstrates that the proposed approach is better than these existing imputation methods in terms of classification accuracy and root mean square error (RMSE) at different missing ratios. Xiaofeng Zhu 0001, Shichao Zhang 0001, Zhi Jin 0001, Zili Zhang 0001, Zhuoming Xu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2010 | Applying Multi-objective Evolutionary Algorithms to QoS-Aware Web Service Composition
Li Li 0006, Peng Cheng 0011, Ling Ou, Zili Zhang 0001 |
ADMA (2) | 4 |
| 2010 | Genetic Algorithm-Based Multi-objective Optimisation for QoS-Aware Web Services Composition
Li Li 0006, Pengyi Yang, Ling Ou, Zili Zhang 0001, Peng Cheng 0011 |
KSEM | 4 |
| 2010 | Chinese Named Entity Recognition Based on Hierarchical Hybrid Model
Zhihua Liao, Zili Zhang 0001 |
PRICAI | 2 |
| 2010 | A multi-filter enhanced genetic ensemble system for gene selection and sample classification of microarray dataabstractBACKGROUND: Feature selection techniques are critical to the analysis of high dimensional datasets. This is especially true in gene selection from microarray data which are commonly with extremely high feature-to-sample ratio. In addition to the essential objectives such as to reduce data noise, to reduce data redundancy, to improve sample classification accuracy, and to improve model generalization property, feature selection also helps biologists to focus on the selected genes to further validate their biological hypotheses. RESULTS: In this paper we describe an improved hybrid system for gene selection. It is based on a recently proposed genetic ensemble (GE) system. To enhance the generalization property of the selected genes or gene subsets and to overcome the overfitting problem of the GE system, we devised a mapping strategy to fuse the goodness information of each gene provided by multiple filtering algorithms. This information is then used for initialization and mutation operation of the genetic ensemble system. CONCLUSION: We used four benchmark microarray datasets (including both binary-class and multi-class classification problems) for concept proving and model evaluation. The experimental results indicate that the proposed multi-filter enhanced genetic ensemble (MF-GE) system is able to improve sample classification accuracy, generate more compact gene subset, and converge to the selection results more quickly. The MF-GE system is very flexible as various combinations of multiple filters and classifiers can be incorporated based on the data characteristics and the user preferences. Pengyi Yang, Bing Bing Zhou, Zili Zhang 0001, Albert Y. Zomaya |
BMC Bioinform. | 3 |
| 2010 | A clustering based hybrid system for biomarker selection and sample classification of mass spectrometry data
Pengyi Yang, Zili Zhang 0001, Bing Bing Zhou, Albert Y. Zomaya |
Neurocomputing | 2 |
| 2010 | Guest Editors' IntroductionabstractGuest editors' introduction Zili Zhang 0001, Yi-Ping Phoebe Chen |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2007 | A Fuzzy Logic Based Approach for Software TestingabstractHow to provide cost-effective strategies for Software Testing has been one of the research focuses in Software Engineering for a long time. Many researchers in Software Engineering have addressed the effectiveness and quality metric of Software Testing, and many interesting results have been obtained. However, one issue of paramount importance in software testing — the intrinsic imprecise and uncertain relationships within testing metrics — is left unaddressed. To this end, a new quality and effectiveness measurement based on fuzzy logic is proposed. Related issues like the software quality features and fuzzy reasoning for test project similarity measurement are discussed, which can deal with quality and effectiveness consistency between different test projects. Experiments were conducted to verify the proposed measurement using real data from actual software testing projects. Experimental results show that the proposed fuzzy logic based metrics is effective and efficient to measure and evaluate the quality and effectiveness of test projects. Zili Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2007 | Building agent-based hybrid intelligent systems: A case study
Zili Zhang 0001, Chengqi Zhang |
Web Intell. Agent Syst. | 1 |
| 2006 | Decision Aggregation in an Agent-Based Financial Investment Planning System
Zili Zhang 0001 |
MDAI | 1 |
| 2006 | Integrating Insurance Services, Trust and Risk Mechanisms into Multi-agent Systems
Yuk-Hei Lam, Zili Zhang 0001, Kok-Leong Ong |
PRICAI | 2 |
| 2006 | Buying and Selling with Insurance in Open Multi-agent Marketplace
Yuk-Hei Lam, Zili Zhang 0001, Kok-Leong Ong |
PRICAI | 2 |
| 2005 | Supporting Adaptive Learning in Hypertext Environment: A High Level Timed Petri Net Based ApproachabstractOne problem for hypertext-based learning application is to control learning paths for different learning activities. This paper first introduced related concepts of hypertext learning state space and Petri net, then proposed a high level timed Petri Net based approach to provide some kinds of adaptation for learning activities. Examples were given while explaining ways to realizing adaptive instructions. Possible future directions were also discussed at the end of this paper. Shang Gao 0003, Zili Zhang 0001, Igor T. Hawryszkiewycz |
ICALT | 2 |
| 2005 | Supporting Adaptive Learning with High Level Timed Petri Nets
Shang Gao 0003, Zili Zhang 0001, Jason Wells, Igor T. Hawryszkiewycz |
KES (3) | 2 |
| 2004 | A Reputation-Based Trust Model for Agent Societies
Yuk-Hei Lam, Zili Zhang 0001, Kok-Leong Ong |
PRICAI | 2 |
| 2003 | Building Agent-Based Hybrid Intelligent Systems
Zili Zhang 0001, Chengqi Zhang |
HIS | 1 |
| 2003 | A Ring-Based Architectural Model for Middle Agents in Agent-Based System
Chengqi Zhang, Zili Zhang 0001 |
IDEAL | 3 |
| 2002 | An Agent-Based Framework for Petroleum Information Services from Distributed Heterogeneous Data ResourcesabstractFor making good decisions in the area of petroleum production, it is becoming a big problem how to timely gather sufficient and correct information, which may be stored in databases, data files, or on the World Wide Web. In this paper, Gaia methodology and Open Agent Architecture were employed to contribute a framework to solve above problem. The framework consists of three levels, namely, role mode, agent type, and agent instance. The model with five roles is analyzed. Four agent types are designed Six agent instances are developed for constructing the system of petroleum information services. The experimental results show that all agents in the system can work cooperatively to organize and retrieve relevant petroleum information. The successful implementation of the framework shows that agent-based technology can significantly facilitate the construction of complex systems in distributed heterogeneous data resource environment. Chengqi Zhang, Zili Zhang 0001 |
APSEC | 3 |
| 2002 | An Agent-Based Hybrid Intelligent System for Financial Investment Planning
Zili Zhang 0001, Chengqi Zhang |
PRICAI | 1 |
| 2000 | Building an Ontology for Financial Investment
Zili Zhang 0001, Chengqi Zhang, Swee San Ong |
IDEAL | 1 |