Wei Liu 0010

dblp:49/3283-10 · DBLP profile ↗
← Back
37ranked-venue papers
9as first author
16since 2021 · last 2025
0000-0001-8503-4063ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Influence maximization based on discrete particle swarm optimization on multilayer network
Saiwei Wang, Wei Liu 0010, Ling Chen 0005, Shijie Zong
Inf. Syst.2
2025 A Greedy Descent Method for Budget Constrained Continuous Influence Maximization in Online Social Network
abstract
Continuous influence maximization (CIM) in social networks aims to maximize the expected influence spreading by assigning each user a continuous weight reflecting the likelihood and cost for him becoming a seed. Traditional CIM assumes a budget constrain to limit the total cost of all users. However, this assumption does not tenable in practical applications. In practice, it is not necessary to incur costs for all the users. Instead, the budget should be set only for the cost associated with the seed set. In this article, an extended CIM problem of cost distribution under budget (CDB) is defined, which aims to assign different costs to the customers according to their ability to spread influence, ensuring that the cost of each potential seed set does not exceed the budget, while the expected spreading of the product’s influence is maximized. The NP-hardness of CDB and the monotonicity and submodularity of its objective function are investigated. We formulate the CDB problem into a constrained optimization, and present a greedy descent-based algorithm for the problem. In each iteration of the greedy descent method, the influence increment of each node is calculated according to its estimated influence spreading range. The cost distribution is updated along the direction with the maximum increment. The optimal cost distribution can be obtained after several iterations. Precision of the results obtained by the proposed algorithm is analyzed. To avoid the time-consuming simulations, we design an effective algorithm for estimating the seed influence spreading range. Experiment results on real and synthetic networks show that the proposed algorithm can significantly improve the expected influence spreading.
Wei Liu 0010, Ziwei Deng, Yixin Chen 0001, Ling Chen 0005
IEEE Trans. Comput. Soc. Syst.1
2024 Powerful Influencer Identification in Temporal Social Networks Adopting Discrete Wild Geese Swarm Optimization
Wei Liu 0010, Shijie Zong, Ling Chen 0005
ICIC (1)1
2024 Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection Systems
abstract
Recently, Graph Neural Network (GNN)-based vulnerability detection systems have achieved remarkable success. However, the lack of explainability poses a critical challenge to deploy black-box models in security-related domains. For this reason, several approaches have been proposed to explain the decision logic of the detection model by providing a set of crucial statements positively contributing to its predictions. Unfortunately, due to the weakly-robust detection models and suboptimal explanation strategy, they have the danger of revealing spurious correlations and redundancy issue.
Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, David Lo 0001, Lili Bo, Bin Li 0006, Wei Liu 0010
ICSE7
2024 Snopy: Bridging Sample Denoising with Causal Graph Learning for Effective Vulnerability Detection
abstract
Deep Learning (DL) has emerged as a promising means for vulnerability detection due to its ability to automatically derive features from vulnerable code. Unfortunately, current solutions struggle to focus on vulnerability-related parts of vulnerable functions, and tend to exploit spurious correlations for prediction, thus undermining their effectiveness in practice. In this paper, we propose Snopy, a novel DL-based approach, which bridges sample denoising with causal graph learning to capture real vulnerability patterns from vulnerable samples with numerous noise for effective detection. Specifically, Snopy adopts a change-based sample denoising approach to automatically weed out vulnerability-irrelevant code elements in the vulnerable functions without sacrificing the label accuracy. Then, Snopy constructs a novel Causality-Aware Graph Attention Network (CA-GAT) with Feature Caching Scheme (FCS) to learn causal vulnerability features while maintaining efficiency. Experiments on the three public benchmark datasets show that Snopy outperforms the state-of-the-art baselines by an average of 27.22%, 85.89%, and 75.50% in terms of F1-score, respectively.
Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, David Lo 0001, Lili Bo, Bin Li 0006, Xiaolei Liu 0001, Xingwei Lin, Wei Liu 0010
ASE9
2024 EXVul: Toward Effective and Explainable Vulnerability Detection for IoT Devices
abstract
As with anything connected to the internet, Internet of Things (IoT) devices are also subject to severe cybersecurity threats because an adversary could exploit vulnerabilities in their internal software to perform malicious attacks. Despite the promising results of Deep Learning (DL)-based approaches, the lack of well-labeled IoT vulnerability samples available for training and explainability pose a critical challenge to deploy them in practice. In this paper, we propose, a novel DL-based approach for Effective and eXplainable IoT VULnerability detection. Specifically, inspired by recent advances of self-supervised learning in label-expensive tasks, we propose a new combinatorial contrastive loss to combine the strengths of large-scale unlabeled code corpus and limited IoT vulnerability samples. Then, given a binary detection result, provides a set of faithful and stable code statements positively contributing to the model’s predictions as understandable explanations. Experimental results indicate that outperforms state-of-the-art baselines by 33.44%-72.91% and 19.52%-98.78% with respect to the accuracy and F1 score metrics, respectively. For vulnerability explanation, improves over the best-performing baseline explainer PGExplainer by 22.97% in MSP, 49.55% in MSR, and 48.40% in MIoU, demonstrating that the explanations provided by can correctly point out the vulnerable statements relevant to the detected vulnerabilities.
Sicong Cao, Xiaobing Sun 0001, Wei Liu 0010, Di Wu 0050, Jiale Zhang 0001, Yan Li 0002, Tom H. Luan, Longxiang Gao
IEEE Internet Things J.3
2024 Hierarchy-Aware Representation Learning for Industrial IoT Vulnerability Classification
abstract
As with anything connected to the internet, industrial Internet of Things (IIoT) devices are also subject to severe cybersecurity threats because an adversary could exploit vulnerabilities in their internal software to perform malicious attacks. Despite the promising results of deep learning-based approaches, most solutions can only detect the presence of a vulnerability but fail to pinpoint its corresponding type. Recently, TreeVul formalizes the task as a hierarchical multilabel classification problem to predict complete coarse-to-fine vulnerability type hierarchy. Yet, the TreeVul approach is still inaccurate and neglects samples labeled at coarse categories. In this article, we proposeHierVul, a novel hierarchy-aware representation learning approach for IIoT vulnerability classification. Specifically, to make full use of vulnerable samples labeled at any granularity,HierVulconstructs hierarchy-specific extractors as well as classifiers to disentangle level-wise vulnerability features from the code representation learning network backbone, and maximizes their marginal probability in the probability space constrained by the Common Weakness Enumeration tree hierarchy. Furthermore, considering that the distinction between two vulnerability types at the same level of abstraction becomes smaller and smaller as the refinement of classification granularity,HierVulleverages residual connections to add parent-level coarser-grained features to child-level finer-grained features to transfer hierarchical knowledge across levels. The experimental results show thatHierVulachieves 15.25%, 45.16%, and 14.52% relative improvement over TreeVul on Weight F1, Macro F1, and PF, respectively, indicating the effectiveness ofHierVulin the practical scenario.
Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, Wei Liu 0010, Bin Li 0006
IEEE Trans. Ind. Informatics5
2024 Learning to Detect Memory-related Vulnerabilities
abstract
Memory-related vulnerabilities can result in performance degradation or even program crashes, constituting severe threats to the security of modern software. Despite the promising results of deep learning (DL)-based vulnerability detectors, there exist three main limitations: (1) rich contextual program semantics related to vulnerabilities have not yet been fully modeled; (2) multi-granularity vulnerability features in hierarchical code structure are still hard to be captured; and (3) heterogeneous flow information is not well utilized. To address these limitations, in this article, we propose a novel DL-based approach, called MVD+ , to detect memory-related vulnerabilities at the statement-level. Specifically, it conducts both intraprocedural and interprocedural analysis to model vulnerability features, and adopts a hierarchical representation learning strategy, which performs syntax-aware neural embedding within statements and captures structured context information across statements based on a novel Flow-Sensitive Graph Neural Networks, to learn both syntactic and semantic features of vulnerable code. To demonstrate the performance, we conducted extensive experiments against eight state-of-the-art DL-based approaches as well as five well-known static analyzers on our constructed dataset with 6,879 vulnerabilities in 12 popular C/C++ applications. The experimental results confirmed that MVD+ can significantly outperform current state-of-the-art baselines and make a great trade-off between effectiveness and efficiency.
Sicong Cao, Xiaobing Sun 0001, Lili Bo, Rongxin Wu, Bin Li 0006, Xiaoxue Wu 0001, Chuanqi Tao, Tao Zhang 0001, Wei Liu 0010
ACM Trans. Softw. Eng. Methodol.9
2023 Improving Java Deserialization Gadget Chain Mining via Overriding-Guided Object Generation
abstract
Java (de)serialization is prone to causing security-critical vulnerabilities that attackers can invoke existing methods (gadgets) on the application's classpath to construct a gadget chain to perform malicious behaviors. Several techniques have been proposed to statically identify suspicious gadget chains and dynamically generate injection objects for fuzzing. However, due to their incomplete support for dynamic program features (e.g., Java runtime polymorphism) and ineffective injection object generation for fuzzing, the existing techniques are still far from satisfactory. In this paper, we first performed an empirical study to investigate the characteristics of Java deserialization vulnerabilities based on our manually collected 86 publicly known gadget chains. The empirical results show that 1) Java deserialization gadgets are usually exploited by abusing runtime polymorphism, which enables attackers to reuse serializable overridden methods; and 2) attackers usually invoke exploitable overridden methods (gadgets) via dynamic binding to generate injection objects for gadget chain construction. Based on our empirical findings, we propose a novel gadget chain mining approach, GCMiner, which captures both explicit and implicit method calls to identify more gadget chains, and adopts an overriding-guided object generation approach to generate valid injection objects for fuzzing. The evaluation results show that GCMiner significantly outperforms the state-of-the-art techniques, and discovers 56 unique gadget chains that cannot be identified by the baseline approaches.
Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, Lili Bo, Bin Li 0006, Rongxin Wu, Wei Liu 0010, Biao He 0002, Yu Ouyang
ICSE7
2023 Identifying multiple influence sources in social networks based on latent space mapping
abstract
We are currently in a network era which enables us to communicate more widely and more easily via the social networks. Meanwhile, negative information, such as fake news, rumors and computer viruses, often spread in social network. In order to restrain the propagation of such negative influence, we must find its sources in the network. But in real-world applications, we usually only know the scope of the negative influence spreading, and do not know who first propagates the negative influence. However, we can identify the sources of the negative influence based on the information of some observed nodes which are negatively influenced. This is the problem of influence sources locating. To tackle this problem, we present a latent space mapping-based method for identifying the multiple influence sources in the independent cascade model. The method first detects the candidate sources of the observed nodes based on message passing in a reversed network. An algorithm is presented to calculate the activation probability between nodes according to the influence spreading pattern in the independent cascade model. To evaluate each node’s rationality as the propagation source, we use the difference between the length of the path influencing an observed node and its activation time. We define two latent spaces, namely the influence senders and receivers’ latent spaces, and map the nodes into these two latent spaces to form a model describing the influence propagation. An estimation-maximization-based algorithm is proposed to optimize the propagation model. Based on this model, we propose a latent space mapping-based algorithm to identify the influence sources. The probability for each node to be a source is calculated by its positions in the latent spaces. Finally, k nodes with the largest probabilities are selected as the sources. Empirical results demonstrate that the influence sources identified by the proposed method can influence more observed nodes at more accurate time than other methods.
Ling Chen 0005, Yixin Chen 0001, Wei Liu 0010, Caiyan Dai
Inf. Sci.4
2022 Social influence source locating based on network sparsification and stratification
Ling Chen 0005, Yixin Chen 0001, Wei Liu 0010
Expert Syst. Appl.4
2022 Random walk-based algorithm for distance-aware influence maximization on multiple query locations
Ling Chen 0005, Yixin Chen 0001, Bin Li 0006, Wei Liu 0010
Knowl. Based Syst.5
2022 Weighted Collaborative Sparse and L1/2 Low-Rank Regularizations With Superpixel Segmentation for Hyperspectral Unmixing
abstract
In this letter, using the sparse unmixing framework, a weighted collaborative sparse and$L_{1/2}$low-rank regularization with superpixel segmentation method is proposed for hyperspectral unmixing. The method outlined here first uses superpixel segmentation to obtain local homogeneous regions. The reason for this approach is that the shape and size of superpixels are adaptive, which are better for obtaining homogeneous regions than square patches. Next, the weighted collaborative sparse term and$L_{1/2}$low-rank regularization were utilized to exploit the spatial and spectral correlation of each superpixel. In addition, the smoothness between adjacent pixels is enforced by total variation regularization. Finally, the proposed method and several state-of-the-art methods were tested on two simulated data sets and two real data sets. The results demonstrate the superiority of the method proposed here.
Le Sun 0002, Feiyang Wu, Chengxun He, Tianming Zhan, Wei Liu 0010, Daopan Zhang
IEEE Geosci. Remote. Sens. Lett.5
2021 Negative influence blocking maximization with uncertain sources under the independent cascade model
Ling Chen 0005, Yixin Chen 0001, Bin Li 0006, Wei Liu 0010
Inf. Sci.5
2021 Minimizing the seed set cost for influence spreading with the probabilistic guarantee
Ling Chen 0005, Yixin Chen 0001, Bin Li 0006, Wei Liu 0010
Knowl. Based Syst.5
2021 A random walk-based method for detecting essential proteins by integrating the topological and biological features of PPI network
Nahla Mohamed Ahmed, Ling Chen 0005, Bin Li 0006, Wei Liu 0010, Caiyan Dai
Soft Comput.4
2020 An algorithm for influence maximization in competitive social networks with unwanted users
Wei Liu 0010, Ling Chen 0005, Bolun Chen
Appl. Intell.1
2020 Maximum likelihood-based influence maximization in social networks
Wei Liu 0010, Yun Li 0010
Appl. Intell.1
2020 A new algorithm for positive influence maximization in signed networks
Weijia Ju, Ling Chen 0005, Bin Li 0006, Wei Liu 0010, Jun Sheng
Inf. Sci.4
2020 Positive influence maximization in signed social networks under independent cascade model
Jun Sheng, Ling Chen 0005, Yixin Chen 0001, Bin Li 0006, Wei Liu 0010
Soft Comput.5
2019 Link prediction on signed social networks based on latent space mapping
Shensheng Gu, Ling Chen 0005, Bin Li 0006, Wei Liu 0010, Bolun Chen
Appl. Intell.4
2019 Influence maximization on signed networks under independent cascade model
Wei Liu 0010, Byeungwoo Jeon, Ling Chen 0005, Bolun Chen
Appl. Intell.1
2019 A practical algorithm for solving the sparseness problem of short text clustering
abstract
Dirichlet Multinomial Mixture (DMM) models have been successful in clustering short texts. However, the word co-occurrence information that can be captured by these models is limited to the short text corpus itself. If two words have strong relatedness but rarely co-occurring in short texts, these models can not fully capture the semantic relatedness between the two words. In this paper, we propose a novel model by incorporating word-word correlation into DMM, called WDMM. By constructing a sparse graph using word-word relationship, our model expands each short text using their neighboring words in each text that can help to solve the problem of sparseness in short texts. Therefore, the cluster label of each text is not only influenced by its words, but decided by their similar words in this corpus. Experimental results on real-world datasets demonstrated the substantial superiority of our WDMM model over the state-of-the-art methods.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Wei Liu 0010, Xindong Wu 0001
Intell. Data Anal.4
2019 Multimodel Framework for Indoor Localization Under Mobile Edge Computing Environment
abstract
Location estimation technology under the wireless environment has become a vital technology in the field of mobile edge computing. Especially, under the mobile edge of entire networks environment, indoor location estimation is gradually getting the interest research and application topic, due to technical constraints of global positioning system technology for indoor environment and the popularity of the mobile edge computing servers. In this paper, the widely used single-model framework for indoor localization is presented as an introduction, which consists of three stages: 1) sample data collection; 2) model building; and 3) localization estimation. And then, through analyzing of the actual scene of indoor localization, a new framework for indoor localization under mobile edge computing environment, named Multimodel, is proposed from the theoretical perspective. It is mainly based on the observation that the environment of the sample data collection and that of localization data collection may change seriously. In order to make up for the shortcomings of this framework, two combinatorial optimization problems are proposed. Later, we discuss the NP-hardness of them in several different cases. In addition, two heuristic algorithms are given, and the performance of which are illustrated by the corresponding experimental results.
Wenjun Li 0001, Zhenyu Chen 0003, Xingyu Gao 0001, Wei Liu 0010, Jin Wang 0001
IEEE Internet Things J.4
2018 A link prediction algorithm based on low-rank matrix completion
Man Gao, Ling Chen 0005, Bin Li 0006, Wei Liu 0010
Appl. Intell.4
2018 Snapshot ensembles of non-negative matrix factorization for stability of topic modeling
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Wei Liu 0010
Appl. Intell.4
2018 Detect potential relations by link prediction in multi-relational social networks
Ling Chen 0005, Man Gao, Bin Li 0006, Wei Liu 0010, Bolun Chen
Decis. Support Syst.4
2018 Prediction of protein essentiality by the improved particle swarm optimization
Wei Liu 0010, Jin Wang 0001, Ling Chen 0005, Bolun Chen
Soft Comput.1
2017 Projection-based link prediction in a bipartite network
Man Gao, Ling Chen 0005, Bin Li 0006, Yun Li 0010, Wei Liu 0010, Yongcheng Xu
Inf. Sci.5
2016 A Novel Link Prediction Algorithm Based on Spatial Mapping in PPI Network
Qiang-Mei Wu, Wei Liu 0010, Haiyan Hong, Ling Chen 0005
IDEAL2
2016 Sampling-based algorithm for link prediction in temporal networks
Nahla Mohamed Ahmed Ibrahim, Ling Chen 0005, Bin Li 0006, Yun Li 0010, Wei Liu 0010
Inf. Sci.6
2015 A New Protein-Protein Interaction Prediction Algorithm Based on Conditional Random Field
Wei Liu 0010, Ling Chen 0005, Bin Li 0006
ICIC (2)1
2015 Density-based modularity for evaluating community structure in bipartite networks
Yongcheng Xu, Ling Chen 0005, Bin Li 0006, Wei Liu 0010
Inf. Sci.4
2012 Identifying CpG Islands in Genome Using Conditional Random Fields
Wei Liu 0010, Hanwu Chen, Ling Chen 0005
ICIC (1)1
2011 Efficiently Detecting Frequent Patterns in Biological Sequences
abstract
Most of the existing algorithms for mining frequent patterns could produce lots of projected databases and short candidate patterns which could increase the time and memory cost of mining. In order to overcome such shortcoming, we propose two fast and efficient algorithms named SBPM and MSPM for mining frequent patterns in single and multiple biological respectively. We first present the concept of primary pattern, and then use prefix tree for mining frequent primary patterns. A pattern growth approach is also presented to mine all the frequent patterns without producing large amount of irrelevant patterns. Our experimental results show that our algorithms not only improve the performance but also achieve effective mining results.
Wei Liu 0010, Ling Chen 0005
WISA1
2006 Partitioned optimization algorithms for multiple sequence alignment
abstract
Multiple sequence alignment is an important and difficult problem in molecular biology and bioinformatics. In this paper, we propose a partitioning approach that significantly improves the solution time and quality by utilizing the locality structure of the problem. The algorithm solves the multiple sequence alignment in three stages. First, an automated and suboptimal partitioning strategy is used to divide the set of sequences into several subsections. Then a multiple sequence alignment algorithm based on ant colony optimization is used to align the sequences of each subsection. Finally, the alignment of original sequences can be obtained by assembling the result of each subsection. The ant colony algorithm is highly optimized in order to avoid local optimal traps and converge to global optimal efficiently. Experimental results show that the algorithm can significantly reduce the running time and improve the solution quality on large-scale multiple sequence alignment benchmarks.
Yixin Chen 0001, Yi Pan 0001, Wei Liu 0010, Ling Chen 0005
AINA (2)4
2006 A fast parallel algorithm for finding the longest common sequence of multiple biosequences
abstract
BACKGROUND: Searching for the longest common sequence (LCS) of multiple biosequences is one of the most fundamental tasks in bioinformatics. In this paper, we present a parallel algorithm named FAST_LCS to speedup the computation for finding LCS. RESULTS: A fast parallel algorithm for LCS is presented. The algorithm first constructs a novel successor table to obtain all the identical pairs and their levels. It then obtains the LCS by tracing back from the identical character pairs at the last level. Effective pruning techniques are developed to significantly reduce the computational complexity. Experimental results on gene sequences in the tigr database show that our algorithm is optimal and much more efficient than other leading LCS algorithms. CONCLUSION: We have developed one of the fastest parallel LCS algorithms on an MPP parallel computing model. For two sequences X and Y with lengths n and m, respectively, the memory required is max{4*(n+1)+4*(m+1), L}, where L is the number of identical character pairs. The time complexity is O(L) for sequential execution, and O(|LCS(X, Y)|) for parallel execution, where |LCS(X, Y)| is the length of the LCS of X and Y. For n sequences X1, X2, ..., Xn, the time complexity is O(L) for sequential execution, and O(|LCS(X1, X2, ..., Xn)|) for parallel execution. Experimental results support our analysis by showing significant improvement of the proposed method over other leading LCS algorithms.
Yixin Chen 0001, Andrew Wan, Wei Liu 0010
BMC Bioinform.3