VLDB 2026 Research / reviewers in the wild / expert
Yan Liu 0085
dblp:150/4295-85
· DBLP profile ↗
16ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-1386-812XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Personalized interpretable classification
Zengyou He, Lianyu Hu 0001, Mudi Jiang, Yan Liu 0085 |
Knowl. Inf. Syst. | 6 |
| 2025 | BG2VN: benchmark graph generator for vital node recognition
Yan Liu 0085, Zengyou He |
Frontiers Comput. Sci. | 2 |
| 2025 | Community structure testing by counting frequent common neighbor setsabstractThe detection of communities from a graph is a key issue in network science and graph data mining . However, existing community detection algorithms can always partition a given network/graph into different communities/subgraphs, even when no community structure exists. Obviously, it will lead to fruitless efforts and erroneous conclusions if we conduct the community detection procedure on a network without a community structure. Hence, prior to community detection, it is a must to test whether the community structure is present in the target network. Unfortunately, the community structure testing issue is still not revolved and existing solutions have some limitations. Therefore, we present a new test, which is called FCN (Frequent Common Neighbor) test to tackle the community structure testing problem . In FCN test, the number of FCN sets is employed as the test statistic, which will approximately follows a Poisson distribution when the support threshold is sufficiently large under the null hypothesis that the graph is generated according to the Erdős-Rényi model. We compare the proposed FCN test with existing community structure testing methods on both real networks and simulated networks. The experimental results demonstrate the effectiveness and advantage of our method. Zengyou He, Lianyu Hu 0001, Mudi Jiang, Yan Liu 0085 |
Inf. Sci. | 5 |
| 2025 | Clusterability test for categorical data
Lianyu Hu 0001, Mudi Jiang, Yan Liu 0085, Zengyou He |
Knowl. Inf. Syst. | 4 |
| 2025 | Clustering Categorical Data via Multiple Hypothesis TestingabstractCategorical data clustering is a fundamental data mining problem, which has been extensively studied during the past decades. To date, many effective clustering algorithms for categorical data are available in the literature. However, almost all existing categorical data clustering algorithms did not address the issue of the statistical significance of detected clusters. In particular, how to assess the statistical significance of a set of non-overlapping categorical clusters still remains unaddressed. In this article, we formulate the categorical data clustering problem as a multiple hypothesis testing problem, where the null hypothesis is that each attribute is independent of the given partition of clusters. Then, all individual \(p\) -values from different attributes are integrated to obtain a consensus \(p\) -value through statistical meta-analysis. Thereafter, a significance-based clustering algorithm is proposed in which the combined \(p\) -value is efficiently optimized in an indirectly and incremental manner. Experimental results on 25 real-world datasets demonstrate that our method is capable of achieving comparable performance to state-of-the-art categorical data clustering algorithms. Furthermore, our method has a good capability of determining whether there really exists a clustering structure and assessing whether a given set of clusters is statistically significant. Lianyu Hu 0001, Mudi Jiang, Yan Liu 0085, Quan Zou 0001, Zengyou He |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Integrating topology and biological information to predict essential proteins via Shannon entropyabstractIdentifying essential proteins is vital for deciphering the intricacies of disease mechanisms and devising efficacious therapeutic strategies. Over the past several decades, a plethora of algorithms have been proposed, aimed at synthesizing topological and biological information to address the complex challenge of essential protein identification. Nevertheless, a critical examination of the current methodologies reveals certain limitations: (1) the aggregation of diverse features in various methods often requires parameter tuning to maintain balance, potentially introducing instability and increasing complexity in practical scenarios; (2) traditional methods commonly combine various features without in-depth mathematical or physical interpretation, possibly falling short of fully revealing the principles behind the observed phenomena. Hence, we propose a new algorithm for essential protein detection, which is named ITBSE. The basic idea behind this method is to reconstruct the PPI network by removing false positive edges and the subsequent allocation of a protein score. This scoring process integrates both the topological attributes of the reconstructed PPI network and biological data, utilizing the computational framework of Shannon entropy for a comprehensive assessment. To evaluate the effectiveness of our method, we conduct the experiments on real PPI networks and compare with 10 popular methods including DC, BC, CC, LID, PR, DMNC, LAC, NC, PeC and esPOS. The comparison results demonstrate that ITBSE is able to achieve better performance than those competing algorithms. Yan Liu 0085, Zhong Wang 0001, Zengyou He, Hexin Zhang, Jing Qin 0007 |
BIBM | 1 |
| 2024 | Essential protein discovery on weighted PPI networks via statistical information fusionabstractIdentifying essential proteins is crucial for understanding disease mechanisms and developing therapeutic strategies. During the past several decades, numerous algorithms have been introduced to integrate topological and biological information to tackle the challenge of identifying essential proteins. However, existing methods still have some drawbacks: (1) the lack of rigorous mathematical interpretation for the parameters that determine their respective weights when integrating topological and biological information; (2) the lack of flexibility for adding or removing topological and biological information from integration process in real applications. To overcome these limitations, we propose a novel essential protein discovery method, called EPSIF, which assigns weights to PPI network interactions and integrates diverse topological and biological information via a statistical ensemble model. To assess the performance of EPSIF, we conduct experiments on real PPI networks and compare with eight state-of-the-art algorithms, which confirm the effectiveness and flexibility of EPSIF. Yan Liu 0085, Zhong Wang 0001, Zengyou He, Hexin Zhang, Hongwei Wu, Ya-Dong Wang |
BIBM | 1 |
| 2024 | Central node identification via weighted kernel density estimation
Yan Liu 0085, Jun Lou, Lianyu Hu 0001, Zengyou He |
Data Min. Knowl. Discov. | 1 |
| 2023 | Predicting active enhancers with DNA methylation and histone modificationabstractBACKGROUND: Enhancers play a crucial role in gene regulation, and some active enhancers produce noncoding RNAs known as enhancer RNAs (eRNAs) bi-directionally. The most commonly used method for detecting eRNAs is CAGE-seq, but the instability of eRNAs in vivo leads to data noise in sequencing results. Unfortunately, there is currently a lack of research focused on the noise inherent in CAGE-seq data, and few approaches have been developed for predicting eRNAs. Bridging this gap and developing widely applicable eRNA prediction models is of utmost importance. RESULTS: In this study, we proposed a method to reduce false positives in the identification of eRNAs by adjusting the statistical distribution of expression levels. We also developed eRNA prediction models using joint gene expressions, DNA methylation, and histone modification. These models achieved impressive performance with an AUC value of approximately 0.95 for intra-cell prediction and 0.9 for cross-cell prediction. CONCLUSIONS: Our method effectively attenuates the noise generated by stochastic RNA production, resulting in more accurate detection of eRNAs. Furthermore, our eRNA prediction model exhibited significant accuracy in both intra-cell and cross-cell validation, highlighting its robustness and potential application in various cellular contexts. Ximei Luo, Yan Liu 0085, Quan Zou 0001, Ying Zhang 0060, Lei Xu 0047 |
BMC Bioinform. | 4 |
| 2023 | Mining Statistically Significant Communities From Weighted NetworksabstractAs one of the most important issues in data mining and network science, the community detection problem has been extensively investigated during the past decades. Despite of the success achieved by existing methods, how to directly access the statistical significance of an individual community in a weighted network remains unsolved. To address this issue, we present a new method to calculate the analytical p-value of an individual community in weighted networks. The proposed analytical p-value is able to assess the statistical significance that one target community appears in a random weighted graph in a straightforward manner. To verify the effectiveness of the proposed p-value in community evaluation, it is utilized as the objective function in a local search procedure to derive a new community detection algorithm. Experimental results show that the new algorithm is able to achieve comparable performance to those state-of-the-art algorithms for identifying communities from weighted networks. The source codes of our method are available at: https://github.com/chenwenfang/MSSC. Zengyou He, Wenfang Chen, Xiaoqi Wei, Yan Liu 0085 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | On the Statistical Significance of a Community StructureabstractThe community structure typically refers to the existence of a network partition in terms of a set of non-overlapping dense sub-graphs, where each sub-graph is called a community and there are few links between different communities. The detection of community structure is able to provide additional knowledge on the organization mechanism of the network and its characteristics. Despite decades of developments in community detection algorithms, how to determine whether a given community structure is true or not in a statistically sound manner still remains unresolved. In this paper, we present an analytical upper bound on thep-value of a community structure under the configuration model. To demonstrate its effectiveness on community structure validation, we further develop a community detection algorithm in which thep-value upper bound is used as the objective function. Experimental results on both real networks and simulated networks show that our algorithm outperforms prior state-of-the-art community detection methods. Zengyou He, Xiaoqi Wei, Wenfang Chen, Yan Liu 0085 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Decision Tree for SequencesabstractCurrent decision trees such as C4.5 and CART are widely used in different fields due to their simplicity, accuracy and intuitive interpretation. Similar to other popular classifiers, these tree-based classification algorithms are developed for fixed-length vector data and suffer from intrinsic limitations in handling complex data such as sequences. To tackle the discrete sequence classification task, the dominant strategy is to adopt a two-step procedure: first transform the sequential dataset into a vector dataset and then apply existing tree-based classifiers on the new vector data. However, such methods are highly dependent on the feature generation procedure and some features that are critical to the tree construction may be missed. To alleviate these issues, we present a new tree-based sequence classification method, which is able to construct a concise decision tree from the feature space that is composed of all subsequences present in the training sequences. Experimental results on fourteen real datasets show that our method can achieve better performance than those state-of-the-art sequence classification algorithms. The source codes of our method are available at: https://github.com/ZiyaoWu/SeqDT. Zengyou He, Ziyao Wu, Guangyao Xu, Yan Liu 0085, Quan Zou 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Significance-Based Essential Protein DiscoveryabstractThe identification of essential proteins is an important problem in bioinformatics. During the past decades, many centrality measures and algorithms have been proposed to address this issue. However, existing methods still deserve the following drawbacks: (1) the lack of a context-free and readily interpretable quantification of their centrality values; (2) the difficulty of specifying a proper threshold for their centrality values; (3) the incapability of controlling the quality of reported essential proteins in a statistically sound manner. To overcome the limitations of existing solutions, we tackle the essential protein discovery problem from a significance testing perspective. More precisely, the essential protein discovery problem is formulated as a multiple hypothesis testing problem, where the null hypothesis is that each protein is not an essential protein. To quantify the statistical significance of each protein, we present a p-value calculation method in which both the degree and the local clustering coefficient are used as the test statistic and the Erdös-Rényi model is employed as the random graph model. After calculating the p-value for each protein, the false discovery rate is used as the error rate in the multiple testing correction procedure. Our significance-based essential protein discovery method is named as SigEP, which is tested on both simulated networks and real PPI networks. The experimental results show that our method is able to achieve better performance than those competing algorithms. Yan Liu 0085, Quan Zou 0001, Zengyou He |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | Detecting Statistically Significant CommunitiesabstractCommunity detection is a key data analysis problem across different fields. During the past decades, numerous algorithms have been proposed to address this issue. However, most work on community detection does not address the issue of statistical significance. Although some research efforts have been made towards mining statistically significant communities, deriving an analytical solution of$p$-value for one community under the configuration model is still a challenging mission that remains unsolved. The configuration model is a widely used random graph model in community detection, in which the degree of each node is preserved in the generated random networks. To partially fulfill this void, we present a tight upper bound on the$p$-value of a single community under the configuration model, which can be used for quantifying the statistical significance of each community analytically. Meanwhile, we present a local search method to detect statistically significant communities in an iterative manner. Experimental results demonstrate that our method is comparable with the competing methods on detecting statistically significant communities. Zengyou He, Can Zhao 0007, Yan Liu 0085 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | Essential Protein Recognition via Community SignificanceabstractEssential protein plays a vital role in understanding the cellular life. With the advance in high-throughput technologies, a number of protein-protein interaction (PPI) networks have been constructed such that essential proteins can be identified from a system biology perspective. Although a series of network-based essential protein discovery methods have been proposed, these existing methods still have some drawbacks. Recently, it has been shown that the significance-based method SigEP is promising on overcoming the defects that are inherent in currently available essential protein identification methods. However, the SigEP method is developed under the unrealistic Erdös-Rényi (E-R) model and its time complexity is very high. Hence, we propose a new significance-based essential protein recognition method named EPCS in which the essential protein discovery problem is formulated as a community significance testing problem. Experimental results on four PPI networks show that EPCS performs better than nine state-of-the-art essential protein identification methods and the only significance-based essential protein identification method SigEP. Yan Liu 0085, Wenfang Chen, Zengyou He |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Computing exact P-values for community detection
Zengyou He, Can Zhao 0007, Yan Liu 0085 |
Data Min. Knowl. Discov. | 5 |