Lianyu Hu 0001

dblp:238/9285-1 · DBLP profile ↗
← Back
15ranked-venue papers in the field
3as first author
15since 2021 · last 2026
0000-0001-7470-9395ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 6 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 6 (1 first)Database Systems & Data Management · 3
YearPublicationVenuePosition
2026 Personalized interpretable classification
Zengyou He, Lianyu Hu 0001, Mudi Jiang, Yan Liu 0085
Knowl. Inf. Syst.4
2026 Rule-based interpretable sequence clustering
Yushuang Liu, Mudi Jiang, Lianyu Hu 0001, Zengyou He
Knowl. Inf. Syst.6
2026 Clustering With Multiview Explanations
abstract
Explainable clustering has become increasingly important as it presents the clustering results in a manner that can be easily understood by the end-users. However, most existing explainable clustering approaches generate only a single explanation such as a decision tree for the given clustering result, overlooking that multiple valid interpretations may exist. To address this limitation, we propose the CME (Clustering with Multiview Explanations) algorithm. This method constructs a diverse set of candidate decision trees by retaining multiple splitting points at each node, where each splitting point is valid in a statistical sense. The tree similarity based on optimal path matches across corresponding clusters is then used to build a similarity graph. A subset of representative decision trees are selected by solving the minimum dominating set problem defined on the corresponding similarity graph. Experiments on 10 real-world categorical datasets demonstrate that CME provides multiple complementary explanations that may be missed by existing algorithms. Moreover, these additional decision trees can be more accurate and interpretable than those ones identified by baseline methods.
Lianyu Hu 0001, Mudi Jiang, Jun Lou, Zengyou He
IEEE Trans. Knowl. Data Eng.2
2026 Interpretable Sequence Classification via Decision Set
abstract
Sequence classification is a fundamental research issue in data mining and machine learning. However, existing sequence classification methods primarily focus on improving the prediction performance. Although a few methods attempt to enhance the interpretability, they often fail to provide intuitive model explanations and come with high computational costs. To fill this gap, we propose an interpretable sequence classification algorithm based on decision set. Each rule in the decision set is only associated with one discriminative pattern (subsequence) and the classification decision is made based on one best-matched rule. Hence, the proposed method has good interpretablility since the classification decision is solely determined by one simple intuitive rule. Experimental results on real-world data sets demonstrate that our algorithm outperforms the state-ofthe- art interpretable sequence classification methods in terms of both interpretability and classification accuracy.
Mudi Jiang, Lianyu Hu 0001, Zengyou He
IEEE Trans. Knowl. Data Eng.5
2026 Clustering Validation via Sample Pair Co-Cluster Testing
abstract
Clustering validation is a fundamental task in cluster analysis. While many clustering validity indices have been proposed, most existing internal validity indices are typically defined heuristically, lacking solid statistical foundation and interpretation. To address this limitation, we introduce a new internal validity index that employs hypothesis testing to determine whether two samples belong to the same cluster and defines the index as the proportion of correctly identified sample pairs based on their cluster memberships. To demonstrate the advantages of proposed validity index, we conduct experiments on various synthetic and real-world data sets. The experimental results indicate that our validation index can beat both classic and state-of-the-art internal validation indices.
Lianyu Hu 0001, Mudi Jiang, Zengyou He
IEEE Trans. Knowl. Data Eng.2
2025 Hamming encoder: mining discriminative k-mers for discrete sequence classification
Mudi Jiang, Lianyu Hu 0001, Zengyou He
Data Min. Knowl. Discov.3
2025 Interpretable sequence clustering
Xinyi Yang 0006, Mudi Jiang, Lianyu Hu 0001, Zengyou He
Inf. Sci.4
2025 Significance-based interpretable sequence clustering
Zengyou He, Lianyu Hu 0001, Jinfeng He, Mudi Jiang
Inf. Sci.2
2025 Conjunction subspaces test for conformal and selective classification
Zengyou He, Zerun Li, Mudi Jiang, Lianyu Hu 0001
Inf. Sci.6
2025 Community structure testing by counting frequent common neighbor sets
abstract
The detection of communities from a graph is a key issue in network science and graph data mining . However, existing community detection algorithms can always partition a given network/graph into different communities/subgraphs, even when no community structure exists. Obviously, it will lead to fruitless efforts and erroneous conclusions if we conduct the community detection procedure on a network without a community structure. Hence, prior to community detection, it is a must to test whether the community structure is present in the target network. Unfortunately, the community structure testing issue is still not revolved and existing solutions have some limitations. Therefore, we present a new test, which is called FCN (Frequent Common Neighbor) test to tackle the community structure testing problem . In FCN test, the number of FCN sets is employed as the test statistic, which will approximately follows a Poisson distribution when the support threshold is sufficiently large under the null hypothesis that the graph is generated according to the Erdős-Rényi model. We compare the proposed FCN test with existing community structure testing methods on both real networks and simulated networks. The experimental results demonstrate the effectiveness and advantage of our method.
Zengyou He, Lianyu Hu 0001, Mudi Jiang, Yan Liu 0085
Inf. Sci.3
2025 Significance-based decision tree for interpretable categorical data clustering
Lianyu Hu 0001, Mudi Jiang, Zengyou He
Inf. Sci.1
2025 Clusterability test for categorical data
Lianyu Hu 0001, Mudi Jiang, Yan Liu 0085, Zengyou He
Knowl. Inf. Syst.1
2025 Clustering Categorical Data via Multiple Hypothesis Testing
abstract
Categorical data clustering is a fundamental data mining problem, which has been extensively studied during the past decades. To date, many effective clustering algorithms for categorical data are available in the literature. However, almost all existing categorical data clustering algorithms did not address the issue of the statistical significance of detected clusters. In particular, how to assess the statistical significance of a set of non-overlapping categorical clusters still remains unaddressed. In this article, we formulate the categorical data clustering problem as a multiple hypothesis testing problem, where the null hypothesis is that each attribute is independent of the given partition of clusters. Then, all individual \(p\) -values from different attributes are integrated to obtain a consensus \(p\) -value through statistical meta-analysis. Thereafter, a significance-based clustering algorithm is proposed in which the combined \(p\) -value is efficiently optimized in an indirectly and incremental manner. Experimental results on 25 real-world datasets demonstrate that our method is capable of achieving comparable performance to state-of-the-art categorical data clustering algorithms. Furthermore, our method has a good capability of determining whether there really exists a clustering structure and assessing whether a given set of clusters is statistically significant.
Lianyu Hu 0001, Mudi Jiang, Yan Liu 0085, Quan Zou 0001, Zengyou He
ACM Trans. Knowl. Discov. Data1
2024 Central node identification via weighted kernel density estimation
Yan Liu 0085, Jun Lou, Lianyu Hu 0001, Zengyou He
Data Min. Knowl. Discov.4
2024 Random subsequence forests
Zengyou He, Mudi Jiang, Lianyu Hu 0001, Quan Zou 0001
Inf. Sci.4