Zhongmin Yan

dblp:64/6062 · DBLP profile ↗
← Back
20ranked-venue papers in the field
1as first author
14since 2021 · last 2025
0000-0002-5271-5417ORCID · corroborated

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 8Information Retrieval & Web Search · 4 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 2Database Systems & Data Management · 1Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2025 Few-shot partial multi-label learning with credible non-candidate label
abstract
Partial multi-label learning (PML) addresses scenarios where each training sample is associated with multiple candidate labels, but only a subset are ground-truth labels. The primary difficulty in PML is to mitigate the negative impact of noisy labels. Most existing PML methods rely on sufficient samples to train a noise-robust multi-label classifier. However, in practical scenarios, such as privacy-sensitive domains or those with limited data, only a few training samples are typically available for the target task. In this paper, we propose an approach called FsPML-CNL (Few-shot Partial Multi-label Learning with Credible Non-candidate Label) to tackle the PML problem with few-shot training samples. Specifically, FsPML-CNL first utilizes the sample features and feature-prototype similarity in the embedding space to disambiguate candidate labels and to obtain label prototypes. Then, the credible non-candidate label is selected based on label correlation and confidence, and its prototype is incorporated into the training samples to generate new data for boosting supervised information . The noise-tolerant multi-label classifier is finally induced with the original and generated samples, along with the confidence-guided loss. Extensive experiments on public datasets demonstrate that FsPML-CNL outperforms competitive baselines across different settings.
Meng Wang 0001, Zhongmin Yan, Guoxian Yu
Inf. Sci.3
2025 Multimodal contrastive learning with hyperbolic geometry for KG-based game recommendation
Yuliang Shi, Jihu Wang, Han Yu 0001, Xinjun Wang 0003, Zhongmin Yan, Fanyu Kong 0002
Knowl. Inf. Syst.6
2025 Cross-space topological contrastive learning for knowledge graph-aware issue recommendation
Leihong Zhang, Yuliang Shi, Kaiyuan Qi, Xinjun Wang 0003, Zhongmin Yan
Knowl. Inf. Syst.6
2024 Few-shot partial multi-label learning with synthetic features network
Yifan Sun 0012, Guoxian Yu, Zhongmin Yan, Carlotta Domeniconi
Knowl. Inf. Syst.4
2023 Sample and Feature Enhanced Few-Shot Knowledge Graph Completion
Daokun Zhang, Ning Liu 0014, Yonghua Yang, Zhongmin Yan, Hui Li 0048, Li-Zhen Cui 0001
DASFAA (2)6
2023 Causal Discovery by Graph Attention Reinforcement Learning
abstract
Discovery the causal structure graph among a set of variables is a fundamental but difficult task in many empirical sciences. Reinforcement learning based causal discovery from observed data achieves prominent results. However, previous algorithms lack interpretability and efficiency, and ignore the prior knowledge of causal structure. To solve these problems, we propose GARL that leverages graph attention network to embed the structure information and the prior knowledge, and reinforcement learning to search the variable ordering with the best score. GARL takes the structure information and prior knowledge as the computational skeleton of attention to obtain the embedded representation of variables, and then generates variable orderings through the designed ordering model. In addition, the structure information is used to form the DAG corresponding to the variable ordering, which reduces the computational difficulty and improves the efficiency. GARL generates DAGs in the reinforcement learning framework, and uses the score of DAG as the reward to optimize the network structure to search the DAG with the best score. Experimental results on synthetic and real datasets show that our GARL has obvious advantages in multi-node operation efficiency, and competitive results with competitive baselines.
Dezhi Yang, Guoxian Yu, Jun Wang 0035, Zhongmin Yan, Maozu Guo 0001
SDM4
2023 Mixed-Curvature Manifolds Interaction Learning for Knowledge Graph-aware Recommendation
abstract
As auxiliary collaborative signals, the entity connectivity and relation semanticity beneath knowledge graph (KG) triples can alleviate the data sparsity and cold-start issues of recommendation tasks. Thus many works consider obtaining user and item representations via information aggregation on graph-structured data within Euclidean space. However, the scale-free graphs (e.g., KGs) inherently exhibit non-Euclidean geometric topologies, such as tree-like and circle-like structures. The existing recommendation models built in a single type of embedding space do not have enough capacity to embrace various geometric patterns, consequently, resulting in suboptimal performance. To address this limitation, we propose a KG-aware recommendation model with mixed-curvature manifolds interaction learning, namely CurvRec. On the one hand, it aims to preserve various global geometric structures in KG with mixed-curvature manifold spaces as the backbone. On the other hand, we integrate Ricci curvature into graph convolutional networks (GCNs) to capture local geometric structural properties when aggregating neighbor nodes. Besides, to exploit the expressive spatial features in KG, we incorporate interaction learning to ensure the geometric message passing between curved manifolds. Specifically, we adopt curvature-aware geodesic distance metrics to maximize the mutual information between Euclidean space and non-Euclidean spaces. Through extensive experiments, we demonstrate that the proposed CurvRec outperforms state-of-the-art baselines.
Jihu Wang, Yuliang Shi, Han Yu 0001, Xinjun Wang 0003, Zhongmin Yan, Fanyu Kong 0002
SIGIR5
2023 KLECA: knowledge-level-evolution and category-aware personalized knowledge recommendation
Lin Cheng 0007, Yuliang Shi, Lin Li 0013, Han Yu 0001, Xinjun Wang 0003, Zhongmin Yan
Knowl. Inf. Syst.6
2023 Few-shot partial multi-label learning via prototype rectification
Guoxian Yu, Lei Liu 0003, Zhongmin Yan, Carlotta Domeniconi, Xiayan Zhang, Li-Zhen Cui 0001
Knowl. Inf. Syst.4
2022 Few-shot Partial Multi-label Learning with Data Augmentation
abstract
Partial multi-label learning (PML) models the scenario where each training sample is annotated with a set of candidate labels, but only a subset of them corresponds to the ground-truths. The key challenge for PML is how to minimize the negative impact of incorrect labels concealed within the candidate ones. Most existing PML solutions require abundant samples to train a noise-robust multi-label predictor. However, due to privacy, safety or ethic issues, we more often have a handful of training samples for the target task. In this paper, we propose an approach named FsPML-DA (Few-shot Partial Multi-Label Learning with Data Augmentation) to simultaneously estimate label confidence, perform data augmentation and induce multilabel classifier. Specifically, FsPML-DA disambiguates the label confidence vector of each PML sample by jointly modeling the feature and semantic similarity, label credibility of other samples and label co-occurrence. Next, FsPML-DA introduces a synthetic feature network to generate more training samples from pairs of given samples with label confidence values. FsPML-DA then leverages original and generated samples to train a noise-tolerant multi-label classifier. Extensive experiments on benchmark datasets show that FsPML-DA performs better than recent competitive PML baselines and few-shot solutions. FsPML-DA can dislodge noisy labels by mining PML data in a sensible way and the proposed data augmentation strategy effectively combats with the scarcity of few-shot training samples.
Yifan Sun 0012, Guoxian Yu, Zhongmin Yan, Carlotta Domeniconi
ICDM4
2022 Joint Attention Networks with Inherent and Contextual Preference-Awareness for Successive POI Recommendation
abstract
Abstract Nowadays recording and sharing personal lives using mobile devices on the Internet is becoming increasingly popular, and successive POI recommendation is gaining growing attention from academia and industry. In mobile scenarios, multiple influencing factors including the diversity of user preferences, the changeability of user behavior and the dynamic of spatiotemporal context bring great challenges to the POI recommender system. In order to accurately capture both the stable and the contextual preferences of mobile users in dynamic contexts, we propose a fusion framework JANICP (Joint Attention Networks with Inherent and Contextual Preferences) for successive POI recommendation by jointly training an offline/nearline user inherent interest perception model and an online user contextual interest prediction model. The offline model is trained based on the global historical behavior data to achieve stable interest representation, while the online model is trained based on the instantly selected context-sensitive data to achieve dynamic interest perception. An attention aggregation and matching module is used to fully connect the two kinds of preference representations and generate the final POI recommendation. Extensive experiments were conducted on three real datasets and experimental results show that the proposed JANICP outperforms existing state-of-the-art methods.
Haiting Zhong, Wei He 0020, Li-Zhen Cui 0001, Lei Liu 0003, Zhongmin Yan
Data Sci. Eng.5
2021 Few-Shot Partial Multi-Label Learning
abstract
Partial multi-label learning (PML) aims at learning a robust multi-label classifier by training on ambiguous data, where each sample is associated with a set of candidate labels, among which only a subset are valid labels. A basic premise of existing PML solutions is to obtain enough partial multi-label samples for inducing the classification model. However, when dealing with new tasks, we may only have a few PML samples for those tasks. Furthermore, existing few-shot learning approaches assume the support (training) samples are precisely labeled; as such, irrelevant labels in the candidate label set may seriously mislead the meta-learner and thus result in a compromised performance. How to achieve PML with limited few-shot support samples is an important and practical problem, but not yet well studied. In this paper, we propose an approach called FsPML (Few-shot PML) to tackle this problem. Specifically, FsPML first performs adaptive distance metric learning via an embedding network using both sample features and label semantics in the embedding space. Next it rectifies the positive and negative prototypes of each new label of the target task in the embedding space. An unseen example can then be classified via its distances to the positive and to the negative prototypes. Experimental results on widely-used multi-label datasets (MS COCO and NUS-WIDE) demonstrate that our FsPML outperforms competitive baselines across different settings, and it can quickly generalize to new tasks with fewer training samples.
Guoxian Yu, Lei Liu 0003, Zhongmin Yan, Carlotta Domeniconi, Li-Zhen Cui 0001
ICDM4
2021 DEKR: Description Enhanced Knowledge Graph for Machine Learning Method Recommendation
abstract
The huge number of machine learning (ML) methods has resulted in significant information overload. Faced with an overwhelming number of ML methods, it is challenging to select appropriate ones for the given dataset and task. In general, the names of ML methods or datasets are rather condensed, thus lacking specific explanations, while the rich latent relationships between ML entities are not fully explored. In this paper, we propose a description-enhanced machine learning knowledge graph-based approach - DEKR - to help recommend appropriate ML methods for given ML datasets. The proposed knowledge graph (KG) not only includes the connections between entities but also contains the descriptions of the dataset and method entities. DEKR fuses the structural information with the description information of entities in the knowledge graph. It is a deep hybrid recommendation framework, which incorporates the knowledge graph-based and text-based methods, overcoming the limitations of previous knowledge graph-based recommendation systems that ignore the description information. There are two key components of DEKR: 1) a graph neural network aggregating information from multi-order neighbors with attention to enrich the seed (i.e. dataset or method) node's own representation, and 2) a deep collaborative filtering network based on the description text to obtain the linear and nonlinear interactions of description features. Through extensive experiments, we demonstrated the efficiency of DEKR, which outperforms the current state-of-the-art baselines by a large margin.
Xianshuai Cao, Yuliang Shi, Han Yu 0001, Jihu Wang, Xinjun Wang 0003, Zhongmin Yan
SIGIR6
2021 Cost-effective multi-instance multilabel active learning
abstract
Multi-instance multi-label (MIML) Active Learning (M2AL) aims to improve the learner while reducing the cost as much as possible by querying informative labels of complex bags composed of diverse instances. Existing M2AL solutions suffer high query costs for scrutinizing all relevant labels of MIML samples, querying excessive bag–label or instance–label pairs. To address these issues, a Cost-effective M2AL solution (CM2AL) is presented. CM2AL first selects the most informative bag–label pairs by leveraging uncertainty, label correlations, label space sparsity, and informativeness from queried instances of the bag, and thus avoids scrutinizing all labels. Next, it queries the most probably positive instance–label pairs of the selected bag–label pair. Particularly, if the feedback is positive, the bag is positively annotated with the label. For negative feedback, it further leverages the label of the neighborhood bags and the label of the nearby instances of queried instances of this bag, if the suggested labels from bag- and instance-levels disagree, CM2AL temporally gives up querying this bag–label pair and moves to another most informativeness one; otherwise, it takes the agreed label to annotate the bag, which further saves the cost by avoiding the excessive query. Extensive experiments on MIML data sets from diverse domains show that CM2AL can more reduce the cost while managing a better performance than state-of-the-art methods, the collaboration between bags and instances contributes to the saved cost.
Cong Su, Zhongmin Yan, Guoxian Yu
Int. J. Intell. Syst.2
2020 Collaborative Filtering: Graph Neural Network with Attention
Yanli Guo, Zhongmin Yan
WISA2
2019 Power Demand Response Incentive Pricing Model
abstract
Power demand response aims to clip the peak and fill the valley of power load by interruptible load management. Its closely related to the safety and economic benefit of the power system. At present, various provinces in China have carried out the pilot work of interruptible load management. The power companies signed contracts with power users to stipulate power users to adjust the power load consumption during peak hours or in emergency situations. The singed contracts should be fair for all power users. At the same time, the specific information about the signed contracts should be protected for privacy preservation, which is a typical federated learning scenario. A fair and privacy-preserving incentive pricing model should be proposed to stimulate power uses to participate in thee interruptible load management. To this end, a power demand response incentive pricing model is proposed. Based on the multi-attribute sealed auction game, the bidding model between the power company and the power users is established. In the single bidding model, the risk of the power user's demand response is evaluated firstly. The power users could be classified into different group according to their power load characteristics. Then a user classification selection algorithm' is proposed, which enables both the power company's revenue and power user risk to be considered, and enables the selected users to achieve balanced peak clipping. A fair mechanism based on integrals is proposed to assure the fairness between all power users. The case simulations demonstrate that effectiveness of the proposed incentive pricing model.
Kun Zhang 0013, Yuliang Shi, Yuecan Liu, Zhongmin Yan
IEEE BigData4
2015 Matching Reviews to Object Based on 2-Stage CRF
Qingzhong Li, Dequan Wang, Yanhui Ding, Congli Liu, Zhongmin Yan
APWeb6
2015 An Effective Hybrid Fraud Detection Method
abstract
The rapid growth of data makes it possible for us to study human behavior patterns. Knowing the patterns of human behavior is of great use to help us detect the unusual fraud human behavior. Existing fraud detection methods can be divided into two categories: pattern based and outlier detection based methods. However, because of the sparsity and complex granularity of big data, these methods have high false positive in fraud detection. In this paper, we propose an effective hybrid fraud detection method. We propose SSIsomap which improves isomap to cluster behaviors into behavior classes and propose SimLOF which improves LOF to conduct outlier detection, then we use Dempster-Shafer evidence Theory for combining behavior pattern evidence and outlier evidence, which yields a degree of belief of fraud to the new coming claim. The experiment result shows our method has significantly higher accuracy than exsiting methods in medical insurance fraud detection.
Chenfei Sun, Qingzhong Li, Li-Zhen Cui 0001, Zhongmin Yan, Hui Li 0048
KSEM4
2012 A Novel URL Assignment Model Based on Multi-objective Decision Making Method
abstract
With the tremendous growth of the Web, it has become a huge challenge for the single-process crawlers to locate the resources that are precise and relevant to some topics in an appropriate amount of time, so it is increasingly important to use the parallel crawler. However, due to the parallelism of crawlers, one headache problem we have to face is how to distribute the URLs to crawlers to make the parallel system work coordinately and thereby make sure that the Web pages fetched are of high quality. In this paper, a novel URL assignment model for the parallel crawler is described, which is based on multi-objective decision making method and considers multiple factors synthetically such as load balance, overlap and so on. Extensive experiments test and validate our techniques.
Qiuyan Huang, Qingzhong Li, Zhongmin Yan
WISA3
2010 MI-WDIS: web data integration system for market intelligence
abstract
As an important supporting technology of Market Intelligence (MI), Web data integration is facing new challenges, such as the integrity of data acquisition, the quality of data extraction and data consolidation. To solve such problems, we propose an MI-oriented web data integration system (MI-WDIS), which achieves excellent performances in integrating Surface Web and Deep Web data with much less manual work. Based on MI-WDIS, we have developed a platform for intelligent analysis of job data. The platform collects tens of thousands of job data daily and provides personalized services for job seekers through diversified channels. Besides, it provides other advanced services, including intelligence analysis, automatic monitoring and alerting, for various organizations, such as enterprises, training institutions and recruitment agencies.
Zhongmin Yan, Qingzhong Li, Shidong Zhang, Zhaohui Peng, Yongquan Dong, Yanhui Ding, Xiuxing Xu
CIKM1