VLDB 2026 Research / reviewers in the wild / expert
Li-Yan Liu
dblp:269/8825
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2023
0000-0002-3409-7133ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Safe co-training for semi-supervised regressionabstractCo-training is a popular semi-supervised learning method. The learners exchange pseudo-labels obtained from different views to reduce the accumulation of errors. One of the key issues is how to ensure the quality of pseudo-labels. However, the pseudo-labels obtained during the co-training process may be inaccurate. In this paper, we propose a safe co-training (SaCo) algorithm for regression with two new characteristics. First, the safe labeling technique obtains pseudo-labels that are certified by both views to ensure their reliability. It differs from popular techniques of using two views to assign pseudo-labels to each other. Second, the label dynamic adjustment strategy updates the previous pseudo-labels to keep them up-to-date. These pseudo-labels are predicted using the augmented training data. Experiments are conducted on twelve datasets commonly used for regression testing. Results show that SaCo is superior to other co-training style regression algorithms and state-of-the-art semi-supervised regression algorithms. Li-Yan Liu, Hong Yu 0007, Fan Min 0001 |
Intell. Data Anal. | 1 |
| 2022 | Safe Multi-view Co-training for Semi-supervised RegressionabstractCo-training is a popular disagreement-based semi-supervised learning method. Learners of different views mutually select reliable unlabeled instances to augment the labeled dataset. Existing co-training style algorithms have cumbersome procedures for selecting confident instances. Furthermore, the pseudo-labels assigned to selected unlabeled instances are not always reliable. In this paper, we propose a safe co-training regression algorithm for multi-view scenarios with two characteristics. An instance selection strategy based on the consistency assumption aims to improve the efficiency of selecting confident unlabeled instances. This strategy makes full use of the information provided by a committee to measure the confidence of unlabeled instances. A safe labeling technique in an ensemble manner is introduced to improve the quality of pseudo-labels. The safe pseudo-labels not only integrate information provided by the committee, but also take into account the part of the receiver. The results over twenty datasets prove the superiority of the proposed algorithm against other state-of-the-art semi-supervised regression algorithms. Li-Yan Liu, Fan Min 0001 |
DSAA | 1 |
| 2022 | Semi-supervised Regression with Data Partitioning and Feature MappingabstractSemi-supervised regression attempts to utilize as much unlabeled data as possible and as little labeled data as possible to improve model performance. Methods based on data partitioning can improve regression performance from the perspective of data distribution. However, most partitioning methods only consider the correlation between the data, but not the relationship between the regressor and the data. In this study, a strategy of dividing the data and then regressing is proposed, where data partitioning is based on the relationship between the regression and data. Accordingly, an intuitive and effective algorithm named SRPF, i.e. Semi-supervised regression based on data partitioning and feature mapping, is proposed. First, we divide the labeled dataset into two disjoint subsets based on the relationship between the predicted and actual values of separation regressor. Second, we label these two subsets as distinct classes and use feature mapping to map data features to higher dimensions to better distinguish the data. Third, we construct a partitioner to determine which subset an unlabeled instance belongs to, and then use the regressor on the corresponding data set to make predictions. Finally, during the iterative training process, a self-training method is used to enrich labeled samples. Experiments are conducted on 15 well-known datasets compared to state-of-the-art algorithms. The results show that our method outperforms them in most datasets. Li-Yan Liu, Jia-Hui Zhang, Fan Min 0001 |
DSAA | 1 |