VLDB 2026 Research / reviewers in the wild / expert
Guowen Yuan
dblp:204/2946
· DBLP profile ↗
10ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0002-8005-4433ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Data mining · 100% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › dimensionality reduction
feature selection |
1.8 | 5 | 2020 | Semi-Supervised Feature Selection via Sparse Rescaled Linear Square Regression · IEEE Trans. Knowl. Data Eng. 2020 Semi-Supervised Feature Selection with Adaptive Discriminant Analysis · AAAI 2019 Discriminative Semi-Supervised Feature Selection via Rescaled Least Squares Regression-Supplement · AAAI 2018 |
Data mining › dimensionality reduction › feature selection › weakly supervised feature selection
semi-supervised feature selection |
1.4 | 4 | 2020 | Semi-Supervised Feature Selection via Sparse Rescaled Linear Square Regression · IEEE Trans. Knowl. Data Eng. 2020 Semi-Supervised Feature Selection with Adaptive Discriminant Analysis · AAAI 2019 Discriminative Semi-Supervised Feature Selection via Rescaled Least Squares Regression-Supplement · AAAI 2018 |
Natural language and speech › Information extraction and text analysis › data annotation
text annotation |
0.7 | 1 | 2023 | CEMA - Cost-Efficient Machine-Assisted Document Annotations · AAAI 2023 |
Data mining
dimensionality reduction |
0.4 | 1 | 2019 | Semi-Supervised Feature Selection with Adaptive Discriminant Analysis · AAAI 2019 |
Data mining › dimensionality reduction › feature selection
feature ranking |
0.3 | 1 | 2018 | A Stratified Feature Ranking Method for Supervised Feature Selection · AAAI 2018 |
Data mining › dimensionality reduction › feature selection
supervised feature selection |
0.3 | 1 | 2018 | A Stratified Feature Ranking Method for Supervised Feature Selection · AAAI 2018 |
Mathematical optimization › least squares
regularized least squares |
0.1 | 1 | 2020 | Semi-Supervised Feature Selection via Sparse Rescaled Linear Square Regression · IEEE Trans. Knowl. Data Eng. 2020 |
Mathematical optimization › statistical estimation › regression
sparse regression |
0.1 | 1 | 2017 | Semi-supervised Feature Selection via Rescaled Linear Regression · IJCAI 2017 |
Methods — techniques the papers use, named apart from their topics
cost estimation · 1.3active learning · 1.3sparse rescaled linear square regression · 0.9l2,p-norm regularization · 0.9rescaled linear regression · 0.6least squares regression · 0.6iterative optimization · 0.4adaptive similarity matrix · 0.4ε-dragging · 0.3subspace feature clustering · 0.3stratified ranking · 0.3rescaled least squares regression · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | CEMA - Cost-Efficient Machine-Assisted Document AnnotationsabstractWe study the problem of semantically annotating textual documents that are complex in the sense that the documents are long, feature rich, and domain specific. Due to their complexity, such annotation tasks require trained human workers, which are very expensive in both time and money. We propose CEMA, a method for deploying machine learning to assist humans in complex document annotation. CEMA estimates the human cost of annotating each document and selects the set of documents to be annotated that strike the best balance between model accuracy and human cost. We conduct experiments on complex annotation tasks in which we compare CEMA against other document selection and annotation strategies. Our results show that CEMA is the most cost-efficient solution for those tasks. Guowen Yuan, Ben Kao, Tien-Hsuan Wu |
AAAI | 1 |
| 2022 | Semisupervised Feature Selection With Sparse Discriminative Least Squares RegressionabstractIn big data time, selecting informative features has become an urgent need. However, due to the huge cost of obtaining enough labeled data for supervised tasks, researchers have turned their attention to semisupervised learning, which exploits both labeled and unlabeled data. In this article, we propose a sparse discriminative semisupervised feature selection (SDSSFS) method. In this method, the$\epsilon $-dragging technique for the supervised task is extended to the semisupervised task, which is used to enlarge the distance between classes in order to obtain a discriminative solution. The flexible$\ell _{2,p}$norm is implicitly used as regularization in the new model. Therefore, we can obtain a more sparse solution by setting smaller$p$. An iterative method is proposed to simultaneously learn the regression coefficients and$\epsilon $-dragging matrix and predicting the unknown class labels. Experimental results on ten real-world datasets show the superiority of our proposed method. Chen Wang 0032, Xiaojun Chen 0006, Guowen Yuan, Feiping Nie 0001, Min Yang 0007 |
IEEE Trans. Cybern. | 3 |
| 2021 | Semantic Search and Summarization of Judgments Using Topic ModelingabstractOnline legal document libraries, such as WorldLII, are indispensable tools for legal professionals to conduct legal research. We study how topic modeling techniques can be applied to such platforms to facilitate searching of court judgments. Specifically, we improve search effectiveness by matching judgments to queries at semantics level rather than at keyword level. Also, we design a system that summarizes a retrieved judgment by highlighting a small number of paragraphs that are semantically most relevant to the user query. This summary serves two purposes: (1) It explains to the user why the machine finds the retrieved judgment relevant to the user’s query, and (2) it helps the user quickly grasp the most salient points of the judgment, which significantly reduces the amount of time needed by the user to go through the returned search results. We further enhance our system by integrating domain knowledge provided by legal experts. The knowledge includes the features and aspects that are most important for a given category of judgments. Users can then view a judgement’s summary focusing on particular aspects only. We illustrate the effectiveness of our techniques with a user evaluation experiment on the HKLII platform. The results show that our methods are highly effective. Tien-Hsuan Wu, Ben Kao, Felix Chan, Anne S. Y. Cheung, Michael M. K. Cheung, Guowen Yuan, Yongxi Chen |
JURIX | 6 |
| 2020 | Integrating Domain Knowledge in AI-Assisted Criminal Sentencing of Drug Trafficking CasesabstractJudgment prediction is the task of predicting various outcomes of legal cases of which sentencing prediction is one of the most important yet difficult challenges. We study the applicability of machine learning (ML) techniques in predicting prison terms of drug trafficking cases. In particular, we study how legal domain knowledge can be integrated with ML models to construct highly accurate predictors. We illustrate how our criminal sentence predictors can be applied to address four important issues in legal knowledge management, which include (1) discovery of model drifts in legal rules, (2) identification of critical features in legal judgments, (3) fairness in machine predictions, and (4) explainability of machine predictions. Tien-Hsuan Wu, Ben Kao, Anne S. Y. Cheung, Michael M. K. Cheung, Yongxi Chen, Guowen Yuan, Reynold Cheng |
JURIX | 7 |
| 2020 | Semi-Supervised Feature Selection via Sparse Rescaled Linear Square RegressionabstractWith the rapid increase of the data size, it has increasing demands for selecting features by exploiting both labeled and unlabeled data. In this paper, we propose a novel semi-supervised embedded feature selection method. The new method extends the least square regression model by rescaling the regression coefficients in the least square regression with a set of scale factors, which is used for evaluating the importance of features. An iterative algorithm is proposed to optimize the new model. It has been proved that solving the new model is equivalent to solving a sparse model with a flexible and adaptable ℓ2;pnorm regularization. Moreover, the optimal solution of scale factors provides a theoretical explanation for why we can use {||w1||2, . . .,||wd||2} to evaluate the importance of features. Experimental results on eight benchmark data sets show the superior performance of the proposed method. Xiaojun Chen 0006, Guowen Yuan, Feiping Nie 0001, Zhong Ming 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2019 | Semi-Supervised Feature Selection with Adaptive Discriminant AnalysisabstractIn this paper, we propose a novel Adaptive Discriminant Analysis for semi-supervised feature selection, namely SADA. Instead of computing fixed similarities before performing feature selection, SADA simultaneously learns an adaptive similarity matrix S and a projection matrix W with an iterative method. In each iteration, S is computed from the projected distance with the learned W and W is computed with the learned S. Therefore, SADA can learn better projection matrix W by weakening the effect of noise features with the adaptive similarity matrix. Experimental results on 4 data sets show the superiority of SADA compared to 5 semisupervised feature selection methods. Weichan Zhong, Xiaojun Chen 0006, Guowen Yuan, Yiqin Li, Feiping Nie 0001 |
AAAI | 3 |
| 2018 | A Stratified Feature Ranking Method for Supervised Feature SelectionabstractMost feature selection methods usually select the highest rank features which may be highly correlated with each other. In this paper, we propose a Stratified Feature Ranking (SFR) method for supervised feature selection. In the new method, a Subspace Feature Clustering (SFC) is proposed to identify feature clusters, and a stratified feature ranking method is proposed to rank the features such that the high rank features are lowly correlated. Experimental results show the superiority of SFR. Renjie Chen 0004, Xiaojun Chen 0006, Guowen Yuan, Wenya Sun, Qingyao Wu |
AAAI | 3 |
| 2018 | Discriminative Semi-Supervised Feature Selection via Rescaled Least Squares Regression-SupplementabstractIn this paper, we propose a Discriminative Semi-Supervised Feature Selection (DSSFS) method. In this method, a ε-dragging technique is introduced to the Rescaled Linear Square Regression in order to enlarge the distances between different classes. An iterative method is proposed to simultaneously learn the regression coefficients, ε-draggings matrix and predicting the unknown class labels. Experimental results show the superiority of DSSFS. Guowen Yuan, Xiaojun Chen 0006, Chen Wang 0032, Feiping Nie 0001, Liping Jing |
AAAI | 1 |
| 2018 | Local Adaptive Projection Framework for Feature Selection of Labeled and Unlabeled DataabstractMost feature selection methods first compute a similarity matrix by assigning a fixed value to pairs of objects in the whole data or to pairs of objects in a class or by computing the similarity between two objects from the original data. The similarity matrix is fixed as a constant in the subsequent feature selection process. However, the similarities computed from the original data may be unreliable, because they are affected by noise features. Moreover, the local structure within classes cannot be recovered if the similarities between the pairs of objects in a class are equal. In this paper, we propose a novel local adaptive projection (LAP) framework. Instead of computing fixed similarities before performing feature selection, LAP simultaneously learns an adaptive similarity matrix and a projection matrix with an iterative method. In each iteration, is computed from the projected distance with the learned and W is computed with the learned . Therefore, LAP can learn better projection matrix by weakening the effect of noise features with the adaptive similarity matrix. A supervised feature selection with LAP (SLAP) method and an unsupervised feature selection with LAP (ULAP) method are proposed. Experimental results on eight data sets show the superiority of SLAP compared with seven supervised feature selection methods and the superiority of ULAP compared with five unsupervised feature selection methods. Xiaojun Chen 0006, Guowen Yuan, Feiping Nie 0001, Xiaojun Chang, Joshua Zhexue Huang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Semi-supervised Feature Selection via Rescaled Linear RegressionabstractWith the rapid increase of complex and high-dimensional sparse data, demands for new methods to select features by exploiting both labeled and unlabeled data have increased. Least regression based feature selection methods usually learn a projection matrix and evaluate the importances of features using the projection matrix, which is lack of theoretical explanation. Moreover, these methods cannot find both global and sparse solution of the projection matrix. In this paper, we propose a novel semi-supervised feature selection method which can learn both global and sparse solution of the projection matrix. The new method extends the least square regression model by rescaling the regression coefficients in the least square regression with a set of scale factors, which are used for ranking the features. It has shown that the new model can learn global and sparse solution. Moreover, the introduction of scale factors provides a theoretical explanation for why we can use the projection matrix to rank the features. A simple yet effective algorithm with proved convergence is proposed to optimize the new model. Experimental results on eight real-life data sets show the superiority of the method. Xiaojun Chen 0006, Guowen Yuan, Feiping Nie 0001, Joshua Zhexue Huang |
IJCAI | 2 |