VLDB 2026 Research / reviewers in the wild / expert
Xin Guo 0003
dblp:17/1430-3
· DBLP profile ↗
6ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0002-7465-9356ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Kernel, tree and ensemble methods · 51% Learning theory · 23% Deep learning architectures and training · 12% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods
kernel methods |
1.1 | 2 | 2025 | Kernel-based L_2-Boosting with Structure Constraints · J. Mach. Learn. Res. 2025 Sparsity and Error Analysis of Empirical Feature-Based Regularization Schemes · J. Mach. Learn. Res. 2016 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel machines
kernel boosting |
0.9 | 1 | 2025 | Kernel-based L_2-Boosting with Structure Constraints · J. Mach. Learn. Res. 2025 |
Mathematical optimization
statistical learning theory |
0.9 | 1 | 2025 | Kernel-based L_2-Boosting with Structure Constraints · J. Mach. Learn. Res. 2025 |
Machine learning › Learning theory
information-theoretic learning |
0.4 | 1 | 2020 | Distributed Minimum Error Entropy Algorithms · J. Mach. Learn. Res. 2020 |
Machine learning › Learning theory › loss function
minimum error entropy |
0.4 | 1 | 2020 | Distributed Minimum Error Entropy Algorithms · J. Mach. Learn. Res. 2020 |
Machine learning › Efficient and distributed learning
divide-and-conquer learning |
0.3 | 1 | 2017 | Distributed Learning with Regularized Least Squares · J. Mach. Learn. Res. 2017 |
Machine learning › Deep learning architectures and training
regularization |
0.2 | 1 | 2016 | Sparsity and Error Analysis of Empirical Feature-Based Regularization Schemes · J. Mach. Learn. Res. 2016 |
Machine learning › Deep learning architectures and training › regularization
sparse regularization |
0.2 | 1 | 2016 | Sparsity and Error Analysis of Empirical Feature-Based Regularization Schemes · J. Mach. Learn. Res. 2016 |
Machine learning › Learning paradigms
semi-supervised learning |
0.1 | 1 | 2020 | Distributed Minimum Error Entropy Algorithms · J. Mach. Learn. Res. 2020 |
Machine learning › Learning theory
nonparametric regression |
0.1 | 1 | 2017 | Distributed Learning with Regularized Least Squares · J. Mach. Learn. Res. 2017 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space |
0.1 | 1 | 2017 | Distributed Learning with Regularized Least Squares · J. Mach. Learn. Res. 2017 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
least squares regression |
0.1 | 1 | 2016 | Sparsity and Error Analysis of Empirical Feature-Based Regularization Schemes · J. Mach. Learn. Res. 2016 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression |
0.1 | 1 | 2016 | Sparsity and Error Analysis of Empirical Feature-Based Regularization Schemes · J. Mach. Learn. Res. 2016 |
Methods — techniques the papers use, named apart from their topics
truncation · 1.7talagrand concentration inequality · 1.7boosting · 1.7u-statistics · 0.4error decomposition · 0.4divide-and-conquer · 0.4integral operator · 0.3error bound analysis · 0.3reproducing kernel hilbert space · 0.2concave regularization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Kernel-based L_2-Boosting with Structure ConstraintsabstractDeveloping efficient kernel methods for regression is popular in the past two decades. In this paper, utilizing boosting on kernel-based weak learners, we propose a novel kernel-based learning algorithm called kernel-based re-scaled boosting with truncation, dubbed as KReBooT. The proposed KReBooT benefits in controlling the structure and producing sparse estimators, and is near overfitting resistant. We conduct both theoretical analysis and numerical simulations to illustrate the excellent performance of KReBooT. Theoretically, we prove that KReBooT can achieve the optimal numerical convergence rate for nonlinear approximation. Furthermore, using a variant of Talagrand's concentration inequality, we provide fast learning rates for KReBooT, which is a new record of boosting-type algorithms. Numerically, we carry out several simulations to show the promising performance of KReBooT in terms of its good generalization, near over-fitting resistance and structure constraints. Yao Wang 0003, Xin Guo 0003, Shaobo Lin |
J. Mach. Learn. Res. | 2 |
| 2022 | Online gradient descent algorithms for functional data learning
Xiaming Chen, Bohao Tang, Xin Guo 0003 |
J. Complex. | 4 |
| 2020 | Distributed Minimum Error Entropy AlgorithmsabstractMinimum Error Entropy (MEE) principle is an important approach in Information Theoretical Learning (ITL). It is widely applied and studied in various fields for its robustness to noise. In this paper, we study a reproducing kernel-based distributed MEE algorithm, DMEE, which is designed to work with both fully supervised data and semi-supervised data. The divide-and-conquer approach is employed, so there is no inter-node communication overhead. Similar as other distributed algorithms, DMEE significantly reduces the computational complexity and memory requirement on single computing nodes. With fully supervised data, our proved learning rates equal the minimax optimal learning rates of the classical pointwise kernel-based regressions. Under the semi-supervised learning scenarios, we show that DMEE exploits unlabeled data effectively, in the sense that first, under the settings with weak regularity assumptions, additional unlabeled data significantly improves the learning rates of DMEE. Second, with sufficient unlabeled data, labeled data can be distributed to many more computing nodes, that each node takes only O(1) labels, without spoiling the learning rates in terms of the number of labels. This conclusion overcomes the saturation phenomenon in unlabeled data size. It parallels a recent results for regularized least squares (Lin and Zhou, 2018), and suggests that an inflation of unlabeled data is a solution to the MEE learning problems with decentralized data source for the concerns of privacy protection. Our work refers to pairwise learning and non-convex loss. The theoretical analysis is achieved by distributed U-statistics and error decomposition techniques in integral operators. Xin Guo 0003, Ting Hu 0002, Qiang Wu 0003 |
J. Mach. Learn. Res. | 1 |
| 2019 | Search for K: Assessing Five Topic-Modeling Approaches to 120, 000 Canadian ArticlesabstractTopic modeling has been an important field in natural language processing (NLP) and recently witnessed great methodological advances. Yet, the development of topic modeling is still, if not increasingly, challenged by two critical issues. First, despite intense efforts toward nonparametric/post-training methods, the search for the optimal number of topics K remains a fundamental question in topic modeling and warrants input from domain experts. Second, with the development of more sophisticated models, topic modeling is now ironically been treated as a black box and it becomes increasingly difficult to tell how research findings are informed by data, model specifications, or inference algorithms. Based on about 120,000 newspaper articles retrieved from three major Canadian newspapers (Globe and Mail, Toronto Star, and National Post) since 1977, we employ five methods with different model specifications and inference algorithms (Latent Semantic Analysis, Latent Dirichlet Allocation, Principal Component Analysis, Factor Analysis, Nonnegative Matrix Factorization) to identify discussion topics. The optimal topics are then assessed using three measures: coherence statistics, held-out likelihood (loss), and graph-based dimensionality selection. Mixed findings from this research complement advances in topic modeling and provide insights into the choice of optimal topics in social science research. Qiang Fu 0022, Yufan Zhuang, Jiaxin Gu, Yushu Zhu, Huihui Qin, Xin Guo 0003 |
IEEE BigData | 6 |
| 2017 | Distributed Learning with Regularized Least SquaresabstractWe study distributed learning with the least squares regularization scheme in a reproducing kernel Hilbert space (RKHS). By a divide-and-conquer approach, the algorithm partitions a data set into disjoint data subsets, applies the least squares regularization scheme to each data subset to produce an output function, and then takes an average of the individual output functions as a final global estimator or predictor. We show with error bounds and learning rates in expectation in both the $L^2$-metric and RKHS-metric that the global output function of this distributed learning is a good approximation to the algorithm processing the whole data in one single machine. Our derived learning rates in expectation are optimal and stated in a general setting without any eigenfunction assumption. The analysis is achieved by a novel second order decomposition of operator differences in our integral operator approach. Even for the classical least squares regularization scheme in the RKHS associated with a general kernel, we give the best learning rate in expectation in the literature. Shaobo Lin, Xin Guo 0003, Ding-Xuan Zhou |
J. Mach. Learn. Res. | 2 |
| 2016 | Sparsity and Error Analysis of Empirical Feature-Based Regularization SchemesabstractWe consider a learning algorithm generated by a regularization scheme with a concave regularizer for the purpose of achieving sparsity and good learning rates in a least squares regression setting. The regularization is induced for linear combinations of empirical features, constructed in the literatures of kernel principal component analysis and kernel projection machines, based on kernels and samples. In addition to the separability of the involved optimization problem caused by the empirical features, we carry out sparsity and error analysis, giving bounds in the norm of the reproducing kernel Hilbert space, based on a priori conditions which do not require assumptions on sparsity in terms of any basis or system. In particular, we show that as the concave exponent $q$ of the concave regularizer increases to $1$, the learning ability of the algorithm improves. Some numerical simulations for both artificial and real MHC-peptide binding data involving the $\ell^q$ regularizer and the SCAD penalty are presented to demonstrate the sparsity and error analysis. Xin Guo 0003, Ding-Xuan Zhou |
J. Mach. Learn. Res. | 1 |