Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ganggang Xu

dblp:166/9950 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Learning theory · 46% Probabilistic and Bayesian machine learning · 38% Representation and self-supervised learning · 9%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › hypothesis testing
nonparametric hypothesis testing
0.812024
Nonparametric Inference under B-bits Quantization · J. Mach. Learn. Res. 2024
Machine learning › Learning theory
statistical learning theory
0.812024
Nonparametric Inference under B-bits Quantization · J. Mach. Learn. Res. 2024
Coding theory › source coding
quantization
0.812024
Nonparametric Inference under B-bits Quantization · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
0.512021
Row-clustering of a Point Process-valued Matrix · NeurIPS 2021
Data mining
clustering
0.512021
Row-clustering of a Point Process-valued Matrix · NeurIPS 2021
Data mining › clustering › model-based clustering
mixture model clustering
0.512021
Row-clustering of a Point Process-valued Matrix · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
functional principal component analysis
0.412020
Semi-parametric Learning of Structured Temporal Point Processes · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.412020
Semi-parametric Learning of Structured Temporal Point Processes · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
log-gaussian cox process
0.412020
Semi-parametric Learning of Structured Temporal Point Processes · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
non-parametric methods
0.412020
Semi-parametric Learning of Structured Temporal Point Processes · J. Mach. Learn. Res. 2020
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel ridge regression
0.312018
Optimal Tuning for Divide-and-conquer Kernel Ridge Regression with Massive Data · ICML 2018
Machine learning › Learning theory
nonparametric regression
0.312018
Optimal Tuning for Divide-and-conquer Kernel Ridge Regression with Massive Data · ICML 2018
Machine learning › Learning theory › model selection
tuning parameter selection
0.312018
Optimal Tuning for Divide-and-conquer Kernel Ridge Regression with Massive Data · ICML 2018
Data mining › statistical analysis
functional data analysis
0.112021
Row-clustering of a Point Process-valued Matrix · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

asymptotic analysis · 2.0nonparametric testing · 1.5functional principal component analysis · 1.0expectation-solution algorithm · 1.0spline models · 0.8spline model · 0.8conditional likelihood maximization · 0.4generalized cross-validation · 0.3distributed generalized cross-validation · 0.3
YearPublicationVenuePosition
2024 Nonparametric Inference under B-bits Quantization
abstract
Statistical inference based on lossy or incomplete samples is often needed in research areas such as signal/image processing, medical image storage, remote sensing, signal transmission. In this paper, we propose a nonparametric testing procedure based on samples quantized to $B$ bits through a computationally efficient algorithm. Under mild technical conditions, we establish the asymptotic properties of the proposed test statistic and investigate how the testing power changes as $B$ increases. In particular, we show that if $B$ exceeds a certain threshold, the proposed nonparametric testing procedure achieves the classical minimax rate of testing (Shang and Cheng, 2015) for spline models. We further extend our theoretical investigations to a nonparametric linearity test and an adaptive nonparametric test, expanding the applicability of the proposed methods. Extensive simulation studies {together with a real-data analysis} are used to demonstrate the validity and effectiveness of the proposed tests.
Kexuan Li, Ruiqi Liu 0003, Ganggang Xu, Zuofeng Shang
J. Mach. Learn. Res.3
2021 Row-clustering of a Point Process-valued Matrix
abstract
Structured point process data harvested from various platforms poses new challenges to the machine learning community. To cluster repeatedly observed marked point processes, we propose a novel mixture model of multi-level marked point processes for identifying potential heterogeneity in the observed data. Specifically, we study a matrix whose entries are marked log-Gaussian Cox processes and cluster rows of such a matrix. An efficient semi-parametric Expectation-Solution (ES) algorithm combined with functional principal component analysis (FPCA) of point processes is proposed for model estimation. The effectiveness of the proposed framework is demonstrated through simulation studies and real data analyses.
Lihao Yin, Ganggang Xu, Huiyan Sang
NeurIPS2
2020 Semi-parametric Learning of Structured Temporal Point Processes
abstract
We propose a general framework of using a multi-level log-Gaussian Cox process to model repeatedly observed point processes with complex structures; such type of data has become increasingly available in various areas including medical research, social sciences, economics, and finance due to technological advances. A novel nonparametric approach is developed to efficiently and consistently estimate the covariance functions of the latent Gaussian processes at all levels. To predict the functional principal component scores, we propose a consistent estimation procedure by maximizing the conditional likelihood of super-positions of point processes. We further extend our procedure to the bivariate point process case in which potential correlations between the processes can be assessed. Asymptotic properties of the proposed estimators are investigated, and the effectiveness of our procedures is illustrated through a simulation study and an application to a stock trading dataset.
Ganggang Xu, Jiangze Bian, Timothy R. Burch, Sandro C. Andrade, Jingfei Zhang
J. Mach. Learn. Res.1
2018 Optimal Tuning for Divide-and-conquer Kernel Ridge Regression with Massive Data
abstract
Divide-and-conquer is a powerful approach for large and massive data analysis. In the nonparameteric regression setting, although various theoretical frameworks have been established to achieve optimality in estimation or hypothesis testing, how to choose the tuning parameter in a practically effective way is still an open problem. In this paper, we propose a data-driven procedure based on divide-and-conquer for selecting the tuning parameters in kernel ridge regression by modifying the popular Generalized Cross-validation (GCV, Wahba, 1990). While the proposed criterion is computationally scalable for massive data sets, it is also shown under mild conditions to be asymptotically optimal in the sense that minimizing the proposed distributed-GCV (dGCV) criterion is equivalent to minimizing the true global conditional empirical loss of the averaged function estimator, extending the existing optimality results of GCV to the divide-and-conquer framework.
Ganggang Xu, Zuofeng Shang, Guang Cheng 0003
ICML1