VLDB 2026 Research / reviewers in the wild / expert
Xunye Tian
dblp:360/6608
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Learning theory · 57% Kernel, tree and ensemble methods · 29% Probabilistic and Bayesian machine learning · 14% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
hypothesis testing |
1.7 | 2 | 2025 | DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence Testing · NeurIPS 2025 Anchor-based Maximum Discrepancy for Relative Similarity Testing · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › kernel design
deep kernel learning |
0.9 | 1 | 2025 | Anchor-based Maximum Discrepancy for Relative Similarity Testing · NeurIPS 2025 |
Machine learning › Learning theory › hypothesis testing
independence testing |
0.9 | 1 | 2025 | DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence Testing · NeurIPS 2025 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.9 | 1 | 2025 | DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence Testing · NeurIPS 2025 |
Machine learning › Learning theory › hypothesis testing › two-sample testing
kernel two-sample test |
0.9 | 1 | 2025 | DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence Testing · NeurIPS 2025 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel learning
multiple kernel learning |
0.9 | 1 | 2025 | DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence Testing · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
submodular selection · 0.9statistical testing · 0.9maximum discrepancy · 0.9kernel diversity · 0.9deep kernel · 0.9asymptotic analysis · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Anchor-based Maximum Discrepancy for Relative Similarity TestingabstractThe relative similarity testing aims to determine which of the distributions, $P$ or $Q$, is closer to an anchor distribution $U$. Existing kernel-based approaches often test the relative similarity with a fixed kernel in a manually specified alternative hypothesis, e.g., $Q$ is closer to $U$ than $P$. Although kernel selection is known to be important to kernel-based testing methods, the manually specified hypothesis poses a significant challenge for kernel selection in relative similarity testing: Once the hypothesis is specified first, we can always find a kernel such that the hypothesis is rejected. This challenge makes relative similarity testing ill-defined when we want to select a good kernel after the hypothesis is specified. In this paper, we cope with this challenge via learning a proper hypothesis and a kernel simultaneously, instead of learning a kernel after manually specifying the hypothesis. We propose an anchor-based maximum discrepancy (AMD), which defines the relative similarity as the maximum discrepancy between the distances of $(U, P)$ and $(U, Q)$ in a space of deep kernels. Based on AMD, our testing incorporates two phases. In Phase I, we estimate the AMD over the deep kernel space and infer the potential hypothesis. In Phase II, we assess the statistical significance of the potential hypothesis, where we propose a unified testing framework to derive thresholds for tests over different possible hypotheses from Phase I. Lastly, we validate our method theoretically and demonstrate its effectiveness via extensive experiments on benchmark datasets. Codes are publicly available at: https://github.com/tmlr-group/AMD. Zhijian Zhou, Liuhua Peng, Xunye Tian, Feng Liu 0003 |
NeurIPS | 3 |
| 2025 | DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence TestingabstractTo adapt kernel two-sample and independence testing to complex structured data, aggregation of multiple kernels is frequently employed to boost testing power compared to single-kernel tests. However, we observe a phenomenon that directly maximizing multiple kernel-based statistics may result in highly similar kernels that capture highly overlapping information, limiting the effectiveness of aggregation.
To address this, we propose an aggregated statistic that explicitly incorporates kernel diversity based on the covariance between different kernels. Moreover, we identify a fundamental challenge: a trade-off between the diversity among kernels and the test power of individual kernels, i.e., the selected kernels should be both effective and diverse. This motivates a testing framework with selection inference, which leverages information from the training phase to select kernels with strong individual performance from the learned diverse kernel pool. We provide rigorous theoretical statements and proofs to show the consistency on the test power and control of Type-I error, along with asymptotic analysis of the proposed statistics. Lastly, we conducted extensive empirical experiments demonstrating the superior performance of our proposed approach across various benchmarks for both two-sample and independence testing. Zhijian Zhou, Xunye Tian, Liuhua Peng, Antonin Schrab, Danica J. Sutherland, Feng Liu 0003 |
NeurIPS | 2 |
| 2025 | A Unified Data Representation Learning for Non-parametric Two-sample TestingabstractLearning effective data representations has been crucial in non-parametric two-sample testing. Common approaches will first split data into training and test sets and then learn data representations purely on the training set. However, recent theoretical studies have shown that, as long as the sample indexes are not used during the learning process, the whole data can be used to learn data representations, meanwhile ensuring control of Type-I errors. The above fact motivates us to use the test set (but without sample indexes) to facilitate the data representation learning in the testing. To this end, we propose a representation-learning two-sample testing (RL-TST) framework. RL-TST first performs purely self-supervised representation learning on the entire dataset to capture inherent representations (IRs) that reflect the underlying data manifold. A discriminative model is then trained on these IRs to learn discriminative representations (DRs), enabling the framework to leverage both the rich structural information from IRs and the discriminative power of DRs. Extensive experiments demonstrate that RL-TST outperforms representative approaches by simultaneously using data manifold information in the test set and enhancing test power via finding the DRs with the training set. Xunye Tian, Liuhua Peng, Zhijian Zhou, Mingming Gong, Arthur Gretton, Feng Liu 0003 |
UAI | 1 |