Xianyang Zhang

dblp:286/4121 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Computer networks · 3 · 3 first-author · 2 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Powerful large scale inference in high dimensional mediation analysis
abstract
In genome-wide epigenetic studies, determining how exposures (e.g., Single Nucleotide Polymorphisms) affect outcomes (e.g., gene expression) through intermediate variables, such as DNA methylation, is a key challenge. Mediation analysis provides a framework to identify these causal pathways; however, testing for mediation effects involves a complex composite null hypothesis. Existing methods, such as Sobel's test or the Max-P test, are often underpowered in this context because they rely on null distributions determined under only a subset of the null space and are not optimized for the multiple testing burden inherent in high-dimensional data. To address these limitations, we introduce MLFDR (Mediation Analysis using Local False Discovery Rates), a novel method for high-dimensional mediation analysis. MLFDR leverages local false discovery rates, calculated from the coefficients of structural equation models, to construct an optimal rejection region. We demonstrate theoretically and through simulation that MLFDR asymptotically controls the false discovery rate and achieves superior statistical power compared to recent high-dimensional mediation methods. In real data applications, MLFDR identified 20%-50% more significant mediators than existing methods, demonstrating its ability to uncover biological signals missed by conventional approaches.
Asmita Roy, Xianyang Zhang
PLoS Comput. Biol.2
2026 Correction: BMDD: A probabilistic framework for accurate imputation of zero-inflated microbiome sequencing data
abstract
[This corrects the article DOI: 10.1371/journal.pcbi.1013124.].
Huijuan Zhou, Jun Chen 0040, Xianyang Zhang
PLoS Comput. Biol.3
2025 A Likelihood Based Approach for Watermark Detection
abstract
Watermarking techniques embed statistical signals within content generated by large language models to help trace its source. Although existing methods perform well on long texts, their effectiveness significantly decreases for shorter texts. We introduce a statistical detection approach that improves the power of watermark detection, particularly in shorter texts. Our method leverages both the watermark key sequence and the next token probabilities (NTPs) to determine whether a text is generated by a large language model. We demonstrate the optimality of our approach and analyze its power properties. We also investigate an approach to estimating NTPs and extend our method to scenarios where texts face potential attacks such as substitutions, insertions, or deletions. We validate the effectiveness of our technique using texts generated by Meta-Llama-3-8B from Meta and Mistral-7B-v0.1 from Mistral AI, utilizing prompts extracted from Google’s C4 dataset. In scenarios without attacks and with short text lengths, our method demonstrates approximately 65% power improvement compared to the baseline method on average. We release all code publicly at \url{https://github.com/doccstat/llm-watermark-adaptive.}
Xingchi Li 0002, Guanxun Li, Xianyang Zhang
AISTATS3
2025 Generalization Bounds and Model Complexity for Kolmogorov-Arnold Networks
abstract
Kolmogorov–Arnold Network (KAN) is a network structure recently proposed in Liu et al. (2024) that offers improved interpretability and a more parsimonious design in many science-oriented tasks compared to multi-layer perceptrons. This work provides a rigorous theoretical analysis of KAN by establishing generalization bounds for KAN equipped with activation functions that are either represented by linear combinations of basis functions or lying in a low-rank Reproducing Kernel Hilbert Space (RKHS). In the first case, the generalization bound accommodates various choices of basis functions in forming the activation functions in each layer of KAN and is adapted to different operator norms at each layer. For a particular choice of operator norms, the bound scales with the $l_1$ norm of the coefficient matrices and the Lipschitz constants for the activation functions, and it has no dependence on combinatorial parameters (e.g., number of nodes) outside of logarithmic factors. Moreover, our result does not require the boundedness assumption on the loss function and, hence, is applicable to a general class of regression-type loss functions. In the low-rank case, the generalization bound scales polynomially with the underlying ranks as well as the Lipschitz constants of the activation functions in each layer. These bounds are empirically investigated for KANs trained with stochastic gradient descent on simulated and real data sets. The numerical results demonstrate the practical relevance of these bounds.
Xianyang Zhang, Huijuan Zhou
ICLR1
2025 BMDD: A probabilistic framework for accurate imputation of zero-inflated microbiome sequencing data
abstract
Microbiome sequencing data are inherently sparse and compositional, with excessive zeros arising from biological absence or insufficient sampling. These zeros pose significant challenges for downstream analyses, particularly those that require log-transformation. We introduce BMDD (BiModal Dirichlet Distribution), a novel probabilistic modeling framework for accurate imputation of microbiome sequencing data. Unlike existing imputation approaches that assume unimodal abundance, BMDD captures the bimodal abundance distribution of the taxa via a mixture of Dirichlet priors. It uses variational inference and a scalable expectation-maximization algorithm for efficient imputation. Through simulations and real microbiome datasets, we demonstrate that BMDD outperforms competing methods in reconstructing true abundances and improves the performance of differential abundance analysis. Through multiple posterior samples, BMDD enables robust inference by accounting for uncertainty in zero imputation. Our method offers a principled and computationally efficient solution for analyzing high-dimensional, zero-inflated microbiome sequencing data and is broadly applicable in microbial biomarker discovery and host-microbiome interaction studies.
Huijuan Zhou, Jun Chen 0040, Xianyang Zhang
PLoS Comput. Biol.3
2025 Bayesian Cramér-Rao Bound Estimation With Score-Based Models
abstract
The Bayesian Cramér-Rao bound (CRB) provides a lower bound on the mean square error of any Bayesian estimator under mild regularity conditions. It can be used to benchmark the performance of statistical estimators, and provides a principled metric for system design and optimization. However, the Bayesian CRB depends on the underlying prior distribution, which is often unknown for many problems of interest. This work introduces a new data-driven estimator for the Bayesian CRB using score matching, i.e., a statistical estimation technique that models the gradient of a probability distribution from a given set of training data. The performance of the proposed estimator is analyzed in both the classical parametric modeling regime and the neural network modeling regime. In both settings, we develop novel non-asymptotic bounds on the score matching error and our Bayesian CRB estimator based on the results from empirical process theory, including classical bounds and recently introduced techniques for characterizing neural networks. We illustrate the performance of the proposed estimator with two application examples: a signal denoising problem and a dynamic phase offset estimation problem with applications in communication systems.
Evan Scope Crafts, Xianyang Zhang, Bo Zhao 0002
IEEE Trans. Inf. Theory2
2024 Soft-constrained Schrödinger Bridge: a Stochastic Control Approach
abstract
Schrödinger bridge can be viewed as a continuous-time stochastic control problem where the goal is to find an optimally controlled diffusion process whose terminal distribution coincides with a pre-specified target distribution. We propose to generalize this problem by allowing the terminal distribution to differ from the target but penalizing the Kullback-Leibler divergence between the two distributions. We call this new control problem soft-constrained Schrödinger bridge (SSB). The main contribution of this work is a theoretical derivation of the solution to SSB, which shows that the terminal distribution of the optimally controlled process is a geometric mixture of the target and some other distribution. This result is further extended to a time series setting. One application is the development of robust generative diffusion models. We propose a score matching-based algorithm for sampling from geometric mixtures and showcase its use via a numerical example for the MNIST data set.
Jhanvi Garg, Xianyang Zhang
AISTATS2
2024 Score Matching with Deep Neural Networks: A Non-Asymptotic Analysis
abstract
Score matching is a statistical approach for estimating the score (the gradient of the log-density) of a probability distribution from samples. It has found a number of applications, including in generative modeling, where it serves as a key component of the state-of-the-art diffusion modeling framework. The goal of this work is to provide non-asymptotic bounds on the score matching risk in the setting where the score model is a deep neural network. Here key challenges include the fact that the score model is vector-valued and that the score matching loss depends on the Jacobian of the score model. Our approach integrates results from empirical process theory, including classical bounds and recently introduced techniques for bounding covering numbers of neural network models, with novel covering results to address these challenges. The resulting bound has logarithmic dependence on the network width, allowing the network size to grow exponentially with the number of training samples without compromising the bound.
Evan Scope Crafts, Xianyang Zhang
ITW2
2024 Segmenting Watermarked Texts From Language Models
abstract
Watermarking is a technique that involves embedding nearly unnoticeable statistical signals within generated content to help trace its source. This work focuses on a scenario where an untrusted third-party user sends prompts to a trusted language model (LLM) provider, who then generates a text from their LLM with a watermark. This setup makes it possible for a detector to later identify the source of the text if the user publishes it. The user can modify the generated text by substitutions, insertions, or deletions. Our objective is to develop a statistical method to detect if a published text is LLM-generated from the perspective of a detector. We further propose a methodology to segment the published text into watermarked and non-watermarked sub-strings. The proposed approach is built upon randomization tests and change point detection techniques. We demonstrate that our method ensures Type I and Type II error control and can accurately identify watermarked sub-strings by finding the corresponding change point locations. To validate our technique, we apply it to texts generated by several language models with prompts extracted from Google's C4 dataset and obtain encouraging numerical results. We release all code publicly at https://github.com/doccstat/llm-watermark-cpd.
Xingchi Li 0002, Guanxun Li, Xianyang Zhang
NeurIPS3
2024 A Robot Navigation System Based on Improved A-Star Algorithm
abstract
Regarding global path planning for search and rescue robots, the traditional method of A* algorithm is slow and has many turning points along the intended route, which hinders its searching speed. A new and enhanced A* algorithm is presented in this paper, incorporating the Floyd-trajectory optimization algorithm. The search directions of the traditional A* algorithm were refined from eight to five. Next, we improved the cost estimation function by introducing a weight value to balance the estimated heuristic function value with the actual cost value, thereby enhancing the algorithm’s search efficiency. Finally, we applied the Floyd path optimization algorithm to eliminate redundant nodes from the route. The use of both the Floyd and A* algorithms in combination has been proven through simulations and experiments to improve search results by 40% compared to solely using the A* algorithm. This integration effectively reduces the search scope, diminishes the number of turning points by 63.8%, improves path smoothness, and effectively shortens the planned path length.
Xiangli Bu, Guangxing Li, Bo Tong, Xianyang Zhang
Int. J. Pattern Recognit. Artif. Intell.4
2023 Sequential Gradient Descent and Quasi-Newton's Method for Change-Point Analysis
abstract
One common approach to detecting change-points is minimizing a cost function over possible numbers and locations of change-points. The framework includes several well-established procedures, such as the penalized likelihood and minimum description length. Such an approach requires finding the cost value repeatedly over different segments of the data set, which can be time-consuming when (i) the data sequence is long and (ii) obtaining the cost value involves solving a non-trivial optimization problem. This paper introduces a new sequential updating method (SE) to find the cost value effectively. The core idea is to update the cost value using the information from previous steps without re-optimizing the objective function. The new method is applied to change-point detection in generalized linear models and penalized regression. Numerical studies show that the new approach can be orders of magnitude faster than the Pruned Exact Linear Time (PELT) method without sacrificing estimation accuracy.
Xianyang Zhang, Trisha Dawn
AISTATS1
2023 A general framework for powerful confounder adjustment in omics association studies
abstract
MOTIVATION: Genomic data are subject to various sources of confounding, such as demographic variables, biological heterogeneity, and batch effects. To identify genomic features associated with a variable of interest in the presence of confounders, the traditional approach involves fitting a confounder-adjusted regression model to each genomic feature, followed by multiplicity correction. RESULTS: This study shows that the traditional approach is suboptimal and proposes a new two-dimensional false discovery rate control framework (2DFDR+) that provides significant power improvement over the conventional method and applies to a wide range of settings. 2DFDR+ uses marginal independence test statistics as auxiliary information to filter out less promising features, and FDR control is performed based on conditional independence test statistics in the remaining features. 2DFDR+ provides (asymptotically) valid inference from samples in settings where the conditional distribution of the genomic variables given the covariate of interest and the confounders is arbitrary and completely unknown. Promising finite sample performance is demonstrated via extensive simulations and real data applications. AVAILABILITY AND IMPLEMENTATION: R codes and vignettes are available at https://github.com/asmita112358/tdfdr.np.
Asmita Roy, Jun Chen 0040, Xianyang Zhang
Bioinform.3
2022 dICC: distance-based intraclass correlation coefficient for metagenomic reproducibility studies
abstract
SUMMARY: Due to the sparsity and high dimensionality, microbiome data are routinely summarized into pairwise distances capturing the compositional differences. Many biological insights can be gained by analyzing the distance matrix in relation to some covariates. A microbiome sampling method that characterizes the inter-sample relationship more reproducibly is expected to yield higher statistical power. Traditionally, the intraclass correlation coefficient (ICC) has been used to quantify the degree of reproducibility for a univariate measurement using technical replicates. In this work, we extend the traditional ICC to distance measures and propose a distance-based ICC (dICC). We derive the asymptotic distribution of the sample-based dICC to facilitate statistical inference. We illustrate dICC using a real dataset from a metagenomic reproducibility study. AVAILABILITY AND IMPLEMENTATION: dICC is implemented in the R CRAN package GUniFrac. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jun Chen 0040, Xianyang Zhang
Bioinform.2
2021 Joint Pushing, Pricing, and Recommendation for Cache-enabled Radio Access Networks
abstract
Proactive pushing can exploit the spectrum underutilized during the off-peak time to push popular content files, thereby significantly improving the spectrum efficiency. Moreover, in a communication system that the virtual network operator (VNO) has to buy spectrum from the base station to conduct file transmission, proactive pushing has been recognized as a promising technology to improve the income of the VNO. However, the appropriate pushing schemes and the achievable income of the VNO are unclear yet. In this paper, joint pushing, pricing, and recommendation (JPPR) schemes are presented for cache-enabled radio access networks. We aim to investigate recommendation-based pushing policy to maximize the average income of the VNO. We establish a Markov chain model, which derives the average income of the VNO. Based on this, we formulate an optimization problem to achieve the maximum average income. We further convert the optimization problem into an equivalent linear programming problem. Moreover, a greedy algorithm is applied to solve the problem with lower computational complexity. Finally, simulation results show the significant income gains that can be achieved by JPPR schemes compared with the system without the JPPR schemes.
Xianyang Zhang, Haiming Hui, Wei Chen 0002, Zhu Han 0001
GLOBECOM1
2021 Joint Recommendation and Pricing for Cache-Aided RAN with Malicious Users: A Game Theoretic Method
abstract
As mobile data traffic has explosively grown during the past decades, pushing popular contents to small cells has been proposed to deal with the growing data demands. To improve the cache hit ratio, the recommender system is employed to recommend cached contents when the requests are not hit by the cache. However, how to persuade users to accept recommended files remains an open problem. We conceive a method that the network operators can give a discount on the traffic cost of the recommended contents. However, some users may maliciously request unpopular contents to get the discount, which will reduce the profit of the virtual network operator (VNO). In order to punish the malicious behaviour, we the VNOcan reduce the recommendation probability to these users. Meanwhile, these malicious users will reduce the malicious probability to increase the revenue. To study the interactions between the profit of the VNO and the revenue of the users, we formulate a non-cooperative game to find the Nash equilibrium (NE) of the recommendation probability of the VNO and the malicious probability of the users. Simulation results indicate that the VNO's profit and the user's revenue can be significantly increased with the proposed system compared with the system without joint recommendation and pricing schemes.
Xianyang Zhang, Haiming Hui, Wei Chen 0002, Zhu Han 0001
GLOBECOM1
2021 D-MANOVA: fast distance-based multivariate analysis of variance for large-scale microbiome association studies
abstract
SUMMARY: PERMANOVA (permutational multivariate analysis of variance based on distances) has been widely used for testing the association between the microbiome and a covariate of interest. Statistical significance is established by permutation, which is computationally intensive for large sample sizes. As large-scale microbiome studies, such as American Gut Project (AGP), become increasingly popular, a computationally efficient version of PERMANOVA is much needed. To achieve this end, we derive the asymptotic distribution of the PERMANOVA pseudo-F statistic and provide analytical P-value calculation based on chi-square approximation. We show that the asymptotic P-value is close to the PERMANOVA P-value even under a moderate sample size. Moreover, it is more accurate and an order-of-magnitude faster than the permutation-free method MDMR. We demonstrated the use of our procedure D-MANOVA on the AGP dataset. AVAILABILITY AND IMPLEMENTATION: D-MANOVA is implemented by the dmanova function in the CRAN package GUniFrac. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jun Chen 0040, Xianyang Zhang
Bioinform.2
2020 Negative Correlation Between Virus-Related Content Popularity and Epidemic Spread
abstract
The coronavirus disease 2019 (COVID-19) has recently attracted extensive attention due to its serious impact on public health worldwide. In this paper, we study and verify that the popularity of virus-related content has a negative correlation with the epidemic spread by means of statistical analysis. Inspired by this result, a practical solution of recommender system is proposed for pushing virus-related content, aiming to gain insight about the newly discovered virus for people and thus reduce the epidemic spread to the utmost extent. First, we formulate the optimization of recommendation policy subject to quality of experience (QoE) loss constraints as a finite-horizon Constrained Markov Decision Problem (CMDP). To solve this problem, then, we present both enumeration and heuristic methods, from perspectives of achieving optimal recommendation policy and reducing computational complexity, respectively. Finally, our simulations validate the benefit of our solution by showing that to recommend virus-related content following our strategy does help slow down the spread of the epidemic.
Xianyang Zhang, Di Han 0001, Zhanyuan Xie, Xin Guo 0008, Haiming Wang 0002, Zhu Han 0001, Wei Chen 0002
GLOBECOM1