EDBT 2026 Demo / reviewers in the wild / expert
Yuling Jiao
dblp:136/7658
· DBLP profile ↗
36ranked-venue papers
6as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 2 since 2021Theory of computation · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation PerspectiveabstractThe Transformer model is widely used in various application areas of machine learning, such as natural language processing. This paper investigates the approximation of the Hölder continuous function class $\mathcal{H}_{Q}^{\beta}\left([0,1]^{d\times n},\mathbb{R}^{d\times n}\right)$ by Transformers and constructs several Transformers that can overcome the curse of dimensionality. These Transformers consist of one self-attention layer with one head and the softmax function as the activation function, along with several feedforward layers. For example, to achieve an approximation accuracy of $\epsilon$, if the activation functions of the feedforward layers in the Transformer are ReLU and floor, only $\mathcal{O}\left(\log\frac{1}{\epsilon}\right)$ layers of feedforward layers are needed, with widths of these layers not exceeding $\mathcal{O}\left(\frac{1}{\epsilon^{2/\beta}}\log\frac{1}{\epsilon}\right)$. If other activation functions are allowed in the feedforward layers, the width of the feedforward layers can be further reduced to a constant. These results demonstrate that Transformers have a strong expressive capability. The construction in this paper is based on the Kolmogorov-Arnold Superposition Theorem and does not require the concept of contextual mapping, hence our proof is more intuitively clear compared to previous Transformer approximation works. Additionally, the translation technique proposed in this paper helps to apply the previous approximation results of feedforward neural networks to Transformer research. Yuling Jiao, Yanming Lai, Yang Wang 0020, Bokai Yan |
J. Mach. Learn. Res. | 1 |
| 2025 | Adv-SSL: Adversarial Self-Supervised Representation Learning with Theoretical GuaranteesabstractLearning transferable data representations from abundant unlabeled data remains a central challenge in machine learning. Although numerous self-supervised learning methods have been proposed to address this challenge, a significant class of these approaches aligns the covariance or correlation matrix with the identity matrix. Despite impressive performance across various downstream tasks, these methods often suffer from biased sample risk, leading to substantial optimization shifts in mini-batch settings and complicating theoretical analysis. In this paper, we introduce a novel \underline{\bf Adv}ersarial \underline{\bf S}elf-\underline{\bf S}upervised Representation \underline{\bf L}earning (Adv-SSL) for unbiased transfer learning with no additional cost compared to its biased counterparts. Our approach not only outperforms the existing methods across multiple benchmark datasets but is also supported by comprehensive end-to-end theoretical guarantees. Our analysis reveals that the minimax optimization in Adv-SSL encourages representations to form well-separated clusters in the embedding space, provided there is sufficient upstream unlabeled data. As a result, our method achieves strong classification performance even with limited downstream labels, shedding new light on few-shot learning. Chenguang Duan, Yuling Jiao, Huazhen Lin, Wensen Ma, Jerry Zhijian Yang |
NeurIPS | 2 |
| 2025 | DRM Revisited: A Complete Error AnalysisabstractIt is widely known that the error analysis for deep learning involves approximation, statistical, and optimization errors. However, it is challenging to combine them together due to overparameterization. In this paper, we address this gap by providing a comprehensive error analysis of the Deep Ritz Method (DRM). Specifically, we investigate a foundational question in the theoretical analysis of DRM under the overparameterized regime: given a target precision level, how can one determine the appropriate number of training samples, the key architectural parameters of the neural networks, the step size for the projected gradient descent optimization procedure, and the requisite number of iterations, such that the output of the gradient descent process closely approximates the true solution of the underlying partial differential equation to the specified precision? Yuling Jiao, Ruoxuan Li, Peiying Wu, Jerry Zhijian Yang, Pingwen Zhang |
J. Mach. Learn. Res. | 1 |
| 2025 | Convergence analysis of deep Ritz method with over-parameterization
Zhao Ding, Yuling Jiao, Xiliang Lu, Peiying Wu, Jerry Zhijian Yang |
Neural Networks | 2 |
| 2025 | Deep contrastive representation learning for supervised tasks
Chenguang Duan, Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Fusheng Zhou |
Pattern Recognit. | 2 |
| 2025 | Semi-Supervised Deep Sobolev Regression: Estimation and Variable Selection by ReQU Neural NetworkabstractWe propose SDORE, asemi-superviseddeep Sobolevregressor, for the nonparametric estimation of the underlying regression function and its gradient. SDORE employs deep ReQU neural networks to minimize the empirical risk with gradient norm regularization, allowing the approximation of the regularization term by unlabeled data. Our study includes a thorough analysis of the convergence rates of SDORE in$L^{2}$-norm, achieving the minimax optimality. Further, we establish a convergence rate for the associated plug-in gradient estimator, even in the presence of significant domain shift. These theoretical findings offer valuable insights for selecting regularization parameters and determining the size of the neural network, while showcasing the provable advantage of leveraging unlabeled data in semi-supervised learning. To the best of our knowledge, SDORE is the first provable neural network-based approach that simultaneously estimates the regression function and its gradient, with diverse applications such as nonparametric variable selection. The effectiveness of SDORE is validated through an extensive range of numerical simulations. Zhao Ding, Chenguang Duan, Yuling Jiao, Jerry Zhijian Yang |
IEEE Trans. Inf. Theory | 3 |
| 2025 | Schrödinger-Föllmer SamplerabstractSampling from probability distributions is a critical problem in statistics and machine learning, particularly in Bayesian inference, where direct integration over the posterior distribution is often infeasible, making sampling from the posterior essential for inference. This paper introduces the Schrödinger-Föllmer sampler (SFS), a novel approach for sampling from potentially unnormalized distributions. The SFS leverages the Schrödinger-Föllmer diffusion process on the unit interval, incorporating a time-dependent drift term that evolves the distribution from a degenerate form at time zero to the target distribution at time one. Unlike existing Markov chain Monte Carlo methods that rely on ergodicity, SFS operates independently of ergodicity. Computationally, SFS is straightforward to implement using the Euler-Maruyama discretization. In our theoretical analysis, we derive non-asymptotic error bounds for the SFS sampling distribution in the Wasserstein distance, subject to reasonable conditions. Numerical experiments demonstrate that SFS generates higher-quality samples than several established methods. Jian Huang 0003, Yuling Jiao, Lican Kang, Jin Liu 0011 |
IEEE Trans. Inf. Theory | 2 |
| 2025 | Model Free Prediction With Uncertainty AssessmentabstractDeep nonparametric regression, characterized by the utilization of deep neural networks to learn target functions, has emerged as a focus of research attention in recent years. Despite considerable progress in understanding convergence rates, the absence of asymptotic properties hinders rigorous statistical inference. To address this gap, we propose a novel framework that transforms the deep estimation paradigm into a platform conducive to conditional mean estimation, leveraging the conditional diffusion model. Theoretically, we develop an end-to-end convergence rate for the conditional diffusion model and establish the asymptotic normality of the generated samples. Consequently, we are equipped to construct confidence regions, facilitating robust statistical inference. Furthermore, through numerical experiments, we empirically validate the efficacy of our proposed methodology. Yuling Jiao, Lican Kang, Jin Liu 0011, Heng Peng, Heng Zuo |
IEEE Trans. Inf. Theory | 1 |
| 2025 | Error Analysis of Three-Layer Neural Network Trained With PGD for Deep Ritz MethodabstractMachine learning is a rapidly advancing field with diverse applications across various domains. One prominent area of research is the utilization of deep learning techniques for solving partial differential equations(PDEs). In this work, we specifically focus on employing a three-layer tanh neural network within the framework of the deep Ritz method(DRM) to solve second-order elliptic equations with three different types of boundary conditions. We perform projected gradient descent(PDG) to train the three-layer network and we establish its global convergence. To the best of our knowledge, we are the first to provide a comprehensive error analysis of using overparameterized networks to solve PDE problems, as our analysis simultaneously includes estimates for approximation error, generalization error, and optimization error. We present error bound in terms of the sample size n and our work provides guidance on how to set the network depth, width, step size, and number of iterations for the projected gradient descent algorithm. Importantly, our assumptions in this work are classical and we do not require any additional assumptions on the solution of the equation. This ensures the broad applicability and generality of our results. Yuling Jiao, Yanming Lai, Yang Wang 0020 |
IEEE Trans. Inf. Theory | 1 |
| 2024 | Neural Network Approximation for Pessimistic Offline Reinforcement LearningabstractDeep reinforcement learning (RL) has shown remarkable success in specific offline decision-making scenarios, yet its theoretical guarantees are still under development. Existing works on offline RL theory primarily emphasize a few trivial settings, such as linear MDP or general function approximation with strong assumptions and independent data, which lack guidance for practical use. The coupling of deep learning and Bellman residuals makes this problem challenging, in addition to the difficulty of data dependence. In this paper, we establish a non-asymptotic estimation error of pessimistic offline RL using general neural network approximation with C-mixing data regarding the structure of networks, the dimension of datasets, and the concentrability of data coverage, under mild assumptions. Our result shows that the estimation error consists of two parts: the first converges to zero at a desired rate on the sample size with partially controllable concentrability, and the second becomes negligible if the residual constraint is tight. This result demonstrates the explicit efficiency of deep adversarial offline RL frameworks. We utilize the empirical process tool for C-mixing sequences and the neural network approximation theory for the Holder class to achieve this. We also develop methods to bound the Bellman estimation error caused by function approximation with empirical Bellman constraint perturbations. Additionally, we present a result that lessens the curse of dimensionality using data with low intrinsic dimensionality and function classes with low complexity. Our estimation provides valuable insights into the development of deep offline RL and guidance for algorithm model design. Yuling Jiao, Li Shen 0008, Haizhao Yang, Xiliang Lu |
AAAI | 2 |
| 2024 | Non-asymptotic Approximation Error Bounds of Parameterized Quantum CircuitsabstractUnderstanding the power of parameterized quantum circuits (PQCs) in accomplishing machine learning tasks is one of the most important questions in quantum machine learning. In this paper, we focus on the PQC expressivity for general multivariate function classes. Previously established Universal Approximation Theorems for PQCs are either nonconstructive or assisted with parameterized classical data processing, making it hard to justify whether the expressive power comes from the classical or quantum parts. We explicitly construct data re-uploading PQCs for approximating multivariate polynomials and smooth functions and establish the first non-asymptotic approximation error bounds for such functions in terms of the number of qubits, the quantum circuit depth and the number of trainable parameters of the PQCs. Notably, we show that for multivariate polynomials and multivariate smooth functions, the quantum circuit size and the number of trainable parameters of our proposed PQCs can be smaller than the deep ReLU neural networks. We further demonstrate the approximation capability of PQCs via numerical experiments. Our results pave the way for designing practical PQCs that can be implemented on near-term quantum devices with limited resources. Qiuhao Chen, Yuling Jiao, Xiliang Lu, Jerry Zhijian Yang |
NeurIPS | 3 |
| 2024 | A Gaussian mixture distribution-based adaptive sampling method for physics-informed neural networks
Yuling Jiao, Xiliang Lu, Jerry Zhijian Yang |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | An error analysis for deep binary classification with sigmoid loss
Changshi Li, Yuling Jiao, Jerry Zhijian Yang |
Inf. Sci. | 2 |
| 2024 | Deep Nonparametric Quantile Regression under Covariate ShiftabstractThis work focuses on addressing the challenges posed by covariate shift in nonparametric quantile regression using deep neural networks. We propose a two-stage pre-training reweighted method that leverages importance weighting to mitigate the effects of distribution shift. In the first stage, density ratios are estimated with a neural network by minimizing least squares. In the second stage, a deep neural network estimator is obtained using pre-training weights. Theoretical analysis is provided, offering non-asymptotic error bounds for the unweighted, reweighted, and pre-training reweighted estimators. We consider scenarios with both bounded and unbounded density ratios. Notably, we employ a novel proof technique to bound the generalization error, characterized by the size and weights bound of ReLU neural networks. This enables us to establish fast rates of convergence under the adaptive self-calibration condition, distinguishing our approach from those relying on local Rademacher complexity techniques. Additionally, we derive the approximation error with weight bounds for ReLU neural networks approximating the Hölder class. Our theoretical findings provide valuable insights for the pre-training process and highlight the efficacy of reweighted techniques. Numerical experiments are conducted to further validate the theoretical findings and demonstrate the effectiveness of our proposed method. Xingdong Feng, Yuling Jiao, Lican Kang, Caixing Wang |
J. Mach. Learn. Res. | 3 |
| 2024 | Gaussian Interpolation FlowsabstractGaussian denoising has emerged as a powerful method for constructing simulation-free continuous normalizing flows for generative modeling. Despite their empirical successes, theoretical properties of these flows and the regularizing effect of Gaussian denoising have remained largely unexplored. In this work, we aim to address this gap by investigating the well-posedness of simulation-free continuous normalizing flows built on Gaussian denoising. Through a unified framework termed Gaussian interpolation flow, we establish the Lipschitz regularity of the flow velocity field, the existence and uniqueness of the flow, and the Lipschitz continuity of the flow map and the time-reversed flow map for several rich classes of target distributions. This analysis also sheds light on the auto-encoding and cycle consistency properties of Gaussian interpolation flows. Additionally, we study the stability of these flows in source distributions and perturbations of the velocity field, using the quadratic Wasserstein distance as a metric. Our findings offer valuable insights into the learning techniques employed in Gaussian interpolation flows for generative modeling, providing a solid theoretical foundation for end-to-end error analyses of learning Gaussian interpolation flows with empirical observations. Yuan Gao 0044, Jian Huang 0003, Yuling Jiao |
J. Mach. Learn. Res. | 3 |
| 2024 | Nonparametric Estimation of Non-Crossing Quantile Regression Process with Deep ReQU Neural NetworksabstractWe propose a penalized nonparametric approach to estimating the quantile regression process (QRP) in a nonseparable model using rectifier quadratic unit (ReQU) activated deep neural networks and introduce a novel penalty function to enforce non-crossing of quantile regression curves. We establish the non-asymptotic excess risk bounds for the estimated QRP and derive the mean integrated squared error for the estimated QRP under mild smoothness and regularity conditions. To establish these non-asymptotic risk and estimation error bounds, we also develop a new error bound for approximating $C^s$ smooth functions with $s >1$ and their derivatives using ReQU activated neural networks. This is a new approximation result for ReQU networks and is of independent interest and may be useful in other problems. Our numerical experiments demonstrate that the proposed method is competitive with or outperforms two existing methods, including methods using reproducing kernels and random forests for nonparametric quantile regression. Guohao Shen, Yuling Jiao, Joel L. Horowitz, Jian Huang 0003 |
J. Mach. Learn. Res. | 2 |
| 2024 | Deep Dimension Reduction for Supervised Representation LearningabstractThe goal of supervised representation learning is to construct effective data representations for prediction. Among all the characteristics of an ideal nonparametric representation of high-dimensional complex data, sufficiency, low dimensionality and disentanglement are some of the most essential ones. We propose a deep dimension reduction approach to learning representations with these characteristics. The proposed approach is a nonparametric generalization of the sufficient dimension reduction method. We formulate the ideal representation learning task as that of finding a nonparametric representation that minimizes an objective function characterizing conditional independence and promoting disentanglement at the population level. We then estimate the target representation at the sample level nonparametrically using deep neural networks. We show that the estimated deep nonparametric representation is consistent in the sense that its excess risk converges to zero. Our extensive numerical experiments using simulated and real benchmark data demonstrate that the proposed methods have better performance than several existing dimension reduction methods and the standard deep learning models in the context of classification and regression. Jian Huang 0003, Yuling Jiao, Jin Liu 0011, Zhou Yu 0006 |
IEEE Trans. Inf. Theory | 2 |
| 2023 | Fast Excess Risk Rates via Offset Rademacher ComplexityabstractBased on the offset Rademacher complexity, this work outlines a systematical framework for deriving sharp excess risk bounds in statistical learning without Bernstein condition. In addition to recovering fast rates in a unified way for some parametric and nonparametric supervised learning models with minimum identifiability assumptions, we also obtain new and improved results for LAD (sparse) linear regression and deep logistic regression with deep ReLU neural networks, respectively. Chenguang Duan, Yuling Jiao, Lican Kang, Xiliang Lu, Jerry Zhijian Yang |
ICML | 2 |
| 2023 | Invariant and Sufficient Supervised Representation LearningabstractImproving the generalization of neural networks under domain shift is an important and challenging task in computer vision. Obtaining an invariant representation across domains is a benchmark method in the literature. In this paper, we propose an invariant and sufficient supervised representation learning (ISSRL) approach to learn a domain invariant representation which is also preserving information used for downstream tasks. To this end, we formulate ISSRL by finding a nonlinear map$\boldsymbol{g}$such that$Y\perp X\vert \boldsymbol{g}(X)$and$(Y,\boldsymbol{g}(X))\perp D$at the population level, where D is the label of the domains and$(X, Y)$is the paired data sampled from domains with label. We use distance correlation to characterize the (conditional) independence. At the sample level, we construct a novel loss function through an unbiased empirical version of distance correlation. We train the representation map by parameterizing it with deep neural networks. Both simulation study and real data evaluation show that ISSRL outperforms the state-of-the-art on out-of-distribution generalization. The PyTorch code for ISSRL is available at https://github.com/CaC033/ISSRL. Junyu Zhu, Changshi Li, Yuling Jiao, Jin Liu 0011, Xiliang Lu |
IJCNN | 4 |
| 2023 | PALM: a powerful and adaptive latent model for prioritizing risk variants with functional annotationsabstractMOTIVATION: The findings from genome-wide association studies (GWASs) have greatly helped us to understand the genetic basis of human complex traits and diseases. Despite the tremendous progress, much effects are still needed to address several major challenges arising in GWAS. First, most GWAS hits are located in the non-coding region of human genome, and thus their biological functions largely remain unknown. Second, due to the polygenicity of human complex traits and diseases, many genetic risk variants with weak or moderate effects have not been identified yet. RESULTS: To address the above challenges, we propose a powerful and adaptive latent model (PALM) to integrate cell-type/tissue-specific functional annotations with GWAS summary statistics. Unlike existing methods, which are mainly based on linear models, PALM leverages a tree ensemble to adaptively characterize non-linear relationship between functional annotations and the association status of genetic variants. To make PALM scalable to millions of variants and hundreds of functional annotations, we develop a functional gradient-based expectation-maximization algorithm, to fit the tree-based non-linear model in a stable manner. Through comprehensive simulation studies, we show that PALM not only controls false discovery rate well, but also improves statistical power of identifying risk variants. We also apply PALM to integrate summary statistics of 30 GWASs with 127 cell type/tissue-specific functional annotations. The results indicate that PALM can identify more risk variants as well as rank the importance of functional annotations, yielding better interpretation of GWAS results. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/YangLabHKUST/PALM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiashun Xiao, Mingxuan Cai, Yuling Jiao, Jin Liu 0011, Can Yang 0002 |
Bioinform. | 4 |
| 2023 | Over-parameterized Deep Nonparametric Regression for Dependent Data with Its Applications to Reinforcement LearningabstractIn this paper, we provide statistical guarantees for over-parameterized deep nonparametric regression in the presence of dependent data. By decomposing the error, we establish non-asymptotic error bounds for deep estimation, which is achieved by effectively balancing the approximation and generalization errors. We have derived an approximation result for Hölder functions with constrained weights. Additionally, the generalization error is bounded by the weight norm, allowing for a neural network parameter number that is much larger than the training sample size. Furthermore, we address the issue of the curse of dimensionality by assuming that the samples originate from distributions with low intrinsic dimensions. Under this assumption, we are able to overcome the challenges posed by high-dimensional spaces. By incorporating an additional error propagation mechanism, we derive oracle inequalities for the over-parameterized deep fitted $Q$-iteration. Xingdong Feng, Yuling Jiao, Lican Kang, Baqun Zhang |
J. Mach. Learn. Res. | 2 |
| 2022 | Approximation with CNNs in Sobolev Space: with Applications to ClassificationabstractWe derive a novel approximation error bound with explicit prefactor for Sobolev-regular functions using deep convolutional neural networks (CNNs). The bound is non-asymptotic in terms of the network depth and filter lengths, in a rather flexible way. For Sobolev-regular functions which can be embedded into the H\"older space, the prefactor of our error bound depends on the ambient dimension polynomially instead of exponentially as in most existing results, which is of independent interest. We also establish a new approximation result when the target function is supported on an approximate lower-dimensional manifold. We apply our results to establish non-asymptotic excess risk bounds for classification using CNNs with convex surrogate losses, including the cross-entropy loss, the hinge loss (SVM), the logistic loss, the exponential loss and the least squares loss. We show that the classification methods with CNNs can circumvent the curse of dimensionality if input data is supported on a neighborhood of a low-dimensional manifold. Guohao Shen, Yuling Jiao, Jian Huang 0003 |
NeurIPS | 2 |
| 2022 | An Error Analysis of Generative Adversarial Networks for Learning DistributionsabstractThis paper studies how well generative adversarial networks (GANs) learn probability distributions from finite samples. Our main results establish the convergence rates of GANs under a collection of integral probability metrics defined through Hölder classes, including the Wasserstein distance as a special case. We also show that GANs are able to adaptively learn data distributions with low-dimensional structures or have Hölder densities, when the network architectures are chosen properly. In particular, for distributions concentrated around a low-dimensional set, we show that the learning rates of GANs do not depend on the high ambient dimension, but on the lower intrinsic dimension. Our analysis is based on a new oracle inequality decomposing the estimation error into the generator and discriminator approximation error and the statistical error, which may be of independent interest. Jian Huang 0003, Yuling Jiao, Zhen Li 0074, Shiao Liu, Yang Wang 0020, Yunfei Yang 0002 |
J. Mach. Learn. Res. | 2 |
| 2022 | PSNA: A pathwise semismooth Newton algorithm for sparse recovery with optimal local convergence and oracle properties
Jian Huang 0003, Yuling Jiao, Xiliang Lu, Yueyong Shi, Qinglong Yang |
Signal Process. | 2 |
| 2021 | Deep Generative Learning via Schrödinger BridgeabstractWe propose to learn a generative model via entropy interpolation with a Schr{ö}dinger Bridge. The generative learning task can be formulated as interpolating between a reference distribution and a target distribution based on the Kullback-Leibler divergence. At the population level, this entropy interpolation is characterized via an SDE on [0,1] with a time-varying drift term. At the sample level, we derive our Schr{ö}dinger Bridge algorithm by plugging the drift term estimated by a deep score estimator and a deep density ratio estimator into the Euler-Maruyama method. Under some mild smoothness assumptions of the target distribution, we prove the consistency of both the score estimator and the density ratio estimator, and then establish the consistency of the proposed Schr{ö}dinger Bridge approach. Our theoretical results guarantee that the distribution learned by our approach converges to the target distribution. Experimental results on multimodal synthetic data and benchmark data support our theoretical findings and indicate that the generative model via Schr{ö}dinger Bridge is comparable with state-of-the-art GANs, suggesting a new formulation of generative learning. We demonstrate its usefulness in image interpolation and image inpainting. Gefei Wang, Yuling Jiao, Yang Wang 0020, Can Yang 0002 |
ICML | 2 |
| 2021 | Non-asymptotic Error Bounds for Bidirectional GANsabstractWe derive nearly sharp bounds for the bidirectional GAN (BiGAN) estimation error under the Dudley distance between the latent joint distribution and the data joint distribution with appropriately specified architecture of the neural networks used in the model. To the best of our knowledge, this is the first theoretical guarantee for the bidirectional GAN learning approach. An appealing feature of our results is that they do not assume the reference and the data distributions to have the same dimensions or these distributions to have bounded support. These assumptions are commonly assumed in the existing convergence analysis of the unidirectional GANs but may not be satisfied in practice. Our results are also applicable to the Wasserstein bidirectional GAN if the target distribution is assumed to have a bounded support. To prove these results, we construct neural network functions that push forward an empirical distribution to another arbitrary empirical distribution on a possibly different-dimensional space. We also develop a novel decomposition of the integral probability metric for the error analysis of bidirectional GANs. These basic theoretical results are of independent interest and can be applied to other related learning problems. Shiao Liu, Yunfei Yang 0002, Jian Huang 0003, Yuling Jiao, Yang Wang 0020 |
NeurIPS | 4 |
| 2021 | Distributed quantile regression for massive heterogeneous data
Aijun Hu, Yuling Jiao, Yueyong Shi, Yuanshan Wu |
Neurocomputing | 2 |
| 2020 | CoMM-S2: a collaborative mixed model using summary statistics in transcriptome-wide association studiesabstractMOTIVATION: Although genome-wide association studies (GWAS) have deepened our understanding of the genetic architecture of complex traits, the mechanistic links that underlie how genetic variants cause complex traits remains elusive. To advance our understanding of the underlying mechanistic links, various consortia have collected a vast volume of genomic data that enable us to investigate the role that genetic variants play in gene expression regulation. Recently, a collaborative mixed model (CoMM) was proposed to jointly interrogate genome on complex traits by integrating both the GWAS dataset and the expression quantitative trait loci (eQTL) dataset. Although CoMM is a powerful approach that leverages regulatory information while accounting for the uncertainty in using an eQTL dataset, it requires individual-level GWAS data and cannot fully make use of widely available GWAS summary statistics. Therefore, statistically efficient methods that leverages transcriptome information using only summary statistics information from GWAS data are required. RESULTS: In this study, we propose a novel probabilistic model, CoMM-S2, to examine the mechanistic role that genetic variants play, by using only GWAS summary statistics instead of individual-level GWAS data. Similar to CoMM which uses individual-level GWAS data, CoMM-S2 combines two models: the first model examines the relationship between gene expression and genotype, while the second model examines the relationship between the phenotype and the predicted gene expression from the first model. Distinct from CoMM, CoMM-S2 requires only GWAS summary statistics. Using both simulation studies and real data analysis, we demonstrate that even though CoMM-S2 utilizes GWAS summary statistics, it has comparable performance as CoMM, which uses individual-level GWAS data. AVAILABILITY AND IMPLEMENTATION: The implement of CoMM-S2 is included in the CoMM package that can be downloaded from https://github.com/gordonliu810822/CoMM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi Yang 0028, Xingjie Shi, Yuling Jiao, Jian Huang 0003, Can Yang 0002, Jin Liu 0011 |
Bioinform. | 3 |
| 2020 | A Semismooth Newton Algorithm for High-Dimensional Nonconvex Sparse LearningabstractThe smoothly clipped absolute deviation (SCAD) and the minimax concave penalty (MCP)-penalized regression models are two important and widely used nonconvex sparse learning tools that can handle variable selection and parameter estimation simultaneously and thus have potential applications in various fields, such as mining biological data in high-throughput biomedical studies. Theoretically, these two models enjoy the oracle property even in the high-dimensional settings, where the number of predictors p may be much larger than the number of observations n . However, numerically, it is quite challenging to develop fast and stable algorithms due to their nonconvexity and nonsmoothness. In this article, we develop a fast algorithm for SCAD- and MCP-penalized learning problems. First, we show that the global minimizers of both models are roots of the nonsmooth equations. Then, a semismooth Newton (SSN) algorithm is employed to solve the equations. We prove that the SSN algorithm converges locally and superlinearly to the Karush-Kuhn-Tucker (KKT) points. The computational complexity analysis shows that the cost of the SSN algorithm per iteration is O(np) . Combined with the warm-start technique, the SSN algorithm can be very efficient and accurate. Simulation studies and a real data example suggest that our SSN algorithm, with comparable solution accuracy with the coordinate descent (CD) and the difference of convex (DC) proximal Newton algorithms, is more computationally efficient. Yueyong Shi, Jian Huang 0003, Yuling Jiao, Qinglong Yang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Deep Generative Learning via Variational Gradient FlowabstractWe propose a framework to learn deep generative models via \textbf{V}ariational \textbf{Gr}adient Fl\textbf{ow} (VGrow) on probability spaces. The evolving distribution that asymptotically converges to the target distribution is governed by a vector field, which is the negative gradient of the first variation of the $f$-divergence between them. We prove that the evolving distribution coincides with the pushforward distribution through the infinitesimal time composition of residual maps that are perturbations of the identity map along the vector field. The vector field depends on the density ratio of the pushforward distribution and the target distribution, which can be consistently learned from a binary classification problem. Connections of our proposed VGrow method with other popular methods, such as VAE, GAN and flow-based methods, have been established in this framework, gaining new insights of deep generative learning. We also evaluated several commonly used divergences, including Kullback-Leibler, Jensen-Shannon, Jeffreys divergences as well as our newly discovered “logD” divergence which serves as the objective function of the logD-trick GAN. Experimental results on benchmark datasets demonstrate that VGrow can generate high-fidelity images in a stable and efficient manner, achieving competitive performance with state-of-the-art GANs. Yuan Gao 0044, Yuling Jiao, Yang Wang 0020, Yao Wang 0003, Can Yang 0002, Shunkang Zhang |
ICML | 2 |
| 2019 | VIMCO: variational inference for multiple correlated outcomes in genome-wide association studiesabstractMOTIVATION: In genome-wide association studies (GWASs) where multiple correlated traits have been measured on participants, a joint analysis strategy, whereby the traits are analyzed jointly, can improve statistical power over a single-trait analysis strategy. There are two questions of interest to be addressed when conducting a joint GWAS analysis with multiple traits. The first question examines whether a genetic loci is significantly associated with any of the traits being tested. The second question focuses on identifying the specific trait(s) that is associated with the genetic loci. Since existing methods primarily focus on the first question, this article seeks to provide a complementary method that addresses the second question. RESULTS: We propose a novel method, Variational Inference for Multiple Correlated Outcomes (VIMCO) that focuses on identifying the specific trait that is associated with the genetic loci, when performing a joint GWAS analysis of multiple traits, while accounting for correlation among the multiple traits. We performed extensive numerical studies and also applied VIMCO to analyze two datasets. The numerical studies and real data analysis demonstrate that VIMCO improves statistical power over single-trait analysis strategies when the multiple traits are correlated and has comparable performance when the traits are not correlated. AVAILABILITY AND IMPLEMENTATION: The VIMCO software can be downloaded from: https://github.com/XingjieShi/VIMCO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xingjie Shi, Yuling Jiao, Yi Yang 0028, Ching Yu Cheng, Can Yang 0002, Jin Liu 0011 |
Bioinform. | 2 |
| 2018 | A Constructive Approach to $L_0$ Penalized RegressionabstractWe propose a constructive approach to estimating sparse, high-dimensional linear regression models. The approach is a computational algorithm motivated from the KKT conditions for the $\ell_0$-penalized least squares solutions. It generates a sequence of solutions iteratively, based on support detection using primal and dual information and root finding. We refer to the algorithm as SDAR for brevity. Under a sparse Riesz condition on the design matrix and certain other conditions, we show that with high probability, the $\ell_2$ estimation error of the solution sequence decays exponentially to the minimax error bound in $O(\log(R\sqrt{J}))$ iterations, where $J$ is the number of important predictors and $R$ is the relative magnitude of the nonzero target coefficients; and under a mutual coherence condition and certain other conditions, the $\ell_{\infty}$ estimation error decays to the optimal error bound in $O(\log(R))$ iterations. Moreover the SDAR solution recovers the oracle least squares estimator within a finite number of iterations with high probability if the sparsity level is known. Computational complexity analysis shows that the cost of SDAR is $O(np)$ per iteration. We also consider an adaptive version of SDAR for use in practical applications where the true sparsity level is unknown. Simulation studies demonstrate that SDAR outperforms Lasso, MCP and two greedy methods in accuracy and efficiency. Jian Huang 0003, Yuling Jiao, Xiliang Lu |
J. Mach. Learn. Res. | 2 |
| 2017 | Iterative Soft/Hard Thresholding With Homotopy Continuation for Sparse RecoveryabstractIn this note, we analyze an iterative soft/hard thresholding algorithm with homotopy continuation for recovering a sparse signal x†from noisy data of a noise level ε. Under suitable regularity and sparsity conditions, we design a path, along which the algorithm can find a solution x*, which admits a sharp reconstruction error ||x* - x†|| ℓ∞ = O(ε) with an iteration complexity O((ln ε)/(ln γ)np), where n and p are problem dimensionality and γ ε (0,1) controls the length of the path. Numerical examples are given to illustrate its performance. Yuling Jiao, Bangti Jin, Xiliang Lu |
IEEE Signal Process. Lett. | 1 |
| 2016 | Stripe Noise Separation and Removal in Remote Sensing Images by Consideration of the Global Sparsity and Local Variational PropertiesabstractRemote sensing images are often contaminated by varying degrees of stripes, which severely affects the visual quality and subsequent application of the data. Unlike with conventional methods, we achieve the destriping by separating the stripe component based on a full analysis of the various stripe properties. Under an optimization framework, an ℓ0-norm-based regularization is used to characterize the global sparse distribution of the stripes. In addition, difference-based constraints are adopted to describe the local smoothness and discontinuity in the along-stripe and across-stripe directions, respectively. The alternating direction method of multipliers is applied to solve and accelerate the model optimization. Experiments with both simulated and real data demonstrate the effectiveness of the proposed model, in terms of both qualitative and quantitative perspectives. Xinxin Liu 0002, Xiliang Lu, Huanfeng Shen, Qiangqiang Yuan, Yuling Jiao, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2015 | A multi-parameter regularization model for image restoration
Qibin Fan, Yuling Jiao |
Signal Process. | 3 |
| 2013 | Hybrid regularization image deblurring in the presence of impulsive noise
Fenge Chen, Yuling Jiao, Guorui Ma, Qianqing Qin |
J. Vis. Commun. Image Represent. | 2 |