EDBT 2026 Demo / reviewers in the wild / expert
Jianqing Fan
dblp:33/2768
· DBLP profile ↗
25ranked-venue papers
11as first author
15since 2021 · last 2026
0000-0003-3250-7677ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 11 since 2021Theory of computation · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Asymptotic Theory of Eigenvectors for Latent Embeddings With Generalized Laplacian MatricesabstractLaplacian matrices are widely used in practice to capture latent structural information in data, ranging from graphs to manifolds. Their normalization forms naturally induce dependencies among matrix entries, and such dependencies are known to pose significant challenges for advances in random matrix theory (RMT). Motivated by this, we introduce a general class of generalized Laplacian matrices, which includes both the standard Laplacian and random adjacency matrices as special cases, and we develop a new framework—Asymptotic Theory of Eigenvectors for latent embeddings with Generalized Laplacian matrices (ATE-GL)—for studying their spectral properties. Our theory is driven by two key ingredients: the use of generalized quadratic vector equations to handle dependency in RMT, and refined high-order asymptotic expansions for empirical spiked eigenvectors and eigenvalues based on local laws. These results lead to asymptotic normality for both spiked eigenvectors and eigenvalues, enabling precise statistical inference and uncertainty quantification for a broad class of applications involving generalized Laplacian matrices. We also discuss two motivating applications of the ATE-GL framework and demonstrate its effectiveness through numerical examples. Jianqing Fan, Jinchi Lv, Fan Yang 0106, Diwen Yu |
IEEE Trans. Inf. Theory | 1 |
| 2025 | Benign Overfitting in Out-of-Distribution Generalization of Linear ModelsabstractBenign overfitting refers to the phenomenon where an over-parameterized model fits the training data perfectly, including noise in the data, but still generalizes well to the unseen test data. While prior work provides some theoretical understanding of this phenomenon under the in-distribution setup, modern machine learning often operates in a more challenging Out-of-Distribution (OOD) regime, where the target (test) distribution can be rather different from the source (training) distribution. In this work, we take an initial step towards understanding benign overfitting in the OOD regime by focusing on the basic setup of over-parameterized linear models under covariate shift. We provide non-asymptotic guarantees proving that benign overfitting occurs in standard ridge regression, even under the OOD regime when the target covariance satisfies certain structural conditions. We identify several vital quantities relating to source and target covariance, which govern the performance of OOD generalization. Our result is sharp, which provably recovers prior in-distribution benign overfitting guarantee (Tsigler & Bartlett, 2023), as well as under-parameterized OOD guarantee (Ge et al., 2024) when specializing to each setup. Moreover, we also present theoretical results for a more general family of target covariance matrix, where standard ridge regression only achieves a slow statistical rate of $\mathcal{O}(1/\sqrt{n})$ for the excess risk, while Principal Component Regression (PCR) is guaranteed to achieve the fast rate $\mathcal{O}(1/n)$, where $n$ is the number of samples. Shange Tang, Jiayun Wu, Jianqing Fan, Chi Jin 0001 |
ICLR | 3 |
| 2025 | A Provable Initialization and Robust Clustering Method for General Mixture ModelsabstractClustering is a fundamental tool in statistical machine learning in the presence of heterogeneous data. Most recent results focus primarily on optimal mislabeling guarantees when data are distributed around centroids with sub-Gaussian errors. Yet, the restrictive sub-Gaussian model is often invalid in practice, since various real-world applications exhibit heavy-tail distributions around the centroids or suffer from possible adversarial attacks that call for robust clustering with a robust data-driven initialization. In this paper, we present initialization and subsequent clustering methods that provably guarantee near-optimal mislabeling for general mixture models when the number of clusters and data dimensions are finite. We first introduce a hybrid clustering technique with a novel multivariate trimmed mean type centroid estimate to produce mislabeling guarantees under a weak initialization condition for general error distributions around the centroids. A matching lower bound is derived, up to factors depending on the number of clusters. In addition, our approach also produces similar mislabeling guarantees even in the presence of adversarial outliers. Our results reduce to the sub-Gaussian case in finite dimensions when errors follow sub-Gaussian distributions. To solve the problem thoroughly, we also present novel data-driven robust initialization techniques and show that, with probabilities approaching one, these initial centroid estimates are sufficiently good for the subsequent clustering algorithm to achieve the optimal mislabeling rates. Furthermore, we demonstrate that Lloyd’s algorithm is suboptimal for more than two clusters even when errors are Gaussian and for two clusters when error distributions have heavy tails. Both simulated data and real data examples further support our robust initialization procedure and clustering algorithm. Soham Jana, Jianqing Fan, Sanjeev R. Kulkarni |
IEEE Trans. Inf. Theory | 2 |
| 2024 | Minimax-optimal reward-agnostic exploration in reinforcement learningabstractThis paper studies reward-agnostic exploration in reinforcement learning (RL) — a scenario where the learner is unware of the reward functions during the exploration stage — and designs an algorithm that improves over the state of the art. More precisely, consider a finite-horizon inhomogeneous Markov decision process with $S$ states, $A$ actions, and horizon length $H$, and suppose that there are no more than a polynomial number of given reward functions of interest. By collecting an order of $\frac{SAH^3}{\varepsilon^2}$ sample episodes (up to log factor) without guidance of the reward information, our algorithm is able to find $\varepsilon$-optimal policies for all these reward functions, provided that $\varepsilon$ is sufficiently small. This forms the first reward-agnostic exploration scheme in this context that achieves provable minimax optimality. Furthermore, once the sample size exceeds $\frac{S^2AH^3}{\varepsilon^2}$ episodes (up to log factor), our algorithm is able to yield $\varepsilon$ accuracy for arbitrarily many reward functions (even when they are adversarially designed), a task commonly dubbed as “reward-free exploration.” The novelty of our algorithm design draws on insights from offline RL: the exploration scheme attempts to maximize a critical reward-agnostic quantity that dictates the performance of offline RL, while the policy learning paradigm leverages ideas from sample-optimal offline RL paradigms. Gen Li 0005, Yuling Yan, Yuxin Chen 0002, Jianqing Fan |
COLT | 4 |
| 2024 | On the Provable Advantage of Unsupervised PretrainingabstractUnsupervised pretraining, which learns a useful representation using a large amount of unlabeled data to facilitate the learning of downstream tasks, is a critical component of modern large-scale machine learning systems. Despite its tremendous empirical success, the rigorous theoretical understanding of why unsupervised pretraining generally helps remains rather limited---most existing results are restricted to particular methods or approaches for unsupervised pretraining with specialized structural assumptions. This paper studies a generic framework,
where the unsupervised representation learning task is specified by an abstract class of latent variable models $\Phi$ and the downstream task is specified by a class of prediction functions $\Psi$. We consider a natural approach of using Maximum Likelihood Estimation (MLE) for unsupervised pretraining and Empirical Risk Minimization (ERM) for learning downstream tasks. We prove that, under a mild ``informative'' condition, our algorithm achieves an excess risk of $\\tilde{\\mathcal{O}}(\sqrt{\mathcal{C}\_\Phi/m} + \sqrt{\mathcal{C}\_\Psi/n})$ for downstream tasks, where $\mathcal{C}\_\Phi, \mathcal{C}\_\Psi$ are complexity measures of function classes $\Phi, \Psi$, and $m, n$ are the number of unlabeled and labeled data respectively. Comparing to the baseline of $\tilde{\mathcal{O}}(\sqrt{\mathcal{C}\_{\Phi \circ \Psi}/n})$ achieved by performing supervised learning using only the labeled data, our result rigorously shows the benefit of unsupervised pretraining when $m \gg n$ and $\mathcal{C}\_{\Phi\circ \Psi} > \mathcal{C}\_\Psi$. This paper further shows that our generic framework covers a wide range of approaches for unsupervised pretraining, including factor models, Gaussian mixture models, and contrastive learning. Jiawei Ge 0003, Shange Tang, Jianqing Fan, Chi Jin 0001 |
ICLR | 3 |
| 2024 | Maximum Likelihood Estimation is All You Need for Well-Specified Covariate ShiftabstractA key challenge of modern machine learning systems is to achieve Out-of-Distribution (OOD) generalization---generalizing to target data whose distribution differs from that of source data. Despite its significant importance, the fundamental question of ``what are the most effective algorithms for OOD generalization'' remains open even under the standard setting of covariate shift.
This paper addresses this fundamental question by proving that, surprisingly, classical Maximum Likelihood Estimation (MLE) purely using source data (without any modification) achieves the *minimax* optimality for covariate shift under the *well-specified* setting. That is, *no* algorithm performs better than MLE in this setting (up to a constant factor), justifying MLE is all you need.
Our result holds for a very rich class of parametric models, and does not require any boundedness condition on the density ratio. We illustrate the wide applicability of our framework by instantiating it to three concrete examples---linear regression, logistic regression, and phase retrieval. This paper further complement the study by proving that, under the *misspecified setting*, MLE is no longer the optimal choice, whereas Maximum Weighted Likelihood Estimator (MWLE) emerges as minimax optimal in certain scenarios. Jiawei Ge 0003, Shange Tang, Jianqing Fan, Cong Ma 0001, Chi Jin 0001 |
ICLR | 3 |
| 2024 | Global Convergence in Training Large-Scale TransformersabstractDespite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously analyzes the convergence properties of gradient flow in training Transformers with weight decay regularization. First, we construct the mean-field limit of large-scale Transformers, showing that as the model width and depth go to infinity, gradient flow converges to the Wasserstein gradient flow, which is represented by a partial differential equation. Then, we demonstrate that the gradient flow reaches a global minimum consistent with the PDE solution when the weight decay regularization parameter is sufficiently small. Our analysis is based on a series of novel mean-field techniques that adapt to Transformers. Compared with existing tools for deep networks (Lu et al., 2020) that demand homogeneity and global Lipschitz smoothness, we utilize a refined analysis assuming only $\textit{partial homogeneity}$ and $\textit{local Lipschitz smoothness}$. These new techniques may be of independent interest. Yuan Cao 0006, Yihan He, Mengdi Wang 0001, Han Liu 0001, Jason M. Klusowski, Jianqing Fan |
NeurIPS | 8 |
| 2024 | Optimal Aggregation of Prediction Intervals under Unsupervised Domain ShiftabstractAs machine learning models are increasingly deployed in dynamic environments, it becomes paramount to assess and quantify uncertainties associated with distribution shifts.
A distribution shift occurs when the underlying data-generating process changes, leading to a deviation in the model's performance.
The prediction interval, which captures the range of likely outcomes for a given prediction, serves as a crucial tool for characterizing uncertainties induced by their underlying distribution.
In this paper, we propose methodologies for aggregating prediction intervals to obtain one with minimal width and adequate coverage on the target domain under unsupervised domain shift, under which we have labeled samples from a related source domain and unlabeled covariates from the target domain.
Our analysis encompasses scenarios where the source and the target domain are related via i) a bounded density ratio, and ii) a measure-preserving transformation.
Our proposed methodologies are computationally efficient and easy to implement. Beyond illustrating the performance of our method through real-world datasets, we also delve into the theoretical details. This includes establishing rigorous theoretical guarantees, coupled with finite sample bounds, regarding the coverage and width of our prediction intervals. Our approach excels in practical applications and is underpinned by a solid theoretical framework, ensuring its reliability and effectiveness across diverse contexts. Jiawei Ge 0003, Debarghya Mukherjee, Jianqing Fan |
NeurIPS | 3 |
| 2024 | One-Layer Transformer Provably Learns One-Nearest Neighbor In ContextabstractTransformers have achieved great success in recent years. Interestingly, transformers have shown particularly strong in-context learning capability -- even without fine-tuning, they are still able to solve unseen tasks well purely based on task-specific prompts. In this paper, we study the capability of one-layer transformers in learning the one-nearest neighbor prediction rule. Under a theoretical framework where the prompt contains a sequence of labeled training data and unlabeled test data, we show that, although the loss function is nonconvex, when trained with gradient descent, a single softmax attention layer can successfully learn to behave like a one-nearest neighbor classifier. Our result gives a concrete example on how transformers can be trained to implement nonparametric machine learning algorithms, and sheds light on the role of softmax attention in transformer models. Yuan Cao 0006, Yihan He, Han Liu 0001, Jason M. Klusowski, Jianqing Fan, Mengdi Wang 0001 |
NeurIPS | 7 |
| 2024 | Uncertainty Quantification of MLE for Entity Ranking with CovariatesabstractWe study statistical estimation and inference for the ranking problems based on pairwise comparisons with additional covariate information. In specific, in this paper, we study a Covariate-Assisted Ranking Estimation (CARE) model in a systematic way, that extends the well-known Bradley-Terry-Luce (BTL) model by incorporating the covariate information. We impose natural identifiability conditions, derive the statistical rates for the MLE under a sparse comparison graph, and obtain its asymptotic distribution. Moreover, we validate our theoretical results through large-scale numerical studies. Jianqing Fan, Jikai Hou, Mengxin Yu |
J. Mach. Learn. Res. | 1 |
| 2023 | The Efficacy of Pessimism in Asynchronous Q-LearningabstractThis paper is concerned with the asynchronous form of Q-learning, which applies a stochastic approximation scheme to Markovian data samples. Motivated by the recent advances in offline reinforcement learning, we develop an algorithmic framework that incorporates the principle of pessimism into asynchronous Q-learning, which penalizes infrequently-visited state-action pairs based on suitable lower confidence bounds (LCBs). This framework leads to, among other things, improved sample efficiency and enhanced adaptivity in the presence of near-expert data. Our approach permits the observed data in some important scenarios to cover only partial state-action space, which is in stark contrast to prior theory that requires uniform coverage of all state-action pairs. When coupled with the idea of variance reduction, asynchronous Q-learning with LCB penalization achieves near-optimal sample complexity, provided that the target accuracy level is small enough. In comparison, prior works were suboptimal in terms of the dependency on the effective horizon even when i.i.d. sampling is permitted. Our results deliver the first theoretical support for the use of pessimism principle in the presence of Markovian non-i.i.d. data. Yuling Yan, Gen Li 0005, Yuxin Chen 0002, Jianqing Fan |
IEEE Trans. Inf. Theory | 4 |
| 2022 | Constructing Phrase-level Semantic Labels to Form Multi-Grained Supervision for Image-Text RetrievalabstractExisting research for image text retrieval mainly relies on sentence-level supervision to distinguish matched and mismatched sentences for a query image. However, semantic mismatch between an image and sentences usually happens in finer grain, i.e., phrase level. In this paper, we explore to introduce additional phrase-level supervision for the better identification of mismatched units in the text. In practice, multi-grained semantic labels are automatically constructed for a query image in both sentence-level and phrase-level. We construct text scene graphs for the matched sentences and extract entities and triples as the phrase-level labels. In order to integrate both supervision of sentence-level and phrase-level, we propose Semantic Structure Aware Multimodal Transformer (SSAMT) for multi-modal representation learning. Inside the SSAMT, we utilize different kinds of attention mechanisms to enforce interactions of multi-grained semantic units in both sides of vision and language. For the training, we propose multi-scale matching from both global and local perspectives, and penalize mismatched phrases. Experimental results on MS-COCO and Flickr30K show the effectiveness of our approach compared to some state-of-the-art models. Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Haijun Shan, Xuanjing Huang 0001, Jianqing Fan |
ICMR | 7 |
| 2021 | Sample-Efficient Reinforcement Learning for Linearly-Parameterized MDPs with a Generative ModelabstractThe curse of dimensionality is a widely known issue in reinforcement learning (RL). In the tabular setting where the state space $\mathcal{S}$ and the action space $\mathcal{A}$ are both finite, to obtain a near optimal policy with sampling access to a generative model, the minimax optimal sample complexity scales linearly with $|\mathcal{S}|\times|\mathcal{A}|$, which can be prohibitively large when $\mathcal{S}$ or $\mathcal{A}$ is large. This paper considers a Markov decision process (MDP) that admits a set of state-action features, which can linearly express (or approximate) its probability transition kernel. We show that a model-based approach (resp.$~$Q-learning) provably learns an $\varepsilon$-optimal policy (resp.$~$Q-function) with high probability as soon as the sample size exceeds the order of $\frac{K}{(1-\gamma)^{3}\varepsilon^{2}}$ (resp.$~$$\frac{K}{(1-\gamma)^{4}\varepsilon^{2}}$), up to some logarithmic factor. Here $K$ is the feature dimension and $\gamma\in(0,1)$ is the discount factor of the MDP. Both sample complexity bounds are provably tight, and our result for the model-based approach matches the minimax lower bound. Our results show that for arbitrarily large-scale MDP, both the model-based approach and Q-learning are sample-efficient when $K$ is relatively small, and hence the title of this paper. Yuling Yan, Jianqing Fan |
NeurIPS | 3 |
| 2021 | Curriculum Learning for Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) is a task where an agent navigates in an embodied indoor environment under human instructions. Previous works ignore the distribution of sample difficulty and we argue that this potentially degrade their agent performance. To tackle this issue, we propose a novel curriculum- based training paradigm for VLN tasks that can balance human prior knowledge and agent learning progress about training samples. We develop the principle of curriculum design and re-arrange the benchmark Room-to-Room (R2R) dataset to make it suitable for curriculum training. Experiments show that our method is model-agnostic and can significantly improve the performance, the generalizability, and the training efficiency of current state-of-the-art navigation agents without increasing model complexity. Jiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie Peng |
NeurIPS | 3 |
| 2021 | Hoeffding's Inequality for General Markov Chains and Its Applications to Statistical LearningabstractThis paper establishes Hoeffding's lemma and inequality for bounded functions of general-state-space and not necessarily reversible Markov chains. The sharpness of these results is characterized by the optimality of the ratio between variance proxies in the Markov-dependent and independent settings. The boundedness of functions is shown necessary for such results to hold in general. To showcase the usefulness of the new results, we apply them for non-asymptotic analyses of MCMC estimation, respondent-driven sampling and high-dimensional covariance matrix estimation on time series data with a Markovian nature. In addition to statistical problems, we also apply them to study the time-discounted rewards in econometric models and the multi-armed bandit problem with Markovian rewards arising from the field of machine learning. Jianqing Fan, Bai Jiang, Qiang Sun 0007 |
J. Mach. Learn. Res. | 1 |
| 2018 | Statistical Sparse Online Regression: A Diffusion Approximation PerspectiveabstractIn this paper, we propose to adopt the diffusion approximation techniques to study online regression. The diffusion approximation techniques allow us to characterize the exact dynamics of the online regression process. As a consequence, we obtain the optimal statistical rate of convergence up to a logarithmic factor of the streaming sample size. Using the idea of trajectory averaging, we further improve the rate of convergence by eliminating the logarithmic factor. Lastly, we propose a two-step algorithm for sparse online regression: a burn-in step using offline learning and a refinement step using a variant of truncated stochastic gradient descent. Under appropriate assumptions, we show the proposed algorithm produces near optimal sparse estimators. Numerical experiments lend further support to our obtained theory. Jianqing Fan, Wenyan Gong, Chris Junchi Li, Qiang Sun 0007 |
AISTATS | 1 |
| 2017 | An $\ell_{\infty}$ Eigenvector Perturbation Bound and Its Application
Jianqing Fan, Yiqiao Zhong |
J. Mach. Learn. Res. | 1 |
| 2016 | Guarding against Spurious Discoveries in High DimensionsabstractMany data mining and statistical machine learning algorithms have been developed to select a subset of covariates to associate with a response variable. Spurious discoveries can easily arise in high-dimensional data analysis due to enormous possibilities of such selections. How can we know statistically our discoveries better than those by chance? In this paper, we define a measure of goodness of spurious fit, which shows how good a response variable can be fitted by an optimally selected subset of covariates under the null model, and propose a simple and effective LAMM algorithm to compute it. It coincides with the maximum spurious correlation for linear models and can be regarded as a generalized maximum spurious correlation. We derive the asymptotic distribution of such goodness of spurious fit for generalized linear models and $L_1$ regression. Such an asymptotic distribution depends on the sample size, ambient dimension, the number of variables used in the fit, and the covariance information. It can be consistently estimated by multiplier bootstrapping and used as a benchmark to guard against spurious discoveries. It can also be applied to model selection, which considers only candidate models with goodness of fits better than those by spurious fits. The theory and method are convincingly illustrated by simulated examples and an application to the binary outcomes from German Neuroblastoma Trials. Jianqing Fan, Wen-Xin Zhou |
J. Mach. Learn. Res. | 1 |
| 2013 | Distributions of angles in random packing on spheres
T. Tony Cai, Jianqing Fan, Tiefeng Jiang |
J. Mach. Learn. Res. | 2 |
| 2011 | Adaptively and Spatially Estimating the Hemodynamic Response Functions in fMRI
Jiaping Wang, Hongtu Zhu, Jianqing Fan, Kelly S. Giovanello, Weili Lin |
MICCAI (2) | 3 |
| 2011 | Nonconcave Penalized Likelihood With NP-DimensionalityabstractPenalized likelihood methods are fundamental to ultra-high dimensional variable selection. How high dimensionality such methods can handle remains largely unknown. In this paper, we show that in the context of generalized linear models, such methods possess model selection consistency with oracle properties even for dimensionality of Non-Polynomial (NP) order of sample size, for a class of penalized likelihood approaches using folded-concave penalty functions, which were introduced to ameliorate the bias problems of convex penalty functions. This fills a long-standing gap in the literature where the dimensionality is allowed to grow slowly with the sample size. Our results are also applicable to penalized likelihood with the L(1)-penalty, which is a convex function at the boundary of the class of folded-concave penalty functions under consideration. The coordinate optimization is implemented for finding the solution paths, whose performance is evaluated by a few simulation examples and the real data analysis. Jianqing Fan, Jinchi Lv |
IEEE Trans. Inf. Theory | 1 |
| 2009 | Ultrahigh Dimensional Feature Selection: Beyond The Linear Model
Jianqing Fan, Richard Samworth, Yichao Wu |
J. Mach. Learn. Res. | 1 |
| 2007 | Selection and validation of normalization methods for c-DNA microarrays using within-array replicationsabstractMOTIVATION: Normalization of microarray data is essential for multiple-array analyses. Several normalization protocols have been proposed based on different biological or statistical assumptions. A fundamental problem arises whether they have effectively normalized arrays. In addition, for a given array, the question arises how to choose a method to most effectively normalize the microarray data. RESULTS: We propose several techniques to compare the effectiveness of different normalization methods. We approach the problem by constructing statistics to test whether there are any systematic biases in the expression profiles among duplicated spots within an array. The test statistics involve estimating the genewise variances. This is accomplished by using several novel methods, including empirical Bayes methods for moderating the genewise variances and the smoothing methods for aggregating variance information. P-values are estimated based on a normal or chi approximation. With estimated P-values, we can choose a most appropriate method to normalize a specific array and assess the extent to which the systematic biases due to the variations of experimental conditions have been removed. The effectiveness and validity of the proposed methods are convincingly illustrated by a carefully designed simulation study. The method is further illustrated by an application to human placenta cDNAs comprising a large number of clones with replications, a customized microarray experiment carrying just a few hundred genes on the study of the molecular roles of Interferons on tumor, and the Agilent microarrays carrying tens of thousands of total RNA samples in the MAQC project on the study of reproducibility, sensitivity and specificity of the data. AVAILABILITY: Code to implement the method in the statistical package R is available from the authors. Jianqing Fan |
Bioinform. | 1 |
| 2003 | Data-analytic approaches to the estimation of Value-at-RiskabstractValue-at-risk measures the worst loss to be expected of a portfolio over a given time horizon at a given confidence level. Calculation of VaR frequently involves estimating the volatility of return processes and quantiles of standardized returns. In this paper, several semiparametric techniques are introduced to estimate the volatilities. In addition, both parametric and nonparametric techniques are proposed to estimate the quantiles of standardized return processes. The newly proposed techniques also have the flexibility to adapt automatically to the changes in the dynamics of market prices over time. The combination of newly proposed techniques for estimating volatility and standardized quantiles yields several new techniques for evaluating multiple period VaR. The performance of the newly proposed VaR estimators is evaluated and compared with some of existing methods. Our simulation results and empirical studies endorse the newly proposed time-dependent semiparametric approach for estimating VaR. Jianqing Fan, Juan Gu |
CIFEr | 1 |
| 2002 | Wavelet deconvolutionabstractThis paper studies the issue of optimal deconvolution density estimation using wavelets. The approach taken here can be considered as orthogonal series estimation in the more general context of the density estimation. We explore the asymptotic properties of estimators based on thresholding of estimated wavelet coefficients. Minimax rates of convergence under the integrated square loss are studied over Besov classes B/sub /spl sigma/pq/ of functions for both ordinary smooth and supersmooth convolution kernels. The minimax rates of convergence depend on the smoothness of functions to be deconvolved and the decay rate of the characteristic function of convolution kernels. It is shown that no linear deconvolution estimators can achieve the optimal rates of convergence in the Besov spaces with p<2 when the convolution kernel is ordinary smooth and super smooth. If the convolution kernel is ordinary smooth, then linear estimators can be improved by using thresholding wavelet deconvolution estimators which are asymptotically minimax within logarithmic terms. Adaptive minimax properties of thresholding wavelet deconvolution estimators are also discussed. Jianqing Fan, Ja-Yong Koo |
IEEE Trans. Inf. Theory | 1 |