VLDB 2026 Research / reviewers in the wild / expert
Raymond K. W. Wong
dblp:57/7726
· DBLP profile ↗
17ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0001-9342-3755ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Learning theory · 28% Reinforcement learning · 24% Probabilistic and Bayesian machine learning · 17% | |
| Theoretical computer science
6 papers |
Mathematical optimization · 62% Algorithms and data structures · 38% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
off-policy evaluation |
1.6 | 2 | 2025 | A Principled Path to Fitted Distributional Evaluation · NeurIPS 2025 A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models · ICML 2024 |
Mathematical optimization › continuous optimization › matrix optimization › matrix recovery
matrix completion |
1.5 | 3 | 2024 | A Pairwise Pseudo-likelihood Approach for Matrix Completion with Informative Missingness · NeurIPS 2024 Median Matrix Completion: from Embarrassment to Optimality · ICML 2020 Matrix Completion with Noisy Entries and Outliers · J. Mach. Learn. Res. 2017 |
Machine learning › Reinforcement learning › off-policy evaluation
fitted q-evaluation |
1.0 | 2 | 2025 | A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models · ICML 2024 A Principled Path to Fitted Distributional Evaluation · NeurIPS 2025 |
Algorithms and data structures
kernel methods |
1.0 | 1 | 2026 | Flexible Functional Treatment Effect Estimation · J. Mach. Learn. Res. 2026 |
Algorithms and data structures › kernel methods
kernel ridge regression |
1.0 | 1 | 2026 | Flexible Functional Treatment Effect Estimation · J. Mach. Learn. Res. 2026 |
Machine learning › Optimization for machine learning
convergence analysis |
0.9 | 1 | 2025 | A Principled Path to Fitted Distributional Evaluation · NeurIPS 2025 |
Machine learning › Reinforcement learning › off-policy evaluation
distributional off-policy evaluation |
0.9 | 1 | 2025 | A Principled Path to Fitted Distributional Evaluation · NeurIPS 2025 |
Machine learning › Learning theory › statistical estimation › minimax estimation
minimax rates |
0.8 | 1 | 2024 | A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models · ICML 2024 |
Machine learning › Learning theory › statistical estimation
nonparametric estimation |
0.8 | 1 | 2024 | A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
pseudolikelihood estimation |
0.8 | 1 | 2024 | A Pairwise Pseudo-likelihood Approach for Matrix Completion with Informative Missingness · NeurIPS 2024 |
Machine learning › Learning theory
statistical learning theory |
0.8 | 1 | 2024 | A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models · ICML 2024 |
Machine learning › Learning theory
sample complexity |
0.7 | 2 | 2019 | Provably Accurate Double-Sparse Coding · J. Mach. Learn. Res. 2019 A Provable Approach for Double-Sparse Coding · AAAI 2018 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding |
0.7 | 2 | 2019 | Provably Accurate Double-Sparse Coding · J. Mach. Learn. Res. 2019 A Provable Approach for Double-Sparse Coding · AAAI 2018 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.7 | 1 | 2023 | Directed Cyclic Graph for Causal Discovery from Multivariate Functional Data · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › uncertainty estimation
bayesian uncertainty quantification |
0.7 | 1 | 2023 | Directed Cyclic Graph for Causal Discovery from Multivariate Functional Data · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
0.7 | 1 | 2023 | Directed Cyclic Graph for Causal Discovery from Multivariate Functional Data · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning
functional data |
0.7 | 1 | 2023 | Directed Cyclic Graph for Causal Discovery from Multivariate Functional Data · NeurIPS 2023 |
Machine learning › Optimization for machine learning › sparse learning
group sparsity |
0.7 | 1 | 2023 | Implicit Regularization for Group Sparsity · ICLR 2023 |
Machine learning › Optimization for machine learning
implicit regularization |
0.7 | 1 | 2023 | Implicit Regularization for Group Sparsity · ICLR 2023 |
Machine learning › Deep learning architectures and training
autoencoder |
0.5 | 1 | 2021 | Benefits of Jointly Training Autoencoders: An Improved Neural Tangent Kernel Analysis · IEEE Trans. Inf. Theory 2021 |
Machine learning › Optimization for machine learning › convergence guarantees
gradient descent convergence |
0.5 | 1 | 2021 | Benefits of Jointly Training Autoencoders: An Improved Neural Tangent Kernel Analysis · IEEE Trans. Inf. Theory 2021 |
Natural language and speech › Machine translation
joint training |
0.5 | 1 | 2021 | Benefits of Jointly Training Autoencoders: An Improved Neural Tangent Kernel Analysis · IEEE Trans. Inf. Theory 2021 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.5 | 1 | 2021 | Benefits of Jointly Training Autoencoders: An Improved Neural Tangent Kernel Analysis · IEEE Trans. Inf. Theory 2021 |
Machine learning › Learning theory
over-parameterization |
0.5 | 1 | 2021 | Benefits of Jointly Training Autoencoders: An Improved Neural Tangent Kernel Analysis · IEEE Trans. Inf. Theory 2021 |
Machine learning › Learning theory
statistical guarantees |
0.5 | 1 | 2021 | Matrix Completion with Model-free Weighting · ICML 2021 |
Machine learning and data management
matrix completion |
0.5 | 1 | 2021 | Matrix Completion with Model-free Weighting · ICML 2021 |
Mathematical optimization › statistical estimation
robust estimation |
0.4 | 1 | 2020 | Median Matrix Completion: from Embarrassment to Optimality · ICML 2020 |
Mathematical optimization › sparse learning
dictionary learning |
0.4 | 1 | 2019 | Provably Accurate Double-Sparse Coding · J. Mach. Learn. Res. 2019 |
Mathematical optimization › continuous optimization › matrix optimization › matrix recovery › matrix completion
noisy matrix completion |
0.3 | 1 | 2017 | Matrix Completion with Noisy Entries and Outliers · J. Mach. Learn. Res. 2017 |
Mathematical optimization › continuous optimization › matrix optimization › matrix recovery › matrix completion
robust matrix completion |
0.3 | 1 | 2017 | Matrix Completion with Noisy Entries and Outliers · J. Mach. Learn. Res. 2017 |
Methods — techniques the papers use, named apart from their topics
gradient descent · 1.8fitted q-evaluation · 1.6regularization · 1.5pairwise pseudo-likelihood · 1.5low-rank estimation · 1.5convex optimization · 1.5group sparsity regularization · 1.3representer theorem · 1.0kernel ridge regression · 1.0distributional reinforcement learning · 0.9ratio function · 0.8nonparametric regression · 0.8bayesian inference · 0.7importance weighting · 0.5pseudo data · 0.4parallel computing · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flexible Functional Treatment Effect EstimationabstractWe study treatment effect estimation with functional treatments where the average potential outcome functional is a function of functions, in contrast to continuous treatment effect estimation where the target is a function of real numbers. By considering a flexible scalar-on-function marginal structural model, a weight-modified kernel ridge regression (WMKRR) is adopted for estimation. The weights are constructed by directly minimizing the uniform balancing error resulting from a decomposition of the WMKRR estimator, instead of being estimated under a particular treatment selection model. Despite the complex structure of the uniform balancing error derived under WMKRR, finite-dimensional convex algorithms can be applied to efficiently solve for the proposed weights thanks to a representer theorem. The optimal convergence rate is shown to be attainable by the proposed WMKRR estimator without any smoothness assumption on the true weight function. Corresponding empirical performance is demonstrated by a simulation study and a real data application. Raymond K. W. Wong, Kwun Chuen Gary Chan |
J. Mach. Learn. Res. | 2 |
| 2025 | Distributional Off-policy Evaluation with Bellman Residual MinimizationabstractWe study distributional off-policy evaluation (OPE), of which the goal is to learn the distribution of the return for a target policy using offline data generated by a different policy. The theoretical foundation of many existing work relies on the supremum-extended statistical distances such as supremum-Wasserstein distance, which are hard to estimate. In contrast, we study the more manageable expectation-extended statistical distances and provide a novel theoretical justification on their validity for learning the return distribution. Based on this attractive property, we propose a new method called Energy Bellman Residual Minimizer (EBRM) for distributional OPE. We provide corresponding in-depth theoretical analyses. We establish a finite-sample error bound for the EBRM estimator under the realizability assumption. Furthermore, we introduce a variant of our method based on a multi-step extension which improves the error bound for non-realizable settings. Notably, unlike prior distributional OPE methods, the theoretical guarantees of our method do not require the completeness assumption. Sungee Hong, Zhengling Qi, Raymond K. W. Wong |
AISTATS | 3 |
| 2025 | A Principled Path to Fitted Distributional EvaluationabstractIn reinforcement learning, distributional off-policy evaluation (OPE) focuses on estimating the return distribution of a target policy using offline data collected under a different policy. This work focuses on extending the widely used fitted Q-evaluation---developed for expectation-based reinforcement learning---to the distributional OPE setting. We refer to this extension as fitted distributional evaluation (FDE). While only a few related approaches exist, there remains no unified framework for designing FDE methods. To fill this gap, we present a set of guiding principles for constructing theoretically grounded FDE methods. Building on these principles, we develop several new FDE methods with convergence analysis and provide theoretical justification for existing methods, even in non-tabular environments. Extensive experiments, including simulations on linear quadratic regulators and Atari games, demonstrate the superior performance of the FDE methods. Sungee Hong, Zhengling Qi, Raymond K. W. Wong |
NeurIPS | 4 |
| 2024 | A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric ModelsabstractIn this paper, we delve into the statistical analysis of the fitted Q-evaluation (FQE) method, which focuses on estimating the value of a target policy using offline data generated by some behavior policy. We provide a comprehensive theoretical understanding of FQE estimators under both parametric and non-parametric models on the Q-function. Specifically, we address three key questions related to FQE that remain largely unexplored in the current literature: (1) Is the optimal convergence rate for estimating the policy value regarding the sample size $n$ ($n^{−1/2}$) achievable for FQE under a nonparametric model with a fixed horizon ($T$ )? (2) How does the error bound depend on the horizon T ? (3) What is the role of the probability ratio function in improving the convergence of FQE estimators? Specifically, we show that under the completeness assumption of Q-functions, which is mild in the non-parametric setting, the estimation errors for policy value using both parametric and non-parametric FQE estimators can achieve an optimal rate in terms of n. The corresponding error bounds in terms of both $n$ and $T$ are also established. With an additional realizability assumption on ratio functions, the rate of estimation errors can be improved from $T^{ 1.5}/\sqrt{n}$ to $T /\sqrt{n}$, which matches the sharpest known bound in the current literature under the tabular setting. Zhengling Qi, Raymond K. W. Wong |
ICML | 3 |
| 2024 | A Pairwise Pseudo-likelihood Approach for Matrix Completion with Informative MissingnessabstractWhile several recent matrix completion methods are developed to deal with non-uniform observation probabilities across matrix entries, very few allow the missingness to depend on the mostly unobserved matrix measurements, which is generally ill-posed. We aim to tackle a subclass of these ill-posed settings, characterized by a flexible separable observation probability assumption that can depend on the matrix measurements. We propose a regularized pairwise pseudo-likelihood approach for matrix completion and prove that the proposed estimator can asymptotically recover the low-rank parameter matrix up to an identifiable equivalence class of a constant shift and scaling, at a near-optimal asymptotic convergence rate of the standard well-posed (non-informative missing) setting, while effectively mitigating the impact of informative missingness. The efficacy of our method is validated via numerical experiments, positioning it as a robust tool for matrix completion to mitigate data bias. Jiangyuan Li, Raymond K. W. Wong, Kwun Chuen Gary Chan |
NeurIPS | 3 |
| 2023 | Implicit Regularization for Group Sparsity
Jiangyuan Li, Thanh Van Nguyen, Chinmay Hegde, Raymond K. W. Wong |
ICLR | 4 |
| 2023 | Directed Cyclic Graph for Causal Discovery from Multivariate Functional DataabstractDiscovering causal relationship using multivariate functional data has received a significant amount of attention very recently. In this article, we introduce a functional linear structural equation model for causal structure learning when the underlying graph involving the multivariate functions may have cycles. To enhance interpretability, our model involves a low-dimensional causal embedded space such that all the relevant causal information in the multivariate functional data is preserved in this lower-dimensional subspace. We prove that the proposed model is causally identifiable under standard assumptions that are often made in the causal discovery literature. To carry out inference of our model, we develop a fully Bayesian framework with suitable prior specifications and uncertainty quantification through posterior summaries. We illustrate the superior performance of our method over existing methods in terms of causal graph estimation through extensive simulation studies. We also demonstrate the proposed method using a brain EEG dataset. Saptarshi Roy, Raymond K. W. Wong |
NeurIPS | 2 |
| 2022 | Extending the Use of MDL for High-Dimensional Problems: Variable Selection, Robust Fitting, and Additive ModelingabstractIn the signal processing and statistics literature, the minimum description length (MDL) principle is a popular tool for choosing model complexity. Successful examples include signal denoising and variable selection in linear regression, for which the corresponding MDL solutions often enjoy consistent properties and produce very promising empirical results. This paper demonstrates that MDL can be extended naturally to the high-dimensional setting, where the number of predictors p is larger than the number of observations n. It first considers the case of linear regression, then allows for outliers in the data, and lastly extends to the robust fitting of nonparametric additive models. Results from numerical experiments are presented to demonstrate the efficiency and effectiveness of the MDL approach. Raymond K. W. Wong, Thomas C. M. Lee |
ICASSP | 2 |
| 2021 | Matrix Completion with Model-free WeightingabstractIn this paper, we propose a novel method for matrix completion under general non-uniform missing structures. By controlling an upper bound of a novel balancing error, we construct weights that can actively adjust for the non-uniformity in the empirical risk without explicitly modeling the observation probabilities, and can be computed efficiently via convex optimization. The recovered matrix based on the proposed weighted empirical risk enjoys appealing theoretical guarantees. In particular, the proposed method achieves stronger guarantee than existing work in terms of the scaling with respect to the observation probabilities, under asymptotically heterogeneous missing settings (where entry-wise observation probabilities can be of different orders). These settings can be regarded as a better theoretical model of missing patterns with highly varying probabilities. We also provide a new minimax lower bound under a class of heterogeneous settings. Numerical experiments are also provided to demonstrate the effectiveness of the proposed method. Raymond K. W. Wong, Xiaojun Mao, Kwun Chuen Gary Chan |
ICML | 2 |
| 2021 | Benefits of Jointly Training Autoencoders: An Improved Neural Tangent Kernel AnalysisabstractDeep neural networks can achieve impressive performance in the regime where they are massively over-parameterized. Consequently, over the past year, there has been a growing interest in analyzing optimization and generalization properties of over-parameterized networks. However, the majority of existing work only applies to supervised learning. The role of over-parameterization in the unsupervised setting has by contrast gained far less attention. In this paper, we study the inductive bias of gradient descent for two-layer over-parameterized autoencoders with ReLU activation. We first provide theoretical evidence for the memorization phenomena observed in recent work using the property that infinitely wide neural networks under gradient descent evolve as linear models. We also analyze the gradient dynamics of the autoencoders in the finite-width setting. Starting from a randomly initialized autoencoder network, we rigorously prove the linear convergence of gradient descent in two weakly-trained and jointly-trained regimes. Our results indicate the considerable benefits of joint training over weak training in finding global optima, achieving a dramatic decrease in the required level of over-parameterization. Finally, we analyze the case of weight-tied autoencoders and prove that in the over-parameterized setting, training such networks from randomly initialized points leads to certain unexpected degeneracies. Thanh Van Nguyen, Raymond K. W. Wong, Chinmay Hegde |
IEEE Trans. Inf. Theory | 2 |
| 2020 | Median Matrix Completion: from Embarrassment to OptimalityabstractIn this paper, we consider matrix completion with absolute deviation loss and obtain an estimator of the median matrix. Despite several appealing properties of median, the non-smooth absolute deviation loss leads to computational challenge for large-scale data sets which are increasingly common among matrix completion problems. A simple solution to large-scale problems is parallel computing. However, embarrassingly parallel fashion often leads to inefficient estimators. Based on the idea of pseudo data, we propose a novel refinement step, which turns such inefficient estimators into a rate (near-)optimal matrix completion procedure. The refined estimator is an approximation of a regularized least median estimator, and therefore not an ordinary regularized empirical risk estimator. This leads to a non-standard analysis of asymptotic behaviors. Empirical results are also provided to confirm the effectiveness of the proposed method. Weidong Liu 0005, Xiaojun Mao, Raymond K. W. Wong |
ICML | 3 |
| 2019 | On the Dynamics of Gradient Descent for AutoencodersabstractWe provide a series of results for unsupervised learning with autoencoders. Specifically, we study shallow two-layer autoencoder architectures with shared weights. We focus on three generative models for data that are common in statistical machine learning: (i) the mixture-of-gaussians model, (ii) the sparse coding model, and (iii) the sparsity model with non-negative coefficients. For each of these models, we prove that under suitable choices of hyperparameters, architectures, and initialization, autoencoders learned by gradient descent can successfully recover the parameters of the corresponding model. To our knowledge, this is the first result that rigorously studies the dynamics of gradient descent for weight-sharing autoencoders. Our analysis can be viewed as theoretical evidence that shallow autoencoder modules indeed can be used as feature learning mechanisms for a variety of data models, and may shed insight on how to train larger stacked architectures with autoencoders as basic building blocks. Thanh Van Nguyen, Raymond K. W. Wong, Chinmay Hegde |
AISTATS | 2 |
| 2019 | Provably Accurate Double-Sparse CodingabstractSparse coding is a crucial subroutine in algorithms for various signal processing, deep learning, and other machine learning applications. The central goal is to learn an overcomplete dictionary that can sparsely represent a given input dataset. However, a key challenge is that storage, transmission, and processing of the learned dictionary can be untenably high if the data dimension is high. In this paper, we consider the double-sparsity model introduced by Rubinstein et al. (2010b) where the dictionary itself is the product of a fixed, known basis and a data-adaptive sparse component. First, we introduce a simple algorithm for double-sparse coding that can be amenable to efficient implementation via neural architectures. Second, we theoretically analyze its performance and demonstrate asymptotic sample complexity and running time benefits over existing (provable) approaches for sparse coding. To our knowledge, our work introduces the first computationally efficient algorithm for double-sparse coding that enjoys rigorous statistical guarantees. Finally, we corroborate our theory with several numerical experiments on simulated data, suggesting that our method may be useful for problem sizes encountered in practice. Thanh Van Nguyen, Raymond K. W. Wong, Chinmay Hegde |
J. Mach. Learn. Res. | 2 |
| 2019 | Locally linear embedding with additive noise
Justin Wang, Raymond K. W. Wong, Thomas C. M. Lee |
Pattern Recognit. Lett. | 2 |
| 2018 | A Provable Approach for Double-Sparse CodingabstractSparse coding is a crucial subroutine in algorithms for various signal processing, deep learning, and other machine learning applications. The central goal is to learn an overcomplete dictionary that can sparsely represent a given dataset. However, storage, transmission, and processing of the learned dictionary can be untenably high if the data dimension is high. In this paper, we consider the double-sparsity model introduced by Rubinstein, Zibulevsky, and Elad (2010) where the dictionary itself is the product of a fixed, known basis and a data-adaptive sparse component. First, we introduce a simple algorithm for double-sparse coding that can be amenable to efficient implementation via neural architectures. Second, we theoretically analyze its performance and demonstrate asymptotic sample complexity and running time benefits over existing (provable) approaches for sparse coding. To our knowledge, our work introduces the first computationally efficient algorithm for double-sparse coding that enjoys rigorous statistical guarantees. Finally, we support our analysis via several numerical experiments on simulated data, confirming that our method can indeed be useful in problem sizes encountered in practical applications. Thanh Van Nguyen, Raymond K. W. Wong, Chinmay Hegde |
AAAI | 2 |
| 2017 | Matrix Completion with Noisy Entries and OutliersabstractThis paper considers the problem of matrix completion when the observed entries are noisy and contain outliers. It begins with introducing a new optimization criterion for which the recovered matrix is defined as its solution. This criterion uses the celebrated Huber function from the robust statistics literature to downweigh the effects of outliers. A practical algorithm is developed to solve the optimization involved. This algorithm is fast, straightforward to implement, and monotonic convergent. Furthermore, the proposed methodology is theoretically shown to be stable in a well defined sense. Its promising empirical performance is demonstrated via a sequence of simulation experiments, including image inpainting. Raymond K. W. Wong, Thomas C. M. Lee |
J. Mach. Learn. Res. | 1 |
| 2010 | Structural break estimation of noisy sinusoidal signals
Raymond K. W. Wong, Randy C. S. Lai, Thomas C. M. Lee |
Signal Process. | 1 |