VLDB 2026 Research / reviewers in the wild / expert
Ke Alexander Wang
dblp:238/0269
· DBLP profile ↗
7ranked-venue papers
3as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 40% Probabilistic and Bayesian machine learning · 28% Transfer learning and domain adaptation · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 50% GPUs and heterogeneous computing · 50% |
Topics — the 19 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.9 | 2 | 2021 | SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes · ICML 2021 Exact Gaussian Processes on a Million Data Points · NeurIPS 2019 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.7 | 1 | 2023 | Sequence Modeling with Multiresolution Convolutional Memory · ICML 2023 |
Machine learning › Deep learning architectures and training › convolutional neural network › dilated convolution
dilated causal convolution |
0.7 | 1 | 2023 | Sequence Modeling with Multiresolution Convolutional Memory · ICML 2023 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.7 | 1 | 2023 | Sequence Modeling with Multiresolution Convolutional Memory · ICML 2023 |
Machine learning › Transfer learning and domain adaptation › instance weighting
importance weighting |
0.6 | 1 | 2022 | Is Importance Weighting Incompatible with Interpolating Classifiers? · ICLR 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.6 | 1 | 2022 | Is Importance Weighting Incompatible with Interpolating Classifiers? · ICLR 2022 |
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design |
0.5 | 1 | 2021 | Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual Information · ICML 2021 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.5 | 1 | 2021 | Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual Information · ICML 2021 |
Machine learning › Learning theory › nonparametric regression › kernel regression
kernel interpolation |
0.5 | 1 | 2021 | SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes · ICML 2021 |
Machine learning › Representation and self-supervised learning
mutual information |
0.5 | 1 | 2021 | Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual Information · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
scalable gaussian process |
0.5 | 1 | 2021 | SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes · ICML 2021 |
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks |
0.4 | 1 | 2020 | Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit Constraints · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › physics-informed neural network
hamiltonian neural network |
0.4 | 1 | 2020 | Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit Constraints · NeurIPS 2020 |
Machine learning › Deep learning architectures and training
physics-informed neural network |
0.4 | 1 | 2020 | Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit Constraints · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
exact inference |
0.4 | 1 | 2019 | Exact Gaussian Processes on a Million Data Points · NeurIPS 2019 |
GPUs and heterogeneous computing
multi-GPU computing |
0.4 | 1 | 2019 | Exact Gaussian Processes on a Million Data Points · NeurIPS 2019 |
Parallel and multicore computing
parallel computing |
0.4 | 1 | 2019 | Exact Gaussian Processes on a Million Data Points · NeurIPS 2019 |
Image and video processing › image filtering › edge-preserving filtering
bilateral filtering |
0.1 | 1 | 2021 | SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes · ICML 2021 |
Graph algorithms and graph theory
shortest path |
0.1 | 1 | 2021 | Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual Information · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
structured kernel interpolation · 1.0mutual information · 1.0gaussian process · 1.0evolution strategies · 1.0CUDA · 1.0wavelet decomposition · 0.7multiresolution analysis · 0.7importance weighting · 0.6lagrange multipliers · 0.4explicit constraints · 0.4linear conjugate gradients · 0.4kernel matrix multiplication · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Sequence Modeling with Multiresolution Convolutional MemoryabstractEfficiently capturing the long-range patterns in sequential data sources salient to a given task---such as classification and generative modeling---poses a fundamental challenge. Popular approaches in the space tradeoff between the memory burden of brute-force enumeration and comparison, as in transformers, the computational burden of complicated sequential dependencies, as in recurrent neural networks, or the parameter burden of convolutional networks with many or large filters. We instead take inspiration from wavelet-based multiresolution analysis to define a new building block for sequence modeling, which we call a MultiresLayer. The key component of our model is the multiresolution convolution, capturing multiscale trends in the input sequence. Our MultiresConv can be implemented with shared filters across a dilated causal convolution tree. Thus it garners the computational advantages of convolutional networks and the principled theoretical motivation of wavelet decompositions. Our MultiresLayer is straightforward to implement, requires significantly fewer parameters, and maintains at most a $O(N \log N)$ memory footprint for a length $N$ sequence. Yet, by stacking such layers, our model yields state-of-the-art performance on a number of sequence classification and autoregressive density estimation tasks using CIFAR-10, ListOps, and PTB-XL datasets. Jiaxin Shi, Ke Alexander Wang, Emily B. Fox |
ICML | 2 |
| 2022 | Is Importance Weighting Incompatible with Interpolating Classifiers?
Ke Alexander Wang, Niladri S. Chatterji, Saminul Haque, Tatsunori B. Hashimoto |
ICLR | 1 |
| 2021 | SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian ProcessesabstractState-of-the-art methods for scalable Gaussian processes use iterative algorithms, requiring fast matrix vector multiplies (MVMs) with the co-variance kernel. The Structured Kernel Interpolation (SKI) framework accelerates these MVMs by performing efficient MVMs on a grid and interpolating back to the original space. In this work, we develop a connection between SKI and the permutohedral lattice used for high-dimensional fast bilateral filtering. Using a sparse simplicial grid instead of a dense rectangular one, we can perform GP inference exponentially faster in the dimension than SKI. Our approach, Simplex-GP, enables scaling SKI to high dimensions, while maintaining strong predictive performance. We additionally provide a CUDA implementation of Simplex-GP, which enables significant GPU acceleration of MVM based inference. Sanyam Kapoor, Marc Finzi, Ke Alexander Wang, Andrew Gordon Wilson |
ICML | 3 |
| 2021 | Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual InformationabstractIn many real world problems, we want to infer some property of an expensive black-box function f, given a budget of T function evaluations. One example is budget constrained global optimization of f, for which Bayesian optimization is a popular method. Other properties of interest include local optima, level sets, integrals, or graph-structured information induced by f. Often, we can find an algorithm A to compute the desired property, but it may require far more than T queries to execute. Given such an A, and a prior distribution over f, we refer to the problem of inferring the output of A using T evaluations as Bayesian Algorithm Execution (BAX). To tackle this problem, we present a procedure, InfoBAX, that sequentially chooses queries that maximize mutual information with respect to the algorithm’s output. Applying this to Dijkstra’s algorithm, for instance, we infer shortest paths in synthetic and real-world graphs with black-box edge costs. Using evolution strategies, we yield variants of Bayesian optimization that target local, rather than global, optima. On these problems, InfoBAX uses up to 500 times fewer queries to f than required by the original algorithm. Our method is closely connected to other Bayesian optimal experimental design procedures such as entropy search methods and optimal sensor placement using Gaussian processes. Willie Neiswanger, Ke Alexander Wang, Stefano Ermon |
ICML | 2 |
| 2020 | Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit ConstraintsabstractReasoning about the physical world requires models that are endowed with the right inductive biases to learn the underlying dynamics. Recent works improve generalization for predicting trajectories by learning the Hamiltonian or Lagrangian of a system rather than the differential equations directly. While these methods encode the constraints of the systems using generalized coordinates, we show that embedding the system into Cartesian coordinates and enforcing the constraints explicitly with Lagrange multipliers dramatically simplifies the learning problem. We introduce a series of challenging chaotic and extended-body systems, including systems with $N$-pendulums, spring coupling, magnetic fields, rigid rotors, and gyroscopes, to push the limits of current approaches. Our experiments show that Cartesian coordinates with explicit constraints lead to a 100x improvement in accuracy and data efficiency. Marc Finzi, Ke Alexander Wang, Andrew Gordon Wilson |
NeurIPS | 2 |
| 2019 | DC2: A Divide-and-conquer Algorithm for Large-scale Kernel Learning with Application to ClusteringabstractDivide-and-conquer is a general strategy to deal with large scale problems. It is typically applied to generate ensemble instances, which potentially limits the problem size it can handle. Additionally, the data are often divided by random sampling which may be suboptimal. To address these concerns, we propose the DC2algorithm. Instead of ensemble instances, we produce structure-preserving signature pieces to be assembled and conquered. DC2achieves the efficiency of sampling-based large scale kernel methods while enabling parallel multicore or clustered computation. The data partition and subsequent compression are unified by recursive random projections. Empirically dividing the data by random projections induces smaller mean squared approximation errors than conventional random sampling. The power of DC2is demonstrated by our clustering algorithm rpfCluster+, which is as accurate as some fastest approximate spectral clustering algorithms while maintaining a running time close to that of K-means clustering. Analysis on DC2when applied to spectral clustering shows that the loss in clustering accuracy due to data division and reduction is upper bounded by the data approximation error which would vanish with recursive random projections. Due to its easy implementation and flexibility, we expect DC2to be applicable to general large scale learning problems. Ke Alexander Wang, Xinran Bian, Pan Liu 0014, Donghui Yan |
IEEE BigData | 1 |
| 2019 | Exact Gaussian Processes on a Million Data PointsabstractGaussian processes (GPs) are flexible non-parametric models, with a capacity that grows with the available data. However, computational constraints with standard inference procedures have limited exact GPs to problems with fewer than about ten thousand training points, necessitating approximations for larger datasets. In this paper, we develop a scalable approach for exact GPs that leverages multi-GPU parallelization and methods like linear conjugate gradients, accessing the kernel matrix only through matrix multiplication. By partitioning and distributing kernel matrix multiplies, we demonstrate that an exact GP can be trained on over a million points, a task previously thought to be impossible with current computing hardware. Moreover, our approach is generally applicable, without constraints to grid data or specific kernel classes. Enabled by this scalability, we perform the first-ever comparison of exact GPs against scalable GP approximations on datasets with $10^4 \!-\! 10^6$ data points, showing dramatic performance improvements. Ke Alexander Wang, Geoff Pleiss, Jacob R. Gardner, Stephen Tyree, Kilian Q. Weinberger, Andrew Gordon Wilson |
NeurIPS | 1 |