Ke Alexander Wang

dblp:238/0269 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 40% Probabilistic and Bayesian machine learning · 28% Transfer learning and domain adaptation · 7%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 50% GPUs and heterogeneous computing · 50%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.922021
SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes · ICML 2021
Exact Gaussian Processes on a Million Data Points · NeurIPS 2019
Machine learning › Deep learning architectures and training
convolutional neural network
0.712023
Sequence Modeling with Multiresolution Convolutional Memory · ICML 2023
Machine learning › Deep learning architectures and training › convolutional neural network › dilated convolution
dilated causal convolution
0.712023
Sequence Modeling with Multiresolution Convolutional Memory · ICML 2023
Machine learning › Deep learning architectures and training
sequence modeling
0.712023
Sequence Modeling with Multiresolution Convolutional Memory · ICML 2023
Machine learning › Transfer learning and domain adaptation › instance weighting
importance weighting
0.612022
Is Importance Weighting Incompatible with Interpolating Classifiers? · ICLR 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
Is Importance Weighting Incompatible with Interpolating Classifiers? · ICLR 2022
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design
0.512021
Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual Information · ICML 2021
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.512021
Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual Information · ICML 2021
Machine learning › Learning theory › nonparametric regression › kernel regression
kernel interpolation
0.512021
SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes · ICML 2021
Machine learning › Representation and self-supervised learning
mutual information
0.512021
Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual Information · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
scalable gaussian process
0.512021
SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes · ICML 2021
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks
0.412020
Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit Constraints · NeurIPS 2020
Machine learning › Deep learning architectures and training › physics-informed neural network
hamiltonian neural network
0.412020
Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit Constraints · NeurIPS 2020
Machine learning › Deep learning architectures and training
physics-informed neural network
0.412020
Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit Constraints · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
exact inference
0.412019
Exact Gaussian Processes on a Million Data Points · NeurIPS 2019
GPUs and heterogeneous computing
multi-GPU computing
0.412019
Exact Gaussian Processes on a Million Data Points · NeurIPS 2019
Parallel and multicore computing
parallel computing
0.412019
Exact Gaussian Processes on a Million Data Points · NeurIPS 2019
Image and video processing › image filtering › edge-preserving filtering
bilateral filtering
0.112021
SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes · ICML 2021
Graph algorithms and graph theory
shortest path
0.112021
Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual Information · ICML 2021

Methods — techniques the papers use, named apart from their topics

structured kernel interpolation · 1.0mutual information · 1.0gaussian process · 1.0evolution strategies · 1.0CUDA · 1.0wavelet decomposition · 0.7multiresolution analysis · 0.7importance weighting · 0.6lagrange multipliers · 0.4explicit constraints · 0.4linear conjugate gradients · 0.4kernel matrix multiplication · 0.4
YearPublicationVenuePosition
2023 Sequence Modeling with Multiresolution Convolutional Memory
abstract
Efficiently capturing the long-range patterns in sequential data sources salient to a given task---such as classification and generative modeling---poses a fundamental challenge. Popular approaches in the space tradeoff between the memory burden of brute-force enumeration and comparison, as in transformers, the computational burden of complicated sequential dependencies, as in recurrent neural networks, or the parameter burden of convolutional networks with many or large filters. We instead take inspiration from wavelet-based multiresolution analysis to define a new building block for sequence modeling, which we call a MultiresLayer. The key component of our model is the multiresolution convolution, capturing multiscale trends in the input sequence. Our MultiresConv can be implemented with shared filters across a dilated causal convolution tree. Thus it garners the computational advantages of convolutional networks and the principled theoretical motivation of wavelet decompositions. Our MultiresLayer is straightforward to implement, requires significantly fewer parameters, and maintains at most a $O(N \log N)$ memory footprint for a length $N$ sequence. Yet, by stacking such layers, our model yields state-of-the-art performance on a number of sequence classification and autoregressive density estimation tasks using CIFAR-10, ListOps, and PTB-XL datasets.
Jiaxin Shi, Ke Alexander Wang, Emily B. Fox
ICML2
2022 Is Importance Weighting Incompatible with Interpolating Classifiers?
Ke Alexander Wang, Niladri S. Chatterji, Saminul Haque, Tatsunori B. Hashimoto
ICLR1
2021 SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes
abstract
State-of-the-art methods for scalable Gaussian processes use iterative algorithms, requiring fast matrix vector multiplies (MVMs) with the co-variance kernel. The Structured Kernel Interpolation (SKI) framework accelerates these MVMs by performing efficient MVMs on a grid and interpolating back to the original space. In this work, we develop a connection between SKI and the permutohedral lattice used for high-dimensional fast bilateral filtering. Using a sparse simplicial grid instead of a dense rectangular one, we can perform GP inference exponentially faster in the dimension than SKI. Our approach, Simplex-GP, enables scaling SKI to high dimensions, while maintaining strong predictive performance. We additionally provide a CUDA implementation of Simplex-GP, which enables significant GPU acceleration of MVM based inference.
Sanyam Kapoor, Marc Finzi, Ke Alexander Wang, Andrew Gordon Wilson
ICML3
2021 Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual Information
abstract
In many real world problems, we want to infer some property of an expensive black-box function f, given a budget of T function evaluations. One example is budget constrained global optimization of f, for which Bayesian optimization is a popular method. Other properties of interest include local optima, level sets, integrals, or graph-structured information induced by f. Often, we can find an algorithm A to compute the desired property, but it may require far more than T queries to execute. Given such an A, and a prior distribution over f, we refer to the problem of inferring the output of A using T evaluations as Bayesian Algorithm Execution (BAX). To tackle this problem, we present a procedure, InfoBAX, that sequentially chooses queries that maximize mutual information with respect to the algorithm’s output. Applying this to Dijkstra’s algorithm, for instance, we infer shortest paths in synthetic and real-world graphs with black-box edge costs. Using evolution strategies, we yield variants of Bayesian optimization that target local, rather than global, optima. On these problems, InfoBAX uses up to 500 times fewer queries to f than required by the original algorithm. Our method is closely connected to other Bayesian optimal experimental design procedures such as entropy search methods and optimal sensor placement using Gaussian processes.
Willie Neiswanger, Ke Alexander Wang, Stefano Ermon
ICML2
2020 Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit Constraints
abstract
Reasoning about the physical world requires models that are endowed with the right inductive biases to learn the underlying dynamics. Recent works improve generalization for predicting trajectories by learning the Hamiltonian or Lagrangian of a system rather than the differential equations directly. While these methods encode the constraints of the systems using generalized coordinates, we show that embedding the system into Cartesian coordinates and enforcing the constraints explicitly with Lagrange multipliers dramatically simplifies the learning problem. We introduce a series of challenging chaotic and extended-body systems, including systems with $N$-pendulums, spring coupling, magnetic fields, rigid rotors, and gyroscopes, to push the limits of current approaches. Our experiments show that Cartesian coordinates with explicit constraints lead to a 100x improvement in accuracy and data efficiency.
Marc Finzi, Ke Alexander Wang, Andrew Gordon Wilson
NeurIPS2
2019 DC2: A Divide-and-conquer Algorithm for Large-scale Kernel Learning with Application to Clustering
abstract
Divide-and-conquer is a general strategy to deal with large scale problems. It is typically applied to generate ensemble instances, which potentially limits the problem size it can handle. Additionally, the data are often divided by random sampling which may be suboptimal. To address these concerns, we propose the DC2algorithm. Instead of ensemble instances, we produce structure-preserving signature pieces to be assembled and conquered. DC2achieves the efficiency of sampling-based large scale kernel methods while enabling parallel multicore or clustered computation. The data partition and subsequent compression are unified by recursive random projections. Empirically dividing the data by random projections induces smaller mean squared approximation errors than conventional random sampling. The power of DC2is demonstrated by our clustering algorithm rpfCluster+, which is as accurate as some fastest approximate spectral clustering algorithms while maintaining a running time close to that of K-means clustering. Analysis on DC2when applied to spectral clustering shows that the loss in clustering accuracy due to data division and reduction is upper bounded by the data approximation error which would vanish with recursive random projections. Due to its easy implementation and flexibility, we expect DC2to be applicable to general large scale learning problems.
Ke Alexander Wang, Xinran Bian, Pan Liu 0014, Donghui Yan
IEEE BigData1
2019 Exact Gaussian Processes on a Million Data Points
abstract
Gaussian processes (GPs) are flexible non-parametric models, with a capacity that grows with the available data. However, computational constraints with standard inference procedures have limited exact GPs to problems with fewer than about ten thousand training points, necessitating approximations for larger datasets. In this paper, we develop a scalable approach for exact GPs that leverages multi-GPU parallelization and methods like linear conjugate gradients, accessing the kernel matrix only through matrix multiplication. By partitioning and distributing kernel matrix multiplies, we demonstrate that an exact GP can be trained on over a million points, a task previously thought to be impossible with current computing hardware. Moreover, our approach is generally applicable, without constraints to grid data or specific kernel classes. Enabled by this scalability, we perform the first-ever comparison of exact GPs against scalable GP approximations on datasets with $10^4 \!-\! 10^6$ data points, showing dramatic performance improvements.
Ke Alexander Wang, Geoff Pleiss, Jacob R. Gardner, Stephen Tyree, Kilian Q. Weinberger, Andrew Gordon Wilson
NeurIPS1