Maximilian Lam

dblp:173/5115 · also Max Lam · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Network and information security
2 papers
Privacy and data protection · 70% Cryptographic protocols and secure computation · 30%
Theoretical computer science
1 paper
Coding theory · 67% Combinatorics and discrete mathematics · 33%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 85% High-performance computing · 15%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genomics
genome-wide association study
1.122023
GWAS quality score for evaluating associated regions in GWAS analyses · Bioinform. 2023
RICOPILI: Rapid Imputation for COnsortias PIpeLIne · Bioinform. 2020
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval
0.812024
GPU-based Private Information Retrieval for On-Device Machine Learning Inference · ASPLOS (1) 2024
Privacy and data protection › privacy-preserving machine learning
privacy-preserving machine learning inference
0.812024
GPU-based Private Information Retrieval for On-Device Machine Learning Inference · ASPLOS (1) 2024
Cryptographic protocols and secure computation
private information retrieval
0.812024
GPU-based Private Information Retrieval for On-Device Machine Learning Inference · ASPLOS (1) 2024
GPUs and heterogeneous computing
GPU computing
0.812024
GPU-based Private Information Retrieval for On-Device Machine Learning Inference · ASPLOS (1) 2024
Bioinformatics and computational biology
genomics
0.712023
GWAS quality score for evaluating associated regions in GWAS analyses · Bioinform. 2023
Bioinformatics and computational biology › population genetics
linkage disequilibrium
0.712023
GWAS quality score for evaluating associated regions in GWAS analyses · Bioinform. 2023
Privacy and data protection
anonymity
0.512021
Gradient Disaggregation: Breaking Privacy in Federated Learning by Reconstructing the User Participant Matrix · ICML 2021
Privacy and data protection › privacy-preserving machine learning
federated learning privacy
0.512021
Gradient Disaggregation: Breaking Privacy in Federated Learning by Reconstructing the User Participant Matrix · ICML 2021
Bioinformatics and computational biology › genomics › genotyping
genotype imputation
0.412020
RICOPILI: Rapid Imputation for COnsortias PIpeLIne · Bioinform. 2020
Combinatorics and discrete mathematics
card shuffling
0.312018
Speeding Up Distributed Machine Learning Using Codes · IEEE Trans. Inf. Theory 2018
Coding theory › error-correcting codes
coded computation
0.312018
Speeding Up Distributed Machine Learning Using Codes · IEEE Trans. Inf. Theory 2018
Coding theory › error-correcting codes › coded computation › coded distributed computing
distributed matrix multiplication
0.312018
Speeding Up Distributed Machine Learning Using Codes · IEEE Trans. Inf. Theory 2018
High-performance computing
scientific computing systems
0.112020
RICOPILI: Rapid Imputation for COnsortias PIpeLIne · Bioinform. 2020
Machine learning › Efficient and distributed learning
distributed training
0.112018
Speeding Up Distributed Machine Learning Using Codes · IEEE Trans. Inf. Theory 2018

Methods — techniques the papers use, named apart from their topics

PIR-ML co-design · 2.3quality control · 0.9polygenic risk scoring · 0.9meta-analysis · 0.9straggler mitigation · 0.7coding theory · 0.7secure aggregation · 0.5gradient inference · 0.5
YearPublicationVenuePosition
2024 GPU-based Private Information Retrieval for On-Device Machine Learning Inference
abstract
On-device machine learning (ML) inference can enable the use of private user data on user devices without revealing them to remote servers. However, a pure on-device solution to private ML inference is impractical for many applications that rely on embedding tables that are too large to be stored on-device. In particular, recommendation models typically use multiple embedding tables each on the order of 1--10 GBs of data, making them impractical to store on-device. To overcome this barrier, we propose the use of private information retrieval (PIR) to efficiently and privately retrieve embeddings from servers without sharing any private information. As off-the-shelf PIR algorithms are usually too computationally intensive to directly use for latency-sensitive inference tasks, we 1) propose novel GPU-based acceleration of PIR, and 2) co-design PIR with the downstream ML application to obtain further speedup. Our GPU acceleration strategy improves system throughput by more than 20× over an optimized CPU PIR implementation, and our PIR-ML co-design provides an over 5× additional throughput improvement at fixed model quality. Together, for various on-device ML applications such as recommendation and language modeling, our system on a single V100 GPU can serve up to 100,000 queries per second---a > 100× throughput improvement over a CPU-based baseline---while maintaining model accuracy.
Maximilian Lam, Jeff Johnson 0004, Wenjie Xiong 0001, Kiwan Maeng, Udit Gupta 0001, Yang Li 0183, Liangzhen Lai, Ilias Leontiadis, Minsoo Rhu, Hsien-Hsin S. Lee, Vijay Janapa Reddi, Gu-Yeon Wei, David Brooks 0001, G. Edward Suh
ASPLOS (1)1
2023 GWAS quality score for evaluating associated regions in GWAS analyses
abstract
MOTIVATION: The number of significantly associated regions reported in genome-wide association studies (GWAS) for polygenic traits typically increases with sample size. A traditional tool for quality control and identification of significant regions has been a visual inspection of how significant and correlated genetic variants cluster within a region. However, while inspecting hundreds of regions, this subjective method can misattribute significance to some loci or neglect others that are significant. RESULTS: The GWAS quality score (GQS) identifies suspicious regions and prevents erroneous interpretations with an objective, quantitative and automated method. The GQS assesses all measured single nucleotide polymorphisms (SNPs) that are linked by inheritance to each other [linkage disequilibrium (LD)] and compares the significance of trait association of each SNP to its LD value for the reported index SNP. A GQS value of 1.0 ascribes a high level of confidence to the entire region and its underlying gene(s), while GQS values <1.0 indicate the need to closely inspect the outliers. We applied the GQS to published and non-published genome-wide summary statistics and report suspicious regions requiring secondary inspection while supporting the majority of reported regions from large-scale published meta-analyses. AVAILABILITY AND IMPLEMENTATION: The GQS code/scripts can be cloned from GitHub (https://github.com/Xswapnil/GQS/). The analyst can use whole-genome summary statistics to estimate GQS for each defined region. We also provide an online tool (http://35.227.18.38/) that gives access to the GQS. The quantitative measure of quality attributes by GQS and its visualization is an objective method that enhances the confidence of each genomic hit. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Swapnil Awasthi, Chia-Yen Chen, Maximilian Lam, Stephan Ripke, C. Anthony Altar
Bioinform.3
2021 Precision Batching: Bitserial Decomposition for Efficient Neural Network Inference on GPUs
abstract
We present PrecisionBatching, a quantized inference algorithm for speeding up neural network inference on traditional hardware platforms at low bitwidths. PrecisionBatching is based on the following insights: 1) neural network inference with low batch sizes on traditional hardware architectures (e.g: GPUs) is memory bound, 2) activation precision is critical to improving quantized model quality and 3) matrix-vector multiplication can be decomposed into binary matrix-matrix multiplications, enabling quantized inference with higher precision activations at the cost of more arithmetic operations. Combining these three insights, PrecisionBatching enables inference at extreme quantization levels (< 8 bits) by shifting a memory bound problem to a compute bound problem and achieves higher compute efficiency and runtime speedup at fixed accuracy thresholds against standard quantized inference methods. Across a variety of applications (MNIST, language modeling, natural language inference, reinforcement learning) and neural network architectures (fully connected, RNN, LSTM), PrecisionBatching yields end-to-end speedups of over 8× on a GPU within a < 1 - 5% error margin of the full precision baseline, outperforming traditional 8-bit quantized inference by over 1.5 × - 2× at the same error tolerance.
Maximilian Lam, Zachary Yedidia, Colby R. Banbury, Vijay Janapa Reddi
PACT1
2021 Gradient Disaggregation: Breaking Privacy in Federated Learning by Reconstructing the User Participant Matrix
abstract
We show that aggregated model updates in federated learning may be insecure. An untrusted central server may disaggregate user updates from sums of updates across participants given repeated observations, enabling the server to recover privileged information about individual users’ private training data via traditional gradient inference attacks. Our method revolves around reconstructing participant information (e.g: which rounds of training users participated in) from aggregated model updates by leveraging summary information from device analytics commonly used to monitor, debug, and manage federated learning systems. Our attack is parallelizable and we successfully disaggregate user updates on settings with up to thousands of participants. We quantitatively and qualitatively demonstrate significant improvements in the capability of various inference attacks on the disaggregated updates. Our attack enables the attribution of learned properties to individual users, violating anonymity, and shows that a determined central server may undermine the secure aggregation protocol to break individual users’ data privacy in federated learning.
Maximilian Lam, Gu-Yeon Wei, David Brooks 0001, Vijay Janapa Reddi, Michael Mitzenmacher
ICML1
2020 RICOPILI: Rapid Imputation for COnsortias PIpeLIne
abstract
SUMMARY: Genome-wide association study (GWAS) analyses, at sufficient sample sizes and power, have successfully revealed biological insights for several complex traits. RICOPILI, an open-sourced Perl-based pipeline was developed to address the challenges of rapidly processing large-scale multi-cohort GWAS studies including quality control (QC), imputation and downstream analyses. The pipeline is computationally efficient with portability to a wide range of high-performance computing environments. RICOPILI was created as the Psychiatric Genomics Consortium pipeline for GWAS and adopted by other users. The pipeline features (i) technical and genomic QC in case-control and trio cohorts, (ii) genome-wide phasing and imputation, (iv) association analysis, (v) meta-analysis, (vi) polygenic risk scoring and (vii) replication analysis. Notably, a major differentiator from other GWAS pipelines, RICOPILI leverages on automated parallelization and cluster job management approaches for rapid production of imputed genome-wide data. A comprehensive meta-analysis of simulated GWAS data has been incorporated demonstrating each step of the pipeline. This includes all the associated visualization plots, to allow ease of data interpretation and manuscript preparation. Simulated GWAS datasets are also packaged with the pipeline for user training tutorials and developer work. AVAILABILITY AND IMPLEMENTATION: RICOPILI has a flexible architecture to allow for ongoing development and incorporation of newer available algorithms and is adaptable to various HPC environments (QSUB, BSUB, SLURM and others). Specific links for genomic resources are either directly provided in this paper or via tutorials and external links. The central location hosting scripts and tutorials is found at this URL: https://sites.google.com/a/broadinstitute.org/RICOPILI/home. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Maximilian Lam, Swapnil Awasthi, Hunna J. Watson, Jackie Goldstein, Georgia Panagiotaropoulou, Vassily Trubetskoy, Robert Karlsson, Oleksander Frei, Chun-Chieh Fan, Ward De Witte, Nina R. Mota, Niamh Mullins, Kim Brügger, Sang Hong Lee, Naomi R. Wray, Nora Skarabis, Benjamin M. Neale, Mark J. Daly, Manuel Mattheisen, Raymond Walters, Stephan Ripke
Bioinform.1
2019 Cataloging the visible universe through Bayesian inference in Julia at petascale
Jeffrey Regier, Keno Fischer, Kiran Pamnany, Andreas Noack 0001, Jarrett Revels, Maximilian Lam, Steve Howard, Ryan Giordano, David Schlegel, Jon D. McAuliffe, Rollin C. Thomas, Prabhat
J. Parallel Distributed Comput.6
2018 Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
abstract
It has been experimentally observed that distributed implementations of mini-batch stochastic gradient descent (SGD) algorithms exhibit speedup saturation and decaying generalization ability beyond a particular batch-size. In this work, we present an analysis hinting that high similarity between concurrently processed gradients may be a cause of this performance degradation. We introduce the notion of gradient diversity that measures the dissimilarity between concurrent gradient updates, and show its key role in the convergence and generalization performance of mini-batch SGD. We also establish that heuristics similar to DropConnect, Langevin dynamics, and quantization, are provably diversity-inducing mechanisms, and provide experimental evidence indicating that these mechanisms can indeed enable the use of larger batches without sacrificing accuracy and lead to faster training in distributed learning. For example, in one of our experiments, for a convolutional neural network to reach 95% training accuracy on MNIST, using the diversity-inducing mechanism can reduce the training time by 30% in the distributed setting.
Ashwin Pananjady, Maximilian Lam, Dimitris S. Papailiopoulos, Kannan Ramchandran, Peter L. Bartlett
AISTATS3
2018 Cataloging the Visible Universe Through Bayesian Inference at Petascale
abstract
Astronomical catalogs derived from wide-field imaging surveys are an important tool for understanding the Universe. We construct an astronomical catalog from 55 TB of imaging data using Celeste, a Bayesian variational inference code written entirely in the high-productivity programming language Julia. Using over 1.3 million threads on 650,000 Intel Xeon Phi cores of the Cori Phase II supercomputer, Celeste achieves a peak rate of 1.54 DP PFLOP/s. Celeste is able to jointly optimize parameters for 188M stars and galaxies, loading and processing 178 TB across 8192 nodes in 14.6 minutes. To achieve this, Celeste exploits parallelism at multiple levels (cluster, node, and thread) and accelerates I/O through Cori's Burst Buffer. Julia's native performance enables Celeste to employ high-level constructs without resorting to hand-written or generated low-level code (C/C++/Fortran), and yet achieve petascale performance.
Jeffrey Regier, Kiran Pamnany, Keno Fischer, Andreas Noack 0001, Maximilian Lam, Jarrett Revels, Steve Howard, Ryan Giordano, David Schlegel, Jon D. McAuliffe, Rollin C. Thomas, Prabhat
IPDPS5
2018 Speeding Up Distributed Machine Learning Using Codes
abstract
Codes are widely used in many engineering applications to offerrobustnessagainstnoise. In large-scale systems, there are several types of noise that can affect the performance of distributed machine learning algorithms—straggler nodes, system failures, or communication bottlenecks—but there has been little interaction cutting across codes, machine learning, and distributed systems. In this paper, we provide theoretical insights on howcodedsolutions can achieve significant gains compared with uncoded ones. We focus on two of the most basic building blocks of distributed learning algorithms:matrix multiplicationanddata shuffling. For matrix multiplication, we use codes to alleviate the effect of stragglers and show that if the number of homogeneous workers is$n$, and the runtime of each subtask has an exponential tail, coded computation can speed up distributed matrix multiplication by a factor of$\log n$. For data shuffling, we use codes to reduce communication bottlenecks, exploiting the excess in storage. We show that when a constant fraction$\alpha $of the data matrix can be cached at each worker, and$n$is the number of workers,coded shufflingreduces the communication cost by a factor of$\left({\alpha + \frac {1}{n}}\right)\gamma (n)$compared with uncoded shuffling, where$\gamma (n)$is the ratio of the cost of unicasting$n$messages to$n$users to multicasting a common message (of the same size) to$n$users. For instance,$\gamma (n) \simeq n$if multicasting a message to$n$users is as cheap as unicasting a message to one user. We also provide experimental results, corroborating our theoretical gains of the coded algorithms.
Kangwook Lee 0001, Maximilian Lam, Ramtin Pedarsani, Dimitris S. Papailiopoulos, Kannan Ramchandran
IEEE Trans. Inf. Theory2
2016 Speeding up distributed machine learning using codes
abstract
Distributed machine learning algorithms that are widely run on modern large-scale computing platforms face several types of randomness, uncertainty and system “noise.” These include stragglers1, system failures, maintenance outages, and communication bottlenecks. In this work, we view distributed machine learning algorithms through a coding-theoretic lens, and show how codes can equip them with robustness against this system noise. Motivated by their importance and universality, we focus on two of the most basic building blocks of distributed learning algorithms: data shuffling and matrix multiplication. In data shuffling, we use codes to reduce communication bottlenecks: when a constant fraction of the data can be cached at each worker node, and n is the number of workers, coded shuffling reduces the communication cost by up to a factor Θ(n) over uncoded shuffling. For matrix multiplication, we use codes to alleviate the effects of stragglers, also known as the straggler problem. We show that if the number of workers is n, and the runtime of each subtask has an exponential tail, the optimal coded matrix multiplication is Θ(log n) times faster than the uncoded matrix multiplication or the optimal task replication scheme.
Kangwook Lee 0001, Maximilian Lam, Ramtin Pedarsani, Dimitris S. Papailiopoulos, Kannan Ramchandran
ISIT2