David Bruns-Smith

dblp:191/0533 · also David A. Bruns-Smith · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0000-9863-3882ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 3Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning theory · 27% Kernel, tree and ensemble methods · 27% Probabilistic and Bayesian machine learning · 23%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Reconfigurable computing and FPGAs · 50% Cloud and datacenter computing · 27% Hardware accelerators and domain-specific architectures · 23%
Theoretical computer science
1 paper
Mathematical optimization · 50% Algorithms and data structures · 50%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.912025
Ridge Boosting is Both Robust and Efficient · NeurIPS 2025
Machine learning › Trustworthy machine learning › robustness
distributional robustness
0.912025
Ridge Boosting is Both Robust and Efficient · NeurIPS 2025
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.912025
Ridge Boosting is Both Robust and Efficient · NeurIPS 2025
Machine learning › Learning theory › statistical estimation › semiparametric estimation
semiparametric efficiency
0.912025
Ridge Boosting is Both Robust and Efficient · NeurIPS 2025
Machine learning › Learning theory
statistical estimation
0.912025
Ridge Boosting is Both Robust and Efficient · NeurIPS 2025
Reconfigurable computing and FPGAs
FPGA accelerator
0.822020
Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020
FPGA Accelerated INDEL Realignment in the Cloud · HPCA 2019
Cloud and datacenter computing › cloud deployment
cloud FPGA deployment
0.522020
FPGA Accelerated INDEL Realignment in the Cloud · HPCA 2019
Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.512021
Model-Free and Model-Based Policy Evaluation when Causality is Uncertain · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal sensitivity analysis
0.512021
Model-Free and Model-Based Policy Evaluation when Causality is Uncertain · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › causal inference
latent confounders
0.512021
Model-Free and Model-Based Policy Evaluation when Causality is Uncertain · ICML 2021
Machine learning › Reinforcement learning
off-policy evaluation
0.512021
Model-Free and Model-Based Policy Evaluation when Causality is Uncertain · ICML 2021
Hardware accelerators and domain-specific architectures › spatial architecture
dataflow accelerator
0.412020
Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020
Mathematical optimization › continuous optimization
convex optimization
0.312025
Ridge Boosting is Both Robust and Efficient · NeurIPS 2025
Algorithms and data structures › kernel methods
kernel ridge regression
0.312025
Ridge Boosting is Both Robust and Efficient · NeurIPS 2025
Machine learning › Reinforcement learning › robust reinforcement learning
robust markov decision process
0.112021
Model-Free and Model-Based Policy Evaluation when Causality is Uncertain · ICML 2021
Bioinformatics and computational biology › genomics
genomic data analysis
0.112020
Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020
Reconfigurable computing and FPGAs › cloud FPGA
FPGA-as-a-service
0.112020
Genesis: A Hardware Acceleration Framework for Genomic Data Analysis · ISCA 2020
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis
0.112019
FPGA Accelerated INDEL Realignment in the Cloud · HPCA 2019

Methods — techniques the papers use, named apart from their topics

semiparametric efficiency theory · 1.7kernel ridge regression · 1.7boosting · 1.7on-chip scratchpad · 0.9non-blocking APIs · 0.9extended SQL · 0.9dataflow architecture · 0.9hardware-software co-design · 0.8FPGA-as-a-service · 0.8worst-case bounds · 0.5robust MDPs · 0.5
YearPublicationVenuePosition
2025 Ridge Boosting is Both Robust and Efficient
abstract
Estimators in statistics and machine learning must typically trade off between efficiency, having low variance for a fixed target, and distributional robustness, such as \textit{multiaccuracy}, or having low bias over a range of possible targets. In this paper, we consider a simple estimator, \emph{ridge boosting}: starting with any initial predictor, perform a single boosting step with (kernel) ridge regression. Surprisingly, we show that ridge boosting simultaneously achieves both efficiency and distributional robustness: for target distribution shifts that lie within an RKHS unit ball, this estimator maintains low bias across all such shifts and has variance at the semiparametric efficiency bound for each target. In addition to bridging otherwise distinct research areas, this result has immediate practical value. Since ridge boosting uses only data from the source distribution, researchers can train a single model to obtain both robust and efficient estimates for multiple target estimands at the same time, eliminating the need to fit separate semiparametric efficient estimators for each target. We assess this approach through simulations and an application estimating the age profile of retirement income.
David Bruns-Smith, Zhongming Xie, Avi Feller
NeurIPS1
2022 Outcome Assumptions and Duality Theory for Balancing Weights
abstract
We study balancing weight estimators, which reweight outcomes from a source population to estimate missing outcomes in a target population. These estimators minimize the worst-case error by making an assumption about the outcome model. In this paper, we show that this outcome assumption has two immediate implications. First, we can replace the minimax optimization problem for balancing weights with a simple convex loss over the assumed outcome function class. Second, we can replace the commonly-made overlap assumption with a more appropriate quantitative measure, the minimum worst-case bias. Finally, we show conditions under which the weights remain robust when our assumptions on the outcomes are wrong.
David Bruns-Smith, Avi Feller
AISTATS1
2021 Model-Free and Model-Based Policy Evaluation when Causality is Uncertain
abstract
When decision-makers can directly intervene, policy evaluation algorithms give valid causal estimates. In off-policy evaluation (OPE), there may exist unobserved variables that both impact the dynamics and are used by the unknown behavior policy. These “confounders” will introduce spurious correlations and naive estimates for a new policy will be biased. We develop worst-case bounds to assess sensitivity to these unobserved confounders in finite horizons when confounders are drawn iid each period. We demonstrate that a model-based approach with robust MDPs gives sharper lower bounds by exploiting domain knowledge about the dynamics. Finally, we show that when unobserved confounders are persistent over time, OPE is far more difficult and existing techniques produce extremely conservative bounds.
David Bruns-Smith
ICML1
2020 Genesis: A Hardware Acceleration Framework for Genomic Data Analysis
abstract
In this paper, we describe our vision to accelerate algorithms in the domain of genomic data analysis by proposing a framework called Genesis (genome analysis) that contains an interface and an implementation of a system that processes genomic data efficiently. This framework can be deployed in the cloud and exploit the FPGAs-as-a-service paradigm to provide cost-efficient secondary DNA analysis. We propose conceptualizing genomic reads and associated read attributes as a very large relational database and using extended SQL as a domain-specific language to construct queries that form various data manipulation operations. To accelerate such queries, we design a Genesis hardware library which consists of primitive hardware modules that can be composed to construct a dataflow architecture specialized for those queries. As a proof of concept for the Genesis framework, we present the architecture and the hardware implementation of several genomic analysis stages in the secondary analysis pipeline corresponding to the best known software analysis toolkit, GATK4 workflow proposed by the Broad Institute. We walk through the construction of genomic data analysis operations using a sequence of SQL-style queries and show how Genesis hardware library modules can be utilized to construct the hardware pipelines designed to accelerate such queries. We exploit parallelism and data reuse by utilizing a dataflow architecture along with the use of on-chip scratchpads as well as non-blocking APIs to manage the accelerators, allowing concurrent execution of the accelerator and the host. Our accelerated system deployed on the cloud FPGA performs up to 19.3× better than GATK4 running on a commodity multi-core Xeon server and obtains up to 15× better cost savings. We believe that if a software algorithm can be mapped onto a hardware library to utilize the underlying accelerator(s) using an already-standardized software interface such as SQL, while allowing the efficient mapping of such interface to primitive hardware modules as we have demonstrated here, it will expedite the acceleration of domainspecific algorithms and allow the easy adaptation of algorithm changes.
Tae Jun Ham, David Bruns-Smith, Brendan Sweeney, Yejin Lee 0001, Seong Hoon Seo, U. Gyeong Song, Young H. Oh, Krste Asanovic, Jae W. Lee, Lisa Wu Wills
ISCA2
2019 FPGA Accelerated INDEL Realignment in the Cloud
abstract
The amount of data being generated in genomics is predicted to be between 2 and 40 exabytes per year for the next decade, making genomic analysis the new frontier and the new challenge for precision medicine. This paper explores targeted deployment of hardware accelerators in the cloud to improve the runtime and throughput of immensescale genomic data analyses. In particular, INDEL (INsertion/DELetion) realignment is a critical operation that enables diagnostic testings of cancer through error correction prior to variant calling. It is the slowest part of the somatic (cancer) genomic analysis pipeline, the alignment refinement pipeline, and represents roughly one-third of the execution time of timesensitive diagnostics for acute cancer patients. To accelerate genomic analysis, this paper describes a hardware accelerator for INDEL realignment (IR), and a hardware-software framework leveraging FPGAs-as-a-service in the cloud. We chose to implement genomics analytics on FPGAs because genomic algorithms are still rapidly evolving (e.g. the de facto standard “GATK Best Practices” has had five releases since January of this year). We chose to deploy genomics accelerators in the cloud to reduce capital expenditure and to provide a more quantitative performance and cost analysis. We built and deployed a sea of IR accelerators using our hardware-software accelerator development framework on AWS EC2 F1 instances. We show that our IR accelerator system performed 81× better than multi-threaded genomic analysis software while being 32× more cost efficient.
Lisa Wu Wills, David Bruns-Smith, Frank A. Nothaft, Qijing Huang 0001, Sagar Karandikar, Johnny Le, Andrew Lin, Howard Mao, Brendan Sweeney, Krste Asanovic, David A. Patterson 0001, Anthony D. Joseph
HPCA2
2019 Enhancing Network Visibility and Security through Tensor Analysis
Muthu Manikandan Baskaran, Thomas Henretty, James R. Ezick, Richard A. Lethin, David Bruns-Smith
Future Gener. Comput. Syst.5