VLDB 2026 Research / reviewers in the wild / expert
Bartolomeo Stellato
dblp:192/2914
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0003-4684-7111ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Optimization for machine learning · 46% Learning theory · 32% Reinforcement learning · 15% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 44% Hardware accelerators and domain-specific architectures · 33% Reconfigurable computing and FPGAs · 22% | |
| Theoretical computer science
4 papers |
Mathematical optimization · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning
learned optimizer |
1.6 | 2 | 2025 | Data-Driven Performance Guarantees for Classical and Learned Optimizers · J. Mach. Learn. Res. 2025 Learning to Warm-Start Fixed-Point Optimization Algorithms · J. Mach. Learn. Res. 2024 |
Machine learning › Learning theory
generalization bounds |
1.1 | 2 | 2025 | Data-Driven Performance Guarantees for Classical and Learned Optimizers · J. Mach. Learn. Res. 2025 Learning to Warm-Start Fixed-Point Optimization Algorithms · J. Mach. Learn. Res. 2024 |
Machine learning › Learning theory › generalization bounds
PAC-Bayes bounds |
1.1 | 2 | 2025 | Data-Driven Performance Guarantees for Classical and Learned Optimizers · J. Mach. Learn. Res. 2025 Learning to Warm-Start Fixed-Point Optimization Algorithms · J. Mach. Learn. Res. 2024 |
Program synthesis and code generation
code generation with language models |
0.9 | 1 | 2025 | AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs? · NeurIPS 2025 |
High-performance computing › performance optimization
numerical program optimization |
0.9 | 1 | 2025 | AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs? · NeurIPS 2025 |
High-performance computing
performance optimization |
0.9 | 1 | 2025 | AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs? · NeurIPS 2025 |
Mathematical optimization › continuous optimization › nonlinear optimization › quadratic programming
convex quadratic programming |
0.8 | 1 | 2024 | Multi-Issue Butterfly Architecture for Sparse Convex Quadratic Programming · MICRO 2024 |
Mathematical optimization › optimization
warm-start optimization |
0.8 | 1 | 2024 | Learning to Warm-Start Fixed-Point Optimization Algorithms · J. Mach. Learn. Res. 2024 |
Hardware accelerators and domain-specific architectures › domain-specific accelerator
optimization accelerator |
0.7 | 1 | 2023 | RSQP: Problem-specific Architectural Customization for Accelerated Convex Quadratic Optimization · ISCA 2023 |
Hardware accelerators and domain-specific architectures
quadratic programming solver |
0.7 | 1 | 2023 | RSQP: Problem-specific Architectural Customization for Accelerated Convex Quadratic Optimization · ISCA 2023 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable accelerator |
0.7 | 1 | 2023 | RSQP: Problem-specific Architectural Customization for Accelerated Convex Quadratic Optimization · ISCA 2023 |
Robotics › Motion planning and robot control › robot control › optimal control
bang-bang control |
0.5 | 1 | 2021 | Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies · NeurIPS 2021 |
Machine learning › Reinforcement learning
continuous control |
0.5 | 1 | 2021 | Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies · NeurIPS 2021 |
Machine learning › Optimization for machine learning
gradient-based optimization |
0.5 | 1 | 2021 | Accelerating Quadratic Optimization with Reinforcement Learning · NeurIPS 2021 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.5 | 1 | 2021 | Accelerating Quadratic Optimization with Reinforcement Learning · NeurIPS 2021 |
Machine learning › Reinforcement learning › policy learning
policy parameterization |
0.5 | 1 | 2021 | Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies · NeurIPS 2021 |
Machine learning › Optimization for machine learning › convex optimization
quadratic programming |
0.5 | 1 | 2021 | Accelerating Quadratic Optimization with Reinforcement Learning · NeurIPS 2021 |
Mathematical optimization
continuous optimization |
0.3 | 1 | 2025 | Data-Driven Performance Guarantees for Classical and Learned Optimizers · J. Mach. Learn. Res. 2025 |
Reconfigurable computing and FPGAs
FPGA prototyping |
0.2 | 1 | 2024 | Multi-Issue Butterfly Architecture for Sparse Convex Quadratic Programming · MICRO 2024 |
Mathematical optimization › continuous optimization
convex optimization |
0.2 | 1 | 2023 | RSQP: Problem-specific Architectural Customization for Accelerated Convex Quadratic Optimization · ISCA 2023 |
Mathematical optimization › continuous optimization › nonlinear optimization
quadratic programming |
0.2 | 1 | 2023 | RSQP: Problem-specific Architectural Customization for Accelerated Convex Quadratic Optimization · ISCA 2023 |
Methods — techniques the papers use, named apart from their topics
statistical learning theory · 1.7sample convergence bound · 1.7language model agent · 1.7benchmarking · 1.7pipelined spatial architecture · 1.5neural network · 1.5fixed-point iteration · 1.5elimination tree · 1.5ADMM · 1.5mixed integer linear programming · 1.3lossless string compression · 1.3reinforcement learning · 0.5imitation learning · 0.5OSQP · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?abstractDespite progress in language model (LM) capabilities, evaluations have thus far focused on models' performance on tasks that humans have previously solved, including in programming (SWE-Bench) and mathematics (FrontierMath). We therefore propose testing models' ability to design and implement algorithms in an open-ended benchmark: We task LMs with writing code that efficiently solves computationally challenging problems in computer science, physics, and mathematics. Our AlgoTune benchmark consists of 120 tasks collected from domain experts and a framework for validating and timing LM-synthesized solution code, which is compared to reference implementations from popular open-source packages.In addition, we develop a baseline LM agent, AlgoTuner, and evaluate its performance across a suite of frontier models.AlgoTuner achieves an average 1.58x speedup against reference solvers, including methods from packages such as SciPy, scikit-learn and CVXPY.However, we find that current models fail to discover algorithmic innovations, instead preferring surface-level optimizations. We hope that AlgoTune catalyzes the development of LM agents exhibiting creative problem solving beyond state-of-the-art human performance. Ori Press, Brandon Amos, Yikai Wu 0001, Samuel K. Ainsworth, Dominik Krupke, Patrick Kidger, Touqir Sajed, Bartolomeo Stellato, Jisun Park 0003, Nathanael Bosch, Eli Meril, Albert Steppi, Arman Zharmagambetov, Fangzhao Zhang, David Pérez-Piñeiro, Alberto Mercurio, Ni Zhan 0002, Talor Abramovich, Kilian Lieret, Shirley Huang, Matthias Bethge, Ofir Press |
NeurIPS | 9 |
| 2025 | Data-Driven Performance Guarantees for Classical and Learned OptimizersabstractWe introduce a data-driven approach to analyze the performance of continuous optimization algorithms using generalization guarantees from statistical learning theory. We study classical and learned optimizers to solve families of parametric optimization problems. We build generalization guarantees for classical optimizers, using a sample convergence bound, and for learned optimizers, using the Probably Approximately Correct (PAC)-Bayes framework. To train learned optimizers, we use a gradient-based algorithm to directly minimize the PAC-Bayes upper bound. Numerical experiments in signal processing, control, and meta-learning showcase the ability of our framework to provide strong generalization guarantees for both classical and learned optimizers given a fixed budget of iterations. For classical optimizers, our bounds which hold with high probability are much tighter than those that worst-case guarantees provide. For learned optimizers, our bounds outperform the empirical outcomes observed in their non-learned counterparts. Rajiv Sambharya, Bartolomeo Stellato |
J. Mach. Learn. Res. | 2 |
| 2024 | Multi-Issue Butterfly Architecture for Sparse Convex Quadratic ProgrammingabstractConvex quadratic optimization solvers are extensively utilized in various domains; however, achieving optimal performance in diverse situations remains a significant challenge due to the sparse nature of objective and constraint matrices. General-purpose architectures struggle with hardware utilization when performing critical sparse matrix operations, such as factorization and multiplication. To address this issue, we introduce a pipelined spatial architecture, Multi-Issue Butterfly (MIB), which supports all primitive scalar, vector, and matrix operations required by the Alternating Direction Method of Multipliers (ADMM) based solver algorithm. The proposed architecture features a butterfly computational network with innovative working modes for each node, controlled by runtime instructions. We developed a companion scheduling method for matrix operations based on their sparsity patterns. For factorization, an elimination tree guides the network instructions reordering to avoid data hazards caused by computation dependencies. For matrix-vector multiplication, data prefetching resolves structural hazards caused by read and write conflicts to register files. Instructions without hazards are issued simultaneously to increase pipeline throughput and function unit utilization. We evaluate the proposed architecture using FPGA prototypes, representing the first fully FPGA-based generic QP solver. Our assessment includes extensive performance and efficiency bench-marks across 100 QP problems from five application domains. Compared to the same algorithm variation running on CPU backends, our prototype achieves a geometric mean of$30.5\times$end-to-end speedup,$127.0 \times$greater energy efficiency, and$16.5\times$less runtime jitter. In comparison to GPU backends, the prototype attains a geometric mean of$4.3\times$faster end-to-end speedup,$21.7\times$higher energy efficiency, and$33.4\times$less runtime jitter. Maolin Wang 0002, Ian McInerney, Bartolomeo Stellato, Fengbin Tu, Stephen P. Boyd, Hayden Kwok-Hay So, Kwang-Ting Cheng |
MICRO | 3 |
| 2024 | Learning to Warm-Start Fixed-Point Optimization AlgorithmsabstractWe introduce a machine-learning framework to warm-start fixed-point optimization algorithms. Our architecture consists of a neural network mapping problem parameters to warm starts, followed by a predefined number of fixed-point iterations. We propose two loss functions designed to either minimize the fixed-point residual or the distance to a ground truth solution. In this way, the neural network predicts warm starts with the end-to-end goal of minimizing the downstream loss. An important feature of our architecture is its flexibility, in that it can predict a warm start for fixed-point algorithms run for any number of steps, without being limited to the number of steps it has been trained on. We provide PAC-Bayes generalization bounds on unseen data for common classes of fixed-point operators: contractive, linearly convergent, and averaged. Applying this framework to well-known applications in control, statistics, and signal processing, we observe a significant reduction in the number of iterations and solution time required to solve these problems, through learned warm starts. Rajiv Sambharya, Georgina Hall, Brandon Amos, Bartolomeo Stellato |
J. Mach. Learn. Res. | 4 |
| 2023 | RSQP: Problem-specific Architectural Customization for Accelerated Convex Quadratic OptimizationabstractConvex optimization is at the heart of many performance-critical applications across a wide range of domains. Although many high-performance hardware accelerators have been developed for specific optimization problems in the past, designing such accelerator is a challenging task and the resulting computing architecture is often so specific to the targeted application that they can hardly be reused even in a related application within the same domain. To accelerate general-purpose optimization solvers that must operate on diverse user input during run time, an ideal hardware solver should be able to adapt to the provided optimization problem dynamically while achieving high performance and power-efficiency. In this work, a hardware-accelerated general-purpose quadratic program solver, called RSQP, with reconfigurable functional units and data path that facilitate problem-specific customization is presented. RSQP uses a string-based encoding to describe the problem structure with fine granularity. Based on this encoding, functional units and datapath customized to the sparsity pattern of the problem are created by solving a dictionary-based lossless string compression problem and a mixed integer linear program respectively. RSQP has been integrated to accelerate the general-purpose quadratic programming solver OSQP and has been tested using an extensive benchmark with 120 optimization problems from 6 application domains. Through architectural customization, RSQP achieves up to 7× performance improvement over its baseline generic design. Furthermore, when compared with a CPU and a GPU-accelerated implementation, RSQP achieves up to 31.2× and 6.9× end-to-end speedup on these benchmark programs, respectively. Finally, the FPGA accelerator operates at up to 6.6× lower dynamic power consumption and up to 22.7× higher power efficiency over the GPU implementation, making it an attractive solution for power-conscious datacenter applications. Maolin Wang 0002, Ian McInerney, Bartolomeo Stellato, Stephen P. Boyd, Hayden Kwok-Hay So |
ISCA | 3 |
| 2022 | Online Mixed-Integer Optimization in MillisecondsabstractWe propose a method to approximate the solution of online mixed-integer optimization (MIO) problems at very high speed using machine learning. By exploiting the repetitive nature of online optimization, we can greatly speed up the solution time. Our approach encodes the optimal solution into a small amount of information denoted as strategy using the voice of optimization framework. In this way, the core part of the optimization routine becomes a multiclass classification problem that can be solved very quickly. In this work, we extend that framework to real-time and high-speed applications focusing on parametric mixed-integer quadratic optimization. We propose an extremely fast online optimization method consisting of a feedforward neural network evaluation and a linear system solution where the matrix has already been factorized. Therefore, this online approach does not require any solver or iterative algorithm. We show the speed of the proposed method both in terms of total computations required and measured execution time. We estimate the number of floating point operations required to completely recover the optimal solution as a function of the problem dimensions. Compared with state-of-the-art MIO routines, the online running time of our method is very predictable and can be lower than a single matrix factorization time. We benchmark our method against the state-of-the-art solver Gurobi obtaining up to two to three orders of magnitude speedups on examples from fuel cell energy management, sparse portfolio optimization, and motion planning with obstacle avoidance. Summary of Contribution: We propose a technique to approximate the solution of online optimization problems at high speed using machine learning. By exploiting the repetitive nature of online optimization, we learn the mapping between the key problem parameters and an encoding of the optimal solution to greatly speed up the solution time. This allows us to significantly improve the computation time and resources needed to solve online mixed-integer optimization problems. We obtain a simple method with a very low computing time variance, which is crucial in online settings. Dimitris Bertsimas, Bartolomeo Stellato |
INFORMS J. Comput. | 2 |
| 2021 | Accelerating Quadratic Optimization with Reinforcement LearningabstractFirst-order methods for quadratic optimization such as OSQP are widely used for large-scale machine learning and embedded optimal control, where many related problems must be rapidly solved. These methods face two persistent challenges: manual hyperparameter tuning and convergence time to high-accuracy solutions. To address these, we explore how Reinforcement Learning (RL) can learn a policy to tune parameters to accelerate convergence. In experiments with well-known QP benchmarks we find that our RL policy, RLQP, significantly outperforms state-of-the-art QP solvers by up to 3x. RLQP generalizes surprisingly well to previously unseen problems with varying dimension and structure from different applications, including the QPLIB, Netlib LP and Maros-M{\'e}sz{\'a}ros problems. Code, models, and videos are available at https://berkeleyautomation.github.io/rlqp/. Jeffrey Ichnowski, Paras Jain 0001, Bartolomeo Stellato, Goran Banjac, Michael Luo, Francesco Borrelli, Joseph Gonzalez 0001, Ion Stoica, Kenneth Y. Goldberg |
NeurIPS | 3 |
| 2021 | Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli PoliciesabstractReinforcement learning (RL) for continuous control typically employs distributions whose support covers the entire action space. In this work, we investigate the colloquially known phenomenon that trained agents often prefer actions at the boundaries of that space. We draw theoretical connections to the emergence of bang-bang behavior in optimal control, and provide extensive empirical evaluation across a variety of recent RL algorithms. We replace the normal Gaussian by a Bernoulli distribution that solely considers the extremes along each action dimension - a bang-bang controller. Surprisingly, this achieves state-of-the-art performance on several continuous control benchmarks - in contrast to robotic hardware, where energy and maintenance cost affect controller choices. Since exploration, learning, and the final solution are entangled in RL, we provide additional imitation learning experiments to reduce the impact of exploration on our analysis. Finally, we show that our observations generalize to environments that aim to model real-world challenges and evaluate factors to mitigate the emergence of bang-bang solutions. Our findings emphasise challenges for benchmarking continuous control algorithms, particularly in light of potential real-world applications. Tim Seyde, Igor Gilitschenski, Wilko Schwarting, Bartolomeo Stellato, Martin A. Riedmiller, Markus Wulfmeier, Daniela Rus |
NeurIPS | 4 |
| 2021 | The voice of optimization
Dimitris Bertsimas, Bartolomeo Stellato |
Mach. Learn. | 2 |
| 2021 | Machine Learning for Real-Time Heart Disease PredictionabstractHeart-related anomalies are among the most common causes of death worldwide. Patients are often asymptomatic until a fatal event happens, and even when they are under observation, trained personnel is needed in order to identify a heart anomaly. In the last decades, there has been increasing evidence of how Machine Learning can be leveraged to detect such anomalies, thanks to the availability of Electrocardiograms (ECG) in digital format. New developments in technology have allowed to exploit such data to build models able to analyze the patterns in the occurrence of heart beats, and spot anomalies from them. In this work, we propose a novel methodology to extract ECG-related features and predict the type of ECG recorded in real time (less than 30 milliseconds). Our models leverage a collection of almost 40 thousand ECGs labeled by expert cardiologists across different hospitals and countries, and are able to detect 7 types of signals: Normal, AF, Tachycardia, Bradycardia, Arrhythmia, Other or Noisy. We exploit the XGBoost algorithm, a leading machine learning method, to train models achieving out of sample F1 Scores in the range 0.93 - 0.99. To our knowledge, this is the first work reporting high performance across hospitals, countries and recording standards. Dimitris Bertsimas, Luca Mingardi, Bartolomeo Stellato |
IEEE J. Biomed. Health Informatics | 3 |