Michael Luo

dblp:152/0092 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Generative modeling · 39% Optimization for machine learning · 16% Reinforcement learning · 11%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 84% Distributed systems · 8% Parallel and multicore computing · 8%
Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 70% Information retrieval · 30%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
2.432026
Agentix: An Efficient Serving Engine for LLM Agents as General Programs · NSDI 2026
Starburst: A Cost-aware Scheduler for Hybrid Cloud · USENIX ATC 2024
SkyPilot: An Intercloud Broker for Sky Computing · NSDI 2023
Cloud and datacenter computing
serverless computing
1.012026
Agentix: An Efficient Serving Engine for LLM Agents as General Programs · NSDI 2026
Machine learning › Generative modeling › generative model
diverse generation
0.912025
SimpleStrat: Diversifying Language Model Generation with Stratification · NeurIPS 2025
Machine learning › Generative modeling
video generation
0.912025
WorldModelBench: Judging Video Generation Models As World Models · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.812024
Stylus: Automatic Adapter Selection for Diffusion Models · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Stylus: Automatic Adapter Selection for Diffusion Models · NeurIPS 2024
Information retrieval
retrieval models
0.812024
Stylus: Automatic Adapter Selection for Diffusion Models · NeurIPS 2024
Cloud and datacenter computing › cluster resource management and scheduling
hybrid cloud scheduling
0.812024
Starburst: A Cost-aware Scheduler for Hybrid Cloud · USENIX ATC 2024
Query processing and optimization › query optimization
join ordering
0.612022
Balsa: Learning a Query Optimizer Without Expert Demonstrations · SIGMOD Conference 2022
Query processing and optimization › query optimization › learned query optimization
learned query optimizer
0.612022
Balsa: Learning a Query Optimizer Without Expert Demonstrations · SIGMOD Conference 2022
Query processing and optimization
query optimization
0.612022
Balsa: Learning a Query Optimizer Without Expert Demonstrations · SIGMOD Conference 2022
Robotics › Legged, aerial and field robots › field robotics
agricultural robotics
0.512021
Learning Seed Placements and Automation Policies for Polyculture Farming with Companion Plants · ICRA 2021
Machine learning › Generative modeling
autoregressive model
0.512021
Discovering Non-monotonic Autoregressive Orderings with Variational Inference · ICLR 2021
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning
0.512021
RLlib Flow: Distributed Reinforcement Learning is a Dataflow Problem · NeurIPS 2021
Machine learning › Optimization for machine learning
gradient-based optimization
0.512021
Accelerating Quadratic Optimization with Reinforcement Learning · NeurIPS 2021
Machine learning › Optimization for machine learning
hyperparameter optimization
0.512021
Accelerating Quadratic Optimization with Reinforcement Learning · NeurIPS 2021
Machine learning › Efficient and distributed learning › model compression
pruning
0.512021
Learning Seed Placements and Automation Policies for Polyculture Farming with Companion Plants · ICRA 2021
Machine learning › Optimization for machine learning › convex optimization
quadratic programming
0.512021
Accelerating Quadratic Optimization with Reinforcement Learning · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.512021
Discovering Non-monotonic Autoregressive Orderings with Variational Inference · ICLR 2021
Parallel and multicore computing › parallel computation models
distributed dataflow
0.512021
RLlib Flow: Distributed Reinforcement Learning is a Dataflow Problem · NeurIPS 2021
Distributed systems › distributed machine learning
distributed reinforcement learning
0.512021
RLlib Flow: Distributed Reinforcement Learning is a Dataflow Problem · NeurIPS 2021
Machine learning › Reinforcement learning › large-scale reinforcement learning › distributed reinforcement learning
asynchronous reinforcement learning
0.412020
IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target Networks · ICLR 2020
Machine learning › Transfer learning and domain adaptation › instance weighting
importance weighting
0.412020
IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target Networks · ICLR 2020
Natural language and speech › Language models and text generation
instruction following
0.312025
WorldModelBench: Judging Video Generation Models As World Models · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning
sampling
0.312025
SimpleStrat: Diversifying Language Model Generation with Stratification · NeurIPS 2025
Cloud and datacenter computing › job scheduling › economic scheduling
cost-aware scheduling
0.212024
Starburst: A Cost-aware Scheduler for Hybrid Cloud · USENIX ATC 2024
Machine learning › Reinforcement learning
policy learning
0.112021
Learning Seed Placements and Automation Policies for Polyculture Farming with Companion Plants · ICRA 2021

Methods — techniques the papers use, named apart from their topics

prompt keyword matching · 1.5embedding · 1.5program compilation · 1.0large language model · 1.0actor model · 1.0judger fine-tuning · 0.9human preference annotation · 0.9simulation-based learning · 0.6deep reinforcement learning · 0.6variational inference · 0.5simulation · 0.5reinforcement learning · 0.5learned policy · 0.5dataflow programming · 0.5OSQP · 0.5
YearPublicationVenuePosition
2026 Agentix: An Efficient Serving Engine for LLM Agents as General Programs
Michael Luo, Xiaoxiang Shi, Colin Cai, Tianjun Zhang, Justin Wong, Yanping Huang, Joseph Gonzalez 0001, Ion Stoica
NSDI1
2025 WorldModelBench: Judging Video Generation Models As World Models
abstract
Video generation models have rapidly progressed, positioning themselves as video world models capable of supporting decision-making applications like robotics and autonomous driving. However, current benchmarks fail to rigorously evaluate these claims, focusing only on general video quality, ignoring important factors to world models such as physics adherence.To bridge this gap, we propose WorldModelBench, a benchmark designed to evaluate the world modeling capabilities of video generation models in application-driven domains. WorldModelBench offers two key advantages: (1) Against to nuanced world modeling violations: By incorporating instruction-following and physics-adherence dimensions, WorldModelBench detects subtle violations, such as irregular changes in object size that breach the mass conservation law—issues overlooked by prior benchmarks. (2) Aligned with large-scale human preferences: We crowd-source 67K human labels to accurately measure 14 frontier models. Using our high-quality human labels, we further fine-tune an accurate judger to automate the evaluation procedure, achieving 9.9% lower error in predicting world modeling violations than GPT-4o with 2B parameters. In addition, we demonstrate that training to align human annotations by maximizing the rewards from the judger noticeably improve the world modeling capability. The dataset is hosted in HuggingFace at https://huggingface.co/datasets/Efficient-Large-Model/worldmodelbench. The code to run evaluation is available at https://github.com/WorldModelBench-Team/WorldModelBench.
Dacheng Li, Yunhao Fang, Yukang Chen, Shuo Yang 0011, Shiyi Cao, Justin Wong, Michael Luo, Xiaolong Wang 0004, Hongxu Yin, Joseph Gonzalez 0001, Ion Stoica, Song Han 0003, Yao Lu 0006
NeurIPS7
2025 SimpleStrat: Diversifying Language Model Generation with Stratification
abstract
Generating diverse responses from large language models (LLMs) is crucial for applications such as adversarial testing, search, and synthetic data generation, where diversity provides distinct answers across generations. Previous approaches rely solely on increasing the temperature, sacrificing quality. Furthermore, the model's next-token probabilities may not be representative of the true answer distribution. To combat these challenges, we propose SimpleStrat, an alternative that uses the language model itself to partition the solution space into strata from which to sample. To measure resampling diversity, we introduce CoverageQA, a dataset of underspecified questions with multiple equally plausible answers. We propose measuring resampling diversity as the KL Divergence between the response distribution and the uniform distribution over valid ground truth answers and use recall as an alternative when assessing proprietary models. On CoverageQA, SimpleStrat improves diversity across all temperatures, showing orthogonal benefits. Quantifiably, we achieve as much as 4X better recall when applied to GPT-4o, and an average reduction in KL divergence by 0.36 when applied to Llama 3. Furthermore, we show that SimpleStrat achieves more resampling diversity at temperature T=0 than scaling temperature to T=1 on creative writing, an open-ended domain. Implementation and dataset available at https://github.com/jwong8314/simplestrat.
Justin Wong, Yury Orlovskiy, Alexander Shypula, Michael Luo, Sanjit A. Seshia, Joseph Gonzalez 0001
NeurIPS4
2024 Stylus: Automatic Adapter Selection for Diffusion Models
abstract
Beyond scaling base models with more data or parameters, fine-tuned adapters provide an alternative way to generate high fidelity, custom images at reduced costs. As such, adapters have been widely adopted by open-source communities, accumulating a database of over 100K adapters—most of which are highly customized with insufficient descriptions. To generate high quality images, this paper explores the problem of matching the prompt to a Stylus of relevant adapters, built on recent work that highlight the performance gains of composing adapters. We introduce Stylus, which efficiently selects and automatically composes task-specific adapters based on a prompt's keywords. Stylus outlines a three-stage approach that first summarizes adapters with improved descriptions and embeddings, retrieves relevant adapters, and then further assembles adapters based on prompts' keywords by checking how well they fit the prompt. To evaluate Stylus, we developed StylusDocs, a curated dataset featuring 75K adapters with pre-computed adapter embeddings. In our evaluation on popular Stable Diffusion checkpoints, Stylus achieves greater CLIP/FID Pareto efficiency and is twice as preferred, with humans and multimodal models as evaluators, over the base model.
Michael Luo, Justin Wong, Brandon Trabucco, Yanping Huang, Joseph Gonzalez 0001, Ruslan Salakhutdinov, Ion Stoica
NeurIPS1
2024 Starburst: A Cost-aware Scheduler for Hybrid Cloud
Michael Luo, Siyuan Zhuang, Suryaprakash Vengadesan, Romil Bhardwaj, Eric J. Friedman, Scott Shenker, Ion Stoica
USENIX ATC1
2023 SkyPilot: An Intercloud Broker for Sky Computing
Zongheng Yang, Zhanghao Wu, Michael Luo, Wei-Lin Chiang, Romil Bhardwaj, Woosuk Kwon, Siyuan Zhuang, Sifei Luan 0001, Gautam Mittal, Scott Shenker, Ion Stoica
NSDI3
2022 Balsa: Learning a Query Optimizer Without Expert Demonstrations
abstract
Query optimizers are a performance-critical component in every database system. Due to their complexity, optimizers take experts months to write and years to refine. In this work, we demonstrate for the first time that learning to optimize queries without learning from an expert optimizer is both possible and efficient. We present Balsa, a query optimizer built by deep reinforcement learning. Balsa first learns basic knowledge from a simple, environment-agnostic simulator, followed by safe learning in real execution. On the Join Order Benchmark, Balsa matches the performance of two expert query optimizers, both open-source and commercial, with two hours of learning, and outperforms them by up to 2.8× in workload runtime after a few more hours. Balsa thus opens the possibility of automatically learning to optimize in future compute environments where expert-designed optimizers do not exist.
Zongheng Yang, Wei-Lin Chiang, Sifei Luan 0001, Gautam Mittal, Michael Luo, Ion Stoica
SIGMOD Conference5
2021 Discovering Non-monotonic Autoregressive Orderings with Variational Inference
Brandon Trabucco, Dong Huk Park, Michael Luo, Sheng Shen 0001, Trevor Darrell, Yang Gao 0029
ICLR4
2021 Learning Seed Placements and Automation Policies for Polyculture Farming with Companion Plants
abstract
Polyculture farming is a sustainable farming technique based on synergistic interactions between differing plant types that make them more resistant to diseases and pests and better able to retain water. Reduced uniformity can reduce use of pesticides, fertilizer, and water, but is more labor intensive and more challenging to automate. We describe a scaled physical testbed (1.5m×3.0m) that uses a high resolution camera and soil sensors to monitor polyculture plants to facilitate tuning of plant growth, companion effects, and irrigation parameters for a first-order garden simulator. We use this simulator to develop a novel seed placement algorithm that increases coverage and diversity, and a learned pruning policy. In simulation experiments, the seed placement algorithm yields 60% more coverage and 10% more diversity than random seed placement and the learned pruning policy runs 1000X faster than a procedural lookahead policy to achieve high leaf coverage and plant diversity on adversarial gardens that include plant species with diverse growth rates. These models and policies provide the groundwork for a fully-automated system under development. Code, datasets and supplementary material can be found at https://github.com/BerkeleyAutomation/AlphaGarden/.
Yahav Avigal, Anna Deza, Sebastian Oehme, Mark Presten, Mark Theis, Jackson Chui, Paul Shao, Atsunobu Kotani, Satvik Sharma, Rishi Parikh, Michael Luo, Sandeep Mukherjee, Stefano Carpin, Joshua Viers, Stavros G. Vougioukas, Kenneth Y. Goldberg
ICRA13
2021 Accelerating Quadratic Optimization with Reinforcement Learning
abstract
First-order methods for quadratic optimization such as OSQP are widely used for large-scale machine learning and embedded optimal control, where many related problems must be rapidly solved. These methods face two persistent challenges: manual hyperparameter tuning and convergence time to high-accuracy solutions. To address these, we explore how Reinforcement Learning (RL) can learn a policy to tune parameters to accelerate convergence. In experiments with well-known QP benchmarks we find that our RL policy, RLQP, significantly outperforms state-of-the-art QP solvers by up to 3x. RLQP generalizes surprisingly well to previously unseen problems with varying dimension and structure from different applications, including the QPLIB, Netlib LP and Maros-M{\'e}sz{\'a}ros problems. Code, models, and videos are available at https://berkeleyautomation.github.io/rlqp/.
Jeffrey Ichnowski, Paras Jain 0001, Bartolomeo Stellato, Goran Banjac, Michael Luo, Francesco Borrelli, Joseph Gonzalez 0001, Ion Stoica, Kenneth Y. Goldberg
NeurIPS5
2021 RLlib Flow: Distributed Reinforcement Learning is a Dataflow Problem
abstract
Researchers and practitioners in the field of reinforcement learning (RL) frequently leverage parallel computation, which has led to a plethora of new algorithms and systems in the last few years. In this paper, we re-examine the challenges posed by distributed RL and try to view it through the lens of an old idea: distributed dataflow. We show that viewing RL as a dataflow problem leads to highly composable and performant implementations. We propose RLlib Flow, a hybrid actor-dataflow programming model for distributed RL, and validate its practicality by porting the full suite of algorithms in RLlib, a widely adopted distributed RL library. Concretely, RLlib Flow provides 2-9$\times$ code savings in real production code and enables the composition of multi-agent algorithms not possible by end users before. The open-source code is available as part of RLlib at https://github.com/ray-project/ray/tree/master/rllib.
Eric Liang, Zhanghao Wu, Michael Luo, Sven Mika, Joseph Gonzalez 0001, Ion Stoica
NeurIPS3
2020 IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target Networks
Michael Luo, Jiahao Yao, Richard Liaw, Eric Liang, Ion Stoica
ICLR1