VLDB 2026 Research / reviewers in the wild / expert
Alex Gu
dblp:285/4734
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-4814-0796ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Preserving privacy, enabling collaboration: Decentralized learning framework for multi-class orthopedic imagingabstractMedical imaging data are inherently distributed across healthcare institutions and subject to strict privacy regulations, limiting the feasibility of centralized model training. In orthopedic imaging, further challenges arise from heterogeneous diagnostic tasks, implant categories, and label spaces that differ across institutions. Existing decentralized approaches, including federated and swarm learning, reduce direct data sharing but typically rely on repeated parameter synchronization and assume partially aligned label spaces, restricting their scalability in heterogeneous clinical environments. To address these limitations, we propose OrthoATD.Net, a decentralized learning framework for collaborative orthopedic image analysis that operates without raw-data sharing or iterative parameter synchronization. The framework combines independent local training with synchronization-free representation sharing, enabling knowledge integration across fully disjoint label spaces. We evaluate OrthoATD.Net across six heterogeneous orthopedic nodes comprising 43,976 X-ray images and 30 implant and diagnostic classes, using identical Vision Transformer backbones and leakage-controlled evaluation protocols. Over three independent runs, the framework achieves a mean accuracy of 97.32±0.03% and a macro F1-score of 96.45±0.07%. Within this heterogeneous disjoint-label setting, relative to the strongest decentralized baseline (Ditto-adapted, 92.93%), it improves accuracy by 4.39 and macro F1-score by 5.84 percentage points, and consistently outperforms NonIID-SL (91.18%), FedPer-adapted (90.82%), centralized learning (89.27%), FedLD (87.25%), and ATD (71.21%) under identical experimental conditions. Multi-seed statistical validation with significance testing, leave-one-node-out generalization analysis, and membership-inference attack analysis further demonstrate the robustness, reproducibility, and practical viability of the framework. The primary contribution of OrthoATD.Net is enabling synchronization-free collaborative learning across heterogeneous clinical nodes with fully disjoint label spaces rather than establishing a universal performance advantage over centralized learning. These findings suggest that synchronization-free representation sharing can serve as an effective and scalable alternative to conventional decentralized learning for heterogeneous orthopedic imaging tasks while preserving data locality, providing a promising basis for privacy-aware, scalable collaborative orthopedic artificial intelligence across distributed healthcare environments. Haider A. Alwzwazy, Alex Gu, Mustafa Dukhan, Zehui Zhao, Laith Alzubaidi |
Artif. Intell. Medicine | 2 |
| 2025 | LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeabstractLarge Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from academia and industry. However, as new and improved LLMs are developed, existing evaluation benchmarks (e.g., HumanEvla, MBPP) are no longer sufficient for assessing their capabilities suffering from data contamination, overfitting, saturation, and focus on merely code generation. In this work, we propose LiveCodeBench, a comprehensive and contamination-free evaluation of LLMs for code, which collects new problems over time from contests across three competition platforms, Leetcode, Atcoder, and Codeforces. Notably, our benchmark also focuses on a broader range of code-related capabilities, such as self-repair, code execution, and test output prediction, beyond just code generation. Currently, LiveCodeBench hosts over six hundred coding problems that were published between May 2023 and Aug 2024. We evaluate over 50 LLMs on LiveCodeBench (LCB for brevity) presenting the largest evaluation study of code LLMs on competition problems. Based on the study, we present novel empirical findings on contamination, overfitting, and holistic evaluations. We demonstrate that time-segmented evaluations serve as a robust approach to evade contamination; they are successful at detecting contamination across a wide range of open and closed models including GPT-4O, Claude, Deepseek, and Codestral. Next, we highlight overfitting and saturation of traditional coding benchmarks like HumanEvla and demonstrate LCB allows more reliable evaluations. Finally, our holistic evaluation scenarios allow for measuring the different capabilities of programming agents in isolation. Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida I. Wang, Armando Solar-Lezama, Koushik Sen, Ion Stoica |
ICLR | 3 |
| 2025 | Mixture of Parrots: Experts improve memorization more than reasoningabstractThe Mixture-of-Experts (MoE) architecture enables a significant increase in the total number of model parameters with minimal computational overhead.
However, it is not clear what performance tradeoffs, if any, exist between MoEs and standard dense transformers.
In this paper,
we show that as we increase the number of experts (while fixing the number of active parameters), the memorization performance consistently increases while the reasoning capabilities saturate.
We begin by analyzing the theoretical limitations of MoEs at reasoning. We prove that there exist graph problems that cannot be solved by any number of experts of a certain width; however, the same task can be easily solved by a dense model with a slightly larger width.
On the other hand, we find that on memory-intensive tasks, MoEs can effectively leverage a small number of active parameters with a large number of experts to memorize the data.
We empirically validate these findings on synthetic graph problems and memory-intensive closed book retrieval tasks.
Lastly, we pre-train a series of MoEs and dense transformers and evaluate them on commonly used benchmarks in math and natural language.
We find that increasing the number of experts helps solve knowledge-intensive tasks, but fails to yield the same benefits for reasoning tasks. Samy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu, Nikhil Vyas 0001, Nikhil Anand, David Alvarez-Melis, Yuanzhi Li, Sham M. Kakade, Eran Malach |
ICLR | 4 |
| 2025 | Solving Inequality Proofs with Large Language ModelsabstractInequality proving, crucial across diverse scientific and mathematical fields, tests advanced reasoning skills such as discovering tight bounds and strategic theorem application. This makes it a distinct, demanding frontier for large language models (LLMs), offering insights beyond general mathematical problem-solving. Progress in this area is hampered by existing datasets that are often scarce, synthetic, or rigidly formal. We address this by proposing an informal yet verifiable task formulation, recasting inequality proving into two automatically checkable subtasks: bound estimation and relation prediction. Building on this, we release IneqMath, an expert-curated dataset of Olympiad-level inequalities, including a test set and training corpus enriched with step-wise solutions and theorem annotations. We also develop a novel LLM-as-judge evaluation suite, combining a final-answer judge with four specialized step-wise judges designed to detect common reasoning flaws. A systematic evaluation of 29 leading LLMs on IneqMath reveals a surprising reality: even top models like o1 achieve less than 10% overall accuracy under step-wise scrutiny; this is a drop of up to 65.5% from their accuracy considering only final answer equivalence. This discrepancy exposes fragile deductive chains and a critical gap for current LLMs between merely finding an answer and constructing a rigorous proof. Scaling model size and increasing test-time computation yield limited gains in overall proof correctness. Instead, our findings highlight promising research directions such as theorem-guided reasoning and self-refinement. Jiayi Sheng, Luna Lyu, Jikai Jin, Tanglin Xia, Alex Gu, James Zou 0001, Pan Lu |
NeurIPS | 5 |
| 2024 | CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionabstractWe present Code Reasoning, Understanding, and eXecution Evaluation, a benchmark consisting of 800 Python functions (3-13 lines). Each function comes with an input-output pair, leading to two natural tasks: input prediction and output prediction. First, we propose a general recipe for generating our execution benchmark by sampling from a model, which can be used for more challenging versions of the benchmark if needed. Second, we evaluate twenty code models on our benchmark and discover that many recent high-scoring models on HumanEval show no improvements on our benchmark. Third, we show that simple CoT and fine-tuning schemes can improve performance on our benchmark but remain far from solving it. The best setup, GPT-4 with chain of thought (CoT), achieves a pass@1 of 75% and 81% on input and output prediction, respectively. In contrast, Code Llama 34B achieves a pass@1 of 50% and 46% on input and output prediction. When it comes to reasoning about code, GPT-4 has a huge edge over other models but still fails consistently on some surprisingly simple Python programs. Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, Sida I. Wang |
ICML | 1 |
| 2024 | Language Agnostic Code EmbeddingsabstractSaiteja Utpala, Alex Gu, Pin-Yu Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Saiteja Utpala, Alex Gu |
NAACL-HLT | 2 |
| 2023 | LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic ProversabstractLogical reasoning, i.e., deductively inferring the truth value of a conclusion from a set of premises, is an important task for artificial intelligence with wide potential impacts on science, mathematics, and society.While many prompting-based strategies have been proposed to enable Large Language Models (LLMs) to do such reasoning more effectively, they still appear unsatisfactory, often failing in subtle and unpredictable ways.In this work, we investigate the validity of instead reformulating such tasks as modular neurosymbolic programming, which we call LINC: Logical Inference via Neurosymbolic Computation.In LINC, the LLM acts as a semantic parser, translating premises and conclusions from natural language to expressions in first-order logic.These expressions are then offloaded to an external theorem prover, which symbolically performs deductive inference.Leveraging this approach, we observe significant performance gains on FOLIO and a balanced subset of ProofWriter for three different models in nearly all experimental conditions we evaluate.On ProofWriter, augmenting the comparatively small open-source StarCoder+ (15.5B parameters) with LINC even outperforms GPT-3.5 and GPT-4 with Chain-of-Thought (CoT) prompting by an absolute 38% and 10%, respectively.When used with GPT-4, LINC scores 26% higher than CoT on ProofWriter while performing comparatively on FOLIO.Further analysis reveals that although both methods on average succeed roughly equally often on this dataset, they exhibit distinct and complementary failure modes.We thus provide promising evidence for how logical reasoning over natural language can be tackled through jointly leveraging LLMs alongside symbolic provers.All corresponding code is publicly available. Theo X. Olausson, Alex Gu, Benjamin Lipkin, Cedegao E. Zhang, Armando Solar-Lezama, Josh Tenenbaum, Roger Levy |
EMNLP | 2 |
| 2023 | Min-Max Multi-objective Bilevel Optimization with Applications in Robust Machine Learning
Alex Gu, Songtao Lu, Parikshit Ram, Tsui-Wei Weng |
ICLR | 1 |
| 2023 | LeanDojo: Theorem Proving with Retrieval-Augmented Language ModelsabstractLarge language models (LLMs) have shown promise in proving formal theorems using proof assistants such as Lean. However, existing methods are difficult to reproduce or build on, due to private code, data, and large compute requirements. This has created substantial barriers to research on machine learning methods for theorem proving. This paper removes these barriers by introducing LeanDojo: an open-source Lean playground consisting of toolkits, data, models, and benchmarks. LeanDojo extracts data from Lean and enables interaction with the proof environment programmatically. It contains fine-grained annotations of premises in proofs, providing valuable data for premise selection—a key bottleneck in theorem proving. Using this data, we develop ReProver (Retrieval-Augmented Prover): an LLM-based prover augmented with retrieval for selecting premises from a vast math library. It is inexpensive and needs only one GPU week of training. Our retriever leverages LeanDojo's program analysis capability to identify accessible premises and hard negative examples, which makes retrieval much more effective. Furthermore, we construct a new benchmark consisting of 98,734 theorems and proofs extracted from Lean's math library. It features challenging data split requiring the prover to generalize to theorems relying on novel premises that are never used in training. We use this benchmark for training and evaluation, and experimental results demonstrate the effectiveness of ReProver over non-retrieval baselines and GPT-4. We thus provide the first set of open-source LLM-based theorem provers without any proprietary datasets and release it under a permissive MIT license to facilitate further research. Kaiyu Yang, Aidan M. Swope, Alex Gu, Rahul Chalamala, Peiyang Song 0002, Shixing Yu, Saad Godil, Ryan Prenger, Anima Anandkumar |
NeurIPS | 3 |
| 2021 | Three Operator Splitting with Subgradients, Stochastic Gradients, and Adaptive Learning RatesabstractThree Operator Splitting (TOS) (Davis & Yin, 2017) can minimize the sum of multiple convex functions effectively when an efficient gradient oracle or proximal operator is available for each term. This requirement often fails in machine learning applications: (i) instead of full gradients only stochastic gradients may be available; and (ii) instead of proximal operators, using subgradients to handle complex penalty functions may be more efficient and realistic. Motivated by these concerns, we analyze three potentially valuable extensions of TOS. The first two permit using subgradients and stochastic gradients, and are shown to ensure a $\mathcal{O}(1/\sqrt{t})$ convergence rate. The third extension AdapTOS endows TOS with adaptive step-sizes. For the important setting of optimizing a convex loss over the intersection of convex sets AdapTOS attains universal convergence rates, i.e., the rate adapts to the unknown smoothness degree of the objective. We compare our proposed methods with competing methods on various applications. Alp Yurtsever, Alex Gu, Suvrit Sra |
NeurIPS | 2 |