VLDB 2026 Research / reviewers in the wild / expert
Siddhartha Jain 0001
dblp:81/8212-1
· DBLP profile ↗
15ranked-venue papers
8as first author
7since 2021 · last 2026
0009-0006-3845-6478ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 55% Trustworthy machine learning · 18% Planning, search and constraint satisfaction · 9% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model evaluation |
1.6 | 2 | 2025 | LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025 Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies · EMNLP 2024 |
Natural language and speech › Language models and text generation
test-time scaling |
1.0 | 1 | 2026 | Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model evaluation
benchmark contamination |
0.9 | 1 | 2025 | LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025 |
Natural language and speech › Language models and text generation
reranking |
0.8 | 1 | 2024 | Lightweight reranking for language model generations · ACL (1) 2024 |
Natural language and speech › Language models and text generation
code generation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Program synthesis and code generation
code generation evaluation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Program synthesis and code generation › code generation with language models
multilingual code generation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | Overinterpretation reveals image classification model pathologies · NeurIPS 2021 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | Overinterpretation reveals image classification model pathologies · NeurIPS 2021 |
Bioinformatics and computational biology
immunoinformatics |
0.5 | 1 | 2021 | Machine learning optimization of peptides for presentation by class II MHCs · Bioinform. 2021 |
Bioinformatics and computational biology › immunoinformatics
peptide-MHC binding prediction |
0.5 | 1 | 2021 | Machine learning optimization of peptides for presentation by class II MHCs · Bioinform. 2021 |
Machine learning › Trustworthy machine learning › uncertainty estimation › model uncertainty
ensemble uncertainty |
0.4 | 1 | 2020 | Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles · AAAI 2020 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.4 | 1 | 2020 | Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles · AAAI 2020 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.4 | 1 | 2020 | Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles · AAAI 2020 |
Machine learning › Reinforcement learning
benchmark design |
0.3 | 1 | 2025 | LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025 |
Bioinformatics and computational biology
systems biology |
0.2 | 1 | 2016 | Reconstructing the temporal progression of HIV-1 immune response pathways · Bioinform. 2016 |
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent debate |
0.2 | 1 | 2024 | Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies · EMNLP 2024 |
Natural language and speech › Language models and text generation
self-consistency |
0.2 | 1 | 2024 | Lightweight reranking for language model generations · ACL (1) 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.1 | 1 | 2021 | Overinterpretation reveals image classification model pathologies · NeurIPS 2021 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.1 | 1 | 2020 | Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles · AAAI 2020 |
Computational complexity
constraint satisfaction |
0.1 | 1 | 2011 | A General Nogood-Learning Framework for Pseudo-Boolean Multi-Valued SAT · AAAI 2011 |
Automated reasoning and model checking
satisfiability |
0.1 | 1 | 2011 | A General Nogood-Learning Framework for Pseudo-Boolean Multi-Valued SAT · AAAI 2011 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.3test-time search · 1.0reinforcement learning · 1.0chain-of-thought prompting · 1.0crowdsourcing · 0.9automated scoring · 0.9self-consistency · 0.8pairwise statistics · 0.8compute budget analysis · 0.8chain-of-thought self-consistency · 0.8yeast display assay · 0.5ensemble learning · 0.5deep residual network · 0.5integer programming · 0.2conflict analysis · 0.1bounds consistency · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight ModelsabstractMehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain, Wasi Uddin Ahmad, Somshubra Majumdar, Vahid Noroozi, Boris Ginsburg. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain 0001, Wasi Uddin Ahmad, Somshubra Majumdar, Vahid Noroozi, Boris Ginsburg |
ACL (1) | 4 |
| 2025 | LiveBench: A Challenging, Contamination-Limited LLM BenchmarkabstractTest set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models. Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain 0001, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, Micah Goldblum |
ICLR | 6 |
| 2024 | Lightweight reranking for language model generationsabstractLarge Language Models (LLMs) can exhibit considerable variation in the quality of their sampled outputs.Reranking and selecting the best generation from the sampled set is a popular way of obtaining strong gains in generation quality.In this paper, we present a novel approach for reranking LLM generations.Unlike other techniques that might involve additional inferences or training a specialized reranker, our approach relies on easy to compute pairwise statistics between the generations that have minimal compute overhead.We show that our approach can be formalized as an extension of self-consistency and analyze its performance in that framework, theoretically as well as via simulations.We show strong improvements for selecting the best k generations for code generation tasks as well as robust improvements for the best generation for the tasks of autoformalization, summarization, and translation.While our approach only assumes black-box access to LLMs, we show that additional access to token probabilities can improve performance even further. Siddhartha Jain 0001, Xiaofei Ma 0001, Anoop Deoras, Bing Xiang |
ACL (1) | 1 |
| 2024 | Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning StrategiesabstractA diverse array of reasoning strategies has been proposed to elicit the capabilities of large language models.However, in this paper, we point out that traditional evaluations which focus solely on performance metrics miss a key factor: the increased effectiveness due to additional compute.By overlooking this aspect, a skewed view of strategy efficiency is often presented.This paper introduces a framework that incorporates the compute budget into the evaluation, providing a more informative comparison that takes into account both performance metrics and computational cost.In this budgetaware perspective, we find that complex reasoning strategies often don't surpass simpler baselines purely due to algorithmic ingenuity, but rather due to the larger computational resources allocated.When we provide a simple baseline like chain-of-thought self-consistency with comparable compute resources, it frequently outperforms reasoning strategies proposed in the literature.In this scale-aware perspective, we find that unlike self-consistency, certain strategies such as multi-agent debate or Reflexion can become worse if more compute budget is utilized. Siddhartha Jain 0001, Dejiao Zhang, Baishakhi Ray, Ben Athiwaratkun |
EMNLP | 2 |
| 2023 | Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang 0002, Xiaopeng Li 0002, Wasi Uddin Ahmad, Shiqi Wang 0002, Qing Sun 0013, Mingyue Shang, Sujan K. Gonugondla, Hantian Ding, Nathan Fulton, Arash Farahani, Siddhartha Jain 0001, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati |
ICLR | 16 |
| 2021 | Overinterpretation reveals image classification model pathologiesabstractImage classifiers are typically scored on their test set accuracy, but high accuracy can mask a subtle type of model failure. We find that high scoring convolutional neural networks (CNNs) on popular benchmarks exhibit troubling pathologies that allow them to display high accuracy even in the absence of semantically salient features. When a model provides a high-confidence decision without salient supporting input features, we say the classifier has overinterpreted its input, finding too much class-evidence in patterns that appear nonsensical to humans. Here, we demonstrate that neural networks trained on CIFAR-10 and ImageNet suffer from overinterpretation, and we find models on CIFAR-10 make confident predictions even when 95% of input images are masked and humans cannot discern salient features in the remaining pixel-subsets. We introduce Batched Gradient SIS, a new method for discovering sufficient input subsets for complex datasets, and use this method to show the sufficiency of border pixels in ImageNet for training and testing. Although these patterns portend potential model fragility in real-world deployment, they are in fact valid statistical patterns of the benchmark that alone suffice to attain high test accuracy. Unlike adversarial examples, overinterpretation relies upon unmodified image pixels. We find ensembling and input dropout can each help mitigate overinterpretation. Brandon Carter 0001, Siddhartha Jain 0001, Jonas Mueller 0001, David K. Gifford |
NeurIPS | 2 |
| 2021 | Machine learning optimization of peptides for presentation by class II MHCsabstractSUMMARY: T cells play a critical role in cellular immune responses to pathogens and cancer and can be activated and expanded by Major Histocompatibility Complex (MHC)-presented antigens contained in peptide vaccines. We present a machine learning method to optimize the presentation of peptides by class II MHCs by modifying their anchor residues. Our method first learns a model of peptide affinity for a class II MHC using an ensemble of deep residual networks, and then uses the model to propose anchor residue changes to improve peptide affinity. We use a high throughput yeast display assay to show that anchor residue optimization improves peptide binding. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zheng Dai, Brooke D. Huisman, Brandon Carter 0001, Siddhartha Jain 0001, Michael E. Birnbaum, David K. Gifford |
Bioinform. | 5 |
| 2020 | Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep EnsemblesabstractThe inaccuracy of neural network models on inputs that do not stem from the distribution underlying the training data is problematic and at times unrecognized. Uncertainty estimates of model predictions are often based on the variation in predictions produced by a diverse ensemble of models applied to the same input. Here we describe Maximize Overall Diversity (MOD), an approach to improve ensemble-based uncertainty estimates by encouraging larger overall diversity in ensemble predictions across all possible inputs. We apply MOD to regression tasks including 38 Protein-DNA binding datasets, 9 UCI datasets, and the IMDB-Wiki image dataset. We also explore variants that utilize adversarial training techniques and data density estimation. For out-of-distribution test examples, MOD significantly improves predictive performance and uncertainty calibration without sacrificing performance on test data drawn from same distribution as the training data. We also find that in Bayesian optimization tasks, the performance of UCB acquisition is improved via MOD uncertainty estimates. Siddhartha Jain 0001, Jonas Mueller 0001, David K. Gifford |
AAAI | 1 |
| 2019 | What made you do this? Understanding black-box decisions with sufficient input subsetsabstractLocal explanation frameworks aim to rationalize particular decisions made by a black-box prediction model. Existing techniques are often restricted to a specific type of predictor or based on input saliency, which may be undesirably sensitive to factors unrelated to the model’s decision making process. We instead propose sufficient input subsets that identify minimal subsets of features whose observed values alone suffice for the same decision to be reached, even if all other input feature values are missing. General principles that globally govern a model’s decision-making can also be revealed by searching for clusters of such input patterns across many data points. Our approach is conceptually straightforward, entirely model-agnostic, simply implemented using instance-wise backward selection, and able to produce more concise rationales than existing techniques. We demonstrate the utility of our interpretation method on various neural network models trained on text, image, and genomic data. Brandon Carter 0001, Jonas Mueller 0001, Siddhartha Jain 0001, David K. Gifford |
AISTATS | 3 |
| 2016 | Reconstructing the temporal progression of HIV-1 immune response pathwaysabstractMOTIVATION: Most methods for reconstructing response networks from high throughput data generate static models which cannot distinguish between early and late response stages. RESULTS: We present TimePath, a new method that integrates time series and static datasets to reconstruct dynamic models of host response to stimulus. TimePath uses an Integer Programming formulation to select a subset of pathways that, together, explain the observed dynamic responses. Applying TimePath to study human response to HIV-1 led to accurate reconstruction of several known regulatory and signaling pathways and to novel mechanistic insights. We experimentally validated several of TimePaths' predictions highlighting the usefulness of temporal models. AVAILABILITY AND IMPLEMENTATION: Data, Supplementary text and the TimePath software are available from http://sb.cs.cmu.edu/timepath CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Siddhartha Jain 0001, Joel Arrais, Narasimhan J. Venkatachari, Velpandi Ayyavoo, Ziv Bar-Joseph |
Bioinform. | 1 |
| 2014 | Multitask Learning of Signaling and Regulatory Networks with Application to Studying Human Response to FluabstractReconstructing regulatory and signaling response networks is one of the major goals of systems biology. While several successful methods have been suggested for this task, some integrating large and diverse datasets, these methods have so far been applied to reconstruct a single response network at a time, even when studying and modeling related conditions. To improve network reconstruction we developed MT-SDREM, a multi-task learning method which jointly models networks for several related conditions. In MT-SDREM, parameters are jointly constrained across the networks while still allowing for condition-specific pathways and regulation. We formulate the multi-task learning problem and discuss methods for optimizing the joint target function. We applied MT-SDREM to reconstruct dynamic human response networks for three flu strains: H1N1, H5N1 and H3N2. Our multi-task learning method was able to identify known and novel factors and genes, improving upon prior methods that model each condition independently. The MT-SDREM networks were also better at identifying proteins whose removal affects viral load indicating that joint learning can still lead to accurate, condition-specific, networks. Supporting website with MT-SDREM implementation: http://sb.cs.cmu.edu/mtsdrem. Siddhartha Jain 0001, Anthony Gitter, Ziv Bar-Joseph |
PLoS Comput. Biol. | 1 |
| 2011 | A General Nogood-Learning Framework for Pseudo-Boolean Multi-Valued SATabstractWe formulate a general framework for pseudo-Boolean multi-valued nogood-learning, generalizing conflict analysis performed by modern SAT solvers and its recent extension for disjunctions of multi-valued variables. This framework can handle more general constraints as well as different domain representations, such as interval domains which are commonly used for bounds consistency in constraint programming (CP), and even set variables. Our empirical evaluation shows that our solver, built upon this framework, works robustly across a number of challenging domains. Siddhartha Jain 0001, Ashish Sabharwal, Meinolf Sellmann |
AAAI | 1 |
| 2011 | Large Neighborhood Search for Dial-a-Ride Problems
Siddhartha Jain 0001, Pascal Van Hentenryck |
CP | 1 |
| 2010 | A Complete Multi-valued SAT Solver
Siddhartha Jain 0001, Eoin O'Mahony, Meinolf Sellmann |
CP | 1 |
| 2010 | Upper Bounds on the Number of Solutions of Binary Integer Programs
Siddhartha Jain 0001, Serdar Kadioglu, Meinolf Sellmann |
CPAIOR | 1 |