Siddhartha Jain 0001

dblp:81/8212-1 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
7since 2021 · last 2026
0009-0006-3845-6478ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 55% Trustworthy machine learning · 18% Planning, search and constraint satisfaction · 9%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
1.622025
LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025
Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies · EMNLP 2024
Natural language and speech › Language models and text generation
test-time scaling
1.012026
Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model evaluation
benchmark contamination
0.912025
LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025
Natural language and speech › Language models and text generation
reranking
0.812024
Lightweight reranking for language model generations · ACL (1) 2024
Natural language and speech › Language models and text generation
code generation
0.712023
Multi-lingual Evaluation of Code Generation Models · ICLR 2023
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation
0.712023
Multi-lingual Evaluation of Code Generation Models · ICLR 2023
Program synthesis and code generation
code generation evaluation
0.712023
Multi-lingual Evaluation of Code Generation Models · ICLR 2023
Program synthesis and code generation › code generation with language models
multilingual code generation
0.712023
Multi-lingual Evaluation of Code Generation Models · ICLR 2023
Computer vision › Image recognition and object detection
image classification
0.512021
Overinterpretation reveals image classification model pathologies · NeurIPS 2021
Machine learning › Trustworthy machine learning
interpretability
0.512021
Overinterpretation reveals image classification model pathologies · NeurIPS 2021
Bioinformatics and computational biology
immunoinformatics
0.512021
Machine learning optimization of peptides for presentation by class II MHCs · Bioinform. 2021
Bioinformatics and computational biology › immunoinformatics
peptide-MHC binding prediction
0.512021
Machine learning optimization of peptides for presentation by class II MHCs · Bioinform. 2021
Machine learning › Trustworthy machine learning › uncertainty estimation › model uncertainty
ensemble uncertainty
0.412020
Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles · AAAI 2020
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.412020
Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles · AAAI 2020
Machine learning › Trustworthy machine learning
uncertainty estimation
0.412020
Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles · AAAI 2020
Machine learning › Reinforcement learning
benchmark design
0.312025
LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025
Bioinformatics and computational biology
systems biology
0.212016
Reconstructing the temporal progression of HIV-1 immune response pathways · Bioinform. 2016
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent debate
0.212024
Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies · EMNLP 2024
Natural language and speech › Language models and text generation
self-consistency
0.212024
Lightweight reranking for language model generations · ACL (1) 2024
Machine learning › Trustworthy machine learning
robustness
0.112021
Overinterpretation reveals image classification model pathologies · NeurIPS 2021
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.112020
Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles · AAAI 2020
Computational complexity
constraint satisfaction
0.112011
A General Nogood-Learning Framework for Pseudo-Boolean Multi-Valued SAT · AAAI 2011
Automated reasoning and model checking
satisfiability
0.112011
A General Nogood-Learning Framework for Pseudo-Boolean Multi-Valued SAT · AAAI 2011

Methods — techniques the papers use, named apart from their topics

large language model · 1.3test-time search · 1.0reinforcement learning · 1.0chain-of-thought prompting · 1.0crowdsourcing · 0.9automated scoring · 0.9self-consistency · 0.8pairwise statistics · 0.8compute budget analysis · 0.8chain-of-thought self-consistency · 0.8yeast display assay · 0.5ensemble learning · 0.5deep residual network · 0.5integer programming · 0.2conflict analysis · 0.1bounds consistency · 0.1
YearPublicationVenuePosition
2026 Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
abstract
Mehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain, Wasi Uddin Ahmad, Somshubra Majumdar, Vahid Noroozi, Boris Ginsburg. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain 0001, Wasi Uddin Ahmad, Somshubra Majumdar, Vahid Noroozi, Boris Ginsburg
ACL (1)4
2025 LiveBench: A Challenging, Contamination-Limited LLM Benchmark
abstract
Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models.
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain 0001, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, Micah Goldblum
ICLR6
2024 Lightweight reranking for language model generations
abstract
Large Language Models (LLMs) can exhibit considerable variation in the quality of their sampled outputs.Reranking and selecting the best generation from the sampled set is a popular way of obtaining strong gains in generation quality.In this paper, we present a novel approach for reranking LLM generations.Unlike other techniques that might involve additional inferences or training a specialized reranker, our approach relies on easy to compute pairwise statistics between the generations that have minimal compute overhead.We show that our approach can be formalized as an extension of self-consistency and analyze its performance in that framework, theoretically as well as via simulations.We show strong improvements for selecting the best k generations for code generation tasks as well as robust improvements for the best generation for the tasks of autoformalization, summarization, and translation.While our approach only assumes black-box access to LLMs, we show that additional access to token probabilities can improve performance even further.
Siddhartha Jain 0001, Xiaofei Ma 0001, Anoop Deoras, Bing Xiang
ACL (1)1
2024 Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
abstract
A diverse array of reasoning strategies has been proposed to elicit the capabilities of large language models.However, in this paper, we point out that traditional evaluations which focus solely on performance metrics miss a key factor: the increased effectiveness due to additional compute.By overlooking this aspect, a skewed view of strategy efficiency is often presented.This paper introduces a framework that incorporates the compute budget into the evaluation, providing a more informative comparison that takes into account both performance metrics and computational cost.In this budgetaware perspective, we find that complex reasoning strategies often don't surpass simpler baselines purely due to algorithmic ingenuity, but rather due to the larger computational resources allocated.When we provide a simple baseline like chain-of-thought self-consistency with comparable compute resources, it frequently outperforms reasoning strategies proposed in the literature.In this scale-aware perspective, we find that unlike self-consistency, certain strategies such as multi-agent debate or Reflexion can become worse if more compute budget is utilized.
Siddhartha Jain 0001, Dejiao Zhang, Baishakhi Ray, Ben Athiwaratkun
EMNLP2
2023 Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang 0002, Xiaopeng Li 0002, Wasi Uddin Ahmad, Shiqi Wang 0002, Qing Sun 0013, Mingyue Shang, Sujan K. Gonugondla, Hantian Ding, Nathan Fulton, Arash Farahani, Siddhartha Jain 0001, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati
ICLR16
2021 Overinterpretation reveals image classification model pathologies
abstract
Image classifiers are typically scored on their test set accuracy, but high accuracy can mask a subtle type of model failure. We find that high scoring convolutional neural networks (CNNs) on popular benchmarks exhibit troubling pathologies that allow them to display high accuracy even in the absence of semantically salient features. When a model provides a high-confidence decision without salient supporting input features, we say the classifier has overinterpreted its input, finding too much class-evidence in patterns that appear nonsensical to humans. Here, we demonstrate that neural networks trained on CIFAR-10 and ImageNet suffer from overinterpretation, and we find models on CIFAR-10 make confident predictions even when 95% of input images are masked and humans cannot discern salient features in the remaining pixel-subsets. We introduce Batched Gradient SIS, a new method for discovering sufficient input subsets for complex datasets, and use this method to show the sufficiency of border pixels in ImageNet for training and testing. Although these patterns portend potential model fragility in real-world deployment, they are in fact valid statistical patterns of the benchmark that alone suffice to attain high test accuracy. Unlike adversarial examples, overinterpretation relies upon unmodified image pixels. We find ensembling and input dropout can each help mitigate overinterpretation.
Brandon Carter 0001, Siddhartha Jain 0001, Jonas Mueller 0001, David K. Gifford
NeurIPS2
2021 Machine learning optimization of peptides for presentation by class II MHCs
abstract
SUMMARY: T cells play a critical role in cellular immune responses to pathogens and cancer and can be activated and expanded by Major Histocompatibility Complex (MHC)-presented antigens contained in peptide vaccines. We present a machine learning method to optimize the presentation of peptides by class II MHCs by modifying their anchor residues. Our method first learns a model of peptide affinity for a class II MHC using an ensemble of deep residual networks, and then uses the model to propose anchor residue changes to improve peptide affinity. We use a high throughput yeast display assay to show that anchor residue optimization improves peptide binding. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zheng Dai, Brooke D. Huisman, Brandon Carter 0001, Siddhartha Jain 0001, Michael E. Birnbaum, David K. Gifford
Bioinform.5
2020 Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep Ensembles
abstract
The inaccuracy of neural network models on inputs that do not stem from the distribution underlying the training data is problematic and at times unrecognized. Uncertainty estimates of model predictions are often based on the variation in predictions produced by a diverse ensemble of models applied to the same input. Here we describe Maximize Overall Diversity (MOD), an approach to improve ensemble-based uncertainty estimates by encouraging larger overall diversity in ensemble predictions across all possible inputs. We apply MOD to regression tasks including 38 Protein-DNA binding datasets, 9 UCI datasets, and the IMDB-Wiki image dataset. We also explore variants that utilize adversarial training techniques and data density estimation. For out-of-distribution test examples, MOD significantly improves predictive performance and uncertainty calibration without sacrificing performance on test data drawn from same distribution as the training data. We also find that in Bayesian optimization tasks, the performance of UCB acquisition is improved via MOD uncertainty estimates.
Siddhartha Jain 0001, Jonas Mueller 0001, David K. Gifford
AAAI1
2019 What made you do this? Understanding black-box decisions with sufficient input subsets
abstract
Local explanation frameworks aim to rationalize particular decisions made by a black-box prediction model. Existing techniques are often restricted to a specific type of predictor or based on input saliency, which may be undesirably sensitive to factors unrelated to the model’s decision making process. We instead propose sufficient input subsets that identify minimal subsets of features whose observed values alone suffice for the same decision to be reached, even if all other input feature values are missing. General principles that globally govern a model’s decision-making can also be revealed by searching for clusters of such input patterns across many data points. Our approach is conceptually straightforward, entirely model-agnostic, simply implemented using instance-wise backward selection, and able to produce more concise rationales than existing techniques. We demonstrate the utility of our interpretation method on various neural network models trained on text, image, and genomic data.
Brandon Carter 0001, Jonas Mueller 0001, Siddhartha Jain 0001, David K. Gifford
AISTATS3
2016 Reconstructing the temporal progression of HIV-1 immune response pathways
abstract
MOTIVATION: Most methods for reconstructing response networks from high throughput data generate static models which cannot distinguish between early and late response stages. RESULTS: We present TimePath, a new method that integrates time series and static datasets to reconstruct dynamic models of host response to stimulus. TimePath uses an Integer Programming formulation to select a subset of pathways that, together, explain the observed dynamic responses. Applying TimePath to study human response to HIV-1 led to accurate reconstruction of several known regulatory and signaling pathways and to novel mechanistic insights. We experimentally validated several of TimePaths' predictions highlighting the usefulness of temporal models. AVAILABILITY AND IMPLEMENTATION: Data, Supplementary text and the TimePath software are available from http://sb.cs.cmu.edu/timepath CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Siddhartha Jain 0001, Joel Arrais, Narasimhan J. Venkatachari, Velpandi Ayyavoo, Ziv Bar-Joseph
Bioinform.1
2014 Multitask Learning of Signaling and Regulatory Networks with Application to Studying Human Response to Flu
abstract
Reconstructing regulatory and signaling response networks is one of the major goals of systems biology. While several successful methods have been suggested for this task, some integrating large and diverse datasets, these methods have so far been applied to reconstruct a single response network at a time, even when studying and modeling related conditions. To improve network reconstruction we developed MT-SDREM, a multi-task learning method which jointly models networks for several related conditions. In MT-SDREM, parameters are jointly constrained across the networks while still allowing for condition-specific pathways and regulation. We formulate the multi-task learning problem and discuss methods for optimizing the joint target function. We applied MT-SDREM to reconstruct dynamic human response networks for three flu strains: H1N1, H5N1 and H3N2. Our multi-task learning method was able to identify known and novel factors and genes, improving upon prior methods that model each condition independently. The MT-SDREM networks were also better at identifying proteins whose removal affects viral load indicating that joint learning can still lead to accurate, condition-specific, networks. Supporting website with MT-SDREM implementation: http://sb.cs.cmu.edu/mtsdrem.
Siddhartha Jain 0001, Anthony Gitter, Ziv Bar-Joseph
PLoS Comput. Biol.1
2011 A General Nogood-Learning Framework for Pseudo-Boolean Multi-Valued SAT
abstract
We formulate a general framework for pseudo-Boolean multi-valued nogood-learning, generalizing conflict analysis performed by modern SAT solvers and its recent extension for disjunctions of multi-valued variables. This framework can handle more general constraints as well as different domain representations, such as interval domains which are commonly used for bounds consistency in constraint programming (CP), and even set variables. Our empirical evaluation shows that our solver, built upon this framework, works robustly across a number of challenging domains.
Siddhartha Jain 0001, Ashish Sabharwal, Meinolf Sellmann
AAAI1
2011 Large Neighborhood Search for Dial-a-Ride Problems
Siddhartha Jain 0001, Pascal Van Hentenryck
CP1
2010 A Complete Multi-valued SAT Solver
Siddhartha Jain 0001, Eoin O'Mahony, Meinolf Sellmann
CP1
2010 Upper Bounds on the Number of Solutions of Binary Integer Programs
Siddhartha Jain 0001, Serdar Kadioglu, Meinolf Sellmann
CPAIOR1