VLDB 2026 Research / reviewers in the wild / expert
Tarun Suresh
dblp:348/7104
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 57% Trustworthy machine learning · 24% Planning, search and constraint satisfaction · 8% | |
| Software engineering, system software, and programming languages
3 papers |
Program verification · 81% Debugging and program repair · 9% Program synthesis and code generation · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › decoding
constrained decoding |
1.7 | 2 | 2025 | DINGO: Constrained Inference for Diffusion LLMs · NeurIPS 2025 CRANE: Reasoning with constrained LLM generation · ICML 2025 |
Natural language and speech › Language models and text generation › text generation
structured generation |
1.7 | 2 | 2025 | DINGO: Constrained Inference for Diffusion LLMs · NeurIPS 2025 IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
1.0 | 2 | 2024 | Incremental Randomized Smoothing Certification · ICLR 2024 Relational Verification Leaps Forward with RABBit · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
backtracking |
0.9 | 1 | 2025 | IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking · ICLR 2025 |
Natural language and speech › Language models and text generation › decoding
diffusion language model decoding |
0.9 | 1 | 2025 | DINGO: Constrained Inference for Diffusion LLMs · NeurIPS 2025 |
Natural language and speech › Language models and text generation › controllable text generation
grammar-guided generation |
0.9 | 1 | 2025 | IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking · ICLR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
symbolic reasoning |
0.9 | 1 | 2025 | CRANE: Reasoning with constrained LLM generation · ICML 2025 |
Information retrieval › document retrieval › domain-specific retrieval
code search |
0.9 | 1 | 2025 | CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking · ICLR 2025 |
Information retrieval
ranking |
0.9 | 1 | 2025 | CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking · ICLR 2025 |
Information retrieval
reranking |
0.9 | 1 | 2025 | CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking · ICLR 2025 |
Machine learning › Trustworthy machine learning › robustness
certified robustness |
0.8 | 1 | 2024 | Incremental Randomized Smoothing Certification · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness › certified robustness
randomized smoothing |
0.8 | 1 | 2024 | Incremental Randomized Smoothing Certification · ICLR 2024 |
Program verification › neural network verification
branch-and-bound verification |
0.8 | 1 | 2024 | Relational Verification Leaps Forward with RABBit · NeurIPS 2024 |
Program verification
neural network verification |
0.8 | 1 | 2024 | Relational Verification Leaps Forward with RABBit · NeurIPS 2024 |
Program verification
relational verification |
0.8 | 1 | 2024 | Relational Verification Leaps Forward with RABBit · NeurIPS 2024 |
Natural language and speech › Language models and text generation
code generation |
0.3 | 1 | 2025 | IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking · ICLR 2025 |
Debugging and program repair › fault localization
bug localization |
0.3 | 1 | 2025 | CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking · ICLR 2025 |
Program synthesis and code generation
code generation with language models |
0.3 | 1 | 2025 | CRANE: Reasoning with constrained LLM generation · ICML 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2024 | Incremental Randomized Smoothing Certification · ICLR 2024 |
Machine learning › Efficient and distributed learning › model compression
pruning and quantization |
0.2 | 1 | 2024 | Incremental Randomized Smoothing Certification · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness › adversarial examples
universal adversarial perturbation |
0.2 | 1 | 2024 | Relational Verification Leaps Forward with RABBit · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
hard negative mining · 1.7grammar augmentation · 1.7fine-tuning · 1.7contrastive learning · 1.7constrained decoding algorithm · 1.7symbol-to-position mapping · 0.9grammar-guided decoding · 0.9dynamic programming · 0.9diffusion language model · 0.9automated puzzle generation · 0.9KV cache · 0.9incremental certification · 0.8branch-and-bound · 0.8bound refinement · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT FormulasabstractAnjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh, Huanmi Tan, Zhanke Zhou, Sanmi Koyejo, Ke Wang, Alex Aiken. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Anjiang Wei, Yingjia Wan, Tarun Suresh, Huanmi Tan, Zhanke Zhou, Oluwasanmi Koyejo, Ke Wang 0022, Alex Aiken |
EMNLP | 4 |
| 2025 | CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and RerankingabstractEffective code retrieval plays a crucial role in advancing code generation, bug fixing, and software maintenance, particularly as software systems increase in complexity. While current code embedding models have demonstrated promise in retrieving code snippets for small-scale, well-defined tasks, they often underperform in more demanding real-world applications such as bug localization within GitHub repositories. We hypothesize that a key issue is their reliance on noisy and inconsistent datasets for training, which impedes their ability to generalize to more complex retrieval scenarios. To address these limitations, we introduce CoRNStack, a large-scale, high-quality contrastive training dataset for code that spans multiple programming languages. This dataset is curated using consistency filtering to eliminate noisy positives and is further enriched with mined hard negatives, thereby facilitating more effective learning. We demonstrate that contrastive training of embedding models using CoRNStack leads to state-of-the-art performance across a variety of code retrieval tasks. Furthermore, the dataset can be leveraged for training code reranking models, a largely underexplored area compared to text reranking. Our finetuned code reranking model significantly improves the ranking quality over the retrieved results. Finally, by employing our code retriever and reranker together, we demonstrate significant improvements in function localization for GitHub issues, an important
component of real-world software development. Tarun Suresh, Revanth Gangi Reddy, Zach Nussbaum, Andriy Mulyar, Brandon Duderstadt, Heng Ji 0001 |
ICLR | 1 |
| 2025 | Tamper-Resistant Safeguards for Open-Weight LLMsabstractRapid advances in the capabilities of large language models (LLMs) have raised widespread concerns regarding their potential for malicious use. Open-weight LLMs present unique challenges, as existing safeguards lack robustness to tampering attacks that modify model weights. For example, recent works have demonstrated that refusal and unlearning safeguards can be trivially removed with a few steps of fine-tuning. These vulnerabilities necessitate new approaches for enabling the safe release of open-weight LLMs. We develop a method, called TAR, for building tamper-resistant safeguards into open-weight LLMs such that adversaries cannot remove the safeguards even after hundreds of steps of fine-tuning. In extensive evaluations and red teaming analyses, we find that our method greatly improves tamper-resistance while preserving benign capabilities. Our results demonstrate that progress on tamper-resistance is possible, opening up a promising new avenue to improve the safety and security of open-weight LLMs. Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li 0026, Dan Hendrycks, Mantas Mazeika |
ICLR | 6 |
| 2025 | IterGen: Iterative Semantic-aware Structured LLM Generation with BacktrackingabstractLarge Language Models (LLMs) are widely used for tasks such as natural language and code generation, but their outputs often suffer from issues like hallucination, toxicity, and incorrect results. Current libraries for structured LLM generation rely on left-to-right decoding without support for backtracking, limiting the ability to correct or refine outputs mid-generation.
To address this, we introduce IterGen, a user-friendly library for iterative, grammar-guided LLM generation that enables users to move both forward and backward within the generated output based on grammar symbols.
By leveraging a symbol-to-position mapping and maintaining the key-value (KV) cache state, IterGen ensures efficient and structured generation while allowing for corrections during the process. We demonstrate IterGen's effectiveness in two important applications: reducing privacy leakage in LLM outputs, improving the accuracy of LLM-generated SQL and Vega-Lite queries.
Our code and additional resources are available at https://structuredllm.com. Shubham Ugare, Rohan Gumaste, Tarun Suresh, Gagandeep Singh 0001, Sasa Misailovic |
ICLR | 3 |
| 2025 | CRANE: Reasoning with constrained LLM generationabstractCode generation, symbolic math reasoning, and other tasks require LLMs to produce outputs that are both syntactically and semantically correct. Constrained LLM generation is a promising direction to enforce adherence to formal grammar, but prior works have empirically observed that strict enforcement of formal constraints often diminishes the reasoning capabilities of LLMs. In this work, we first provide a theoretical explanation for why constraining LLM outputs to very restrictive grammars that only allow syntactically valid final answers reduces the reasoning capabilities of the model. Second, we demonstrate that by augmenting the output grammar with carefully designed additional rules, it is always possible to preserve the reasoning capabilities of the LLM while ensuring syntactic and semantic correctness in its outputs. Building on these theoretical insights, we propose a reasoning-augmented constrained decoding algorithm, CRANE, which effectively balances the correctness of constrained generation with the flexibility of unconstrained generation.
Experiments on multiple open-source LLMs and benchmarks show that CRANE significantly outperforms both state-of-the-art
constrained decoding strategies and standard unconstrained decoding, showing up to 10% points accuracy improvement over baselines on challenging symbolic reasoning benchmarks GSM-symbolic and FOLIO. Debangshu Banerjee 0001, Tarun Suresh, Shubham Ugare, Sasa Misailovic, Gagandeep Singh 0001 |
ICML | 2 |
| 2025 | DINGO: Constrained Inference for Diffusion LLMsabstractDiffusion LLMs have emerged as a promising alternative to conventional autoregressive LLMs, offering substantial potential for improving runtime efficiency. However, existing diffusion models fail to provably enforce user-specified formal constraints, such as regular expressions, which makes them unreliable for tasks that require structured outputs, such as fixed-schema JSON generation. Unlike autoregressive models, which generate tokens sequentially, diffusion LLMs predict a block of tokens in parallel. This parallelism makes traditional constrained decoding algorithms, designed to enforce constraints with sequential token prediction, ineffective at preserving the true output distribution. To address this limitation, we propose DINGO, a dynamic programming-based constrained decoding strategy that is both efficient and provably distribution-preserving. DINGO enables sampling of output strings with the highest probability under the model’s predicted distribution while strictly adhering to any user-specified regular expression. On standard symbolic math and JSON generation benchmarks, DINGO achieves up to a $68$\% points of improvement over unconstrained inference. The code is available at [**DINGO**](https://github.com/uiuc-focal-lab/DINGO). Tarun Suresh, Debangshu Banerjee 0001, Shubham Ugare, Sasa Misailovic, Gagandeep Singh 0001 |
NeurIPS | 1 |
| 2024 | Incremental Randomized Smoothing CertificationabstractRandomized smoothing-based certification is an effective approach for obtaining robustness certificates of deep neural networks (DNNs) against adversarial attacks. This method constructs a smoothed DNN model and certifies its robustness through statistical sampling, but it is computationally expensive, especially when certifying with a large number of samples. Furthermore, when the smoothed model is modified (e.g., quantized or pruned), certification guarantees may not hold for the modified DNN, and recertifying from scratch can be prohibitively expensive.
We present the first approach for incremental robustness certification for randomized smoothing, IRS. We show how to reuse the certification guarantees for the original smoothed model to certify an approximated model with very few samples. IRS significantly reduces the computational cost of certifying modified DNNs while maintaining strong robustness guarantees. We experimentally demonstrate the effectiveness of our approach, showing up to 4.1x certification speedup over the certification that applies randomized smoothing of the approximate model from scratch. Shubham Ugare, Tarun Suresh, Debangshu Banerjee 0001, Gagandeep Singh 0001, Sasa Misailovic |
ICLR | 2 |
| 2024 | Relational Verification Leaps Forward with RABBitabstractWe propose RABBit, a Branch-and-Bound-based verifier for verifying relational properties defined over Deep Neural Networks, such as robustness against universal adversarial perturbations (UAP). Existing SOTA complete $L_{\infty}$-robustness verifiers can not reason about dependencies between multiple executions and, as a result, are imprecise for relational verification. In contrast, existing SOTA relational verifiers only apply a single bounding step and do not utilize any branching strategies to refine the obtained bounds, thus producing imprecise results. We develop the first scalable Branch-and-Bound-based relational verifier, RABBit, which efficiently combines branching over multiple executions with cross-executional bound refinement to utilize relational constraints, gaining substantial precision over SOTA baselines on a wide range of datasets and networks. Our code is at https://github.com/uiuc-focal-lab/RABBit. Tarun Suresh, Debangshu Banerjee 0001, Gagandeep Singh 0001 |
NeurIPS | 1 |