EDBT 2026 Demo / reviewers in the wild / expert
Pranjal Aggarwal
dblp:163/0764
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-2962-1535ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 68% Efficient and distributed learning · 15% Trustworthy machine learning · 13% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 78% Data mining · 22% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | Agentic-R1: Distilled Dual-Strategy Reasoning · EMNLP 2025 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.9 | 1 | 2025 | Agentic-R1: Distilled Dual-Strategy Reasoning · EMNLP 2025 |
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models |
0.9 | 1 | 2025 | Agentic-R1: Distilled Dual-Strategy Reasoning · EMNLP 2025 |
Program synthesis and code generation › formal synthesis
verified code generation |
0.9 | 1 | 2025 | AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement · ICML 2025 |
Natural language and speech › Language models and text generation
model routing |
0.8 | 1 | 2024 | AutoMix: Automatically Mixing Language Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › verification
self-verification |
0.8 | 1 | 2024 | AutoMix: Automatically Mixing Language Models · NeurIPS 2024 |
Information retrieval
generative engine optimization |
0.8 | 1 | 2024 | GEO: Generative Engine Optimization · KDD 2024 |
Information retrieval
retrieval models |
0.8 | 1 | 2024 | GEO: Generative Engine Optimization · KDD 2024 |
Information retrieval
search engines |
0.8 | 1 | 2024 | GEO: Generative Engine Optimization · KDD 2024 |
Natural language and speech › Language models and text generation
test-time scaling |
0.7 | 1 | 2023 | Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs · EMNLP 2023 |
Data mining › predictive modeling › classification › multi-label classification
extreme classification |
0.7 | 1 | 2023 | SemSup-XC: Semantic Supervision for Zero and Few-shot Extreme Classification · ICML 2023 |
Program synthesis and code generation
code generation with language models |
0.3 | 1 | 2025 | AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement · ICML 2025 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2023 | SemSup-XC: Semantic Supervision for Zero and Few-shot Extreme Classification · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.6semantic class descriptions · 1.3hybrid matching · 1.3contrastive learning · 1.3tree search · 0.9tool use · 0.9self-improvement · 0.9knowledge distillation · 0.9few-shot self-verification · 0.8black-box optimization · 0.8POMDP · 0.8stopping criterion · 0.7self-consistency · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Agentic-R1: Distilled Dual-Strategy ReasoningabstractCurrent long chain-of-thought (long-CoT) models excel at mathematical reasoning but rely on slow and error-prone natural language traces.Tool-augmented agents address arithmetic via code execution, but often falter on complex logical tasks.We introduce a fine-tuning framework, DualDistill, that distills complementary reasoning strategies from multiple teachers into a unified student model.Using this approach, we train Agentic-R1, which dynamically selects the optimal strategy for each query, invoking tools for arithmetic and algorithmic problems, and using text-based reasoning for abstract ones.Our method improves accuracy across a range of tasks, including both computation-intensive and standard benchmarks, demonstrating the effectiveness of multi-strategy distillation in achieving robust and efficient reasoning.Our project is available at https://github.com/StigLidu/DualDistill. Weihua Du, Pranjal Aggarwal, Sean Welleck, Yiming Yang 0002 |
EMNLP | 2 |
| 2025 | AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and TreefinementabstractAutomated code generation with large language models has gained significant traction, but there remains no guarantee of the correctness of generated code. We aim to use formal verification to provide mathematical guarantees that the generated code is correct. However, generating formally verified code with LLMs is hindered by the scarcity of training data and the complexity of formal proofs. To tackle this challenge, we introduce AlphaVerus, a self-improving framework that bootstraps formally verified code generation by iteratively translating programs from a higher-resource language and leveraging feedback from a verifier. AlphaVerus operates in three phases: exploration of candidate translations, Treefinement -- a novel tree search algorithm for program refinement using verifier feedback, and filtering misaligned specifications and programs to prevent reward hacking. Through this iterative process, AlphaVerus enables the LLaMA-3.1-70B model to generate verified code without human intervention or model finetuning. AlphaVerus shows an ability to generate formally verified solutions for HumanEval and MBPP, laying the groundwork for truly trustworthy code-generation agents. Pranjal Aggarwal, Bryan Parno, Sean Welleck |
ICML | 1 |
| 2025 | Investigating the Reasoning Abilities of Large Language Models for Understanding Spoken Language in Interpersonal Interactions
Pranjal Aggarwal, Ghritachi Mahajani, Pavan Kumar Malasani, Vaibhav Jamadagni, Caroline J. Wendt, Ehsanul Haque Nirjhar, Theodora Chaspari |
INTERSPEECH | 1 |
| 2025 | A New Genetic Algorithm-Based Network for Text Localization in Degraded Social Media ImagesabstractABSTRACT This paper presents a novel model for understanding social image content through text localization. For text localization, we explore maximally stable extremal regions (MSER) for detecting components that work by clustering pixels with similar properties. The output of component detection includes several non‐text components due to the degradations of social media images. To select the best components among many, we explore the genetic algorithm by convolving different kernels with components, which results in a feature matrix that is further fed to EfficientNet for choosing actual text components. Therefore, the proposed model is called genetic algorithm based network for text localization in degraded social media images (TLDSMI). For evaluating text localization, we consider the images of the standard dataset of natural scenes by uploading and downloading from different social media platforms, namely, WhatsApp, Telegram, and Instagram. The effectiveness of our method is shown by testing on original and degraded standard datasets. For example, for the degraded images of different complexities including degradations caused by social media platforms, the proposed method performs well in almost all situations. In addition, the proposed model achieves the best F1‐Score, 0.76, 0.77, 0.70, and 0.78 for the degraded images of CUTE, ICDAR 2013, Total‐Text, and CTW1500, respectively, compared to the state‐of‐the‐art methods. Palaiahnakote Shivakumara, C. Pavan Kumar 0001, Pranjal Aggarwal, Pasupuleti Chandana, M. Basavanna, Umapada Pal 0001 |
IET Image Process. | 3 |
| 2024 | GEO: Generative Engine OptimizationabstractThe advent of large language models (LLMs) has ushered in a new paradigm of search engines that use generative models to gather and summarize information to answer user queries. This emerging technology, which we formalize under the unified framework of generative engines (GEs), can generate accurate and personalized responses, rapidly replacing traditional search engines like Google and Bing. Generative Engines typically satisfy queries by synthesizing information from multiple sources and summarizing them using LLMs. While this shift significantly improvesuser utility and generative search engine traffic, it poses a huge challenge for the third stakeholder -- website and content creators. Given the black-box and fast-moving nature of generative engines, content creators have little to no control over when and how their content is displayed. With generative engines here to stay, we must ensure the creator economy is not disadvantaged. To address this, we introduce Generative Engine Optimization (GEO), the first novel paradigm to aid content creators in improving their content visibility in generative engine responses through a flexible black-box optimization framework for optimizing and defining visibility metrics. We facilitate systematic evaluation by introducing GEO-bench, a large-scale benchmark of diverse user queries across multiple domains, along with relevant web sources to answer these queries. Through rigorous evaluation, we demonstrate that GEO can boost visibility by up to 40% in generative engine responses. Moreover, we show the efficacy of these strategies varies across domains, underscoring the need for domain-specific optimization methods. Our work opens a new frontier in information discovery systems, with profound implications for both developers of generative engines and content creators. Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande |
KDD | 1 |
| 2024 | AutoMix: Automatically Mixing Language ModelsabstractLarge language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively leveraging the options to optimize computational cost and performance remains challenging. In this work, we present AutoMix, an approach that strategically routes queries to larger LMs, based on the approximate correctness of outputs from a smaller LM. Central to AutoMix are two key technical contributions. First, it has a few-shot self-verification mechanism, which estimates the reliability of its own outputs without requiring extensive training. Second, given that self-verification can be noisy, it employs a POMDP based router that can effectively select an appropriately sized model, based on answer confidence. Experiments across five language models and five challenging datasets show that Automix consistently surpasses strong baselines, reducing computational cost by over 50\% for comparable performance. Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Aditya Gupta 0001, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang 0002, Shyam Upadhyay, Manaal Faruqui, Mausam |
NeurIPS | 1 |
| 2023 | Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMsabstractA popular approach for improving the correctness of output from large language models (LLMs) is Self-Consistency -poll the LLM multiple times and output the most frequent solution.Existing Self-Consistency techniques always generate a constant number of samples per question, where a better approach will be to non-uniformly distribute the available budget based on the amount of agreement in the samples generated so far.In response, we introduce Adaptive-Consistency, a cost-efficient, model-agnostic technique that dynamically adjusts the number of samples per question using a lightweight stopping criterion.Our experiments over 17 reasoning and code generation datasets and three LLMs demonstrate that Adaptive-Consistency reduces sample budget by up to 7.9 times with an average accuracy drop of less than 0.1%. 1 Pranjal Aggarwal, Aman Madaan, Yiming Yang 0002, Mausam |
EMNLP | 1 |
| 2023 | SemSup-XC: Semantic Supervision for Zero and Few-shot Extreme ClassificationabstractExtreme classification (XC) involves predicting over large numbers of classes (thousands to millions), with real-world applications like news article classification and e-commerce product tagging. The zero-shot version of this task requires generalization to novel classes without additional supervision. In this paper, we develop SemSup-XC, a model that achieves state-of-the-art zero-shot and few-shot performance on three XC datasets derived from legal, e-commerce, and Wikipedia data. To develop SemSup-XC, we use automatically collected semantic class descriptions to represent classes and facilitate generalization through a novel hybrid matching module that matches input instances to class descriptions using a combination of semantic and lexical similarity. Trained with contrastive learning, SemSup-XC significantly outperforms baselines and establishes state-of-the-art performance on all three datasets considered, gaining up to 12 precision points on zero-shot and more than 10 precision points on one-shot tests, with similar gains for recall@10. Our ablation studies highlight the relative importance of our hybrid matching module and automatically collected class descriptions. Pranjal Aggarwal, Ameet Deshpande, Karthik Narasimhan |
ICML | 1 |