EDBT 2026 Demo / reviewers in the wild / expert
Abhijeet Awasthi
dblp:233/8164
· DBLP profile ↗
11ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 42% Information extraction and text analysis · 22% Knowledge representation and reasoning · 18% | |
| Databases, data mining, and information retrieval
2 papers |
Machine learning and data management · 54% Data integration and cleaning · 46% | |
| Software engineering, system software, and programming languages
3 papers |
Program synthesis and code generation · 100% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
case-based reasoning |
0.7 | 1 | 2023 | Structured Case-Based Reasoning for Inference-Time Adaptation of Text-to-SQL Parsers · AAAI 2023 |
Natural language and speech › Language models and text generation › large language model inference
inference-time adaptation |
0.7 | 1 | 2023 | Structured Case-Based Reasoning for Inference-Time Adaptation of Text-to-SQL Parsers · AAAI 2023 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.7 | 1 | 2023 | Structured Case-Based Reasoning for Inference-Time Adaptation of Text-to-SQL Parsers · AAAI 2023 |
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL |
0.7 | 1 | 2023 | Structured Case-Based Reasoning for Inference-Time Adaptation of Text-to-SQL Parsers · AAAI 2023 |
Program synthesis and code generation › code generation from natural language
text-to-SQL generation |
0.7 | 1 | 2023 | Conditional Tree Matching for Inference-Time Adaptation of Tree Prediction Models · ICML 2023 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
text-to-SQL parsing |
0.6 | 1 | 2022 | Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL Parsers · EMNLP 2022 |
Data integration and cleaning › data generation
synthetic data generation |
0.6 | 1 | 2022 | Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL Parsers · EMNLP 2022 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.5 | 1 | 2021 | Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages Study · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation
multilingual language models |
0.5 | 1 | 2021 | Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages Study · ACL/IJCNLP (1) 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
rule learning |
0.4 | 1 | 2020 | Learning from Rules Generalizing Labeled Exemplars · ICLR 2020 |
Program synthesis and code generation
rule learning |
0.4 | 1 | 2020 | Learning from Rules Generalizing Labeled Exemplars · ICLR 2020 |
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
sequence transduction |
0.4 | 1 | 2019 | Parallel Iterative Edit Models for Local Sequence Transduction · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation
code language models |
0.3 | 1 | 2025 | NextCoder: Robust Adaptation of Code LMs to Diverse Code Edits · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
model adaptation |
0.3 | 1 | 2025 | NextCoder: Robust Adaptation of Code LMs to Diverse Code Edits · ICML 2025 |
Natural language and speech › Language models and text generation › language modeling
low-resource language modeling |
0.1 | 1 | 2021 | Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages Study · ACL/IJCNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 2.9synthetic data generation · 1.7sparse projection · 1.7tree matching · 1.3sinkhorn algorithm · 1.3optimal transport · 1.3retrieve-and-edit · 1.1rule generalization · 0.9subtree similarity · 0.7sequence-to-sequence model · 0.7case-based reasoning · 0.7multilingual pretraining · 0.5language model adaptation · 0.5weak supervision · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NextCoder: Robust Adaptation of Code LMs to Diverse Code EditsabstractSoftware engineering activities frequently involve edits to existing code. However, contemporary code language models (LMs) lack the ability to handle diverse types of code-edit requirements. In this work, we attempt to overcome this shortcoming through (1) a novel synthetic data generation pipeline and (2) a robust model adaptation algorithm. Starting with seed code examples and diverse editing criteria, our pipeline generates high-quality samples comprising original and modified code, along with natural language instructions in different styles and verbosity. Today’s code LMs come bundled with strong abilities, such as code generation and instruction following, which should not be lost due to fine-tuning. To ensure this, we propose a novel adaptation algorithm, SeleKT, that (a) leverages a dense gradient-based step to identify the weights that are most important for code editing, and (b) does a sparse projection onto the base model to avoid overfitting. Using our approach, we obtain a new series of models NextCoder (adapted from QwenCoder-2.5) that achieves strong results on five code-editing benchmarks, outperforming comparable size models and even several larger ones. We show the generality of our approach on two model families DeepSeekCoder and QwenCoder), compare against other fine-tuning approaches, and demonstrate robustness by showing retention of code generation and general problem-solving abilities post adaptation. We opensource the models, synthetic dataset, and implementation at https://aka.ms/nextcoder. Tushar Aggarwal, Swayam Singh, Abhijeet Awasthi, Aditya Kanade 0001, Nagarajan Natarajan |
ICML | 3 |
| 2023 | Structured Case-Based Reasoning for Inference-Time Adaptation of Text-to-SQL ParsersabstractInference-time adaptation methods for semantic parsing are useful for leveraging examples from newly-observed domains without repeated fine-tuning. Existing approaches typically bias the decoder by simply concatenating input-output example pairs (cases) from the new domain at the encoder’s input in a Seq-to-Seq model. Such methods cannot adequately leverage the structure of logical forms in the case examples. We propose StructCBR, a structured case-based reasoning approach, which leverages subtree-level similarity between logical forms of cases and candidate outputs, resulting in better decoder decisions. For the task of adapting Text-to-SQL models to unseen schemas, we show that exploiting case examples in a structured manner via StructCBR offers consistent performance improvements over prior inference-time adaptation methods across five different databases. To the best of our knowledge, we are the first to attempt inference-time adaptation of Text-to-SQL models, and harness trainable structured similarity between subqueries. Abhijeet Awasthi, Soumen Chakrabarti, Sunita Sarawagi |
AAAI | 1 |
| 2023 | Bootstrapping Multilingual Semantic Parsers using Large Language ModelsabstractAbhijeet Awasthi, Nitish Gupta, Bidisha Samanta, Shachi Dave, Sunita Sarawagi, Partha Talukdar. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Abhijeet Awasthi, Nitish Gupta, Bidisha Samanta, Shachi Dave, Sunita Sarawagi, Partha Talukdar |
EACL | 1 |
| 2023 | Conditional Tree Matching for Inference-Time Adaptation of Tree Prediction ModelsabstractWe present CTreeOT, a convergent, differentiable algorithm for matching two trees when each tree is conditioned on some input. Such conditional tree matching is useful for light-weight, few-shot adaptation of tree prediction models without parameter fine-tuning. CTreeOT includes an alignment algorithm that extends the popular Sinkhorn algorithm for matching tree nodes while supporting constraints on tree edges. The algorithm involves alternating between matrix rescaling and message passing updates, and can be efficiently expressed as GPU tensor operations. The second part of CTreeOT is fine-grained relevance-based reweighting of nodes that makes the match scores useful for prediction tasks. We demonstrate the usefulness of CTreeOT for cross-schema adaptation of Text-to-SQL, a popular semantic parsing task. We show that compared to state-of-the-art methods, we achieve significant increase in adaptation accuracy. Harshit Varma, Abhijeet Awasthi, Sunita Sarawagi |
ICML | 2 |
| 2022 | Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL ParsersabstractText-to-SQL parsers typically struggle with databases unseen during the train time.Adapting parsers to new databases is a challenging problem due to the lack of natural language queries in the new schemas.We present REFILL, a framework for synthesizing highquality and textually diverse parallel datasets for adapting a Text-to-SQL parser to a target schema.REFILL learns to retrieve-andedit text queries from the existing schemas and transfers them to the target schema.We show that retrieving diverse existing text, masking their schema-specific tokens, and refilling with tokens relevant to the target schema, leads to significantly more diverse text queries than achievable by standard SQL-to-Text generation methods.Through experiments spanning multiple databases, we demonstrate that fine-tuning parsers on datasets synthesized using REFILL consistently outperforms the prior data-augmentation methods. Abhijeet Awasthi, Ashutosh Sathe, Sunita Sarawagi |
EMNLP | 1 |
| 2021 | Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages StudyabstractYash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi, Partha Talukdar, Sunita Sarawagi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi, Partha P. Talukdar, Sunita Sarawagi |
ACL/IJCNLP (1) | 4 |
| 2021 | Error-Driven Fixed-Budget ASR Personalization for Accented SpeakersabstractWe consider the task of personalizing ASR models while being constrained by a fixed budget on recording speaker specific utterances. Given a speaker and an ASR model, we propose a method of identifying sentences for which the speaker’s utterances are likely to be harder for the given ASR model to recognize. We assume a tiny amount of speaker-specific data to learn phoneme-level error models which help us select such sentences. We show that speaker’s utterances on the sentences selected using our error model indeed have larger error rates when compared to speaker’s utterances on randomly selected sentences. We find that fine-tuning the ASR model on the sentence utterances selected with the help of error models yield higher WER improvements in comparison to fine-tuning on an equal number of randomly selected sentence utterances. Thus, our method provides an efficient way of collecting speaker utterances under budget constraints for personalizing ASR models. Abhijeet Awasthi, Aman Kansal, Sunita Sarawagi, Preethi Jyothi |
ICASSP | 1 |
| 2021 | Teaching Keyword Spotters to Spot New Keywords with Limited ExamplesabstractLearning to recognize new keywords with just a few examples is essential for personalizing keyword spotting (KWS) models to a user's choice of keywords. However, modern KWS models are typically trained on large datasets and restricted to a small vocabulary of keywords, limiting their transferability to a broad range of unseen keywords. Towards easily customizable KWS models, we present KeySEM (Keyword Speech EMbedding), a speech embedding model pre-trained on the task of recognizing a large number of keywords. Speech representations offered by KeySEM are highly effective for learning new keywords from a limited number of examples. Comparisons with a diverse range of related work across several datasets show that our method achieves consistently superior performance with fewer training examples. Although KeySEM was pre-trained only on English utterances, the performance gains also extend to datasets from four other languages indicating that KeySEM learns useful representations well aligned with the task of keyword spotting. Finally, we demonstrate KeySEM's ability to learn new keywords sequentially without requiring to re-train on previously learned keywords. Our experimental observations suggest that KeySEM is well suited to on-device environments where post-deployment learning and ease of customization are often desirable. Abhijeet Awasthi, Kevin Kilgour, Hassan Rom |
Interspeech | 1 |
| 2020 | Learning from Rules Generalizing Labeled Exemplars
Abhijeet Awasthi, Sabyasachi Ghosh, Rasna Goyal, Sunita Sarawagi |
ICLR | 1 |
| 2020 | Black-Box Adaptation of ASR for Accented SpeechabstractWe introduce the problem of adapting a black-box, cloud-based ASR system to speech from a target accent. While leading online ASR services obtain impressive performance on main-stream accents, they perform poorly on sub-populations - we observed that the word error rate (WER) achieved by Google's ASR API on Indian accents is almost twice the WER on US accents. Existing adaptation methods either require access to model parameters or overlay an error-correcting module on output transcripts. We highlight the need for correlating outputs with the original speech to fix accent errors. Accordingly, we propose a novel coupling of an open-source accent-tuned local model with the black-box service where the output from the service guides frame-level inference in the local model. Our fine-grained merging algorithm is better at fixing accent errors than existing word-level combination strategies. Experiments on Indian and Australian accents with three leading ASR models as service, show that we achieve as much as 28% relative reduction in WER over both the local and service models. Kartik Khandelwal, Preethi Jyothi, Abhijeet Awasthi, Sunita Sarawagi |
INTERSPEECH | 3 |
| 2019 | Parallel Iterative Edit Models for Local Sequence TransductionabstractAbhijeet Awasthi, Sunita Sarawagi, Rasna Goyal, Sabyasachi Ghosh, Vihari Piratla. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Abhijeet Awasthi, Sunita Sarawagi, Rasna Goyal, Sabyasachi Ghosh, Vihari Piratla |
EMNLP/IJCNLP (1) | 1 |