Ashutosh Sathe

dblp:332/0994 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 59% Transfer learning and domain adaptation · 15% Representation and self-supervised learning · 15%
Databases, data mining, and information retrieval
2 papers
Data integration and cleaning · 74% Data models and query languages · 26%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.912025
Improving Cross Lingual Transfer by Pretraining with Active Forgetting · EMNLP 2025
Natural language and speech › Language models and text generation
multilingual language models
0.912025
Improving Cross Lingual Transfer by Pretraining with Active Forgetting · EMNLP 2025
Machine learning › Representation and self-supervised learning
pre-training
0.912025
Improving Cross Lingual Transfer by Pretraining with Active Forgetting · EMNLP 2025
Natural language and speech › Language models and text generation › natural language understanding
ambiguity handling
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Language models and text generation › decoding
constrained decoding
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Language models and text generation
decoding
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Language models and text generation › natural language understanding › question answering
text-to-SQL parsing
0.612022
Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL Parsers · EMNLP 2022
Data integration and cleaning › data generation
synthetic data generation
0.612022
Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL Parsers · EMNLP 2022
Data models and query languages › SQL
SQL query generation
0.212023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

plan-based template generation · 1.3constrained infilling · 1.3beam search · 1.3retrieve-and-edit · 1.1fine-tuning · 1.1pre-training · 0.9active forgetting · 0.9
YearPublicationVenuePosition
2025 Improving Cross Lingual Transfer by Pretraining with Active Forgetting
abstract
Large Language Models (LLMs) demonstrate exceptional capabilities in a multitude of NLP tasks. However, the efficacy of such models to languages other than English is often limited. Prior works have shown that encoder-only models such as BERT or XLM-RoBERTa show impressive cross lingual transfer of their capabilities from English to other languages. In this work, we propose a pretraining strategy that uses active forgetting to achieve similar cross lingual transfer in decoder-only LLMs. We show that LLMs pretrained with active forgetting are highly effective when adapting to new and unseen languages. Through extensive experimentation, we find that LLMs pretrained with active forgetting are able to learn better multilingual representations which translates to better performance in many downstream tasks.
Divyanshu Aggarwal, Ashutosh Sathe, Sunayana Sitaram
EMNLP2
2024 MAFIA: Multi-Adapter Fused Inclusive Language Models
abstract
Prachi Jain, Ashutosh Sathe, Varun Gumma, Kabir Ahuja, Sunayana Sitaram. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ashutosh Sathe, Varun Gumma, Kabir Ahuja, Sunayana Sitaram
EACL (1)2
2024 MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
abstract
Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Mohamed Ahmed, Kalika Bali, Sunayana Sitaram. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Kalika Bali, Sunayana Sitaram
NAACL-HLT5
2023 Benchmarking and Improving Text-to-SQL Generation under Ambiguity
abstract
Research in Text-to-SQL conversion has been largely benchmarked against datasets where each text query corresponds to one correct SQL.However, natural language queries over reallife databases frequently involve significant ambiguity about the intended SQL due to overlapping schema names and multiple confusing relationship paths.To bridge this gap, we develop a novel benchmark called AmbiQT with over 3000 examples where each text is interpretable as two plausible SQLs due to lexical and/or structural ambiguity.When faced with ambiguity, an ideal top-k decoder should generate all valid interpretations for possible disambiguation by the user (Elgohary et al., 2021;Zhong et al., 2022).We evaluate several Text-to-SQL systems and decoding algorithms, including those employing state-of-the-art LLMs, and find them to be far from this ideal.The primary reason is that the prevalent beam search algorithm and its variants, treat SQL queries as a string and produce unhelpful token-level diversity in the top-k.We propose LogicalBeam, a new decoding algorithm that navigates the SQL logic space using a blend of plan-based template generation and constrained infilling.Counterfactually generated plans diversify templates while in-filling with a beam-search, that branches solely on schema names, provides value diversity.Log-icalBeam is up to 2.5× more effective than state-of-the-art models at generating all candidate SQLs in the top-k ranked outputs.It also enhances the top-5 Exact and Execution Match Accuracies on SPIDER and Kaggle DBQA 1 .
Adithya Bhaskar, Tushar Tomar, Ashutosh Sathe, Sunita Sarawagi
EMNLP3
2022 Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL Parsers
abstract
Text-to-SQL parsers typically struggle with databases unseen during the train time.Adapting parsers to new databases is a challenging problem due to the lack of natural language queries in the new schemas.We present REFILL, a framework for synthesizing highquality and textually diverse parallel datasets for adapting a Text-to-SQL parser to a target schema.REFILL learns to retrieve-andedit text queries from the existing schemas and transfers them to the target schema.We show that retrieving diverse existing text, masking their schema-specific tokens, and refilling with tokens relevant to the target schema, leads to significantly more diverse text queries than achievable by standard SQL-to-Text generation methods.Through experiments spanning multiple databases, we demonstrate that fine-tuning parsers on datasets synthesized using REFILL consistently outperforms the prior data-augmentation methods.
Abhijeet Awasthi, Ashutosh Sathe, Sunita Sarawagi
EMNLP2