Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tushar Tomar

dblp:339/3377 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 52% Machine translation · 30% Information extraction and text analysis · 17%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › natural language understanding
ambiguity handling
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Language models and text generation › decoding
constrained decoding
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Language models and text generation
decoding
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.712023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
0.612022
Quality Scoring of Source Words in Neural Translation Models · EMNLP 2022
Natural language and speech › Machine translation › machine translation evaluation › translation quality estimation
word-level quality estimation
0.612022
Quality Scoring of Source Words in Neural Translation Models · EMNLP 2022
Data models and query languages › SQL
SQL query generation
0.212023
Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

plan-based template generation · 1.3constrained infilling · 1.3beam search · 1.3language model probability comparison · 0.6
YearPublicationVenuePosition
2023 Benchmarking and Improving Text-to-SQL Generation under Ambiguity
abstract
Research in Text-to-SQL conversion has been largely benchmarked against datasets where each text query corresponds to one correct SQL.However, natural language queries over reallife databases frequently involve significant ambiguity about the intended SQL due to overlapping schema names and multiple confusing relationship paths.To bridge this gap, we develop a novel benchmark called AmbiQT with over 3000 examples where each text is interpretable as two plausible SQLs due to lexical and/or structural ambiguity.When faced with ambiguity, an ideal top-k decoder should generate all valid interpretations for possible disambiguation by the user (Elgohary et al., 2021;Zhong et al., 2022).We evaluate several Text-to-SQL systems and decoding algorithms, including those employing state-of-the-art LLMs, and find them to be far from this ideal.The primary reason is that the prevalent beam search algorithm and its variants, treat SQL queries as a string and produce unhelpful token-level diversity in the top-k.We propose LogicalBeam, a new decoding algorithm that navigates the SQL logic space using a blend of plan-based template generation and constrained infilling.Counterfactually generated plans diversify templates while in-filling with a beam-search, that branches solely on schema names, provides value diversity.Log-icalBeam is up to 2.5× more effective than state-of-the-art models at generating all candidate SQLs in the top-k ranked outputs.It also enhances the top-5 Exact and Execution Match Accuracies on SPIDER and Kaggle DBQA 1 .
Adithya Bhaskar, Tushar Tomar, Ashutosh Sathe, Sunita Sarawagi
EMNLP2
2022 Quality Scoring of Source Words in Neural Translation Models
abstract
Word-level quality scores on input source sentences can provide useful feedback to an enduser when translating into an unfamiliar target language.Recent approaches either require training custom models on synthetic data or repeatedly invoking the translation model.We propose a simple approach based on comparing probabilities from two language models.The basic premise of our method is to reason how well each source word is explained by the generated translation as against the preceding source language words.Our approach provides between 2.2 and 27.1 higher F1 score and is significantly faster than state of the art methods on three language pairs.Also, our method does not require training any new model.We release a public dataset on word omissions and mistranslations on a new language pair. 1
Priyesh Jain, Sunita Sarawagi, Tushar Tomar
EMNLP3