VLDB 2026 Research / reviewers in the wild / expert
Tushar Tomar
dblp:339/3377
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 52% Machine translation · 30% Information extraction and text analysis · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Data models and query languages · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › natural language understanding
ambiguity handling |
0.7 | 1 | 2023 | Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023 |
Natural language and speech › Language models and text generation › decoding
constrained decoding |
0.7 | 1 | 2023 | Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023 |
Natural language and speech › Language models and text generation
decoding |
0.7 | 1 | 2023 | Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL |
0.7 | 1 | 2023 | Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023 |
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation |
0.6 | 1 | 2022 | Quality Scoring of Source Words in Neural Translation Models · EMNLP 2022 |
Natural language and speech › Machine translation › machine translation evaluation › translation quality estimation
word-level quality estimation |
0.6 | 1 | 2022 | Quality Scoring of Source Words in Neural Translation Models · EMNLP 2022 |
Data models and query languages › SQL
SQL query generation |
0.2 | 1 | 2023 | Benchmarking and Improving Text-to-SQL Generation under Ambiguity · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
plan-based template generation · 1.3constrained infilling · 1.3beam search · 1.3language model probability comparison · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Benchmarking and Improving Text-to-SQL Generation under AmbiguityabstractResearch in Text-to-SQL conversion has been largely benchmarked against datasets where each text query corresponds to one correct SQL.However, natural language queries over reallife databases frequently involve significant ambiguity about the intended SQL due to overlapping schema names and multiple confusing relationship paths.To bridge this gap, we develop a novel benchmark called AmbiQT with over 3000 examples where each text is interpretable as two plausible SQLs due to lexical and/or structural ambiguity.When faced with ambiguity, an ideal top-k decoder should generate all valid interpretations for possible disambiguation by the user (Elgohary et al., 2021;Zhong et al., 2022).We evaluate several Text-to-SQL systems and decoding algorithms, including those employing state-of-the-art LLMs, and find them to be far from this ideal.The primary reason is that the prevalent beam search algorithm and its variants, treat SQL queries as a string and produce unhelpful token-level diversity in the top-k.We propose LogicalBeam, a new decoding algorithm that navigates the SQL logic space using a blend of plan-based template generation and constrained infilling.Counterfactually generated plans diversify templates while in-filling with a beam-search, that branches solely on schema names, provides value diversity.Log-icalBeam is up to 2.5× more effective than state-of-the-art models at generating all candidate SQLs in the top-k ranked outputs.It also enhances the top-5 Exact and Execution Match Accuracies on SPIDER and Kaggle DBQA 1 . Adithya Bhaskar, Tushar Tomar, Ashutosh Sathe, Sunita Sarawagi |
EMNLP | 2 |
| 2022 | Quality Scoring of Source Words in Neural Translation ModelsabstractWord-level quality scores on input source sentences can provide useful feedback to an enduser when translating into an unfamiliar target language.Recent approaches either require training custom models on synthetic data or repeatedly invoking the translation model.We propose a simple approach based on comparing probabilities from two language models.The basic premise of our method is to reason how well each source word is explained by the generated translation as against the preceding source language words.Our approach provides between 2.2 and 27.1 higher F1 score and is significantly faster than state of the art methods on three language pairs.Also, our method does not require training any new model.We release a public dataset on word omissions and mistranslations on a new language pair. 1 Priyesh Jain, Sunita Sarawagi, Tushar Tomar |
EMNLP | 3 |