VLDB 2026 Research / reviewers in the wild / expert
Varun Gangal
dblp:178/8576 · also Varun Prashant Gangal
· DBLP profile ↗
14ranked-venue papers
4as first author
7since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Information extraction and text analysis · 30% Language models and text generation · 30% Question answering and dialogue systems · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 100% |
Topics — the 25 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.9 | 2 | 2022 | Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models · AAAI 2022 Detecting and Explaining Causes From Text For a Time Series Event · EMNLP 2017 |
Natural language and speech › Language models and text generation › text generation › knowledge-grounded generation
commonsense generation |
0.6 | 1 | 2022 | Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models · AAAI 2022 |
Natural language and speech › Language models and text generation › text generation
story generation |
0.6 | 1 | 2022 | NAREOR: The Narrative Reordering Problem · AAAI 2022 |
Natural language and speech › Language models and text generation
text generation |
0.6 | 1 | 2022 | Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models · AAAI 2022 |
Computer vision › Vision and language
visual grounding |
0.6 | 1 | 2022 | Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models · AAAI 2022 |
Computer vision › Image recognition and object detection › image classification › hierarchical classification
coarse-to-fine classification |
0.5 | 1 | 2021 | Coarse2Fine: Fine-grained Text Classification on Coarsely-grained Annotated Data · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › text classification
fine-grained text classification |
0.5 | 1 | 2021 | Coarse2Fine: Fine-grained Text Classification on Coarsely-grained Annotated Data · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
text classification |
0.5 | 1 | 2021 | Coarse2Fine: Fine-grained Text Classification on Coarsely-grained Annotated Data · EMNLP (1) 2021 |
Machine learning › Learning paradigms
weakly supervised learning |
0.5 | 1 | 2021 | Coarse2Fine: Fine-grained Text Classification on Coarsely-grained Annotated Data · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › natural language understanding
cloze task |
0.4 | 1 | 2020 | SCDE: Sentence Cloze Dataset with High Quality Distractors From Examinations · ACL 2020 |
Natural language and speech › Information extraction and text analysis › discourse analysis › discourse processing
discourse interpretation |
0.4 | 1 | 2020 | SCDE: Sentence Cloze Dataset with High Quality Distractors From Examinations · ACL 2020 |
Machine learning › Generative modeling › generative model
generative classifier |
0.4 | 1 | 2020 | Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection in Task Oriented Dialog · AAAI 2020 |
Natural language and speech › Language models and text generation
natural language understanding |
0.4 | 1 | 2020 | SCDE: Sentence Cloze Dataset with High Quality Distractors From Examinations · ACL 2020 |
Natural language and speech › Question answering and dialogue systems › intent detection
out-of-domain detection |
0.4 | 1 | 2020 | Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection in Task Oriented Dialog · AAAI 2020 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.4 | 1 | 2020 | Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection in Task Oriented Dialog · AAAI 2020 |
Natural language and speech › Information extraction and text analysis › text mining
stylometry |
0.4 | 1 | 2019 | (Male, Bachelor) and (Female, Ph.D) have different connotations: Parallelly Annotated Stylistic Language Dataset with Multiple Personas · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis › data annotation
text annotation |
0.4 | 1 | 2019 | (Male, Bachelor) and (Female, Ph.D) have different connotations: Parallelly Annotated Stylistic Language Dataset with Multiple Personas · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation › text generation › domain-specific text generation
commentary generation |
0.3 | 1 | 2018 | Learning to Generate Move-by-Move Commentary for Chess Games from Large-Scale Social Forum Data · ACL (1) 2018 |
Machine learning › Generative modeling
multimodal generation |
0.3 | 1 | 2018 | Learning to Generate Move-by-Move Commentary for Chess Games from Large-Scale Social Forum Data · ACL (1) 2018 |
Natural language and speech › Information extraction and text analysis › relation extraction › event relation extraction
causal relation extraction |
0.3 | 1 | 2017 | Detecting and Explaining Causes From Text For a Time Series Event · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis › computational morphology
word formation |
0.3 | 1 | 2017 | Charmanteau: Character Embedding Models For Portmanteau Creation · EMNLP 2017 |
Web and social media mining › social network analysis
centrality measures |
0.2 | 1 | 2016 | Trust and Distrust Across Coalitions: Shapley Value Based Centrality Measures for Signed Networks (Student Abstract Version) · AAAI 2016 |
Web and social media mining › social network analysis
signed network analysis |
0.2 | 1 | 2016 | Trust and Distrust Across Coalitions: Shapley Value Based Centrality Measures for Signed Networks (Student Abstract Version) · AAAI 2016 |
Algorithmic game theory and mechanism design › cooperative game theory › solution concepts
shapley value |
0.2 | 1 | 2016 | Trust and Distrust Across Coalitions: Shapley Value Based Centrality Measures for Signed Networks (Student Abstract Version) · AAAI 2016 |
Natural language and speech › Language models and text generation › text generation › data-to-text generation
concept-to-text generation |
0.2 | 1 | 2022 | Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
t5 · 1.1BART · 1.1shapley value · 0.8transformer · 0.6sequence-to-sequence fine-tuning · 0.6image captioning · 0.6label-conditioned fine-tuning · 0.5iterative weak supervision · 0.5external knowledge resource · 0.5data augmentation · 0.5sentence cloze dataset · 0.4generative classifier · 0.4distractor generation · 0.4annotation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | PANCETTA: Phoneme Aware Neural Completion to Elicit Tongue Twisters AutomaticallyabstractTongue twisters are meaningful sentences that are difficult to pronounce.The process of automatically generating tongue twisters is challenging since the generated utterance must satisfy two conditions at once: phonetic difficulty and semantic meaning.Furthermore, phonetic difficulty is itself hard to characterize and is expressed in tongue twisters through a heterogeneous mix of phenomena such as alliteration and homophony.In this paper, we propose PANCETTA: Phoneme Aware Neural Completion to Elicit Tongue Twisters Automatically.We leverage phoneme representations to capture the notion of phonetic difficulty, and we train language models to generate original tongue twisters on two proposed task settings.To do this, we curate a dataset called TT-Corp, consisting of existing English tongue twisters.Through automatic and human evaluation, as well as qualitative analysis, we show that PANCETTA generates novel, phonetically difficult, fluent, and semantically meaningful tongue twisters. Sedrick Keh, Steven Y. Feng, Varun Gangal, Malihe Alikhani, Eduard H. Hovy |
EACL | 3 |
| 2022 | Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation ModelsabstractWe investigate the use of multimodal information contained in images as an effective method for enhancing the commonsense of Transformer models for text generation. We perform experiments using BART and T5 on concept-to-text generation, specifically the task of generative commonsense reasoning, or CommonGen. We call our approach VisCTG: Visually Grounded Concept-to-Text Generation. VisCTG involves captioning images representing appropriate everyday scenarios, and using these captions to enrich and steer the generation process. Comprehensive evaluation and analysis demonstrate that VisCTG noticeably improves model performance while successfully addressing several issues of the baseline generations, including poor commonsense, fluency, and specificity. Steven Y. Feng, Zhuofu Tao, Malihe Alikhani, Teruko Mitamura, Eduard H. Hovy, Varun Gangal |
AAAI | 7 |
| 2022 | NAREOR: The Narrative Reordering ProblemabstractMany implicit inferences exist in text depending on how it is structured that can critically impact the text's interpretation and meaning. One such structural aspect present in text with chronology is the order of its presentation. For narratives or stories, this is known as the narrative order. Reordering a narrative can impact the temporal, causal, event-based, and other inferences readers draw from it, which in turn can have strong effects both on its interpretation and interestingness. In this paper, we propose and investigate the task of Narrative Reordering (NAREOR) which involves rewriting a given story in a different narrative order while preserving its plot. We present a dataset, NAREORC, with human rewritings of stories within ROCStories in non-linear orders, and conduct a detailed analysis of it. Further, we propose novel task-specific training methods with suitable evaluation metrics. We perform experiments on NAREORC using state-of-the-art models such as BART and T5 and conduct extensive automatic and human evaluations. We demonstrate that although our models can perform decently, NAREOR is a challenging task with potential for further exploration. We also investigate two applications of NAREOR: generation of more interesting variations of stories and serving as adversarial sets for temporal/event-related tasks, besides discussing other prospective ones, such as for pedagogical setups related to language skills like essay writing and applications to medicine involving clinical narratives. Varun Gangal, Steven Y. Feng, Malihe Alikhani, Teruko Mitamura, Eduard H. Hovy |
AAAI | 1 |
| 2022 | PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification Data for Learning Enhanced GenerationabstractA personification is a figure of speech that endows inanimate entities with properties and actions typically seen as requiring animacy. In this paper, we explore the task of personification generation. To this end, we propose PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification data for Learning Enhanced generation. We curate a corpus of personifications called PersonifCorp, together with automatically generated de-personified literalizations of these personifications. We demonstrate the usefulness of this parallel corpus by training a seq2seq model to personify a given literal input. Both automatic and human evaluations show that fine-tuning with PersonifCorp leads to significant gains in personification-related qualities such as animacy and interestingness. A detailed qualitative analysis also highlights key strengths and imperfections of PINEAPPLE over baselines, demonstrating a strong ability to generate diverse and creative personifications that enhance the overall appeal of a sentence. Sedrick Keh, Varun Gangal, Steven Y. Feng, Harsh Jhamtani, Malihe Alikhani, Eduard H. Hovy |
COLING | 3 |
| 2021 | Investigating Robustness of Dialog Models to Popular Figurative Language ConstructsabstractHumans often employ figurative language use in communication, including during interactions with dialog systems.Thus, it is important for real-world dialog systems to be able to handle popular figurative language constructs like metaphor and simile.In this work, we analyze the performance of existing dialog models in situations where the input dialog context exhibits use of figurative language.We observe large gaps in handling of figurative language when evaluating the models on two open domain dialog datasets.When faced with dialog contexts consisting of figurative language, some models show very large drops in performance compared to contexts without figurative language.We encourage future research in dialog modeling to separately analyze and report results on figurative language in order to better test model capabilities relevant to real-world use.Finally, we propose lightweight solutions to help existing models become more robust to figurative language by simply using an external resource to translate figurative language to literal (non-figurative) forms while preserving the meaning to the best extent possible. Harsh Jhamtani, Varun Gangal, Eduard H. Hovy, Taylor Berg-Kirkpatrick |
EMNLP (1) | 2 |
| 2021 | Coarse2Fine: Fine-grained Text Classification on Coarsely-grained Annotated DataabstractExisting text classification methods mainly focus on a fixed label set, whereas many realworld applications require extending to new fine-grained classes as the number of samples per label increases.To accommodate such requirements, we introduce a new problem called coarse-to-fine grained classification, which aims to perform fine-grained classification on coarsely annotated data.Instead of asking for new fine-grained human annotations, we opt to leverage label surface names as the only human guidance and weave in rich pretrained generative language models into the iterative weak supervision strategy.Specifically, we first propose a label-conditioned finetuning formulation to attune these generators for our task.Furthermore, we devise a regularization objective based on the coarse-fine label constraints derived from our problem setting, giving us even further improvements over the prior formulation.Our framework uses the fine-tuned generative models to sample pseudo-training data for training the classifier, and bootstraps on real unlabeled data for model refinement.Extensive experiments and case studies on two real-world datasets demonstrate superior performance over SOTA zeroshot classification baselines. Dheeraj Mekala, Varun Gangal, Jingbo Shang |
EMNLP (1) | 2 |
| 2021 | SAPPHIRE: Approaches for Enhanced Concept-to-Text GenerationabstractWe motivate and propose a suite of simple but effective improvements for concept-to-text generation called SAPPHIRE: Set Augmentation and Post-hoc PHrase Infilling and REcombination.We demonstrate their effectiveness on generative commonsense reasoning, a.k.a. the CommonGen task, through experiments using both BART and T5 models.Through extensive automatic and human evaluation, we show that SAPPHIRE noticeably improves model performance.An in-depth qualitative analysis illustrates that SAPPHIRE effectively addresses many issues of the baseline model generations, including lack of commonsense, insufficient specificity, and poor fluency. Steven Y. Feng, Jessica Huynh, Chaitanya Narisetty, Eduard H. Hovy, Varun Gangal |
INLG | 5 |
| 2020 | Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection in Task Oriented DialogabstractThe task of identifying out-of-domain (OOD) input examples directly at test-time has seen renewed interest recently due to increased real world deployment of models. In this work, we focus on OOD detection for natural language sentence inputs to task-based dialog systems. Our findings are three-fold:First, we curate and release ROSTD (Real Out-of-Domain Sentences From Task-oriented Dialog) - a dataset of 4K OOD examples for the publicly available dataset from (Schuster et al. 2019). In contrast to existing settings which synthesize OOD examples by holding out a subset of classes, our examples were authored by annotators with apriori instructions to be out-of-domain with respect to the sentences in an existing dataset.Second, we explore likelihood ratio based approaches as an alternative to currently prevalent paradigms. Specifically, we reformulate and apply these approaches to natural language inputs. We find that they match or outperform the latter on all datasets, with larger improvements on non-artificial OOD benchmarks such as our dataset. Our ablations validate that specifically using likelihood ratios rather than plain likelihood is necessary to discriminate well between OOD and in-domain data.Third, we propose learning a generative classifier and computing a marginal likelihood (ratio) for OOD detection. This allows us to use a principled likelihood while at the same time exploiting training-time labels. We find that this approach outperforms both simple likelihood (ratio) based and other prior approaches. We are hitherto the first to investigate the use of generative classifiers for OOD detection at test-time. Varun Gangal, Abhinav Arora, Arash Einolghozati, Sonal Gupta |
AAAI | 1 |
| 2020 | SCDE: Sentence Cloze Dataset with High Quality Distractors From ExaminationsabstractWe introduce SCDE, a dataset to evaluate the performance of computational models through sentence prediction.SCDE is a humancreated sentence cloze dataset, collected from public school English examinations.Our task requires a model to fill up multiple blanks in a passage from a shared candidate set with distractors designed by English teachers.Experimental results demonstrate that this task requires the use of non-local, discourse-level context beyond the immediate sentence neighborhood.The blanks require joint solving and significantly impair each other's context.Furthermore, through ablations, we show that the distractors are of high quality and make the task more challenging.Our experiments show that there is a significant performance gap between advanced models (72%) and humans (87%), encouraging future models to bridge this gap. 1 2 Passage: A student's life is never easy.And it is even more difficult if you will have to complete your study in a foreign land. 1The following are some basic things you need to do before even seizing that passport and boarding on the plane.Knowing the country.You shouldn't bother researching the country's hottest tourist spots or historical places.You won't go there as a tourist, but as a student.2 In addition, read about their laws.You surely don't want to face legal problems, especially if you're away from home.3 Don't expect that you can graduate abroad without knowing even the basics of the language.Before leaving your home country, take online lessons to at least master some of their words and sentences.This will be useful in living and studying there.Doing this will also prepare you in communicating with those who can't speak English.Preparing for other needs.Check the conversion of your money to their local currency.4. The Internet of your intended school will be very helpful in findings an apartment and helping you understand local currency.Remember, you're not only carrying your own reputation but your country's reputation as well.If you act foolishly, people there might think that all of your countrymen are foolish as well. 5Candidates: A. Studying their language.B. That would surely be a very bad start for your study abroad program.C. Going with their trends will keep it from being too obvious that you're a foreigner.D. Set up your bank account so you can use it there, get an insurance, and find an apartment.E. It'll be helpful to read the most important points in their history and to read up on their culture.F. A lot of preparations are needed so you can be sure to go back home with a diploma and a bright future waiting for you.G. Packing your clothes.Answers with Reasoning Type: 1→F (Summary) , 2→E (Inference) , 3→A (Paraphrase) , 4→D (WordMatch), 5→B (Inference) (C and G are distractors) Discussion: Blank 3 is the easiest to solve, since "Studying their language" is a near-paraphrase of "Knowing even the basics of the language".Blank 2 needs to be reasoned out by Inference -specifically E can be inferred from the previous sentence.Note however that C is also a possible inference from the previous sentence -it is only after reading the entire context, which seems to be about learning various aspects of a country, that E seems to fit better.Blank 1 needs Summary → it requires understanding several later sentences and abstracting out that they all refer to lots of preparations.Finally, Blank 5 can be mapped to B by inferring that people thinking all your countrymen are foolish is bad, while Blank 4 is a easy WordMatch on apartment to D. The other distractor G, although topically related to preparation for going abroad, does not directly fit into any of the blank contexts Xiang Kong, Varun Gangal, Eduard H. Hovy |
ACL | 2 |
| 2019 | (Male, Bachelor) and (Female, Ph.D) have different connotations: Parallelly Annotated Stylistic Language Dataset with Multiple PersonasabstractDongyeop Kang, Varun Gangal, Eduard Hovy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dongyeop Kang, Varun Gangal, Eduard H. Hovy |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Learning to Generate Move-by-Move Commentary for Chess Games from Large-Scale Social Forum DataabstractHarsh Jhamtani, Varun Gangal, Eduard Hovy, Graham Neubig, Taylor Berg-Kirkpatrick. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Harsh Jhamtani, Varun Gangal, Eduard H. Hovy, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 2 |
| 2017 | Charmanteau: Character Embedding Models For Portmanteau CreationabstractPortmanteaus are a word formation phenomenon where two words are combined to form a new word.We propose character-level neural sequence-tosequence (S2S) methods for the task of portmanteau generation that are end-toend-trainable, language independent, and do not explicitly use additional phonetic information.We propose a noisy-channelstyle model, which allows for the incorporation of unsupervised word lists, improving performance over a standard sourceto-target model.This model is made possible by an exhaustive candidate generation strategy specifically enabled by the features of the portmanteau task.Experiments find our approach superior to a state-of-the-art FST-based baseline with respect to ground truth accuracy and human evaluation. Varun Gangal, Harsh Jhamtani, Graham Neubig, Eduard H. Hovy, Eric Nyberg |
EMNLP | 1 |
| 2017 | Detecting and Explaining Causes From Text For a Time Series EventabstractExplaining underlying causes or effects about events is a challenging but valuable task.We define a novel problem of generating explanations of a time series event by (1) searching cause and effect relationships of the time series with textual data and (2) constructing a connecting chain between them to generate an explanation.To detect causal features from text, we propose a novel method based on the Granger causality of time series between features extracted from text such as N-grams, topics, sentiments, and their composition.The generation of the sequence of causal entities requires a commonsense causative knowledge base with efficient reasoning.To ensure good interpretability and appropriate lexical usage we combine symbolic and neural representations, using a neural reasoning algorithm trained on commonsense causal tuples to predict the next cause step.Our quantitative and human analysis show empirical evidence that our method successfully extracts meaningful causality relationships between time series with textual features and generates appropriate explanation between them. Dongyeop Kang, Varun Gangal, Ang Lu, Eduard H. Hovy |
EMNLP | 2 |
| 2016 | Trust and Distrust Across Coalitions: Shapley Value Based Centrality Measures for Signed Networks (Student Abstract Version)abstractWe propose Shapley Value based centrality measures for signed social networks. We also demonstrate that they lead to improved precision for the troll detection task. Varun Gangal, Abhishek Narwekar, Balaraman Ravindran, Ramasuri Narayanam |
AAAI | 1 |