VLDB 2026 Research / reviewers in the wild / expert
Tanmay Parekh
dblp:209/9604
· DBLP profile ↗
12ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Information extraction and text analysis · 65% Language models and text generation · 31% Transfer learning and domain adaptation · 4% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › event extraction
event detection |
1.7 | 2 | 2025 | DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning · EMNLP 2025 SNaRe: Domain-aware Data Generation for Low-Resource Event Detection · EMNLP 2025 |
Natural language and speech › Language models and text generation › decoding
constrained decoding |
0.9 | 1 | 2025 | DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › event extraction › event detection
zero-shot event detection |
0.9 | 1 | 2025 | DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
event extraction |
0.8 | 1 | 2024 | SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and Preparedness · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis › event extraction
multilingual event extraction |
0.8 | 1 | 2024 | SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and Preparedness · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis › event extraction
event argument extraction |
0.7 | 1 | 2023 | GENEVA: Benchmarking Generalizability for Event Argument Extraction with Hundreds of Event Types and Argument Roles · ACL (1) 2023 |
Natural language and speech › Language models and text generation › controllable text generation
text style transfer |
0.4 | 1 | 2020 | Politeness Transfer: A Tag and Generate Approach · ACL 2020 |
Natural language and speech › Language models and text generation › language modeling
code-switched language modeling |
0.3 | 1 | 2018 | Code-switched Language Models Using Dual RNNs and Same-Source Pretraining · EMNLP 2018 |
Natural language and speech › Language models and text generation
language modeling |
0.3 | 1 | 2018 | Code-switched Language Models Using Dual RNNs and Same-Source Pretraining · EMNLP 2018 |
Natural language and speech › Language models and text generation › neural language model
recurrent neural network language model |
0.3 | 1 | 2018 | Code-switched Language Models Using Dual RNNs and Same-Source Pretraining · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
large language model · 3.3synthetic data generation · 0.9finite-state machine guided constrained decoding · 0.9chain-of-thought reasoning · 0.9tag and generate · 0.4pre-training · 0.3generative model · 0.3dual recurrent neural network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SNaRe: Domain-aware Data Generation for Low-Resource Event DetectionabstractEvent Detection (ED) -the task of identifying event mentions from natural language text -is critical for enabling reasoning in highly specialized domains such as biomedicine, law, and epidemiology.Data generation has proven to be effective in broadening its utility to wider applications without requiring expensive expert annotations.However, when existing generation approaches are applied to specialized domains, they struggle with label noise, where annotations are incorrect, and domain drift, characterized by a distributional mismatch between generated sentences and the target domain.To address these issues, we introduce SNARE, a domain-aware synthetic data generation framework composed of three components: Scout, Narrator, and Refiner.Scout extracts triggers from unlabeled target domain data and curates a high-quality domain-specific trigger list using corpus-level statistics to mitigate domain drift.Narrator, conditioned on these triggers, generates high-quality domainaligned sentences, and Refiner identifies additional event mentions, ensuring high annotation quality.Experimentation on three diverse domain ED datasets reveals how SNARE outperforms the best baseline, achieving average F1 gains of 3-7% in the zero-shot/few-shot settings and 4-20% F1 improvement for multilingual generation.Analyzing the generated trigger hit rate and human evaluation substantiates SNARE's stronger annotation quality and reduced domain drift.We will release our code at https://github.com/PlusLabNLP/SNaRe. Tanmay Parekh, Lucas Bandarkar, Artin Kim, I-Hung Hsu, Kai-Wei Chang 0001, Nanyun Peng 0001 |
EMNLP | 1 |
| 2025 | DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM ReasoningabstractZero-shot Event Detection (ED), the task of identifying event mentions in natural language text without any training data, is critical for document understanding in specialized domains.Understanding the complex event ontology, extracting domain-specific triggers from the passage, and structuring them appropriately overloads and limits the utility of Large Language Models (LLMs) for zero-shot ED.To this end, we propose DICORE, a divergent-convergent reasoning framework that decouples the task of ED using Dreamer and Grounder.Dreamer encourages divergent reasoning through openended event discovery, which helps to boost event coverage.Conversely, Grounder introduces convergent reasoning to align the freeform predictions with the task-specific instructions using finite-state machine guided constrained decoding.Additionally, an LLM-Judge verifies the final outputs to ensure high precision.Through extensive experiments on six datasets across five domains and nine LLMs, we demonstrate how DICORE consistently outperforms prior zero-shot, transfer-learning, and reasoning baselines, achieving 4-7% average F1 gains over the best baseline -establishing DICORE as a strong zero-shot ED framework. Tanmay Parekh, Kartik Mehta, Ninareh Mehrabi, Kai-Wei Chang 0001, Nanyun Peng 0001 |
EMNLP | 1 |
| 2024 | SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and PreparednessabstractTanmay Parekh, Jeffrey Kwan, Jiarui Yu, Sparsh Johri, Hyosang Ahn, Sreya Muppalla, Kai-Wei Chang, Wei Wang, Nanyun Peng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Tanmay Parekh, Jeffrey Kwan, Jiarui Yu, Sparsh Johri, Hyosang Ahn, Sreya Muppalla, Kai-Wei Chang 0001, Wei Wang 0010, Nanyun Peng 0001 |
EMNLP | 1 |
| 2024 | QUDSELECT: Selective Decoding for Questions Under Discussion ParsingabstractQuestion Under Discussion (QUD) is a discourse framework that uses implicit questions to reveal discourse relationships between sentences.In QUD parsing, each sentence is viewed as an answer to a question triggered by an anchor sentence in prior context.The resulting QUD structure is required to conform to several theoretical criteria like answer compatibility (how well the question is answered), making QUD parsing a challenging task.Previous works construct QUD parsers in a pipelined manner (i.e.detect the trigger sentence in context and then generate the question).However, these parsers lack a holistic view of the task and can hardly satisfy all the criteria.In this work, we introduce QUDSELECT, a joint-training framework that selectively decodes the QUD dependency structures considering the QUD criteria.Using instruction-tuning, we train models to simultaneously predict the anchor sentence and generate the associated question.To explicitly incorporate the criteria, we adopt a selective decoding strategy of sampling multiple QUD candidates during inference, followed by selecting the best one with criteria scorers.Our method outperforms the state-of-the-art baseline models by 9% in human evaluation and 4% in automatic evaluation, demonstrating the effectiveness of our framework.Code and data are in https://github.com/ asuvarna31/qudselect. Ashima Suvarna, Xiao Liu 0032, Tanmay Parekh, Kai-Wei Chang 0001, Nanyun Peng 0001 |
EMNLP | 3 |
| 2024 | Contextual Label Projection for Cross-Lingual Structured PredictionabstractTanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang 0001, Nanyun Peng 0001 |
NAACL-HLT | 1 |
| 2024 | Event Detection from Social Media for Epidemic PredictionabstractTanmay Parekh, Anh Mac, Jiarui Yu, Yuxuan Dong, Syed Shahriar, Bonnie Liu, Eric Yang, Kuan-Hao Huang, Wei Wang, Nanyun Peng, Kai-Wei Chang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Tanmay Parekh, Anh Mac, Jiarui Yu, Syed Shahriar, Bonnie Liu, Eric Yang, Kuan-Hao Huang, Wei Wang 0010, Nanyun Peng 0001, Kai-Wei Chang 0001 |
NAACL-HLT | 1 |
| 2023 | GENEVA: Benchmarking Generalizability for Event Argument Extraction with Hundreds of Event Types and Argument RolesabstractRecent works in Event Argument Extraction (EAE) have focused on improving model generalizability to cater to new events and domains.However, standard benchmarking datasets like ACE and ERE cover less than 40 event types and 25 entity-centric argument roles.Limited diversity and coverage hinder these datasets from adequately evaluating the generalizability of EAE models.In this paper, we first contribute by creating a large and diverse EAE ontology.This ontology is created by transforming FrameNet, a comprehensive semantic role labeling (SRL) dataset for EAE, by exploiting the similarity between these two tasks.Then, exhaustive human expert annotations are collected to build the ontology, concluding with 115 events and 220 argument roles, with a significant portion of roles not being entities.We utilize this ontology to further introduce GENEVA, a diverse generalizability benchmarking dataset comprising four test suites, aimed at evaluating models' ability to handle limited data and unseen event type generalization.We benchmark six EAE models from various families.The results show that owing to non-entity argument roles, even the best-performing model can only achieve 39% F1 score, indicating how GENEVA provides new challenges for generalization in EAE.Overall, our large and diverse EAE ontology can aid in creating more comprehensive future resources, while GENEVA is a challenging benchmarking dataset encouraging further research for improving generalizability in EAE. Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang 0001, Nanyun Peng 0001 |
ACL (1) | 1 |
| 2021 | Towards Using Heterogeneous Relation Graphs for End-to-End TTSabstractNeural models for end-to-end text-to-speech (TTS) synthe-sis are increasingly outperforming traditional approaches in statistical parametric speech synthesis. Speech generation in these neural models predominantly relies on using free-form text as the input modality. However, the earlier statistical parametric models were built on encoded phonetic and syn-tactic features. In this work, we explore the possibility of explicitly feeding deterministic linguistic structure to a neural TTS system in the form of Heterogeneous Relational Graphs (HRGs), an expressive formalism capable of representing pho-netic and syntactic information. Specifically, we use Graph Convolutional Networks to learn structurally informed contin-uous representations of the HRGs, which can be seamlessly passed to the encoders of popular neural TTS models like TransformerTTS or Tacotron. Furthermore, our simple HRG based text-to-speech synthesis leverages the syntactic bias in HRGs as demonstrated by improvements in automated met-rics and human evaluation on i) the single speaker dataset LJSpeech; ii) the multi-speaker dataset Arctic; and iii) out-of-domain test sets from the Blizzard challenge.11The code, trained models, and our dataset of HRGs will be released at https://github.com/ars22/GraphNeuralTTS/. Amrith Setlur, Aman Madaan, Tanmay Parekh, Yiming Yang 0002, Alan W. Black |
ASRU | 3 |
| 2020 | Politeness Transfer: A Tag and Generate ApproachabstractAman Madaan, Amrith Setlur, Tanmay Parekh, Barnabas Poczos, Graham Neubig, Yiming Yang, Ruslan Salakhutdinov, Alan W Black, Shrimai Prabhumoye. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Aman Madaan, Amrith Setlur, Tanmay Parekh, Barnabás Póczos, Graham Neubig, Yiming Yang 0002, Ruslan Salakhutdinov, Alan W. Black, Shrimai Prabhumoye |
ACL | 3 |
| 2020 | Understanding Linguistic Accommodation in Code-Switched Human-Machine DialoguesabstractCode-switching is a ubiquitous phenomenon in multilingual communities. Natural language technologies that wish to communicate like humans must therefore adaptively incorporate code-switching techniques when they are deployed in multilingual settings. To this end, we propose a Hindi-English human-machine dialogue system that elicits code-switching conversations in a controlled setting. It uses different code-switching agent strategies to understand how users respond and accommodate to the agent's language choice. Through this system, we collect and release a new dataset CommonDost, comprising of 439 human-machine multilingual conversations. We adapt pre-defined metrics to discover linguistic accommodation from users to agents. Finally, we compare these dialogues with Spanish-English dialogues collected in a similar setting, and analyze the impact of linguistic and socio-cultural factors on code-switching patterns across the two language pairs. Tanmay Parekh, Emily P. Ahn, Yulia Tsvetkov, Alan W. Black |
CoNLL | 1 |
| 2018 | Code-switched Language Models Using Dual RNNs and Same-Source PretrainingabstractThis work focuses on building language models (LMs) for code-switched text.We propose two techniques that significantly improve these LMs: 1) A novel recurrent neural network unit with dual components that focus on each language in the code-switched text separately 2) Pretraining the LM using synthetic text from a generative model estimated using the training data.We demonstrate the effectiveness of our proposed techniques by reporting perplexities on a Mandarin-English task and derive significant reductions in perplexity. Saurabh Garg 0003, Tanmay Parekh, Preethi Jyothi |
EMNLP | 2 |
| 2018 | Dual Language Models for Code Switched Speech RecognitionabstractIn this work, we present a simple and elegant approach to language modeling for bilingual code-switched text.Since codeswitching is a blend of two or more different languages, a standard bilingual language model can be improved upon by using structures of the monolingual language models.We propose a novel technique called dual language models, which involves building two complementary monolingual language models and combining them using a probabilistic model for switching between the two.We evaluate the efficacy of our approach using a conversational Mandarin-English speech corpus.We prove the robustness of our model by showing significant improvements in perplexity measures over the standard bilingual language model without the use of any external information.Similar consistent improvements are also reflected in automatic speech recognition error rates. Saurabh Garg 0003, Tanmay Parekh, Preethi Jyothi |
INTERSPEECH | 2 |