VLDB 2026 Research / reviewers in the wild / expert
Vijeta Deshpande
dblp:329/4901
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Question answering and dialogue systems · 33% Language models and text generation · 33% Trustworthy machine learning · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model training
data selection for language model training |
0.9 | 1 | 2025 | Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models · EMNLP 2025 |
Machine learning › Trustworthy machine learning
fairness and bias |
0.9 | 1 | 2025 | Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models · EMNLP 2025 |
Natural language and speech › Question answering and dialogue systems › dialogue generation › dialogue response generation
response diversity |
0.9 | 1 | 2025 | Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
preference optimization · 0.9data filtering · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language ModelsabstractDiverse language model responses are crucial for creative generation, open-ended tasks, and self-improvement training.We show that common diversity metrics, and even reward models used for preference optimization, systematically bias models toward shorter outputs, limiting expressiveness.To address this, we introduce Diverse, not Short (Diverse-NS), a length-controlled data selection strategy that improves response diversity while maintaining length parity.By generating and filtering preference data that balances diversity, quality, and length, Diverse-NS enables effective training using only 3,000 preference pairs.Applied to LLaMA-3.1-8B and the Olmo-2 family, Diverse-NS substantially enhances lexical and semantic diversity.We show consistent improvement in diversity with minor reduction or gains in response quality on four creative generation tasks: Divergent Associations, Persona Generation, Alternate Uses, and Creative Writing.Surprisingly, experiments with the Olmo-2 model family (7B, and 13B) show that smaller models like Olmo-2-7B can serve as effective "diversity teachers" for larger models.By explicitly addressing length bias, our method efficiently pushes models toward more diverse and expressive outputs 1 . Vijeta Deshpande, Debasmita Ghose, John D. Patterson, Roger E. Beaty, Anna Rumshisky |
EMNLP | 1 |
| 2024 | LocalTweets to LocalHealth: A Mental Health Surveillance Framework Based on Twitter DataabstractPrior research on Twitter (now X) data has provided positive evidence of its utility in developing supplementary health surveillance systems. In this study, we present a new framework to surveil public health, focusing on mental health (MH) outcomes. We hypothesize that locally posted tweets are indicative of local MH outcomes and collect tweets posted from 765 neighborhoods (census block groups) in the USA. We pair these tweets from each neighborhood with the corresponding MH outcome reported by the Center for Disease Control (CDC) to create a benchmark dataset, LocalTweets. With LocalTweets, we present the first population-level evaluation task for Twitter-based MH surveillance systems. We then develop an efficient and effective method, LocalHealth, for predicting MH outcomes based on LocalTweets. When used with GPT3.5, LocalHealth achieves the highest F1-score and accuracy of 0.7429 and 79.78%, respectively, a 59% improvement in F1-score over the GPT3.5 in zero-shot setting. We also utilize LocalHealth to extrapolate CDC’s estimates to proxy unreported neighborhoods, achieving an F1-score of 0.7291. Our work suggests that Twitter data can be effectively leveraged to simulate neighborhood-level MH outcomes. Vijeta Deshpande, Minhwa Lee, Zonghai Yao, Zihao Zhang 0001, Jason Brian Gibbons, Hong Yu 0001 |
LREC/COLING | 1 |
| 2022 | Extracting Biomedical Factual Knowledge Using Pretrained Language Model and Electronic Health Record Context
Zonghai Yao, Zhichao Yang 0001, Vijeta Deshpande, Hong Yu 0001 |
AMIA | 4 |