VLDB 2026 Research / reviewers in the wild / expert
Sonal Gupta
dblp:41/3406
· DBLP profile ↗
26ranked-venue papers
8as first author
11since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Text-to-Sticker: Style Tailoring Latent Diffusion Models for Human Expression
Animesh Sinha, Anmol Kalia, Arantxa Casanova, Elliot Blanchard, David Yan, Winnie Zhang, Tony Nelli, Hardik Shah, Licheng Yu, Mitesh Kumar Singh, Ankit Ramchandani, Maziar Sanjabi, Sonal Gupta, Amy Bearman, Dhruv Mahajan 0001 |
ECCV (70) | 15 |
| 2023 | SpaText: Spatio-Textual Representation for Controllable Image GenerationabstractRecent text-to-image diffusion models are able to generate convincing results of unprecedented quality. However, it is nearly impossible to control the shapes of different regions/objects or their layout in a fine-grained fashion. Previous attempts to provide such controls were hindered by their reliance on a fixed set of labels. To this end, we present SpaText — a new method for text-to-image generation using open-vocabulary scene control. In addition to a global text prompt that describes the entire scene, the user provides a segmentation map where each region of interest is annotated by a free-form natural language description. Due to lack of large-scale datasets that have a detailed textual description for each region in the image, we choose to leverage the current large-scale text-to-image datasets and base our approach on a novel CLIP-based spatio-textual representation, and show its effectiveness on two state-of-the-art diffusion models: pixel-based and latent-based. In addition, we show how to extend the classifier-free guidance method in diffusion models to the multi-conditional case and present an alternative accelerated inference algorithm. Finally, we offer several automatic evaluation metrics and use them, in addition to FID scores and a user study, to evaluate our method and show that it achieves state-of-the-art results on image generation with free-form textual scene control. Omri Avrahami, Thomas Hayes, Oran Gafni, Sonal Gupta, Yaniv Taigman, Devi Parikh, Dani Lischinski, Ohad Fried, Xi Yin 0001 |
CVPR | 4 |
| 2023 | Make-An-Animation: Large-Scale Text-conditional 3D Human Motion GenerationabstractText-guided human motion generation has drawn significant interest because of its impactful applications spanning animation and robotics. Recently, application of diffusion models for motion generation has enabled improvements in the quality of generated motions. However, existing approaches are limited by their reliance on relatively small-scale motion capture data, leading to poor performance on more diverse, in-the-wild prompts. In this paper, we introduce Make-An-Animation, a text-conditioned human motion generation model which learns more diverse poses and prompts from large-scale image-text datasets, enabling significant improvement in performance over prior works. Make-An-Animation is trained in two stages. First, we train on a curated large-scale dataset of (text, static pseudo-pose) pairs extracted from image-text datasets. Second, we fine-tune on motion capture data, adding additional layers to model the temporal dimension. Unlike prior diffusion models for motion generation, Make-An-Animation uses a U-Net architecture similar to recent text-to-video generation models. Human evaluation of motion realism and alignment with input text shows that our model reaches state-of-the-art performance on text-to-motion generation. Generated samples can be viewed at https://azadis.github.io/make-an-animation. Samaneh Azadi, Akbar Shah, Thomas Hayes, Devi Parikh, Sonal Gupta |
ICCV | 5 |
| 2023 | Make-A-Video: Text-to-Video Generation without Text-Video Data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin 0001, Jie An 0002, Songyang Zhang 0004, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, Yaniv Taigman |
ICLR | 12 |
| 2021 | Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-TuningabstractArmen Aghajanyan, Sonal Gupta, Luke Zettlemoyer. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Armen Aghajanyan, Sonal Gupta, Luke Zettlemoyer |
ACL/IJCNLP (1) | 2 |
| 2021 | El Volumen Louder Por Favor: Code-switching in Task-oriented Semantic ParsingabstractArash Einolghozati, Abhinav Arora, Lorena Sainz-Maza Lecanda, Anuj Kumar, Sonal Gupta. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Arash Einolghozati, Abhinav Arora, Lorena Sainz-Maza Lecanda, Sonal Gupta |
EACL | 5 |
| 2021 | MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing BenchmarkabstractHaoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, Yashar Mehdad. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Haoran Li 0007, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, Yashar Mehdad |
EACL | 5 |
| 2021 | Muppet: Massive Multi-task Representations with Pre-FinetuningabstractWe propose pre-finetuning, an additional largescale learning stage between language model pre-training and fine-tuning.Pre-finetuning is massively multi-task learning (around 50 datasets, over 4.8 million total labeled examples), and is designed to encourage learning of representations that generalize better to many different tasks.We show that prefinetuning consistently improves performance for pretrained discriminators (e.g.RoBERTa) and generation models (e.g.BART) on a wide range of tasks (sentence prediction, commonsense reasoning, MRC, etc.), while also significantly improving sample efficiency during fine-tuning.We also show that large-scale multi-tasking is crucial; pre-finetuning can hurt performance when few tasks are used up until a critical point (usually above 15) after which performance improves linearly in the number of tasks. Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen 0002, Luke Zettlemoyer, Sonal Gupta |
EMNLP (1) | 6 |
| 2021 | Better Fine-Tuning by Reducing Representational Collapse
Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal 0001, Luke Zettlemoyer, Sonal Gupta |
ICLR | 6 |
| 2021 | Learning Better Structured Representations Using Low-rank Adaptive Label Smoothing
Asish Ghoshal, Xilun Chen 0002, Sonal Gupta, Luke Zettlemoyer, Yashar Mehdad |
ICLR | 3 |
| 2021 | Getting to Production with Few-shot Natural Language Generation ModelsabstractPeyman Heidari, Arash Einolghozati, Shashank Jain, Soumya Batra, Lee Callender, Ankit Arun, Shawn Mei, Sonal Gupta, Pinar Donmez, Vikas Bhardwaj, Anuj Kumar, Michael White. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Peyman Heidari, Arash Einolghozati, Shashank Jain, Soumya Batra, Lee Callender, Ankit Arun, Shawn Mei, Sonal Gupta, Pinar Donmez, Vikas Bhardwaj, Michael White 0001 |
SIGDIAL | 8 |
| 2020 | Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection in Task Oriented DialogabstractThe task of identifying out-of-domain (OOD) input examples directly at test-time has seen renewed interest recently due to increased real world deployment of models. In this work, we focus on OOD detection for natural language sentence inputs to task-based dialog systems. Our findings are three-fold:First, we curate and release ROSTD (Real Out-of-Domain Sentences From Task-oriented Dialog) - a dataset of 4K OOD examples for the publicly available dataset from (Schuster et al. 2019). In contrast to existing settings which synthesize OOD examples by holding out a subset of classes, our examples were authored by annotators with apriori instructions to be out-of-domain with respect to the sentences in an existing dataset.Second, we explore likelihood ratio based approaches as an alternative to currently prevalent paradigms. Specifically, we reformulate and apply these approaches to natural language inputs. We find that they match or outperform the latter on all datasets, with larger improvements on non-artificial OOD benchmarks such as our dataset. Our ablations validate that specifically using likelihood ratios rather than plain likelihood is necessary to discriminate well between OOD and in-domain data.Third, we propose learning a generative classifier and computing a marginal likelihood (ratio) for OOD detection. This allows us to use a principled likelihood while at the same time exploiting training-time labels. We find that this approach outperforms both simple likelihood (ratio) based and other prior approaches. We are hitherto the first to investigate the use of generative classifiers for OOD detection at test-time. Varun Gangal, Abhinav Arora, Arash Einolghozati, Sonal Gupta |
AAAI | 4 |
| 2020 | Conversational Semantic ParsingabstractArmen Aghajanyan, Jean Maillard, Akshat Shrivastava, Keith Diedrick, Michael Haeger, Haoran Li, Yashar Mehdad, Veselin Stoyanov, Anuj Kumar, Mike Lewis, Sonal Gupta. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Armen Aghajanyan, Jean Maillard, Akshat Shrivastava, Keith Diedrick, Michael Haeger, Haoran Li 0007, Yashar Mehdad, Veselin Stoyanov, Mike Lewis, Sonal Gupta |
EMNLP (1) | 11 |
| 2020 | Low-Resource Domain Adaptation for Compositional Task-Oriented Semantic ParsingabstractTask-oriented semantic parsing is a critical component of virtual assistants, which is responsible for understanding the user's intents (set reminder, play music, etc.).Recent advances in deep learning have enabled several approaches to successfully parse more complex queries (Gupta et al., 2018;Rongali et al., 2020), but these models require a large amount of annotated training data to parse queries on new domains (e.g.reminder, music).In this paper, we focus on adapting taskoriented semantic parsers to low-resource domains, and propose a novel method that outperforms a supervised neural model at a 10-fold data reduction.In particular, we identify two fundamental factors for low-resource domain adaptation: better representation learning and better training techniques.Our representation learning uses BART (Lewis et al., 2020) to initialize our model which outperforms encoder-only pre-trained representations used in previous work.Furthermore, we train with optimization-based meta-learning (Finn et al., 2017) to improve generalization to lowresource domains.This approach significantly outperforms all baseline methods in the experiments on a newly collected multi-domain taskoriented semantic parsing dataset (TOPv2 1 ). Xilun Chen 0002, Asish Ghoshal, Yashar Mehdad, Luke Zettlemoyer, Sonal Gupta |
EMNLP (1) | 5 |
| 2020 | Sound Natural: Content Rephrasing in Dialog SystemsabstractWe introduce a new task of rephrasing for a more natural virtual assistant.Currently, virtual assistants work in the paradigm of intentslot tagging and the slot values are directly passed as-is to the execution engine.However, this setup fails in some scenarios such as messaging when the query given by the user needs to be changed before repeating it or sending it to another user.For example, for queries like 'ask my wife if she can pick up the kids' or 'remind me to take my pills', we need to rephrase the content to 'can you pick up the kids' and 'take your pills'.In this paper, we study the problem of rephrasing with messaging as a use case and release a dataset of 3000 pairs of original query and rephrased query.We show that BART, a pre-trained transformers-based masked language model with auto-regressive decoding, is a strong baseline for the task, and show improvements by adding a copy-pointer and copy loss to it.We analyze different tradeoffs of BART-based and LSTM-based seq2seq models, and propose a distilled LSTM-based seq2seq as the best practical model. Arash Einolghozati, Anchit Gupta, Keith Diedrick, Sonal Gupta |
EMNLP (1) | 4 |
| 2019 | Span-based Hierarchical Semantic Parsing for Task-Oriented DialogabstractPanupong Pasupat, Sonal Gupta, Karishma Mandyam, Rushin Shah, Mike Lewis, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Panupong Pasupat, Sonal Gupta, Karishma Mandyam, Rushin Shah, Mike Lewis, Luke Zettlemoyer |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Semantic Parsing for Task Oriented Dialog using Hierarchical RepresentationsabstractTask oriented dialog systems typically first parse user utterances to semantic frames comprised of intents and slots.Previous work on task oriented intent and slot-filling work has been restricted to one intent per query and one slot label per token, and thus cannot model complex compositional requests.Alternative semantic parsing systems have represented queries as logical forms, but these are challenging to annotate and parse.We propose a hierarchical annotation scheme for semantic parsing that allows the representation of compositional queries, and can be efficiently and accurately parsed by standard constituency parsing models.We release a dataset of 44k annotated queries 1 , and show that parsing models outperform sequence-to-sequence approaches on this dataset. Sonal Gupta, Rushin Shah, Mrinal Mohit, Mike Lewis |
EMNLP | 1 |
| 2015 | Forum77: An Analysis of an Online Health Forum Dedicated to Addiction RecoveryabstractPrescription drug abuse is a pressing public health issue, and people who misuse prescription drugs are turning to online forums for help. Are such forums effective? We analyze the process of opioid withdrawal, recovery and relapse on Forum77, MedHelp.org's online health forum for substance abuse recovery. Applying Prochashka's Transtheoretical Model for behavior change, we develop a taxonomy describing phases of addiction expressed by Forum77 members. We examine activity and linguistic features across the phases USING, WITHDRAWING and RECOVERING. We train statistical classifiers to identify addiction phase, relapse and whether a user was RECOVERING at the time of her last post. Applying our classifiers to 2,848 users, we find that while almost 50% relapse, the prognosis for ending in RECOVERING is favorable. Supplementing our results with users' own accounts of their experiences, we discuss Forum77's efficacy and shortcomings, and implications for future technologies. Diana L. MacLean, Sonal Gupta, Anna Lembke, Christopher D. Manning, Jeffrey Heer |
CSCW | 2 |
| 2015 | Distributed Representations of Words to Guide Bootstrapped Entity ClassifiersabstractBootstrapped classifiers iteratively generalize from a few seed examples or prototypes to other examples of target labels.However, sparseness of language and limited supervision make the task difficult.We address this problem by using distributed vector representations of words to aid the generalization.We use the word vectors to expand entity sets used for training classifiers in a bootstrapped pattern-based entity extraction system.Our experiments show that the classifiers trained with the expanded sets perform better on entity extraction from four online forums, with 30% F 1 improvement on one forum.The results suggest that distributed representations can provide good directions for generalization in a bootstrapping system. Sonal Gupta, Christopher D. Manning |
HLT-NAACL | 1 |
| 2014 | Improved Pattern Learning for Bootstrapped Entity ExtractionabstractBootstrapped pattern learning for entity extraction usually starts with seed entities and iteratively learns patterns and entities from unlabeled text. Patterns are scored by their ability to extract more positive en-tities and less negative entities. A prob-lem is that due to the lack of labeled data, unlabeled entities are either assumed to be negative or are ignored by the existing pat-tern scoring measures. In this paper, we improve pattern scoring by predicting the labels of unlabeled entities. We use var-ious unsupervised features based on con-trasting domain-specific and general text, and exploiting distributional similarity and edit distances to learned entities. Our system outperforms existing pattern scor-ing algorithms for extracting drug-and-treatment entities from four medical fo-rums. 1 Sonal Gupta, Christopher D. Manning |
CoNLL | 1 |
| 2014 | Research and applications: Induced lexico-syntactic patterns improve information extraction from online medical forumsabstractOBJECTIVE: To reliably extract two entity types, symptoms and conditions (SCs), and drugs and treatments (DTs), from patient-authored text (PAT) by learning lexico-syntactic patterns from data annotated with seed dictionaries. BACKGROUND AND SIGNIFICANCE: Despite the increasing quantity of PAT (eg, online discussion threads), tools for identifying medical entities in PAT are limited. When applied to PAT, existing tools either fail to identify specific entity types or perform poorly. Identification of SC and DT terms in PAT would enable exploration of efficacy and side effects for not only pharmaceutical drugs, but also for home remedies and components of daily care. MATERIALS AND METHODS: We use SC and DT term dictionaries compiled from online sources to label several discussion forums from MedHelp (http://www.medhelp.org). We then iteratively induce lexico-syntactic patterns corresponding strongly to each entity type to extract new SC and DT terms. RESULTS: Our system is able to extract symptom descriptions and treatments absent from our original dictionaries, such as 'LADA', 'stabbing pain', and 'cinnamon pills'. Our system extracts DT terms with 58-70% F1 score and SC terms with 66-76% F1 score on two forums from MedHelp. We show improvements over MetaMap, OBA, a conditional random field-based classifier, and a previous pattern learning approach. CONCLUSIONS: Our entity extractor based on lexico-syntactic patterns is a successful and preferable technique for identifying specific entity types in PAT. To the best of our knowledge, this is the first paper to extract SC and DT entities from PAT. We exhibit learning of informal terms often used in PAT but missing from typical dictionaries. Sonal Gupta, Diana L. MacLean, Jeffrey Heer, Christopher D. Manning |
J. Am. Medical Informatics Assoc. | 1 |
| 2013 | Topic Model Diagnostics: Assessing Domain Relevance via Topical AlignmentabstractThe use of topic models to analyze domain-specific texts often requires manual validation of the latent topics to ensure they are meaningful. We introduce a framework to support large-scale assessment of topical relevance. We measure the correspondence between a set of latent topics and a set of reference concepts to quantify four types of topical misalignment: junk, fused, missing, and repeated topics. Our analysis compares 10,000 topic model variants to 200 expert-provided domain concepts, and demonstrates how our framework can inform choices of model parameters, inference algorithms, and intrinsic measures of topical quality. Jason Chuang, Sonal Gupta, Christopher D. Manning, Jeffrey Heer |
ICML (3) | 2 |
| 2011 | Analyzing the Dynamics of Research by Extracting Key Aspects of Scientific Papers
Sonal Gupta, Christopher D. Manning |
IJCNLP | 1 |
| 2010 | Using Closed Captions as Supervision for Video Activity RecognitionabstractRecognizing activities in real-world videos is a difficult problem exacerbated by background clutter, changes in camera angle & zoom, and rapid camera movements. Large corpora of labeled videos can be used to train automated activity recognition systems, but this requires expensive human labor and time. This paper explores how closed captions that naturally accompany many videos can act as weak supervision that allows automatically collecting "labeled" data for activity recognition. We show that such an approach can improve activity retrieval in soccer videos. Our system requires no manual labeling of video clips and needs minimal human supervision. We also present a novel caption classifier that uses additional linguistic information to determine whether a specific comment refers to an ongoing activity. We demonstrate that combining linguistic analysis and automatically trained activity recognizers can significantly improve the precision of video retrieval. Sonal Gupta, Raymond J. Mooney |
AAAI | 1 |
| 2009 | Catching the drift: learning broad matches from clickthrough dataabstractIdentifying similar keywords, known as broad matches, is an important task in online advertising that has become a standard feature on all major keyword advertising platforms. Effective broad matching leads to improvements in both relevance and monetization, while increasing advertisers' reach and making campaign management easier. In this paper, we present a learning-based approach to broad matching that is based on exploiting implicit feedback in the form of advertisement clickthrough logs. Our method can utilize arbitrary similarity functions by incorporating them as features. We present an online learning algorithm, Amnesiac Averaged Perceptron, that is highly efficient yet able to quickly adjust to the rapidly-changing distributions of bidded keywords, advertisements and user behavior. Experimental results obtained from (1) historical logs and (2) live trials on a large-scale advertising platform demonstrate the effectiveness of the proposed algorithm and the overall success of our approach in identifying high-quality broad match mappings. Sonal Gupta, Mikhail Bilenko, Matthew Richardson |
KDD | 1 |
| 2008 | Watch, Listen & Learn: Co-training on Captioned Images and Videos
Sonal Gupta, Joohyun Kim 0002, Kristen Grauman, Raymond J. Mooney |
ECML/PKDD (1) | 1 |