VLDB 2026 Research / reviewers in the wild / expert
Aparna Garimella
dblp:183/5034
· DBLP profile ↗
26ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-3111-0686ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 5 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TabReX: Tabular Referenceless eXplainable EvaluationabstractEvaluating the quality of tables generated by large language models (LLMs) remains an open challenge: existing metrics either flatten tables into text, ignoring structure, or rely on fixed references that limit generalization.We present TABREX, a reference-less, propertydriven framework for evaluating tabular generation via graph-based reasoning.TABREX converts both source text and generated tables into canonical knowledge graphs, aligns them through an LLM-guided matching process, and computes interpretable, rubric-aware scores that quantify structural and factual fidelity.The resulting metric provides controllable tradeoffs between sensitivity and specificity, yielding human-aligned judgments and cell-level error traces.To systematically assess metric robustness, we introduce TABREX-BENCH, a large-scale benchmark spanning six domains and twelve planner-driven perturbation types across three difficulty tiers.Empirical results show that TABREX achieves the highest correlation with expert rankings, remains stable under harder perturbations, and enables finegrained model-vs-prompt analysis establishing a new paradigm for trustworthy, explainable evaluation of structured generation systems. Tejas Anvekar, Junha Park, Aparna Garimella, Vivek Gupta 0001 |
ACL (1) | 3 |
| 2026 | Decisive: Guiding User Decisions with Optimal Preference Elicitation from Unstructured DocumentsabstractAkriti Jain, Anish Mulay, Divyansh Verma, Aishani Pandey, Pritika Ramu, Aparna Garimella. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Akriti Jain 0001, Anish Mulay, Divyansh Verma, Aishani Pandey, Pritika Ramu, Aparna Garimella |
ACL (1) | 6 |
| 2025 | Infogen: Generating Complex Statistical Infographics from DocumentsabstractAkash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Akash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha 0001 |
ACL (1) | 2 |
| 2025 | ADAPTIVE IE: Investigating the Complementarity of Human-AI Collaboration to Adaptively Extract Information on-the-flyabstractInformation extraction (IE) needs vary over time, where a flexible information extraction (IE) system can be useful. Despite this, existing IE systems are either fully supervised, requiring expensive human annotations, or fully unsupervised, extracting information that often do not cater to user’s needs. To address these issues, we formally introduce the task of “IE on-the-fly”, and address the problem using our proposed Adaptive IE framework that uses human-in-the-loop refinement to adapt to changing user questions. Through human experiments on three diverse datasets, we demonstrate that Adaptive IE is a domain-agnostic, responsive, efficient framework for helping users access useful information while quickly reorganizing information in response to evolving information needs. Ishani Mondal, Michelle Yuan, Anandhavelu Natarajan, Aparna Garimella, Francis Ferraro, Andrew Blair-Stanek, Benjamin Van Durme, Jordan L. Boyd-Graber |
COLING | 4 |
| 2025 | Doc2Chart: Intent-Driven Zero-Shot Chart Generation from DocumentsabstractLarge Language Models (LLMs) have demonstrated strong capabilities in transforming text descriptions or tables to data visualizations via instruction-tuning methods.However, it is not straightforward to apply these methods directly for a more real-world use case of visualizing data from long documents based on user-given intents, as opposed to the user pre-selecting the relevant content manually.We introduce the task of intent-based chart generation from documents: given a user-specified intent and document(s), the goal is to generate a chart adhering to the intent and grounded on the document(s) in a zero-shot setting.We propose an unsupervised, two-staged framework in which an LLM first extracts relevant information from the document(s) by decomposing the intent and iteratively validates and refines this data.Next, a heuristic-guided module selects an appropriate chart type before final code generation.To assess the data accuracy of the generated charts, we propose an attribution-based metric that uses a structured textual representation of charts, instead of relying on visual decoding metrics that often fail to capture the chart data effectively.To validate our approach, we curate a dataset comprising of 1,242 tuples from two domains, finance and scientific, in contrast to the existing datasets that are largely limited to parallel text descriptions/ tables and their corresponding charts.We compare our approach with baselines using single-shot chart generation using LLMs and query-based retrieval methods; our method outperforms by upto 9 points and 17 points in terms of chart data accuracy and chart type respectively over the best baselines. Akriti Jain 0001, Pritika Ramu, Aparna Garimella, Apoorv Saxena |
EMNLP | 3 |
| 2024 | DocScript: Document-level Script Event PredictionabstractWe present a novel task of document-level script event prediction, which aims to predict the next event given a candidate list of narrative events in long-form documents. To enable this, we introduce DocSEP, a challenging dataset in two new domains - contractual documents and Wikipedia articles, where timeline events may be paragraphs apart and may require multi-hop temporal and causal reasoning. We benchmark existing baselines and present a novel architecture called DocScript to learn sequential ordering between events at the document scale. Our experimental results on the DocSEP dataset demonstrate that learning longer-range dependencies between events is a key challenge and show that contemporary LLMs such as ChatGPT and FlanT5 struggle to solve this task, indicating their lack of reasoning abilities for understanding causal relationships and temporal sequences within long texts. Puneet Mathur, Vlad I. Morariu, Aparna Garimella, Franck Dernoncourt, Jiuxiang Gu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha, Rajiv Jain |
LREC/COLING | 3 |
| 2024 | Presentations by the Humans and For the Humans: Harnessing LLMs for Generating Persona-Aware Slides from DocumentsabstractIshani Mondal, Shwetha S, Anandhavelu Natarajan, Aparna Garimella, Sambaran Bandyopadhyay, Jordan Boyd-Graber. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ishani Mondal, Shwetha S, Anandhavelu Natarajan, Aparna Garimella, Sambaran Bandyopadhyay, Jordan L. Boyd-Graber |
EACL (1) | 4 |
| 2024 | Is This a Bad Table? A Closer Look at the Evaluation of Table Generation from TextabstractUnderstanding whether a generated table is of good quality is important to be able to use it in creating or editing documents using automatic methods.In this work, we underline that existing measures for table quality evaluation fail to capture the overall semantics of the tables, and sometimes unfairly penalize good tables and reward bad ones.We propose TABEVAL, a novel table evaluation strategy that captures table semantics by first breaking down a table into a list of natural language atomic statements and then compares them with ground truth statements using entailment-based measures.To validate our approach, we curate a dataset comprising of text descriptions for 1,250 diverse Wikipedia tables, covering a range of topics and structures, in contrast to the limited scope of existing datasets.We compare TABEVAL with existing metrics using unsupervised and supervised textto-table generation methods, demonstrating its stronger correlation with human judgments of table quality across four datasets. Pritika Ramu, Aparna Garimella, Sambaran Bandyopadhyay |
EMNLP | 2 |
| 2024 | Zooming in on Zero-Shot Intent-Guided and Grounded Document Generation using LLMsabstractRepurposing existing content on-the-fly to suit author's goals for creating initial drafts is crucial for document creation.We introduce the task of intent-guided and grounded document generation: given a user-specified intent (e.g., section title) and a few reference documents, the goal is to generate section-level multimodal documents spanning text and images, grounded on the given references, in a zero-shot setting.We present a data curation strategy to obtain general-domain samples from Wikipedia, and collect 1,000 Wikipedia sections consisting of textual and image content along with appropriate intent specifications and references.We propose a simple yet effective planningbased prompting strategy Multimodal Plan-And-Write (MM-PAW), to prompt LLMs to generate an intermediate plan with text and image descriptions, to guide the subsequent generation.We compare the performances of MM-PAW and a text-only variant of it with those of zero-shot Chain-of-Thought (CoT) using recent close and open-domain LLMs.Both of them lead to significantly better performances in terms of content relevance, structure, and groundedness to the references, more so in the smaller models (upto 12.5 points ↑ in Rouge 1-F1) than in the larger ones (upto 4 points ↑ R1-F1).They are particularly effective in improving relatively smaller models' performances, to be on par or higher than those of their larger counterparts for this task. Pritika Ramu, Pranshu Gaur, Rishita Emandi, Himanshu Maheshwari, Danish Javed, Aparna Garimella |
INLG | 6 |
| 2024 | IndiBias: A Benchmark Dataset to Measure Social Biases in Language Models for Indian ContextabstractNihar Sahoo, Pranamya Kulkarni, Arif Ahmad, Tanu Goyal, Narjis Asad, Aparna Garimella, Pushpak Bhattacharyya. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nihar R. Sahoo, Pranamya Prashant Kulkarni, Arif Ahmad, Tanu Goyal, Narjis Asad, Aparna Garimella, Pushpak Bhattacharyya |
NAACL-HLT | 6 |
| 2023 | What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and ProhibitionsabstractReviewing and comprehending key obligations, entitlements, and prohibitions in legal contracts can be a tedious task due to their length and domain-specificity.Furthermore, the key rights and duties requiring review vary for each contracting party.In this work, we propose a new task of party-specific extractive summarization for legal contracts to facilitate faster reviewing and improved comprehension of rights and duties.To facilitate this, we curate a dataset comprising of party-specific pairwise importance comparisons annotated by legal experts, covering ∼293K sentence pairs that include obligations, entitlements, and prohibitions extracted from lease agreements.Using this dataset, we train a pairwise importance ranker and propose a pipeline-based extractive summarization system that generates a party-specific contract summary.We establish the need for incorporating domain-specific notion of importance during summarization by comparing our system against various baselines using both automatic and human evaluation methods 1 . Abhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, Rachel Rudinger |
EMNLP | 2 |
| 2023 | kNN-LM Does Not Improve Open-ended Text GenerationabstractIn this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs).These methods, best exemplified by the kNN-LM (Khandelwal et al., 2020), interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix.While the kNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations.Digging deeper, we find that interpolating with a retrieval distribution actually increases perplexity compared to the baseline LM for the majority of tokens in the WikiText-103 test set, even though the overall perplexity is lower due to a smaller number of tokens for which perplexity dramatically decreases after interpolation.However, when decoding a long sequence at inference time, significant improvements on this smaller subset of tokens are washed out by slightly worse predictions on most tokens.Furthermore, we discover that the entropy of the retrieval distribution increases faster than that of the base LM as the generated sequence becomes longer, which indicates that retrieval is less reliable when using model-generated text as queries (i.e., is subject to exposure bias).We hope that our analysis spurs future work on improved decoding algorithms and interpolation strategies for retrieval-augmented language models. Shufan Wang, Yixiao Song, Andrew Drozdov, Aparna Garimella, Varun Manjunatha, Mohit Iyyer |
EMNLP | 4 |
| 2023 | Computable Contracts by Extracting Obligation Logic GraphsabstractThe emergence of contract specific programming languages has struggled to translate into widespread adoption of computable contracts due largely to high conversion costs. In this work, we present the first system for converting natural language contracts into code through the extraction of key entities, relationships, and formulas into a graph representation called the Obligation Logic Graph (OLG). This approach allows the semantic meaning of contract obligations, including dependencies between obligations, to be captured through the OLG and mapped to code downstream. We also introduce OLG extraction as a new joint entity and relation prediction task for legal contracts, and present the Contract-OLG dataset, consisting of 1,876 contract provisions, 18,597 entities and 18,170 relationships. We perform detailed experiments to understand the capabilities of state-of-the-art Transformer and graph-based models at completing these tasks, and identify where there is currently a significant gap between human expert and machine performance, particularly for relation extraction. Sergio Servantez, Nedim Lipka, Alexa F. Siu, Milan Aggarwal, Balaji Krishnamurthy, Aparna Garimella, Kristian J. Hammond, Rajiv Jain |
ICAIL | 6 |
| 2023 | Reflection of Demographic Background on Word UsageabstractAbstract The availability of personal writings in electronic format provides researchers in the fields of linguistics, psychology, and computational linguistics with an unprecedented chance to study, on a large scale, the relationship between language use and the demographic background of writers, allowing us to better understand people across different demographics. In this article, we analyze the relation between language and demographics by developing cross-demographic word models to identify words with usage bias, or words that are used in significantly different ways by speakers of different demographics. Focusing on three demographic categories, namely, location, gender, and industry, we identify words with significant usage differences in each category and investigate various approaches of encoding a word’s usage, allowing us to identify language aspects that contribute to the differences. Our word models using topic-based features achieve at least 20% improvement in accuracy over the baseline for all demographic categories, even for scenarios with classification into 15 categories, illustrating the usefulness of topic-based features in identifying word usage differences. Further, we note that for location and industry, topics extracted from immediate context are the best predictors of word usages, hinting at the importance of word meaning and its grammatical function for these demographics, while for gender, topics obtained from longer contexts are better predictors for word usage. Aparna Garimella, Carmen Banea, Rada Mihalcea |
Comput. Linguistics | 1 |
| 2022 | Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsabstractTransformer-based language models trained on large natural language corpora have been very useful in downstream entity extraction tasks.However, they often result in poor performances when applied to domains that are different from those they are pretrained on.Continued pretraining using unlabeled data from target domains can help improve the performances of these language models on the downstream tasks.However, using all of the available unlabeled data for pretraining can be time-intensive; also, it can be detrimental to the performance of the downstream tasks, if the unlabeled data is not aligned with the data distribution for the target tasks.Previous works employed external supervision in the form of ontologies for selecting appropriate data samples for pretraining, but external supervision can be quite hard to obtain in low-resource domains.In this paper, we introduce effective ways to select data from unlabeled corpora of target domains for language model pretraining to improve the performances in target entity extraction tasks.Our data selection strategies do not require any external supervision.We conduct extensive experiments for the task of named entity recognition (NER) on seven different domains and show that language models pretrained on target domain unlabeled data obtained using our data selection strategies achieve better performances compared to those using data selection strategies in previous works that use external supervision.We also show that these pretrained language models using our data selection strategies outperform those pretrained on all of the available unlabeled target domain data. Aniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu Natarajan |
EMNLP | 3 |
| 2022 | Agent-Specific Deontic Modality Detection in Legal LanguageabstractLegal documents are typically long and written in legalese, which makes it particularly difficult for laypeople to understand their rights and duties.While natural language understanding technologies can be valuable in supporting such understanding in the legal domain, the limited availability of datasets annotated for deontic modalities in the legal domain, due to the cost of hiring experts and privacy issues, is a bottleneck.To this end, we introduce, LEXDE-MOD, a corpus of English contracts annotated with deontic modality expressed with respect to a contracting party or agent along with the modal triggers.We benchmark this dataset on two tasks: (i) agent-specific multi-label deontic modality classification, and (ii) agent-specific deontic modality and trigger span detection using Transformer-based (Vaswani et al., 2017) language models.Transfer learning experiments show that the linguistic diversity of modal expressions in LEXDEMOD generalizes reasonably from lease to employment and rental agreements.A small case study indicates that a model trained on LEXDEMOD can detect red flags with high recall.We believe our work offers a new research direction for deontic modality detection in the legal domain 1 . Abhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, Rachel Rudinger |
EMNLP | 2 |
| 2022 | Investigating Strategies for Clause RecommendationabstractClause recommendation is the problem of recommending a clause to a legal contract, given the context of the contract in question and the clause type to which the clause should belong. With not much prior work being done toward the generation of legal contracts, this problem was proposed as a first step toward the bigger problem of contract generation. As an open-ended text generation problem, the distinguishing characteristics of this problem lie in the nature of legal language as a sublanguage and the considerable similarity of textual content within the clauses of a specific type. This similarity aspect in legal clauses drives us to investigate the importance of similar contracts’ representation for recommending clauses. In our work, we experiment with generating clauses for 15 commonly occurring clause types in contracts expanding upon the previous work on this problem and analyzing clause recommendations in varying settings using information derived from similar contracts. Sagar Joshi, Sumanth Balaji, Jerrin Thomas, Aparna Garimella, Vasudeva Varma |
JURIX | 4 |
| 2021 | EmpathBERT: A BERT-based Framework for Demographic-aware Empathy PredictionabstractAffect preferences vary with user demographics, and tapping into demographic information provides important cues about the users' language preferences.In this paper, we utilize the user demographics, and propose EMPATH-BERT, a demographic-aware framework for empathy prediction based on BERT.Through several comparative experiments, we show that EMPATHBERT surpasses traditional machine learning and deep learning models, and illustrate the importance of user demographics to predict empathy and distress in user responses to stimulative news articles.We also highlight the importance of affect information in the responses by developing affect-aware models to predict user demographic attributes. Bhanu Prakash Reddy Guda, Aparna Garimella, Niyati Chhaya |
EACL | 2 |
| 2021 | DRAG: Director-Generator Language Modelling Framework for Non-Parallel Author Stylized RewritingabstractAuthor stylized rewriting is the task of rewriting an input text in a particular author's style.Recent works in this area have leveraged Transformer-based language models in a denoising autoencoder setup to generate author stylized text without relying on a parallel corpus of data.However, these approaches are limited by the lack of explicit control of target attributes and being entirely data-driven.In this paper, we propose a Director-Generator framework to rewrite content in the target author's style, specifically focusing on certain target attributes.We show that our proposed framework works well even with a limitedsized target author corpus.Our experiments on corpora consisting of relatively small-sized text authored by three distinct authors show significant improvements upon existing works to rewrite input texts in target author's style.Our quantitative and qualitative analyses further show that our model has better meaning retention and results in more fluent generations. Hrituraj Singh, Aparna Garimella, Balaji Vasan Srinivasan |
EACL | 3 |
| 2021 | ClauseRec: A Clause Recommendation Framework for AI-aided Contract AuthoringabstractContracts are a common type of legal document that frequent in several day-to-day business workflows.However, there has been very limited NLP research in processing such documents, and even lesser in generating them.These contracts are made up of clauses, and the unique nature of these clauses calls for specific methods to understand and generate such documents.In this paper, we introduce the task of clause recommendation, as a first step to aid and accelerate the authoring of contract documents.We propose a twostaged pipeline to first predict if a specific clause type is relevant to be added in a contract, and then recommend the top clauses for the given type based on the contract context.We pretrain BERT on an existing library of clauses with two additional tasks and use it for our prediction and recommendation.We experiment with classification methods and similarity-based heuristics for clause relevance prediction, and generation-based methods for clause recommendation, and evaluate the results from various methods on several clause types.We provide analyses on the results, and further outline the advantages and limitations of the various methods for this line of research. Vinay Aggarwal, Aparna Garimella, Balaji Vasan Srinivasan, Anandhavelu Natarajan, Rajiv Jain |
EMNLP (1) | 2 |
| 2021 | AUTOSUMM: Automatic Model Creation for Text SummarizationabstractSharmila Reddy Nangi, Atharv Tyagi, Jay Mundra, Sagnik Mukherjee, Raj Snehal, Niyati Chhaya, Aparna Garimella. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Sharmila Reddy Nangi, Atharv Tyagi, Jay Mundra, Sagnik Mukherjee, Raj Snehal, Niyati Chhaya, Aparna Garimella |
EMNLP (1) | 7 |
| 2020 | "Judge me by my size (noun), do you?" YodaLib: A Demographic-Aware Humor Generation FrameworkabstractThe subjective nature of humor makes computerized humor generation a challenging task.We propose an automatic humor generation framework for filling the blanks in Mad Libs R stories, while accounting for the demographic backgrounds of the desired audience.We collect a dataset consisting of such stories, which are filled in and judged by carefully selected workers on Amazon Mechanical Turk.We build upon the BERT platform to predict location-biased word fillings in incomplete sentences, and we fine-tune BERT to classify location-specific humor in a sentence.We leverage these components to produce YODALIB, a fully-automated Mad Libs style humor generation framework, which selects and ranks appropriate candidate words and sentences in order to generate a coherent and funny story tailored to certain demographics.Our experimental results indicate that YODALIB outperforms a previous semi-automated approach proposed for this task, while also surpassing human annotators in both qualitative and quantitative analyses. Aparna Garimella, Carmen Banea, Nabil Hossain, Rada Mihalcea |
COLING | 1 |
| 2020 | Understanding and Explicitly Measuring Linguistic and Stylistic Properties of Deception via Generation and TranslationabstractMassive digital disinformation is one of the main risks of modern society.Hundreds of models and linguistic analyses have been done to compare and contrast misleading and credible content online.However, most models do not remove the confounding factor of a topic or narrative when training, so the resulting models learn a clear topical separation for misleading versus credible content.We study the feasibility of using two strategies to disentangle the topic bias from the models to understand and explicitly measure linguistic and stylistic properties of content from misleading versus credible content.First, we develop conditional generative models to create news content that is characteristic of different credibility levels.We perform multi-dimensional evaluation of model performance on mimicking both the style and linguistic differences that distinguish news of different credibility using machine translation metrics and classification models.We show that even though generative models are able to imitate both the style and language of the original content, additional conditioning on both the news category and the topic leads to reduced performance.In a second approach, we perform deception style "transfer" by translating deceptive content into the style of credible content and vice versa.Extending earlier studies, we demonstrate that, when conditioned on a topic, deceptive content is shorter, less readable, more biased, and more subjective than credible content, and transferring the style from deceptive to credible content is more challenging than the opposite direction. Emily Saldanha, Aparna Garimella, Svitlana Volkova |
INLG | 2 |
| 2019 | Women's Syntactic Resilience and Men's Grammatical Luck: Gender-Bias in Part-of-Speech Tagging and Dependency ParsingabstractSeveral linguistic studies have shown the prevalence of various lexical and grammatical patterns in texts authored by a person of a particular gender, but models for part-of-speech tagging and dependency parsing have still not adapted to account for these differences.To address this, we annotate the Wall Street Journal part of the Penn Treebank with the gender information of the articles' authors, and build taggers and parsers trained on this data that show performance differences in text written by men and women.Further analyses reveal numerous part-of-speech tags and syntactic relations whose prediction performances benefit from the prevalence of a specific gender in the training data.The results underscore the importance of accounting for gendered differences in syntactic tasks, and outline future venues for developing more accurate taggers and parsers.We release our data to the research community. Aparna Garimella, Carmen Banea, Eduard H. Hovy, Rada Mihalcea |
ACL (1) | 1 |
| 2017 | Demographic-aware word associationsabstractVariations of word associations across different groups of people can provide insights into people's psychologies and their world views.To capture these variations, we introduce the task of demographicaware word associations.We build a new gold standard dataset consisting of word association responses for approximately 300 stimulus words, collected from more than 800 respondents of different gender (male/female) and from different locations (India/United States), and show that there are significant variations in the word associations made by these groups.We also introduce a new demographic-aware word association model based on a neural net skip-gram architecture, and show how computational methods for measuring word associations that specifically account for writer demographics can outperform generic methods that are agnostic to such information. Aparna Garimella, Carmen Banea, Rada Mihalcea |
EMNLP | 1 |
| 2016 | Identifying Cross-Cultural Differences in Word UsageabstractPersonal writings have inspired researchers in the fields of linguistics and psychology to study the relationship between language and culture to better understand the psychology of people across different cultures. In this paper, we explore this relation by developing cross-cultural word models to identify words with cultural bias – i.e., words that are used in significantly different ways by speakers from different cultures. Focusing specifically on two cultures: United States and Australia, we identify a set of words with significant usage differences, and further investigate these words through feature analysis and topic modeling, shedding light on the attributes of language that contribute to these differences. Aparna Garimella, Rada Mihalcea, James W. Pennebaker |
COLING | 1 |