VLDB 2026 Research / reviewers in the wild / expert
Abhilasha Sancheti
dblp:210/2594
· DBLP profile ↗
11ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 29% Representation and self-supervised learning · 21% Language models and text generation · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 11 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
1.6 | 2 | 2025 | On the Mutual Influence of Gender and Occupation in LLM Representations · ACL (1) 2025 On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis › document understanding
legal text analysis |
1.2 | 2 | 2023 | What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and Prohibitions · EMNLP 2023 Agent-Specific Deontic Modality Detection in Legal Language · EMNLP 2022 |
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
0.9 | 1 | 2025 | Less Mature is More Adaptable for Sentence-level Language Modeling · ACL (1) 2025 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.9 | 1 | 2025 | On the Mutual Influence of Gender and Occupation in LLM Representations · ACL (1) 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
relation prediction |
0.8 | 1 | 2024 | On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models · EMNLP 2024 |
Machine learning › Trustworthy machine learning › fairness › bias in language models
social bias in language models |
0.8 | 1 | 2024 | On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models · EMNLP 2024 |
Natural language and speech › Language models and text generation
text summarization |
0.7 | 1 | 2023 | What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and Prohibitions · EMNLP 2023 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
0.6 | 1 | 2022 | Entailment Relation Aware Paraphrase Generation · AAAI 2022 |
Information retrieval
ranking |
0.2 | 1 | 2023 | What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and Prohibitions · EMNLP 2023 |
Machine learning › Deep learning architectures and training
data augmentation |
0.2 | 1 | 2022 | Entailment Relation Aware Paraphrase Generation · AAAI 2022 |
Machine learning › Deep learning architectures and training › data augmentation
paraphrase augmentation |
0.2 | 1 | 2022 | Entailment Relation Aware Paraphrase Generation · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
probing · 1.7name-replacement experiments · 1.5contextualized embeddings · 1.5pairwise importance ranking · 1.3human evaluation · 1.3mean pooling · 0.9fine-tuning · 0.9embedding analysis · 0.9reinforcement learning · 0.6natural language inference · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Mutual Influence of Gender and Occupation in LLM RepresentationsabstractWe examine LLM representations of gender for first names in various occupational contexts to study how occupations and the gender perception of first names in LLMs influence each other mutually.We find that LLMs' first-name gender representations correlate with real-world gender statistics associated with the name, and are influenced by the co-occurrence of stereotypically feminine or masculine occupations.Additionally, we study the influence of firstname gender representations on LLMs in a downstream occupation prediction task and their potential as an internal metric to identify extrinsic model biases.While feminine firstname embeddings often raise the probabilities for female-dominated jobs (and vice versa for male-dominated jobs), reliably using these internal gender representations for bias detection remains challenging.Is Jody male or female? Female MaleJody is a nurse.Is Jody male or female? Female MaleJody is a comedian.Is Jody male or female? Female Male Haozhe An, Connor Baumler, Abhilasha Sancheti, Rachel Rudinger |
ACL (1) | 3 |
| 2025 | Less Mature is More Adaptable for Sentence-level Language ModelingabstractThis work investigates sentence-level models (i.e., models that operate at the sentence-level) to study how sentence representations from various encoders influence downstream task performance, and which syntactic, semantic, and discourse-level properties are essential for strong performance.Our experiments encompass encoders with diverse training regimes and pretraining domains, as well as various pooling strategies applied to multi-sentence input tasks (including sentence ordering, sentiment classification, and natural language inference) requiring coarse-to-fine-grained reasoning.We find that "less mature" representations (e.g., mean-pooled representations from BERT's first or last layer, or representations from encoders with limited fine-tuning) exhibit greater generalizability and adaptability to downstream tasks compared to representations from extensively fine-tuned models (e.g.,, SBERT or Sim-CSE).These findings are consistent across different pretraining seed initializations for BERT.Our probing analysis reveals that syntactic and discourse-level properties are stronger indicators of downstream performance than MTEB scores or decodability.Furthermore, the data and time efficiency of sentence-level models, often outperforming token-level models, underscores their potential for future research. Abhilasha Sancheti, David Dale, Artyom Kozhevnikov, Maha Elbayad |
ACL (1) | 1 |
| 2024 | On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language ModelsabstractWe study the presence of heteronormative biases and prejudice against interracial romantic relationships in large language models by performing controlled name-replacement experiments for the task of relationship prediction.We show that models are less likely to predict romantic relationships for (a) same-gender character pairs than different-gender pairs; and (b) intra/inter-racial character pairs involving Asian names as compared to Black, Hispanic, or White names.We examine the contextualized embeddings of first names and find that gender for Asian names is less discernible than non-Asian names.We discuss the social implications of our findings, underlining the need to prioritize the development of inclusive and equitable technology.2 Except for Hispanic wherein we did not get any names in 5 -10% bin and only 1 name in 25 -50% bin.480 0 -2 2 -5 5 -1 0 1 0 -2 5 2 5 -5 0 5 0 -7 5 7 5 -9 0 9 0 -9 5 9 5 -9 8 9 8 -1 0 0 % Female 0-2 2-5 5-10 10-25 25-50 50-75 75-90 90-95 95-98 98-100 % Female 0.56 0.55 0.59 0.58 0.61 0.59 0.69 0.68 0.64 0.72 0.56 0.52 0.57 0.58 0.59 0.58 0.65 0.66 0.63 0.69 0.61 0.58 0.63 0.60 0.64 0.62 0.70 0.69 0.67 0.72 0.58 0.57 0.59 0.60 0.62 0.59 0.67 0.66 0.63 0.70 0.61 0.60 0.64 0.62 0.65 0.62 0.69 0.69 0.66 0.69 0.59 0.56 0.61 0.60 0.61 0.60 0.66 0.64 0.62 0.67 0.66 0.64 0.67 0.66 0.67 0.64 0.69 0.66 0.65 0.67 0.67 0.65 0.68 0.66 0.68 0.64 0.67 0.67 0.64 0.67 0.64 0.61 0.65 0.63 0.65 0.61 0.65 0.63 0.62 0.65 0.70 0.66 0.70 0.68 0.68 0.66 0.68 0.67 0.65 0.68 Male Neutral Female Male Neutral Female Asian (Recall) 0 -2 2 -5 5 -1 0 1 0 -2 5 2 5 -5 0 5 0 -7 5 7 5 -9 0 9 0 -9 5 9 5 -9 8 9 8 -1 0 0 Abhilasha Sancheti, Haozhe An, Rachel Rudinger |
EMNLP | 1 |
| 2023 | What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and ProhibitionsabstractReviewing and comprehending key obligations, entitlements, and prohibitions in legal contracts can be a tedious task due to their length and domain-specificity.Furthermore, the key rights and duties requiring review vary for each contracting party.In this work, we propose a new task of party-specific extractive summarization for legal contracts to facilitate faster reviewing and improved comprehension of rights and duties.To facilitate this, we curate a dataset comprising of party-specific pairwise importance comparisons annotated by legal experts, covering ∼293K sentence pairs that include obligations, entitlements, and prohibitions extracted from lease agreements.Using this dataset, we train a pairwise importance ranker and propose a pipeline-based extractive summarization system that generates a party-specific contract summary.We establish the need for incorporating domain-specific notion of importance during summarization by comparing our system against various baselines using both automatic and human evaluation methods 1 . Abhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, Rachel Rudinger |
EMNLP | 1 |
| 2023 | SALAD : Source-free Active Label-Agnostic Domain Adaptation for Classification, Segmentation and DetectionabstractWe present a novel method, SALAD, for the challenging vision task of adapting a pre-trained "source" domain network to a "target" domain, with a small budget for annotation in the "target" domain and a shift in the label space. Further, the task assumes that the source data is not available for adaptation, due to privacy concerns or otherwise. We postulate that such systems need to jointly optimize the dual task of (i) selecting fixed number of samples from the target domain for annotation and (ii) transfer of knowledge from the pre-trained network to the target domain. To do this, SALAD consists of a novel Guided Attention Transfer Network (GATN) and an active learning function, HAL. The GATN enables feature distillation from pre-trained network to the target network, complemented with the target samples mined by HALusing transfer-ability and uncertainty criteria. SALAD has three key benefits: (i) it is task-agnostic, and can be applied across various visual tasks such as classification, segmentation and detection; (ii) it can handle shifts in output label space from the pre-trained source network to the target domain; (iii) it does not require access to source data for adaptation. We conduct extensive experiments across 3 visual tasks, viz. digits classification (MNIST, SVHN, VISDA), synthetic (GTA5) to real (CityScapes) image segmentation, and document layout detection (PubLayNet to DSSE). We show that our source-free approach, SALAD, results in an improvement of 0.5%−31.3% (across datasets and tasks) over prior adaptation methods that assume access to large amounts of annotated source data for adaptation. Code is available here. Divya Kothandaraman, Abhilasha Sancheti, Manoj Ghuhan, Tripti Shukla, Dinesh Manocha |
WACV | 3 |
| 2022 | Entailment Relation Aware Paraphrase GenerationabstractWe introduce a new task of entailment relation aware paraphrase generation which aims at generating a paraphrase conforming to a given entailment relation (e.g. equivalent, forward entailing, or reverse entailing) with respect to a given input. We propose a reinforcement learning-based weakly-supervised paraphrasing system, ERAP, that can be trained using existing paraphrase and natural language inference (NLI) corpora without an explicit task-specific corpus. A combination of automated and human evaluations show that ERAP generates paraphrases conforming to the specified entailment relation and are of good quality as compared to the baselines and uncontrolled paraphrasing systems. Using ERAP for augmenting training data for downstream textual entailment task improves performance over an uncontrolled paraphrasing system, and introduces fewer training artifacts, indicating the benefit of explicit control during paraphrasing. Abhilasha Sancheti, Balaji Vasan Srinivasan, Rachel Rudinger |
AAAI | 1 |
| 2022 | Agent-Specific Deontic Modality Detection in Legal LanguageabstractLegal documents are typically long and written in legalese, which makes it particularly difficult for laypeople to understand their rights and duties.While natural language understanding technologies can be valuable in supporting such understanding in the legal domain, the limited availability of datasets annotated for deontic modalities in the legal domain, due to the cost of hiring experts and privacy issues, is a bottleneck.To this end, we introduce, LEXDE-MOD, a corpus of English contracts annotated with deontic modality expressed with respect to a contracting party or agent along with the modal triggers.We benchmark this dataset on two tasks: (i) agent-specific multi-label deontic modality classification, and (ii) agent-specific deontic modality and trigger span detection using Transformer-based (Vaswani et al., 2017) language models.Transfer learning experiments show that the linguistic diversity of modal expressions in LEXDEMOD generalizes reasonably from lease to employment and rental agreements.A small case study indicates that a model trained on LEXDEMOD can detect red flags with high recall.We believe our work offers a new research direction for deontic modality detection in the legal domain 1 . Abhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, Rachel Rudinger |
EMNLP | 1 |
| 2021 | Multi-Style Transfer with Discriminative Feedback on Disjoint CorpusabstractNavita Goyal, Balaji Vasan Srinivasan, Anandhavelu N, Abhilasha Sancheti. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Navita Goyal, Balaji Vasan Srinivasan, Anandhavelu Natarajan, Abhilasha Sancheti |
NAACL-HLT | 4 |
| 2020 | Reinforced Rewards Framework for Text Style Transfer
Abhilasha Sancheti, Kundan Krishna, Balaji Vasan Srinivasan, Anandhavelu Natarajan |
ECIR (1) | 1 |
| 2020 | Goal-driven Command Recommendations for AnalystsabstractRecent times have seen data analytics software applications become an integral part of the decision-making process of analysts. The users of these software applications generate a vast amount of unstructured log data. These logs contain clues to the user’s goals, which traditional recommender systems may find difficult to model implicitly from the log data. With this assumption, we would like to assist the analytics process of a user through command recommendations. We categorize the commands into software and data categories based on their purpose to fulfill the task at hand. On the premise that the sequence of commands leading up to a data command is a good predictor of the latter, we design, develop, and validate various sequence modeling techniques. In this paper, we propose a framework to provide goal-driven data command recommendations to the user by leveraging unstructured logs. We use the log data of a web-based analytics software to train our neural network models and quantify their performance, in comparison to relevant and competitive baselines. We propose a custom loss function to tailor the recommended data commands according to the goal information provided exogenously. We also propose an evaluation metric that captures the degree of goal orientation of the recommendations. We demonstrate the promise of our approach by evaluating the models with the proposed metric and showcasing the robustness of our models in the case of adversarial examples, where the user activity is misaligned with selected goal, through offline evaluation. Samarth Aggarwal, Rohin Garg, Abhilasha Sancheti, Bhanu Prakash Reddy Guda, Iftikhar Ahamath Burhanuddin |
RecSys | 3 |
| 2018 | Harvesting Knowledge from Cultural Heritage Artifacts in Museums of India
Abhilasha Sancheti, Paridhi Maheshwari, Rajat Chaturvedi, Anish V. Monsy, Tanya Goyal, Balaji Vasan Srinivasan |
PAKDD (2) | 1 |