Harika Abburi

dblp:184/2150 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0001-8280-6152ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Toward Robust Generative AI Text Detection: Generalizable Neural Model
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in generating text that closely resembles human writing across wide range of styles and genres. However, such capabilities are prone to potential misuse, such as fake news generation, spam email creation, and misuse in academic assignments. Hence, it is essential to build automated approaches capable of distinguishing between Artificial Intelligence-generated text and human-authored text. In this paper, we proposed a fine-tuning based neural model which is a combination of transformer models, linguistic features and state-of-the-art embedding models. We have also curated a training dataset encompassing diverse samples from different LLMs and domains to fine-tune pretrained language models. We evaluated our model's performance against state-of-the-art methods, and the comparative analysis demonstrates that our approach consistently outperforms other methods across various datasets using established evaluation metrics.
Harika Abburi, Nirmala Pudota, Balaji Veeramani, Edward Bowen, Sanmitra Bhattacharya
ICMLA1
2024 Pre-train. Mixup and Fine-tune: A Simple Strategy to Handle Domain Shift
abstract
Transfer learning leverages models trained on large source datasets to target domains with limited datasets by fine- tuning pre-trained models. These approaches work well with minimal distribution shifts. However, the nature of distribution shift is unknown in real-world applications which may lead to worse performance of models in the field. Domain adaptation approaches handles this issue explicitly by adapting source trained models using discrepancy, adversarial or reconstruction based approaches, but doesn't tackle the limited target data issue. A data augmentation approach such as mixup helps train resilient models with limited datasets, and is recently being considered for its ability to handle distribution shifts. In this work, we investigate how mixup can be used along with transfer learning to improve model performance on target domains with distribution shifts. Our experimental results shows mixup is complementary to transfer learning which we demonstrate by varying the percentage of available target training data with publicly available source and target datasets. Our proposed approach of using mixup along with fine tuning shows improved performance than just fine tuning or mixup across varying percentage of target dataset sizes.
Haider Ilyas, Harika Abburi, Edward Bowen, Balaji Veeramani
ICMLA2
2024 Multi-task learning neural framework for categorizing sexism
Harika Abburi, Pulkit Parikh, Niyati Chhaya, Vasudeva Varma
Comput. Speech Lang.1
2023 An Ensemble-Based Approach for Generative Language Model Attribution
Harika Abburi, Michael Suesserman, Nirmala Pudota, Balaji Veeramani, Edward Bowen, Sanmitra Bhattacharya
WISE1
2022 Leveraging Mental Health Forums for User-level Depression Detection on Social Media
abstract
The number of depression and suicide risk cases on social media platforms is ever-increasing, and the lack of depression detection mechanisms on these platforms is becoming increasingly apparent. A majority of work in this area has focused on leveraging linguistic features while dealing with small-scale datasets. However, one faces many obstacles when factoring into account the vastness and inherent imbalance of social media content. In this paper, we aim to optimize the performance of user-level depression classification to lessen the burden on computational resources. The resulting system executes in a quicker, more efficient manner, in turn making it suitable for deployment. To simulate a platform agnostic framework, we simultaneously replicate the size and composition of social media to identify victims of depression. We systematically design a solution that categorizes post embeddings, obtained by fine-tuning transformer models such as RoBERTa, and derives user-level representations using hierarchical attention networks. We also introduce a novel mental health dataset to enhance the performance of depression categorization. We leverage accounts of depression taken from this dataset to infuse domain-specific elements into our framework. Our proposed methods outperform numerous baselines across standard metrics for the task of depression detection in text.
Sravani Boinepelli, Tathagata Raha, Harika Abburi, Pulkit Parikh, Niyati Chhaya, Vasudeva Varma
LREC3
2021 Fine-Grained Multi-label Sexism Classification Using a Semi-Supervised Multi-level Neural Approach
abstract
Abstract Sexism, a permeate form of oppression, causes profound suffering through various manifestations. Given the increasing number of experiences of sexism shared online, categorizing these recollections automatically can support the battle against sexism, since it can promote successful evaluations by gender studies researchers and government representatives engaged in policy making. In this paper, we examine the fine-grained, multi-label classification of accounts (reports) of sexism. To the best of our knowledge, we consider substantially more categories of sexism than any related prior work through our 23-class problem formulation. Moreover, we present the first semi-supervised work for the multi-label classification of accounts describing any type(s) of sexism. We devise self-training-based techniques tailor-made for the multi-label nature of the problem to utilize unlabeled samples for augmenting the labeled set. We identify high textual diversity with respect to the existing labeled set as a desirable quality for candidate unlabeled instances and develop methods for incorporating it into our approach. We also explore ways of infusing class imbalance alleviation for multi-label classification into our semi-supervised learning, independently and in conjunction with the method involving diversity. In addition to data augmentation methods, we develop a neural model which combines biLSTM and attention with a domain-adapted BERT model in an end-to-end trainable manner. Further, we formulate a multi-level training approach in which models are sequentially trained using categories of sexism of different levels of granularity. Moreover, we devise a loss function that exploits any label confidence scores associated with the data. Several proposed methods outperform various baselines on a recently released dataset for multi-label sexism categorization across several standard metrics.
Harika Abburi, Pulkit Parikh, Niyati Chhaya, Vasudeva Varma
Data Sci. Eng.1
2021 Categorizing Sexism and Misogyny through Neural Approaches
abstract
Sexism, an injustice that subjects women and girls to enormous suffering, manifests in blatant as well as subtle ways. In the wake of growing documentation of experiences of sexism on the web, the automatic categorization of accounts of sexism has the potential to assist social scientists and policymakers in studying and thereby countering sexism. The existing work on sexism classification has certain limitations in terms of the categories of sexism used and/or whether they can co-occur. To the best of our knowledge, this is the first work on the multi-label classification of sexism of any kind(s). 1 We also consider the related task of misogyny classification. While sexism classification is performed on textual accounts describing sexism suffered or observed, misogyny classification is carried out on tweets perpetrating misogyny. We devise a novel neural framework for classifying sexism and misogyny that can combine text representations obtained using models such as Bidirectional Encoder Representations from Transformers with distributional and linguistic word embeddings using a flexible architecture involving recurrent components and optional convolutional ones. Further, we leverage unlabeled accounts of sexism to infuse domain-specific elements into our framework. To evaluate the versatility of our neural approach for tasks pertaining to sexism and misogyny, we experiment with adapting it for misogyny identification. For categorizing sexism, we investigate multiple loss functions and problem transformation techniques to address the multi-label problem formulation. We develop an ensemble approach using a proposed multi-label classification model with potentially overlapping subsets of the category set. Proposed methods outperform several deep-learning as well as traditional machine learning baselines for all three tasks.
Pulkit Parikh, Harika Abburi, Niyati Chhaya, Manish Gupta 0001, Vasudeva Varma
ACM Trans. Web2
2020 Semi-supervised Multi-task Learning for Multi-label Fine-grained Sexism Classification
abstract
Sexism, a form of oppression based on one’s sex, manifests itself in numerous ways and causes enormous suffering. In view of the growing number of experiences of sexism reported online, categorizing these recollections automatically can assist the fight against sexism, as it can facilitate effective analyses by gender studies researchers and government officials involved in policy making. In this paper, we investigate the fine-grained, multi-label classification of accounts (reports) of sexism. To the best of our knowledge, we work with considerably more categories of sexism than any published work through our 23-class problem formulation. Moreover, we propose a multi-task approach for fine-grained multi-label sexism classification that leverages several supporting tasks without incurring any manual labeling cost. Unlabeled accounts of sexism are utilized through unsupervised learning to help construct our multi-task setup. We also devise objective functions that exploit label correlations in the training data explicitly. Multiple proposed methods outperform the state-of-the-art for multi-label sexism classification on a recently released dataset across five standard metrics.
Harika Abburi, Pulkit Parikh, Niyati Chhaya, Vasudeva Varma
COLING1
2020 Fine-grained Multi-label Sexism Classification Using Semi-supervised Learning
Harika Abburi, Pulkit Parikh, Niyati Chhaya, Vasudeva Varma
WISE (2)1
2019 Multi-label Categorization of Accounts of Sexism using a Neural Framework
abstract
Pulkit Parikh, Harika Abburi, Pinkesh Badjatiya, Radhika Krishnan, Niyati Chhaya, Manish Gupta, Vasudeva Varma. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Pulkit Parikh, Harika Abburi, Pinkesh Badjatiya, Radhika Krishnan, Niyati Chhaya, Manish Gupta 0001, Vasudeva Varma
EMNLP/IJCNLP (1)2