Bharathi Raja Chakravarthi

dblp:241/0168 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
13since 2021 · last 2027
0000-0002-4575-7934ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2027 SegTM-SGD: Few-shot emotion classification framework using meta-learning paradigm
Jairam R., Jyothish Lal G., B. Premjith, Bharathi Raja Chakravarthi, Barathi Ganesh H. B.
Comput. Speech Lang.4
2026 INCLUDE: A chain-of-thought based mixture of experts for inclusive language generation
abstract
• Inclusive language generation using Retrieval-Augmented Generation and expert-driven reasoning. • Chain-of-Thought based Mixture of Experts to ensure fairness, neutrality and coherence. • Model-agnostic design with validated bias mitigation across six social dimensions. Developing intelligent inclusive language generation systems that promote inclusivity and mitigate harmful or exclusive terms is a key challenge in advancing Equity, Diversity and Inclusion (EDI) principles. Using inclusive language in communication helps create a respectful, bias free and mutual understanding between the peers, which is also essential for organizations to promote safe and equitable workspaces. Inclusive language involves neutrality, tone sensitivity and fairness across diverse contexts, enabling meaningful engagement without marginalization. Considering this, we propose INCLUDE, an inclusive language generator designed for promoting respectful and equitable communication in workplaces. We curated a non-inclusive vs inclusive pair dataset including real-world workplace, advertisements and HR policy discourse with annotated inclusive rewrites. The proposed framework employs three experts dedicated to inclusiveness, bias and stereotype free and contextual relevance to facilitate the learning of diverse semantics. To optimize these experts, we propose a self-calibration mechanism using meta-prompting guided by a novel multi-dimensional reward function. Extensive evaluations, including metrics LLM based assessments and human in the loop analysis shows that proposed model effectively counters non-inclusive language into contextually appropriate, inclusive and accessible responses while maintaining the original intent.
Bharathi Raja Chakravarthi, Shunmuga Priya Muthusamy Chinnan, Prasanna Kumar Kumaresan, Rahul Ponnusamy
Knowl. Based Syst.1
2025 Benchmarking Hindi Term Extraction in Education: A Dataset and Analysis
abstract
This paper introduces the HTEC HindiTerm Extraction Dataset 2.0, a resourcedesigned to support terminology extractionand classification tasks within the education domain. HTEC 2.0 has been developed with the objective of providing a high-quality benchmark dataset for the evaluation of term recognition and classification methodologies in Hindi educationaldiscourse. The dataset consists of 97 documents sourced from Hindi Wikipedia, covering a diverse range of topics relevant tothe education sector. Within these documents, 1,702 terms have been manuallyannotated where each term is defined as asingle-word or multi-word expression thatconveys a domain-specific meaning. Theannotated terms in HTEC 2.0 are systematically categorized into seven distinct classes.Furthermore, this paper outlines the development of annotation guidelines, detailingthe criteria used to determine term boundaries and category assignments. By offeringa structured dataset with clearly definedterm classifications, HTEC 2.0 serves as avaluable resource for researchers workingon terminology extraction, domain-specificnamed entity recognition, and text classification in Hindi.
Shubhanker Banerjee, Bharathi Raja Chakravarthi, John P. McCrae
LDK2
2025 Gender inclusive language generation framework: A reasoning approach with RAG and CoT
abstract
Language is a dynamic and evolving concept that shapes thought and perception. The increasing reliance on Natural Language Processing models necessitates careful consideration of their alignment with inclusive language practices. However, Large Language Models often perpetuate biases due to training on androcentric and stereotypical data, undermining fairness and inclusivity. To address this, we propose a novel Two-Pass Retrieval Augmented Generation RAG with Chain of Thought framework that first retrieves contextual, unbiased references from a created corpus of inclusive texts and then applies structured, step-by-step reasoning via CoT prompting to enhance inclusivity in LLM output. By systematically retrieving relevant, unbiased references and enforcing structured reasoning, the framework promotes the generation of more inclusive and less biased content. Both LLM and human based evaluation using structured prompts with metrics like gender assumption, gender neutrality and quality and relevance Score are utilized. The text completion and generation tasks demonstrate that the proposed framework reduces gender bias.
Shunmuga Priya Muthusamy Chinnan, Meghann L. Drury-Grogan, Bharathi Raja Chakravarthi
Knowl. Based Syst.3
2025 A Multimodal Approach for Hate and Offensive Content Detection in Tamil: From Corpus Creation to Model Development
abstract
Detecting hate speech on social media platforms is vital to mitigate technology-facilitated violence. Extensive research has targeted widely spoken languages like English, but there is a notable gap in studying hate speech detection in low-resource languages like Tamil. Additionally, with social media platforms now supporting various modalities, including text, speech, and video, effective techniques for hate speech detection in multimodal formats, especially videos, are crucial. However, detecting hate speech in Tamil presents unique challenges due to its morphology and code-mixing nature. This article presents a comprehensive approach for detecting hate speech in Tamil, with a focus on multimodal data. We introduce a new dataset, the MultimodAl Tamil Hate dataset, comprising videos along with their audio and textual transcripts, annotated with four categories of hate speech: offensive, sexist, racist, and casteist. To classify hate speech categories, we leverage transformer-based models. Through a series of experiments, we evaluated the performance of each modality individually and explored their fusion using a multimodal approach. BERT-based models were used for textual analysis to extract informative features, the TimeSformer model was employed for video modality, and Wav2Vec2 was used for audio modality. Specifically, we attained 81.82% accuracy and a 68.65% F1-score for the text modality, 63.63% accuracy and a 50.60% F1-score for the audio modality, and 45.45% accuracy and a 33.64% F1-score for the video modality. By integrating optimal combinations of models from each modality and employing machine learning classifiers, we achieved an accuracy of 81.82% and an F1-score of 66.67% in our hate speech classification task. Our research findings highlight the effectiveness of employing a multimodal approach for hate speech detection in Tamil, showcasing its efficacy in curbing the dissemination of hateful content on social media platforms.
Jayanth Mohan, Spandana Reddy Mekapati, B. Premjith, Jyothish Lal G., Bharathi Raja Chakravarthi
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2024 Dataset for Identification of Homophobia and Transphobia for Telugu, Kannada, and Gujarati
abstract
Users of social media platforms are negatively affected by the proliferation of hate or abusive content. There has been a rise in homophobic and transphobic content in recent years targeting LGBT+ individuals. The increasing levels of homophobia and transphobia online can make online platforms harmful and threatening for LGBT+ persons, potentially inhibiting equality, diversity, and inclusion. We are introducing a new dataset for three languages, namely Telugu, Kannada, and Gujarati. Additionally, we have created an expert-labeled dataset to automatically identify homophobic and transphobic content within comments collected from YouTube. We provided comprehensive annotation rules to educate annotators in this process. We collected approximately 10,000 comments from YouTube for all three languages. Marking the first dataset of these languages for this task, we also developed a baseline model with pre-trained transformers.
Prasanna Kumar Kumaresan, Rahul Ponnusamy, Dhruv Sharma, Paul Buitelaar, Bharathi Raja Chakravarthi
LREC/COLING5
2024 From Laughter to Inequality: Annotated Dataset for Misogyny Detection in Tamil and Malayalam Memes
abstract
In this digital era, memes have become a prevalent online expression, humor, sarcasm, and social commentary. However, beneath their surface lies concerning issues such as the propagation of misogyny, gender-based bias, and harmful stereotypes. To overcome these issues, we introduced MDMD (Misogyny Detection Meme Dataset) in this paper. This article focuses on creating an annotated dataset with detailed annotation guidelines to delve into online misogyny within the Tamil and Malayalam-speaking communities. Through analyzing memes, we uncover the intricate world of gender bias and stereotypes in these communities, shedding light on their manifestations and impact. This dataset, along with its comprehensive annotation guidelines, is a valuable resource for understanding the prevalence, origins, and manifestations of misogyny in various contexts, aiding researchers, policymakers, and organizations in developing effective strategies to combat gender-based discrimination and promote equality and inclusivity. It enables a deeper understanding of the issue and provides insights that can inform strategies for cultivating a more equitable and secure online environment. This work represents a crucial step in raising awareness and addressing gender-based discrimination in the digital space.
Rahul Ponnusamy, Kathiravan Pannerselvam, Saranya Rajiakodi, Prasanna Kumar Kumaresan, Sajeetha Thavareesan, Bhuvaneswari Sivagnanam, Anshid K. A., Susminu S. Kumar, Paul Buitelaar, Bharathi Raja Chakravarthi
LREC/COLING10
2024 Large Language Models for Few-Shot Automatic Term Extraction
Shubhanker Banerjee, Bharathi Raja Chakravarthi, John P. McCrae
NLDB (1)2
2024 Overlapping word removal is all you need: revisiting data imbalance in hope speech detection
abstract
Hope speech detection is a new task for finding and highlighting positive comments or supporting content from user-generated social media comments. For this task, we have used a Shared Task multilingual dataset on Hope Speech Detection for Equality, Diversity, and Inclusion (HopeEDI) for three languages English, code-switched Tamil and Malayalam. In this paper, we present deep learning techniques using context-aware string embeddings for word representations and Recurrent Neural Network (RNN) and pooled document embeddings for text representation. We have evaluated and compared the three models for each language with different approaches. Our proposed methodology works fine and achieved higher performance than baselines. The highest weighted average F-scores of 0.93, 0.58, and 0.84 are obtained on the task organisers{'} final evaluation test set. The proposed models are outperforming the baselines by 3{\%}, 2{\%} and 11{\%} in absolute terms for English, Tamil and Malayalam respectively.
Hariharan RamakrishnaIyer LekshmiAmmal, Manikandan Ravikiran, Gayathri Nisha, Navyasree Balamuralidhar, Adithya Madhusoodanan, Anand Kumar Madasamy, Bharathi Raja Chakravarthi
J. Exp. Theor. Artif. Intell.7
2023 MG2P: An Empirical Study Of Multilingual Training for Manx G2P
Shubhanker Banerjee, Bharathi Raja Chakravarthi, John P. McCrae
LDK2
2023 TrollsWithOpinion: A taxonomy and dataset for predicting domain-specific opinion manipulation in troll memes
Shardul Suryawanshi, Bharathi Raja Chakravarthi, Mihael Arcan, Paul Buitelaar
Multim. Tools Appl.2
2022 Thirumurai: A Large Dataset of Tamil Shaivite Poems and Classification of Tamil Pann
abstract
Thirumurai, also known as Panniru Thirumurai, is a collection of Tamil Shaivite poems dating back to the Hindu revival period between the 6th and the 10th century. These poems are par excellence, in both literary and musical terms. They have been composed based on the ancient, now non-existent Tamil Pann system and can be set to music. We present a large dataset containing all the Thirumurai poems and also attempt to classify the Pann and author of each poem using transformer based architectures. Our work is the first of its kind in dealing with ancient Tamil text datasets, which are severely under-resourced. We explore several Deep Learning-based techniques for solving this challenge effectively and provide essential insights into the problem and how to address it.
Shankar Mahadevan, Rahul Ponnusamy, Prasanna Kumar Kumaresan, Prabakaran Chandran, Ruba Priyadharshini, Sangeetha Sivanesan, Bharathi Raja Chakravarthi
LREC7
2022 Offensive language detection in Tamil YouTube comments by adapters and cross-domain knowledge transfer
Rahul Ponnusamy, Sean Benhur, Shanmuga Vadivel Kogilavani, Adhithiya Ganesan, Deepti Ravi, Gowtham Krishnan Shanmugasundaram, Ruba Priyadharshini, Bharathi Raja Chakravarthi
Comput. Speech Lang.9
2020 Unsupervised Deep Language and Dialect Identification for Short Texts
abstract
Automatic Language Identification (LI) or Dialect Identification (DI) of short texts of closely related languages or dialects, is one of the primary steps in many natural language processing pipelines.Language identification is considered a solved task in many cases; however, in the case of very closely related languages, or in an unsupervised scenario (where the languages are not known in advance), performance is still poor.In this paper, we propose the Unsupervised Deep Language and Dialect Identification (UDLDI) method, which can simultaneously learn sentence embeddings and cluster assignments from short texts.The UDLDI model understands the sentence constructions of languages by applying attention to character relations which helps to optimize the clustering of languages.We have performed our experiments on three shorttext datasets for different language families, each consisting of closely related languages or dialects, with very minimal training sets.Our experimental evaluations on these datasets have shown significant improvement over state-of-the-art unsupervised methods and our model has outperformed state-of-the-art LI and DI systems in supervised settings.
Koustava Goswami, Rajdeep Sarkar, Bharathi Raja Chakravarthi, Theodorus Fransen, John P. McCrae
COLING3
2020 Classification Benchmarks for Under-resourced Bengali Language based on Multichannel Convolutional-LSTM Network
abstract
Exponential growths of social media and micro-blogging sites not only provide platforms for empowering freedom of expressions and individual voices, but also enables people to express anti-social behavior like online harassment, cyberbul-lying, and hate speech. Numerous works have been proposed to utilize these data for social and anti-social behavior analysis, document characterization, and sentiment analysis by predicting the contexts mostly for highly resourced languages like English. However, some languages are under-resources, e.g., South Asian languages like Bengali, Tamil, Assamese, Malayalam, that lack of computational resources for natural language processing. In this paper1, we provide several classification benchmarks for Bengali, an under-resourced language. We prepared three datasets of expressing hate, commonly used topics, and opinions for hate speech detection, document classification, and sentiment analysis. We built the largest Bengali word embedding models to date based on 250 million articles, which we call BengFastText. We perform three experiments, covering document classification, sentiment analysis, and hate speech detection. We incorporate word embeddings into a Multichannel Convolutional-LSTM (MC-LSTM) network for predicting different types of hate speech, document classification, and sentiment analysis. Experiments demonstrate that BengFastText can capture the semantics of words from respective contexts correctly. Evaluations against several baseline embedding models, e.g., Word2Vec and GloVe yield up to 92.30%, 82.25%, and 90.45% F1-scores in case of document classification, sentiment analysis, and hate speech detection, respectively during 5-fold cross-validation tests.
Md. Rezaul Karim 0001, Bharathi Raja Chakravarthi, John P. McCrae, Michael Cochez
DSAA2
2019 Comparison of Different Orthographies for Machine Translation of Under-Resourced Dravidian Languages
abstract
Under-resourced languages are a significant challenge for statistical approaches to machine translation, and recently it has been shown that the usage of training data from closely-related languages can improve machine translation quality of these languages. While languages within the same language family share many properties, many under-resourced languages are written in their own native script, which makes taking advantage of these language similarities difficult. In this paper, we propose to alleviate the problem of different scripts by transcribing the native script into common representation i.e. the Latin script or the International Phonetic Alphabet (IPA). In particular, we compare the difference between coarse-grained transliteration to the Latin script and fine-grained IPA transliteration. We performed experiments on the language pairs English-Tamil, English-Telugu, and English-Kannada translation task. Our results show improvements in terms of the BLEU, METEOR and chrF scores from transliteration and we find that the transliteration into the Latin script outperforms the fine-grained IPA transcription.
Bharathi Raja Chakravarthi, Mihael Arcan, John P. McCrae
LDK1
2019 Leveraging Rule-Based Machine Translation Knowledge for Under-Resourced Neural Machine Translation Models
Daniel Torregrosa, Nivranshu Pasricha, Maraim Masoud, Bharathi Raja Chakravarthi, Juan A. Alonso, Noe Casas, Mihael Arcan
MTSummit (2)4
2018 Improving Wordnets for Under-Resourced Languages Using Machine Translation
abstract
Wordnets are extensively used in natural language processing, but the current approaches for manually building a wordnet from scratch involves large research groups for a long period of time, which are typically not available for under-resourced languages.Even if wordnet-like resources are available for under-resourced languages, they are often not easily accessible, which can alter the results of applications using these resources.Our proposed method presents an expand approach for improving and generating wordnets with the help of machine translation.We apply our methods to improve and extend wordnets for the Dravidian languages, i.e., Tamil, Telugu, Kannada, which are severly under-resourced languages.We report evaluation results of the generated wordnet senses in term of precision for these languages.In addition to that, we carried out a manual evaluation of the translations for the Tamil language, where we demonstrate that our approach can aid in improving wordnet resources for under-resourced Dravidian languages.
Bharathi Raja Chakravarthi, Mihael Arcan, John P. McCrae
GWC1