Raksha Sharma

dblp:46/7472 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 8 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Power doesn't reside in size: A Low Parameter Hybrid Language Model (HLM) for Sentiment Analysis in Code-mixed data
abstract
Code-mixed text-where multiple languages are used within the same utterance-is increasingly common in both spoken and written communication.However, it presents significant challenges for machine learning models due to the interplay of distinct grammatical structures, effectively forming a hybrid language.While fine-tuning large language models (LLMs) such as GPT-3, or Llama-3 on code-mixed data has led to performance improvements, these models still lag behind their monolingual counterparts and incur high computational costs due to the large number of trainable parameters.In this paper, we focus on the task of sentiment detection in code-mixed text and propose a Hybrid Language Model (HLM) that combines a multilingual encoder (e.g., mBERT) with a lightweight decoder (e.g., Sarvam-1) (< 3B parameters).Despite having significantly fewer trainable parameters, HLM achieves sentiment classification performance comparable to that of fine-tuned Large Language Models (LLMs) (> 7B parameters).Furthermore, our results demonstrate that HLM significantly outperforms models trained individually, underscoring its effectiveness for low-resource, codemixed sentiment analysis.
Pavan Sai Balaga, Nagasamudram Karthik, Challa Vishwanath, Raksha Sharma, Rudra Murthy, Ashish R. Mittal
EMNLP4
2025 A TextGCN-Based Decoding Approach for Improving Remote Sensing Image Captioning
abstract
Remote sensing (RS) images are highly valued for their ability to address complex real-world issues such as risk management, security, and meteorology. However, manually captioning these images is challenging and requires specialized knowledge across various domains. This letter presents an approach for automatically describing (captioning) RS images. We propose a novel encoder-decoder (ED) setup that deploys a text graph convolutional network (TextGCN) and multilayer long short-term memory (LSTM). The embeddings generated by TextGCN enhance the decoder’s understanding by capturing the semantic relationships among words at both the sentence and corpus levels. Furthermore, we advance our approach with a comparison-based beam search method to ensure fairness in the search strategy for generating the final caption. We present an extensive evaluation of our approach against various other state-of-the-art ED frameworks. We evaluated our method using seven metrics: BLEU-1 to BLEU-4, METEOR, ROUGE-L, and CIDEr. The results demonstrate that our approach significantly outperforms other state-of-the-art ED methods.
Swadhin Das, Raksha Sharma
IEEE Geosci. Remote. Sens. Lett.2
2024 Enhancing Legal Named Entity Recognition Using RoBERTa-GCN with CRF: A Nuanced Approach for Fine-Grained Entity Recognition
Arihant Jain, Raksha Sharma
ECIR (3)2
2024 Unveiling the Power of Convolutional Neural Networks: A Comprehensive Study on Remote Sensing Image Captioning and Encoder Selection
abstract
Extracting semantic information from remote sensing (RS) images has gained attention for its wide applications in defense, disaster management, and urban planning. Captioning RS images is challenging due to intricate properties like resolutions, color bands, and object types. Generating precise captions requires domain expertise, and manual annotation is time-consuming. The common approach involves using an encoder-decoder-based framework for RS image captioning, where an input image is encoded into a feature vector and decoded into a caption. Selecting the right image encoder is vital for optimizing caption prediction systems in specific domains. While Convolutional Neural Network (CNN) based encoders are acknowledged for extracting crucial image features, it’s important to assess variations in their mechanisms and architectures carefully. This paper thoroughly examines various CNNs to evaluate their effectiveness in RS image captioning. We also explore the performance of two caption generation techniques, viz., greedy search and beam search. The encoders are clustered as good, medium, and bad, with ResNet (CNN) emerging as the preferred choice in the good cluster across all considered datasets. The impact of choosing between beam search and greedy search is minimal. Additionally, we conduct a subjective evaluation of leading models to address limitations associated with purely numerical assessments. The paper is a novel contribution, providing the first-of-its-kind subjective evaluation of CNN-based encoders for the RS image captioning task.
Swadhin Das, Akshat Khandelwal, Raksha Sharma
IJCNN3
2023 Leveraging Small-BERT and Bio-BERT for Abbreviation Identification in Scientific Text
Piyush Miglani, Pranav Vatsal, Raksha Sharma
NLDB3
2022 Named Entity Recognition System for the Biomedical Domain
abstract
The recent advancements in medical science have caused a considerable acceleration in the rate at which new information is being published.The MEDLINE database is growing at 500,000 new citations each year.As a result of this exponential increase, it is not easy to manually keep up with this increasing swell of information.Thus, there is a need for automatic information extraction systems to retrieve and organize information in the biomedical domain.Biomedical Named Entity Recognition is one such fundamental information extraction task, leading to significant information management goals in the biomedical domain.Due to the complex vocabulary (e.g., mRNA) and free nomenclature (e.g., IL2), identifying named entities in the biomedical domain is more challenging than any other domain, hence requires special attention.In this paper, we deploy two novel bi-directional encoder-based systems, viz., BioBERT and RoBERTa to identify named entities in the biomedical text.Due to the domain-specific training of BioBERT, it gives reasonably good performance for the NER task in the biomedical domain.However, the structure of RoBERTa makes it more suitable for the task.We obtain a significant improvement in F-score by RoBERTa over BioBERT.In addition, we present a comparative study on training loss attained with ADAM and LAMB optimizers.
Raghav Sharma, Deependra Singh, Raksha Sharma
FedCSIS3
2022 A deep crystal structure identification system for X-ray diffraction patterns
Abhik Chakraborty, Raksha Sharma
Vis. Comput.2
2021 CrypTop12: A Dataset For Cryptocurrency Price Movement Prediction From Tweets And Historical Prices
abstract
Cryptocurrencies are gaining popularity day by day, and their analysis is a fascinating and demanding research topic. The average daily trading volume of Bitcoin was ${\$}67$ billion in May 2021. A peculiar feature of cryptocurrencies is that they are not generally issued by a central authority, making them insusceptible to any governmental impedance. Cryptocurrency rates are closely related to news and influenced by tweets. However, no available dataset can analyze the crypto market adequately. We present CrypTop12, a benchmark dataset for Cryptocurrency Price Movement Prediction based on tweets and historical prices. We collect over 576K tweets related to the top 12 cryptocurrencies, spanning over 1255 days and refine them to filter the tweets that are most relevant to price fluctuations. We also demonstrate use-cases by providing adapted baseline methods and a quantitative results analysis on our dataset.
Amish Garg, Tanav Shah, Vinay Kumar Jain, Raksha Sharma
ICMLA4
2021 Virus Causes Flu: Identifying Causality in the Biomedical Domain Using an Ensemble Approach with Target-Specific Semantic Embeddings
Raksha Sharma, Girish Keshav Palshikar
NLDB1
2020 See Deeper: Identifying Crystal Structure from X-ray Diffraction Patterns
abstract
X-ray diffraction is a commonly used experimental science to detect the atomic and molecular structure of crystalline material. The process is called X-ray crystallography (XRC). Traditionally, it is done by human experts with some conjecture about what structure the crystalline material is likely to be. However, the study of crystal structure using X-ray diffraction patterns is applicable in many domains, such as chemistry, physics, biology, etc. It is tedious to have manual crystallography of X-ray diffraction patterns to determine a crystal structure with a massive amount of dataset. With the advent of high computational resources, deep learning techniques have taken classification to its peak. Convolution Neural Network (CNN) maps an input image into a high dimensional space and produce a low-cost function for image classification. In this paper, we deploy a variation of the Convolution Neural Network to predict crystal structure from X-ray diffraction patterns. We compare our approach with a wide range of conventional as well as modern Machine Learning based classification techniques for the structure prediction of a crystal. We report a cross-validation accuracy of 95.6% and Micro F1-score of 0.949 with our model for the proposed task which is significantly better than the other reported baseline methods.
Abhik Chakraborty, Raksha Sharma
CW2
2019 Cold Is a Disease and D-cold Is a Drug: Identifying Biological Types of Entities in the Biomedical Domain
Suyash Sangwan, Raksha Sharma, Girish Keshav Palshikar, Asif Ekbal
CICLing (2)2
2018 Identifying Transferable Information Across Domains for Cross-domain Sentiment Classification
abstract
Getting manually labeled data in each domain is always an expensive and a time consuming task.Cross-domain sentiment analysis has emerged as a demanding concept where a labeled source domain facilitates a sentiment classifier for an unlabeled target domain.However, polarity orientation (positive or negative) and the significance of a word to express an opinion often differ from one domain to another domain.Owing to these differences, crossdomain sentiment classification is still a challenging task.In this paper, we propose that words that do not change their polarity and significance represent the transferable (usable) information across domains for cross-domain sentiment classification.We present a novel approach based on χ 2 test and cosine-similarity between context vector of words to identify polarity preserving significant words across domains.Furthermore, we show that a weighted ensemble of the classifiers enhances the cross-domain classification performance.
Raksha Sharma, Pushpak Bhattacharyya, Sandipan Dandapat, Himanshu S. Bhatt
ACL (1)1
2018 An Unsupervised Approach for Cause-Effect Relation Extraction from Biomedical Text
Raksha Sharma, Girish Keshav Palshikar, Sachin Pawar
NLDB1
2017 A Comparison Among Significance Tests and Other Feature Building Methods for Sentiment Analysis: A First Study
Raksha Sharma, Dibyendu Mondal, Pushpak Bhattacharyya
CICLing (2)1
2017 Sentiment Intensity Ranking among Adjectives Using Sentiment Bearing Word Embeddings
abstract
Identification of intensity ordering among polar (positive or negative) words which have the same semantics can lead to a finegrained sentiment analysis.For example, master, seasoned and familiar point to different intensity levels, though they all convey the same meaning (semantics), i.e., expertise: having a good knowledge of.In this paper, we propose a semisupervised technique that uses sentiment bearing word embeddings to produce a continuous ranking among adjectives that share common semantics.Our system demonstrates a strong Spearman's rank correlation of 0.83 with the gold standard ranking.We show that sentiment bearing word embeddings facilitate a more accurate intensity ranking system than other standard word embeddings (word2vec and GloVe).Word2vec is the state-of-the-art for intensity ordering task.
Raksha Sharma, Arpan Somani, Lakshya Kumar, Pushpak Bhattacharyya
EMNLP1
2016 High, Medium or Low? Detecting Intensity Variation Among polar synonyms in WordNet
abstract
For fine-grained sentiment analysis, we need to go beyond zero-one polarity and find a way to compare adjectives (synonyms) that share the same sense.Choice of a word from a set of synonyms, provides a way to select the exact polarityintensity.For example, choosing to describe a person as benevolent rather than kind 1 changes the intensity of the expression.In this paper, we present a sense based lexical resource, where synonyms are assigned intensity levels, viz., high, medium and low.We show that the measure P (s|w) (probability of a sense s given the word w) can derive the intensity of a word within the sense.We observe a statistically significant positive correlation between P (s|w) and intensity of synonyms for three languages, viz., English, Marathi and Hindi.The average correlation scores are 0.47 for English, 0.56 for Marathi and 0.58 for Hindi.
Raksha Sharma, Pushpak Bhattacharyya
GWC1
2015 Adjective Intensity and Sentiment Analysis
abstract
For fine-grained sentiment analysis, we need to go beyond zero-one polarity and find a way to compare adjectives that share a common semantic property. In this paper, we present a semi-supervised ap-proach to assign intensity levels to adjec-tives, viz. high, medium and low, where adjectives are compared when they belong
Raksha Sharma, Astha Agarwal, Pushpak Bhattacharyya
EMNLP1
2013 Detecting Domain Dedicated Polar Words
Raksha Sharma, Pushpak Bhattacharyya
IJCNLP1