Shakshi Sharma

dblp:276/6386 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0001-8091-0781ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2024 AMIR: An Automated Misinformation Rebuttal System - A COVID-19 Vaccination Datasets-Based Exposition
abstract
Misinformation has emerged as a major societal threat in the recent years in general; specifically in the context of the COVID-19 pandemic, it has wrecked havoc, for instance, by fueling vaccine hesitancy. Cost-effective, scalable solutions for combating misinformation are the need of the hour. This work explored how existing information obtained from social media and augmented with more curated fact checked data repositories can be harnessed to facilitate automated rebuttal of misinformation at scale. While the ideas herein can be generalized and reapplied in the broader context of misinformation mitigation using a multitude of information sources and catering to the spectrum of social media platforms, this work serves as a proof of concept, and as such, it is confined in its scope to only rebuttal of tweets, and in the specific context of misinformation regarding COVID-19. It leverages two publicly available datasets, viz. FaCov (fact-checked articles) Sharma et al., 2022 and misleading (social media Twitter) Sharma et al., 2024 data on COVID-19 vaccination.
Shakshi Sharma, Anwitaman Datta, Rajesh Sharma 0002
IEEE Trans. Comput. Soc. Syst.1
2024 (Mis)leading the COVID-19 Vaccination Discourse on Twitter: An Exploratory Study of Infodemic Around the Pandemic
abstract
In this work, we collect a moderate-sized representative corpus of tweets (over 200 000) pertaining to COVID-19 vaccination spanning for a period of seven months (September 2020–March 2021). Following a transfer learning approach, we utilize a pretrained transformer-based XLNet model to classify tweets as misleading or nonmisleading and manually validate the results with random subsets of samples. We leverage this to study and contrast the characteristics of tweets in the corpus that are misleading in nature against non-misleading ones. This exploratory analysis enables us to design features such as sentiments, hashtags, nouns, and pronouns which can, in turn, be exploited for classifying tweets as (non-)misleading using various machine learning (ML) models in an explainable manner. Specifically, several ML models are employed for prediction, with up to 90% accuracy, with the importance of each feature is explained using SHAP Explainable AI (XAI) tool. While the thrust of this work is principally exploratory in nature to obtain insight on the online discourse on COVID-19 vaccination, we conclude the article by outlining how these insights provide the foundations for a more actionable approach to mitigate misinformation. We have made the curated data as well as the accompanying code available so that the research community at large can reproduce, compare against, or build upon this work.
Shakshi Sharma, Rajesh Sharma 0002, Anwitaman Datta
IEEE Trans. Comput. Soc. Syst.1
2023 Misinformation Concierge: A Proof-of-Concept with Curated Twitter Dataset on COVID-19 Vaccination
abstract
We demonstrate the Misinformation Concierge, a proof-of-concept that provides actionable intelligence on misinformation prevalent in social media. Specifically, it uses language processing and machine learning tools to identify subtopics of discourse and discerns non/misleading posts; presents statistical reports for policy-makers to understand the big picture of prevalent misinformation in a timely manner; and recommends rebuttal messages for specific pieces of misinformation, identified from within the corpus of data - providing means to intervene and counter misinformation promptly. The Misinformation Concierge proof-of-concept using a curated dataset is accessible at: https://demo-frontend-uy34.onrender.com/
Shakshi Sharma, Anwitaman Datta, Vigneshwaran Shankaran, Rajesh Sharma 0002
CIKM1
2022 DEAP-FAKED: Knowledge Graph based Approach for Fake News Detection
abstract
Fake News on social media platforms has attracted a lot of attention in recent times, primarily for events related to politics (2016 US Presidential elections), and healthcare (infodemic during COVID-19), to name a few. Various methods have been proposed for detecting Fake News. The approaches span from exploiting techniques related to network analysis, Natural Language Processing (NLP), and the usage of Graph Neural Networks (GNNs). In this work, we propose DEAP-FAKED, a knowleDgE grAPh FAKe nEws Detection framework for identifying Fake News. Our approach combines natural language processing (NLP) and tensor decomposition model to encode news content and embed Knowledge Graph (KG) entities, respectively. A variety of these encodings provides a complementary advantage to our detector. We evaluate our framework using two publicly available datasets containing articles from domains such as politics, business, technology, and healthcare. As part of dataset pre-processing, we also remove the bias, such as the source of the articles, which could impact the performance of the models. DEAP-FAKED obtains an F1-score of 88% and 78% for the two datasets, which is an improvement of ~21 %, and ~3%, respectively, which shows the effectiveness of the approach.
Mohit Mayank, Shakshi Sharma, Rajesh Sharma 0002
ASONAM2
2022 FaCov: COVID-19 Viral News and Rumors Fact-Check Articles Dataset
Shakshi Sharma, Ekanshi Agrawal, Rajesh Sharma 0002, Anwitaman Datta
ICWSM1
2021 Identifying Possible Rumor Spreaders on Twitter: A Weak Supervised Learning Approach
abstract
Online Social Media (OSM) platforms such as Twitter, Facebook are extensively exploited by the users of these platforms for spreading the (mis)information to a large audience effortlessly at a rapid pace. It has been observed that the misinformation can cause panic, fear, and financial loss to society. Thus, it is important to detect and control the misinformation in such platforms before it spreads to the masses. In this work, we focus on rumors, which is one type of misinformation (other types are fake news, hoaxes, etc). One way to control the spread of the rumors is by identifying users who are possibly the rumor spreaders, that is, users who are often involved in spreading the rumors. Due to the lack of availability of rumor spreaders labeled dataset (which is an expensive task), we use publicly available PHEME dataset, which contains rumor and non-rumor tweets information, and then apply a weak supervised learning approach to transform the PHEME dataset into rumor spreaders dataset. We utilize three types of features, that is, user, text, and ego-network features, before applying various supervised learning approaches. In particular, to exploit the inherent network property in this dataset (user-user reply graph), we explore Graph Convolutional Network (GCN), a type of Graph Neural Network (GNN) technique. We compare GCN results with the other approaches: SVM, RF, and LSTM. Extensive experiments performed on the rumor spreaders dataset, where we achieve up to 0.864 value for F1-Score and 0.720 value for AUC-ROC, shows the effectiveness of our methodology for identifying possible rumor spreaders using the GCN technique.
Shakshi Sharma, Rajesh Sharma 0002
IJCNN1
2020 Forecasting Transactional Amount in Bitcoin Network Using Temporal GNN Approach
abstract
Financial institutions such as banks regularly forecast the amount of finances an individual will have in his/her account in the near future. This can help banks in categorizing their customers so that banks can recommend financial products that matches the needs of their customers. In this work, we explored the historical financial transactions for predicting the amount a customer will receive through his/her transacting partners at a specific time. In particular, we use the Bitcoin transactional dataset, which has two main characteristics: i) network, and ii) temporal. This paper contributes by exploiting a specific kind of Graph Neural Network approach called Temporal-Graph Convolutional Network (T-GCN) for predicting the amount of Bitcoins received by a customer at a particular timestamp. The lower errors obtained using T-GCN approach compared to 11 baseline approaches (such as Support Vector Regression (SVR), Random Forest Regression (RFR), Vector Auto-Regressive (VAR), Long Short-Term Memory (LSTM), etc.) clearly demonstrate the effectiveness of T-GCN approach. In addition, our findings reveal that time is an important feature for such kind of predictive tasks.
Shakshi Sharma, Rajesh Sharma 0002
ASONAM1