Shuvam Shiwakoti

dblp:351/5799 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0004-4716-2696ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Agentic AI Framework for Low-Resource Essay Evaluation via Scoring, Explanation, and Debate
Surendrabikram Thapa, Kritesh Rauniyar, Shuvam Shiwakoti, Surabhi Adhikari, Junaid Rashid, Jungeun Kim, Usman Naseem
IEEE Big Data3
2024 THYMES: A Framework for Detecting Suicidal Ideation from Social Media Posts Using Hyperbolic Learning
abstract
Mental health concerns are a critical issue in today’s digital age, posing a threat to both individual and societal well-being and making the identification of at-risk individuals crucial. Analyzing an individual’s social media post history can offer insights into their mental health state and help identify the presence of suicidal ideation. However, the complexity of linguistic and temporal data, along with sparsity and time irregularities, poses a formidable challenge in machine learning. Previous methods in this domain either rely on Euclidean space for processing which does not adequately model the power-law properties of social media posts, or lose information due to the discretization of the time axis. To address these challenges, we propose a novel framework, THYMES, which leverages pre-trained encoders and a rich representation learning paradigm with hyperbolic learning to model power-law features for enhanced sequence modeling. We perform experiments on two datasets and demonstrate that THYMES outperforms previously proposed methods while maintaining classification fairness under heavy data imbalances. Additionally, we qualitatively analyze commonly misclassified samples to reveal the shortcomings of models in this domain.
Surendrabikram Thapa, Mohammad Salman, Siddhant Bikram Shah, Shuvam Shiwakoti, Qi Zhang 0020, Liang Hu 0004, Muhammad Imran Razzak, Usman Naseem
IEEE Big Data4
2024 Analyzing the Dynamics of Climate Change Discourse on Twitter: A New Annotated Corpus and Multi-Aspect Classification
abstract
The discourse surrounding climate change on social media platforms has emerged as a significant avenue for understanding public sentiments, perspectives, and engagement with this critical global issue. The unavailability of publicly available datasets, coupled with ignoring the multi-aspect analysis of climate discourse on social media platforms, has underscored the necessity for further advancement in this area. To address this gap, in this paper, we present an extensive exploration of the intricate realm of climate change discourse on Twitter, leveraging a meticulously annotated ClimaConvo dataset comprising 15,309 tweets. Our annotations encompass a rich spectrum, including aspects like relevance, stance, hate speech, the direction of hate, and humor, offering a nuanced understanding of the discourse dynamics. We address the challenges inherent in dissecting online climate discussions and detail our comprehensive annotation methodology. In addition to annotations, we conduct benchmarking assessments across various algorithms for six tasks: relevance detection, stance detection, hate speech identification, direction and target, and humor analysis. This assessment enhances our grasp of sentiment fluctuations and linguistic subtleties within the discourse. Our analysis extends to exploratory data examination, unveiling tweet distribution patterns, stance prevalence, and hate speech trends. Employing sophisticated topic modeling techniques uncovers underlying thematic clusters, providing insights into the diverse narrative threads woven within the discourse. The findings present a valuable resource for researchers, policymakers, and communicators seeking to navigate the intricacies of climate change discussions. The dataset and resources for this paper are available at https://github.com/shucoll/ClimaConvo.
Shuvam Shiwakoti, Surendrabikram Thapa, Kritesh Rauniyar, Akshyat Shah, Aashish Bhandari, Usman Naseem
LREC/COLING1
2024 MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
abstract
The complexity of text-embedded images presents a formidable challenge in machine learning given the need for multimodal understanding of multiple aspects of expression conveyed by them.While previous research in multimodal analysis has primarily focused on singular aspects such as hate speech and its subclasses, this study expands this focus to encompass multiple aspects of linguistics: hate, targets of hate, stance, and humor.We introduce a novel dataset PrideMM comprising 5,063 text-embedded images associated with the LGBTQ+ Pride movement, thereby addressing a serious gap in existing resources.We conduct extensive experimentation on PrideMM by using unimodal and multimodal baseline methods to establish benchmarks for each task.Additionally, we propose a novel framework MemeCLIP for efficient downstream learning while preserving the knowledge of the pre-trained CLIP model.The results of our experiments show that Meme-CLIP achieves superior performance compared to previously proposed frameworks on two real-world datasets.We further compare the performance of MemeCLIP and zero-shot GPT-4 on the hate classification task.Finally, we discuss the shortcomings of our model by qualitatively analyzing misclassified samples.
Siddhant Bikram Shah, Shuvam Shiwakoti, Maheep Chaudhary, Haohan Wang
EMNLP2
2023 NEHATE: Large-Scale Annotated Data Shedding Light on Hate Speech in Nepali Local Election Discourse
abstract
The use of social media during election campaigns has become increasingly popular. However, the unbridled nature of online discourse can lead to the propagation of hate speech, which has far-reaching implications for the democratic process. Natural Language Processing (NLP) techniques are being used to counteract the spread of hate speech and promote healthy online discourse. Despite the increasing need for NLP techniques to combat hate speech, research on low-resource languages such as Nepali is limited, posing a challenge to the realization of the United Nations’ Leave No One Behind principle, which calls for inclusive development that benefits all individuals and communities, regardless of their backgrounds or circumstances. To bridge this gap, we introduce NEHATE, a large-scale manually annotated dataset of hate speech and its targets in Nepali local election discourse. The dataset comprises 13,505 tweets, annotated for hate speech with further sub-categorization of hate speech into targets such as community, individual, and organization. Benchmarking of the dataset with various algorithms has shown potential for performance improvement. We have made the dataset publicly available at https://github.com/shucoll/NEHate to promote further research and development, while also contributing to the UN SDGs aimed at fostering peaceful, inclusive societies, and justice and strong institutions.
Surendrabikram Thapa, Kritesh Rauniyar, Shuvam Shiwakoti, Sweta Poudel, Usman Naseem, Mehwish Nasim
ECAI3