Swapna S. Gokhale

dblp:g/SwapnaSGokhale · DBLP profile ↗
← Back
11ranked-venue papers in the field
1as first author
4since 2021 · last 2024
0000-0001-8443-8146ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 5 (1 first)Big Data, Cloud & Distributed Data Systems · 4Information Retrieval & Web Search · 2
YearPublicationVenuePosition
2024 Geographical Insights into Suicide Mortality Through Spatial Machine Learning
abstract
Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2improved from 0.59 to 0.67) and local (R2improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.
Viswadeep Lebakula, Swapna S. Gokhale, Anuj J. Kapadia, Jodie Trafton, Alina Peluso
IEEE Big Data2
2022 Comparing Deep and Machine Learning Models for Sentiment and Emotion Classification from Vaccine #sideffects
abstract
The accelerated development of Covid-19 vaccines offered tremendous promise and hope, yet stirred significant trepidation and fear. These conflicting emotions motivated many to turn to social media to share their experiences and side effects during the process of getting vaccinated. This paper analyzes sentiment and emotions from tweets collected using the hashtag #sideffects during the early roll out of the Covid-19 vaccine. Each tweet was labeled according to its sentiment polarity (positive vs. negative), and was assigned one of four emotion labels (joy, gratitude, apprehension, and sadness). Exploratory analysis of the tweets through word cloud visualizations revealed that the negativity of emotions intensified with the severity of side effects. Word and numerical features extracted from the text of the tweets and metadata were used to train conventional machine learning and deep learning models. These models resulted in an accuracy of 81% for binary sentiment classification, and 71 % for multi-label emotion identification. The proposed framework, which yielded competitive performance, may be employed to gain insights into people's thoughts and feelings from vaccine-related conversations. These insights can be helpful in devising communication and education strategies to mitigate vaccine hesitancy.
Aditya Dubey, Swapna S. Gokhale
ASONAM2
2021 Identifying Social Media Content Supporting Proud Boys
abstract
While most conversations on social media are inspiring and uplifting, a fraction of the users do engage in sharing content that supports extremist, radical philosophies and organizations. Such rhetoric, if left unchecked, can propagate virally on these platforms, ultimately escalating to turbulence and violence in the physical, offline spaces. Identifying such content from the volumes of social media feeds is therefore necessary to prevent the damage that it may cause, yet it is infeasible to undertake manually, calling for an automated approach. This paper demonstrates the potential of machine learning to separate social media feeds that are sympathetic to radical extremism using the example of Proud Boys, a contemporary right-wing group. From the tweets collected after Proud Boys protests in August 2020; linguistic, social and auxiliary features are extracted. Significance tests are used to select a subset of these features that contribute towards separating between extremist and normal content. Several machine learning models are trained based on a combination of these features. These models can identify tweets that support Proud Boys with excellent performance metrics. Artificial Neural Networks offer the best accuracy, precision, recall and F1-score. Feature importance, assessed using the Random Forest model indicates that users rely both on the expressive power of the language and other metadata features such as punctuations, mentions and URLs to voice and spread hateful ideology on online platforms.
Md Fahim, Swapna S. Gokhale
IEEE BigData2
2021 Identifying Social Media Content Supporting Proud Boys
abstract
While most conversations on social media are inspiring and uplifting, a fraction of the users do engage in sharing content that supports extremist, radical philosophies and organizations. Such rhetoric, if left unchecked, can propagate virally on these platforms, ultimately escalating to turbulence and violence in the physical, offline spaces. Identifying such content from the volumes of social media feeds is therefore necessary to prevent the damage that it may cause, yet it is infeasible to undertake manually, calling for an automated approach. This paper demonstrates the potential of machine learning to separate social media feeds that are sympathetic to radical extremism using the example of Proud Boys, a contemporary right-wing group. From the tweets collected after Proud Boys protests in August 2020; linguistic, social and auxiliary features are extracted. Significance tests are used to select a subset of these features that contribute towards separating between extremist and normal content. Several machine learning models are trained based on a combination of these features. These models can identify tweets that support Proud Boys with excellent performance metrics. Artificial Neural Networks offer the best accuracy, precision, recall and F1-score. Feature importance, assessed using the Random Forest model indicates that users rely both on the expressive power of the language and other metadata features such as punctuations, mentions and URLs to voice and spread hateful ideology on online platforms.
Md Fahim, Swapna S. Gokhale
IEEE BigData2
2020 Comparing the Impact of Unhealthy Behaviors and Preventive Services on Chronic Health Outcomes
abstract
Chronic health outcomes are a leading cause of death and disability, and also prominent drivers of health care costs. Most chronic health outcomes can be attributed to a few risky behaviors. It is believed that chronic health outcomes and their burdens can be alleviated by the use of clinical preventive services. This paper seeks to assess the relative influence of unhealthy behaviors and preventive services on chronic health outcomes using the 500 Cities data. The approach comprises of three-way clustering of the 500 cities, one each based on health outcomes, preventive services, and unhealthy behaviors, and then measuring the pairwise similarity between the clustering solutions based on health outcomes and preventive services, and health outcomes and unhealthy behaviors. A variant of the Rand Index is defined to assess this clustering similarity. Higher similarity between clusterings based on health outcomes and unhealthy behaviors compared to the clusterings based on health outcomes and preventive services is observed. These findings suggest a greater influence of unhealthy behaviors over preventive services on chronic health outcomes. The paper concludes with a question of whether investing in facilitating healthier lifestyle choices will yield a higher return towards promoting health as opposed to investing in improving access to preventive services.
Swapna S. Gokhale
ASONAM1
2020 Analysis and Classification of Vaccine Dialogue in the Coronavirus Era
abstract
As the coronavirus tears through our global community, the world pins its hopes on the expedient availability of a safe and effective vaccine. In the U.S, however, mere mention of vaccines galvanizes a community that remains steadfastly opposed to them. This paper analyzes the vaccine dialogue on Twitter in the coronavirus era, using the data collected a week after President Trump's announcement of Operation Warp Speed. These tweets are explored in three ways. Informal opinion mining reveals both concerns and support; the anti-vaxx community is vociferous in opposing the vaccine, spreading misinformation, spinning conspiracies and whipping hysteria. Significant hesitation about the safety of the Covid-19 vaccine is also expressed in particular because of its rapid deployment. The pro-vaxx community counters this opposition by pointing to prior successes of immunizations as well as by mocking the anti-vaxx attitudes. A comparison of the social features of the anti-vaxx and pro-vaxx tweets suggests that the anti-vaxx community has gained steam on social media platforms and is better connected than the pro-vaxx community, which may lead to a penetration of discordant information through the online world. Identifying and labeling tweets that sow discordant information is one way to prevent their spread, which is facilitated by our classification framework that can distinguish between the anti-vaxx and pro-vaxx tweets with an accuracy of over 80%. Taken together, our results suggest that unless a concerted effort is made to dispel these myths and misgivings, the fringe anti-vaxx minority is likely to become an outspoken majority by dragging many skeptics into their fold, and hence, hinder herd immunity.
Nijhum Paul, Swapna S. Gokhale
IEEE BigData2
2013 Human sensing for smart cities
abstract
Smart cities are powered by the ability to self-monitor and respond to signals and data feeds from heterogeneous physical sensors. These physical sensors, however, are fraught with interoperability and dependability challenges. Moreover, they also cannot shed light on human emotions and factors that impact smart city initiatives. Yet everyday, millions of city dwellers share their observations, thoughts, feelings, and experiences about their city through social media updates. This paper describes how citizens can serve as human sensors in providing supplementary, alternate, and complementary sources of information for smart cities. It presents a methodology, based on a probabilistic language model, to extract the perceptions that may be relevant to smart city initiatives from social media updates. Geo-tagged tweets collected over a two-month period from New York City are used to illustrate the potential of social media powered human sensors.
Derek Doran, Swapna S. Gokhale, Aldo Dagnino
ASONAM2
2013 A comparison of web robot and human requests
abstract
Sophisticated Web robots sport a wide variety of functionality and visiting characteristics, constituting a significant percentage of the requests serviced by a Web server. Unlike human clients that retrieve information off a site by navigating links and ignoring irrelevant information, Web robots may collect many different types of resources, and employ varying navigation strategies to find the knowledge on the site they desire. Thus, the resource request patterns of their visits are unpredictable and cannot be inferred based on our knowledge of human request patterns. In this paper, we perform an analysis on the types of resources requested by Web robots using recent Web logs from an academic Web server. We study the distribution of response sizes and response codes, the types of resources requested, and popularity of resources for requests from Web robots. Throughout, we contrast our findings against human resource request patterns. We find reasons to suggest that robots severely handicaps the ability of Web server caches to operate with high performance.
Derek Doran, Kevin Morillo, Swapna S. Gokhale
ASONAM3
2012 A classification framework for web robots
abstract
The behavior of modern web robots varies widely when they crawl for different purposes. In this article, we present a framework to classify these web robots from two orthogonal perspectives, namely, their functionality and the types of resources they consume. Applying the classification framework to a year‐long access log from the UConn SoE web server, we present trends that point to significant differences in their crawling behavior.
Derek Doran, Swapna S. Gokhale
J. Assoc. Inf. Sci. Technol.2
2011 Web robot detection techniques: overview and limitations
Derek Doran, Swapna S. Gokhale
Data Min. Knowl. Discov.2
2006 Web server performance analysis
abstract
No abstract available.
Jijun Lu, Swapna S. Gokhale
ICWE2