Snehanshu Saha

dblp:130/3938 · DBLP profile ↗
← Back
31ranked-venue papers
2as first author
26since 2021 · last 2026
0000-0002-8458-604XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PowerQuant: Architecture-Agnostic GPU Power Estimation via Quantile Regression
abstract
Accurate prediction of NVIDIA GPU power consumption remains challenging due to rapid architectural evolution. Existing machine-learning–based power models are tightly coupled to specific GPU architectures and degrade sharply on unseen platforms, requiring retraining and extensive power measurements, which hinder scalability. This paper presents a quantile-regression–based GPU power prediction framework that enables architecture-agnostic power estimation using static analysis-based compile-time CUDA kernel features. The key insight is that architectural changes primarily induce systematic shifts in power scale, while the relative ordering of kernel power demands remains preserved. By learning power quantiles that capture this ordering and mapping them to new GPUs through one-time calibration, the proposed approach mitigates cross-architecture distribution shift. Extensive evaluation across multiple NVIDIA GPU generations shows that, on unseen architectures, the proposed method improves prediction accuracy by up to 30–50% over existing regression models, while maintaining comparable accuracy in in-distribution settings. The resulting low-overhead, generalizable power estimates make the approach practical for power-aware scheduling, energy budgeting, and sustainability-oriented resource management in large HPC systems.
Aditya Challa, Tanish Desai, Gargi Alavani Prabhu, Snehanshu Saha, Santonu Sarkar
HPDC4
2025 Adaptive GPU Power Capping: Balancing Energy Efficiency, Thermal Control and Performance
abstract
As GPUs become increasingly popular in commodity hardware as well as High Performance Computing(HPC) systems, the need for sustainable computing is more critical. This work addresses the challenge of identifying the optimal operating power for GPUs to minimize energy consumption and operational temperature while incurring only minimal performance overhead. We propose a machine learning-based solution that leverages tree-based models to predict the optimal GPU power cap using key system parameters, including GPU utilization, Memory utilization, Temperature, and Frequency. Our experimental results demonstrate that our model can achieve a maximum energy saving of 12. 87% and a temperature reduction of 11. 38%, with only a 2.69% increase in execution time. These findings highlight the potential of our approach to enhance energy efficiency and thermal management in GPU-based systems, paving the way for more sustainable computing practices.
Tanish Desai, Jainam Shah, Gargi Alavani Prabhu, Snehanshu Saha, Santonu Sarkar
HPDC4
2024 Improving Continuous Emotion Annotation in Video Platforms via Physiological Response Profiling
abstract
Many video applications (e.g., gaming, meeting, tutoring) aim to improve the user's interaction experience based on continuously inferred user emotion. To infer user emotion, these apps typically deploy machine learning models, trained with continuously collected emotion ground truth labels. However, as continuous annotations are generally collected as emotion self-reports during video consumption (using an auxiliary device), they incur significant annotation effort. To address this problem, we propose PResUP, a framework that creates users' profile using physiological responses (e.g., GSR or galvanic skin re-sponse) and deploys an LSTM network to identify the opportune probing moments for emotion ground truth (emotion self-report) collection instead of continuous annotation. We evaluate the proposed approach on a large-scale publicly available dataset (CASE) containing the physiological signals of subjects during video consumption. The evaluation of PRes UP reveals that it reduces the probing rate by 30.3 % (on average), detects the opportune probing moments with a TPR (True Positive Rate) of 86.1 %, and yet maintains the quality of the emotion annotations as observed in the continuous self-reports. Furthermore, we evaluated the generalizability of PResUP on another public dataset (K-emocon); which reveals an average probing rate reduction of 25.71 %. These results underscore the efficiency of PResUP in reducing the continuous emotion annotation overhead.
Swarnali Banik, Sougata Sen, Snehanshu Saha, Surjya Ghosh
ACII3
2024 Estimating Power Consumption of GPU Application Using Machine Learning Tool
abstract
As Graphic Processing Units (GPU)s play an increasingly important role in High-Performance Computing (HPC) and data-intensive Machine Learning (ML) tasks, accurate power prediction is essential. Traditional methods, relying on architecture-specific models like DVFS and hardware counters, limit cross-architecture applicability of these models. We propose a static analysis framework that predicts an application's power usage across different NVIDIA GPU architectures without execution. Extensive experiments with state-of-the-art ML approaches show promising results, demonstrating generalizability in predicting power consumption for a newer architecture without the need for complete retraining.11This research is partially supported by the New Faculty Seed Grant of BITS Pilani under Grant No.NFSG/GOA/2023/G0916.
Gargi Alavani Prabhu, Tanish Desai, Sharvil Potdar, Nayan Gogari, Snehanshu Saha, Santonu Sarkar
ICTAI5
2024 DeliverAI: Reinforcement Learning Based Distributed Path-Sharing Network for Food Deliveries
abstract
The Online Food Delivery (OFD) industry, propelled by the recent pandemic, has witnessed substantial growth in the last decade. Major players like Amazon Fresh, GrubHub, UberEats, Postmates, InstaCart, and DoorDash share a common food delivery business model. However, existing methods lack efficiency as deliveries are individually optimized or bundled inefficiently. Recognizing the potential for cost reduction, we model our food delivery problem as a multi-objective optimization, focusing on consumer satisfaction and delivery costs. Taking inspiration from ride-sharing taxis and the prevalent order bundling, we propose DeliverAI - a reinforcement learning-based path-sharing algorithm. Our novel agent interaction scheme dynamically groups deliveries going in the same direction to reduce the total distance traveled while keeping a satisfactory delivery completion time. We test DeliverAI vigorously on a simulation setup using real data from the city of Chicago. Our results show that DeliverAI can reduce the delivery fleet size by 12%, the distance traveled by 13%, and achieve 50% higher fleet utilization compared to the baselines.
Ashman Mehra, Snehanshu Saha, Vaskar Raychoudhury, Archana Mathur
IJCNN2
2024 Self-SLAM: A Self-supervised Learning Based Annotation Method to Reduce Labeling Overhead
Alfiya M. Shaikh, Hrithik Nambiar, Kshitish Ghate, Swarnali Banik, Sougata Sen, Surjya Ghosh, Vaskar Raychoudhury, Niloy Ganguly, Snehanshu Saha
ECML/PKDD (9)9
2024 QuantProb: Generalizing Probabilities along with Predictions for a Pre-trained Classifier
abstract
Quantification of Uncertainty in predictions is a challenging problem. In the classification settings, although deep learning based models generalize well, class probabilities often lack reliability. Calibration errors are used to quantify uncertainty, and several methods exist to minimize calibration error. We argue that between the choice of having a minimum calibration error on original distribution which increases across distortions or having a (possibly slightly higher) calibration error which is constant across distortions, we prefer the latter We hypothesize that the reason for unreliability of deep networks is - The way neural networks are currently trained, the probabilities do not generalize across small distortions. We observe that quantile based approaches can potentially solve this problem. We propose an innovative approach to decouple the construction of quantile representations from the loss function allowing us to compute quantile based probabilities without disturbing the original network. We achieve this by establishing a novel duality property between quantiles and probabilities, and an ability to obtain quantile probabilities from any pre-trained classifier. While post-hoc calibration techniques successfully minimize calibration errors, they do not preserve robustness to distortions. We show that, Quantile probabilities (QuantProb), obtained from Quantile representations, preserve the calibration errors across distortions, since quantile probabilities generalize better than the naive Softmax probabilities.
Aditya Challa, Soma S. Dhavala, Snehanshu Saha
UAI3
2024 DiEvD-SF: Disruptive Event Detection Using Continual Machine Learning With Selective Forgetting
abstract
Detecting disruptive events (DEs), such as riots, protests, and natural calamities, from social media is essential for studying geopolitical dynamics. To automate the process, existing methods rely on classical machine learning (ML) models applied to static datasets, which is counterproductive. To detect DEs from dynamic data streams, this article introduces a novelDiEvD-SFframework, which uses continual machine learning (CML) with selective forgetting. Twitter (currently “X”) is used as a real-time and dynamic data source for validation.DiEvD-SFconsiders the temporal nature of DEs and “selectively forgets” outdated DEs through machine unlearning. To the best of our knowledge, this article is the first to apply CML with selective forgetting to discard outdated DEs and to continue learning about the new DEs. Extensive evaluation using a painstakingly collected Twitter dataset shows that the proposed framework continually identifies new DEs with an average incremental accuracy of 78.942% and successfully forgets old DEs with an average forgetting time of 118.498 seconds, which is better than the state-of-the-art. Additionally, computational analysis is performed to establish the effectiveness of theDiEvD-SFframework by applying various candidate selection strategies.
Aditi Seetha, Satyendra Singh Chouhan, Emmanuel S. Pilli, Vaskar Raychoudhury, Snehanshu Saha
IEEE Trans. Comput. Soc. Syst.5
2024 Last Mile: A Novel, Hotspot-Based Distributed Path-Sharing Network for Food Deliveries
abstract
Delivery of items from the producer to the consumer has experienced significant growth over the past decade and has been greatly fueled by the recent pandemic. Amazon Fresh, GrubHub, UberEats, Postmates, InstaCart, and DoorDash are rapidly growing and are sharing the same business model of consumer items or food delivery. Existing food delivery methods are sub-optimal because each delivery is individually optimized to go directly from the producer to the consumer via the shortest time path. We observe a significant scope for reducing the costs associated with completing deliveries under the current model. For this, we model our food delivery problem as a multi-objective optimization, where consumer satisfaction and delivery costs, both, need to be optimized. Taking inspiration from the success of ride-sharing in the taxi industry, we propose DeliverAI - a reinforcement learning-based path-sharing algorithm. Unlike previous attempts for path-sharing, DeliverAI can provide real-time, time-efficient decision-making using a Reinforcement learning-enabled agent system. Our novel agent interaction scheme leverages path-sharing among deliveries to reduce the total distance traveled while keeping the delivery completion time under check. We generate and test our methodology vigorously on a simulation setup using real data from the city of Chicago. Our results show that DeliverAI can reduce the delivery fleet size by 15%, the distance traveled by 16%, and 50% higher fleet utilization w.r.t point-to-point delivery systems.
Ashman Mehra, Divyanshu Singh, Vaskar Raychoudhury, Archana Mathur, Snehanshu Saha
IEEE Trans. Intell. Transp. Syst.5
2023 Efficient anomaly identification in temporal and non-temporal industrial data using tree based approaches
Jyotirmoy Sarkar, Snehanshu Saha, Santonu Sarkar
Appl. Intell.2
2023 Proposal of SVM Utility Kernel for Breast Cancer Survival Estimation
abstract
The advancement of medical research in the field of cancer prognosis and diagnosis using various modalities has put oncologists under tremendous stress. The complexity and heterogeneity involved in multiple modalities and their significantly varied clinical outcomes make it difficult to analyze the disease and provide the correct treatment. Breast cancer is the major concern among all cancers worldwide, specifically for females. To help oncologists and cancer patients, research for breast cancer survival estimation has been proposed. It ranges from complex deep neural networks to simple and interpretable architectures. We propose a utility kernel for a support vector machine (SVM) in this article. It is a simple yet powerful function, which performs better than other popular machine learning algorithms and deep neural networks in the task of breast cancer survival prediction using the TCGA-BRCA dataset. This study validates the proposed utility kernel using four different modalities (gene expression, copy number variation, clinical, and histopathological tissue images) and their multi-modal combinations. The SVM based on our utility kernel empirically proves its efficacy by achieving the highest value on various performance measures, whereas advanced deep neural networks fail to train on small and highly imbalanced breast cancer data.
Nikhilanand Arya, Archana Mathur, Snehanshu Saha, Sriparna Saha 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Fairly Constricted Multi-objective Particle Swarm Optimization
Anwesh Bhattacharya, Snehanshu Saha, Nithin Nagaraj
ICONIP (4)2
2022 A Fast and Robust Photometric Redshift Forecasting Method Using Lipschitz Adaptive Learning Rate
Snigdha Sen, Snehanshu Saha, Pavan Chakraborty, Krishna Pratap Singh
ICONIP (5)2
2022 P-LSTM: A Novel LSTM Architecture for Glucose Level Prediction Problem
Abhijeet Swain, Vaibhav Ganatra, Snehanshu Saha, Archana Mathur, Rekha Phadke
ICONIP (7)3
2022 HMC-PSO: A Hamiltonian Monte Carlo and Particle Swarm Optimization-Based Optimizer
Omatharv Bharat Vaidya, Rithvik Terence DSouza, Soma S. Dhavala, Snehanshu Saha, Swagatam Das
ICONIP (1)4
2022 LipGene: Lipschitz Continuity Guided Adaptive Learning Rates for Fast Convergence on Microarray Expression Data Sets
abstract
Hyperparameter tuning, specifically tuning of learning rate, can often be a time-consuming process, especially when dealing with large data sets. A mathematical foundation in the choice of learning rate can minimize tuning efforts. We propose the application of a novel adaptive learning rate paradigm, guided by Lipschitz continuity of the loss functions (LipGene), to the task of Gene Expression Inference using shallow neural networks. We utilize Mean Absolute Error and Quantile loss separately for training. Our adaptive learning rate, which is dynamically computed for each epoch, is based on the principle of Lipschitz constant and requires no tuning. Experimentally, we prove that our proposed approach greatly surpasses conventional choices of learning rates in terms of both speed of convergence and generalizability. Advocating the principle of Parsimonious Computing, our method can reduce compute infrastructure required for training by using smaller networks with a minimal compromise on the prediction error.
Tejas Prashanth, Snehanshu Saha, Sumedh Basarkod, Suraj Aralihalli, Soma S. Dhavala, Sriparna Saha 0001, Raviprasad Aduri
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 CARE-Share: A Cooperative and Adaptive Strategy for Distributed Taxi Ride Sharing
abstract
Given the fast growth of on-demand transportation services and ride-sharing platforms, the concept of private vehicle ownership is rapidly declining. Although there are multiple fully-grown ride-sharing systems, they are proprietary and centrally controlled. Facilitating ride-sharing using a localized distributed coordination between the riders and the drivers is in need. However, fully distributed systems deal with a large number of variables and objectives and are often sub-optimal. In this paper, we propose a distributed ride-sharing system with multiple objectives which are often conflicting to each other. Therefore, we model it as a multi-objective optimization problem and solve it using the Ant Colony optimization technique which sports a multi-agent behavior. We critically analyze the spatio-temporal challenges posed by the ride sharing problem and define novel performance metrics to capture the underlying subtlety of the distributed system performance. An in-depth experimentation with recent large-scale single-ride taxi trip data from Chicago shows that our solution can ensure up to 79.65% success rate of ride sharing. We have shown that ride sharing is more successful during non-peak traffic hours due to less contention and a healthy balance in passenger and taxi numbers. Further, it has been observed that ride-sharing always reduces thetotal distance travelledby all the taxis and thetotal number of taxison-road; both of which positively impact road congestion and environment. The results obtained from the experiments are very much comparable to real time behaviour of taxi networks. Finally, a revenue framework is proposed to analyse nuances of the operating environment.
Aishwarya Manjunath, Vaskar Raychoudhury, Snehanshu Saha, Saibal Kar, Anusha Kamath
IEEE Trans. Intell. Transp. Syst.3
2021 Solving the N-Queens and Golomb Ruler Problems Using DQN and an Approximation of the Convergence
Patnala Prudhvi Raj, Snehanshu Saha, Gowri Srinivasa
ICONIP (6)2
2021 d-BTAI: The Dynamic-Binary Tree Based Anomaly Identification Algorithm for Industrial Systems
Jyotirmoy Sarkar, Santonu Sarkar, Snehanshu Saha, Swagatam Das
IEA/AIE (2)3
2021 Implementation of Neural Network Regression Model for Faster Redshift Analysis on Cloud-Based Spark Platform
Snigdha Sen, Snehanshu Saha, Pavan Chakraborty, Krishna Pratap Singh
IEA/AIE (2)2
2021 Prediction of Protein-Protein Interactions using Deep Multi-Modal Representations
abstract
Protein-protein interaction (PPI) plays essential roles in nearly all biological processes of living organisms. Identification of PPI helps in understanding the cellular pathways and complex structure of proteins. Most of the works on the prediction of PPI utilized one type of information, mainly sequence-based. With the recent development in deep learning technologies, capturing the diverse and sparse biological dataset's relevant features is possible. This paper proposes a framework that incorporates a multi-modal representation of proteins to predict the interaction between them. Current work utilizes 3D structure and Gene ontology (GO) information to generate the vector representations of proteins. The task of PPI is divided into two phases: feature generation and prediction. We use deep learning algorithms in both stages. We validate our approach on two datasets: Human and S. cerevisiae. The trained model achieves accuracy values of 97.94% and 95.33% on the human and S. cerevisiae test sets, respectively. The results obtained over these datasets illustrate the superiority of the proposed method as compared to state-of-the-art methods.
Kanchan Jha, Sriparna Saha 0001, Snehanshu Saha
IJCNN3
2021 A Swarm Variant for the Schrödinger Solver
abstract
This paper introduces the application of the Exponentially Averaged Momentum Particle Swarm Optimization (EM-PSO) as a derivative-free optimizer for Neural Networks. It adopts PSO's major advantages such as search space exploration and higher robustness to local minima compared to gradient-descent optimizers such as Adam. Neural network based solvers endowed with gradient optimization are now being used to approximate solutions to Differential Equations. Here, we demonstrate the novelty of EM-PSO in approximating gradients and leveraging the property in solving the Schrödinger equation, for the Particle-in-a-Box problem. We also provide the optimal set of hyper-parameters supported by mathematical proofs, suited for our algorithm11Snehanshu Saha would like to thank the Science and Engineering Research Board (SERB), DST, Government of India, for supporting our research (project reference number: EMR/2016/005687)..
Urvil Nileshbhai Jivani, Omatharv Bharat Vaidya, Anwesh Bhattacharya, Snehanshu Saha
IJCNN4
2021 LipARELU: ARELU Networks aided by Lipschitz Acceleration
abstract
We present LipARELU, a novel framework for training L-hidden layer neural networks, equipped to handle large, adaptive learning rates. The framework is based on smoothness assumptions on the proposed activation function, AREL U. The generalization and approximation abilities of ARELU are discussed in detail. The framework assumes weaker conditions on the loss functions used to train the network. Using the fact that the inverse of the Lipschitz constant of the loss function is an ideal learning rate, we compute Lipschitz Adaptive Learning Rates (LALR) for tractable functions such as the Quadratic Loss (QL) and robust functions such as the Mean Absolute Error (MAE). Theoretical and experimental validation testify strength of our approach on several datasets in regression and classification tasks. The performance of our method is comparable to the current state-of-the-art (SOTA) methods.
Ishita Mediratta, Snehanshu Saha, Shubhad Mathur
IJCNN2
2021 DiffAct: A Unifying Framework for Activation Functions
abstract
The evolution of activation functions in deep neural networks is usually driven by fixed goals and incremental steps toward solving specific problems. We introduce differential equation activations (DiffAct), an umbrella of activation units to improve the process of learning activation functions in complex neural network structures on different data sets. Our stated objective is to enable a feed-forward neural network (FNN) to learn nonlinear activation functional forms from a family of solutions to an ordinary differential equation. In this paper, we also present new activation functions in parabolic forms, as offspring of the family of DiffActs. We demonstrate the effectiveness of the exercise of exploiting DiffAct based units in accomplishing comparable, sometimes superior, performance against single, fixed activation functions. The stated objective is to discover the internal dynamics of such activation functions via a common framework instead of producing State-of-the-Art (SOTA) results. The paper also hypothesizes on the ability to draw effective inferences on simple, 1-hidden layer architectures, wherever possible.
Snehanshu Saha, Archana Mathur, Aditya Pandey, Harshith Arun Kumar
IJCNN1
2021 Dynamic Taxi Ride-Sharing Through Adaptive Request Propagation Using Regional Taxi Demand and Supply
Haoxiang Yu, Vaskar Raychoudhury, Snehanshu Saha
MobiQuitous3
2021 LipschitzLR: Using theoretically computed adaptive learning rates for fast convergence
Rahul Yedida, Snehanshu Saha, Tejas Prashanth
Appl. Intell.2
2020 LALR: Theoretical and Experimental validation of Lipschitz Adaptive Learning Rate in Regression and Neural Networks
abstract
We propose a theoretical framework for an adaptive learning rate policy for the Mean Absolute Error loss function and Quantile loss function and evaluate its effectiveness for regression tasks. The framework is based on the theory of Lipschitz continuity, specifically utilizing the relationship between learning rate and Lipschitz constant of the loss function. Based on experimentation, we have found that the adaptive learning rate policy enables up to 20× faster convergence compared to a constant learning rate policy.
Snehanshu Saha, Tejas Prashanth, Suraj Aralihalli, Sumedh Basarkod, T. S. B. Sudarshan, Soma S. Dhavala
IJCNN1
2020 Parsimonious Computing: A Minority Training Regime for Effective Prediction in Large Microarray Expression Data Sets
abstract
Rigorous mathematical investigation of learning rates used in back-propagation in shallow neural networks has become a necessity. This is because experimental evidence needs to be endorsed by a theoretical background. Such theory may be helpful in reducing the volume of experimental effort to accomplish desired results. We leveraged the functional property of Mean Square Error, which is Lipschitz continuous to compute learning rate in shallow neural networks. We claim that our approach reduces tuning efforts, especially when a significant corpus of data has to be handled. We achieve remarkable improvement in saving computational cost while surpassing prediction accuracy reported in literature. The learning rate, proposed here, is the inverse of the Lipschitz constant. The work results in a novel method for carrying out gene expression inference on large microarray data sets with a shallow architecture constrained by limited computing resources. A combination of random sub-sampling of the dataset, an adaptive Lipschitz constant inspired learning rate and a new activation function, A-ReLU helped accomplish the results reported in the paper.
Shailesh Sridhar, Snehanshu Saha, Azhar Shaikh, Rahul Yedida, Sriparna Saha 0001
IJCNN2
2020 A Dynamic Taxi Ride Sharing System Using Particle Swarm Optimization
abstract
With the rapid growth of on-demand taxi services, like Uber, Lyft, etc., urban public transportation scenario is shifting towards a personalized transportation choice for most commuters. While taxi rides are comfortable and time efficient, they often lead to higher cost and road congestion due to lower overall occupancy than bigger vehicles. One efficient way to improve taxi occupancy is to adopt ride sharing. Existing ride sharing solutions are mostly centralized and proprietary. Moreover, given the wide spatio-temporal variation of incoming ride requests designing a dynamic and distributed shared-ride scheduling system is NP-hard. In this paper, we have proposed a publisher (passengers) and subscriber (taxis) based ride sharing system that provides effective real-time ride scheduling for multiple passengers. A particle swarm based route optimization strategy has been applied to determine the most preferable route for passengers. Empirical analysis using large scale single-user taxi ride records from Chicago Transit Authority, show that, our proposed system, ensures a maximum of 91.74% and 63.29% overall success rates during non-peak and peak hours, respectively.
Shrawani Silwal, Vaskar Raychoudhury, Snehanshu Saha, Md. Osman Gani
MASS3
2017 Early Prediction of LBW Cases via Minimum Error Rate Classifier: A Statistical Machine Learning Approach
abstract
Low Birth weight (LBW) acts as an indicator of sickness in newborn babies. LBW is closely associated with infant mortality as well as various health outcomes later in life. Various studies show strong correlation between maternal health during pregnancy and the child's birth weight. This manuscript exploits machine learning techniques to gain useful information from health indicators of pregnant women for early detection of potential LBW cases. The forecasting problem has been reformulated as a classification problem between LBW and NOT-LBW classes using the Bayes' minimum error rate classifier rendering LBW detection as a binary machine classification problem. Expectedly, the proposed model achieved accuracy of 96.77%. Indian health care data was used to construct decision rules to be extrapolated to predictive health care in smart cities. A screening tool based on the decision model is developed to assist health care professionals in Obstetrics and Gynecology (OBG).
Anisha R. Yarlapati, Sudeepa Roy Dey, Snehanshu Saha
SMARTCOMP3
2016 A Novel Approach to Big Data Veracity Using Crowdsourcing Techniques and Bayesian Predictors
abstract
In today's world data is being generated at a tremendous pace and there have to be enough measures in place to verify the nature of big data. Analysis performed on 'dirty' data may lead to erroneous insights and thereby shaping decisions poorly. The aspect of big data that deals with its correctness is known as big data veracity. Trusting the data acquired goes a long way in implementing decisions from an automated decision-making system and veracity helps to validate the data acquired. In this paper, we present our solution to the big data veracity problem using crowdsourcing techniques. Our solution involves the use of sentiment analysis, which deals with identifying the sentiment expressed in a piece of text. As a proof of concept, we have developed an app that requires users to tag tweets as per the sentiment it evokes in them. Each tweet would therefore get ratified by hundreds of our participants and the sentiment associated to the tweet gets tagged. The tagged emotion was then evaluated against the verified emotion as compared to a verified data set. This analysis was then plotted on a ROC curve and also evaluated against verified data using a Bayesian predictor trained with a trinomial function. As can be seen, an accuracy of 81% was obtained as displayed by the ROC curve and 89% through the Bayesian predictor. Also, a MAP analysis of the Bayesian predictor yields neutral sentiment as the most probable hypothesis. By doing this, we have proven that crowdsourcing of sentiment analysis is a viable solution to the problem of big data veracity and therefore an aid in making better decisions.
Bhoomika Agarwal, Abhiram Ravikumar, Snehanshu Saha
ICMLA3