EDBT 2026 Demo / reviewers in the wild / expert
Hemant Purohit
dblp:77/3904ohio
· DBLP profile ↗
22ranked-venue papers in the field
5as first author
9since 2021 · last 2026
0000-0002-4573-8450ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 8 (3 first)Data Mining & Knowledge Discovery · 7 (1 first)Big Data, Cloud & Distributed Data Systems · 5Other / Interdisciplinary · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SHIELD: A Framework for Online Conversational AI System for Self-Regulated Learning at Emergency Communications CentersabstractAdvancements in conversational artificial intelligence (AI) enable scalable, web-based training environments beyond traditional classroom settings to support learning. In emergency communication centers such as 9-1-1, trainees have limited opportunities to practice diverse call-handling scenarios, as it relies on role-play instruction and constrained instructor availability. We introduce SHIELD (Strengthening Human Intervention in Emergencies through Learning with Data and AI), a framework and online conversational AI system designed to support scenario-based training through interactive simulations, adaptive feedback, and learning analytics. SHIELD integrates AI-generated call simulations with data science techniques to capture fine-grained trainee interaction data, including decision-making behavior, response timing, and corrective actions. These logged interactions are analyzed to provide real-time metacognitive nudges during simulated calls and post-performance feedback through AI-assisted analytics. The framework is informed by principles of self-regulated learning to structure training across planning, execution, and reflection phases. This demo showcases SHIELD as a web-based prototype deployed for emergency call-taker training and discusses design insights from initial testing sessions conducted in collaboration with a public safety communications agency. The system also highlights potential applicability to other high-stakes operational training settings. Ramya Sreekanta Nayaka, Ritesh Somashekar, Kathryn B. Laskey, Linton Wells, Hemant Purohit |
WSDM | 5 |
| 2025 | Scam Shield: Multi-Model Voting and Fine-Tuned LLMs Against Adversarial Attacks
Chen-Wei Chang, Shailik Sarkar, Hossein Salemi, Shutonu Mitra, Hemant Purohit, Fengxiu Zhang, Michin Hong, Jin-Hee Cho, Chang-Tien Lu |
IEEE Big Data | 6 |
| 2024 | Exposing LLM Vulnerabilities: Adversarial Scam Detection and PerformanceabstractCan we trust Large Language Models (LLMs) to accurately predict scam? This paper investigates the vulnerabilities of LLMs when facing adversarial scam messages for the task of scam detection. We addressed this issue by creating a comprehensive dataset with fine-grained labels of scam messages, including both original and adversarial scam messages. The dataset extended traditional binary classes for the scam detection task into more nuanced scam types. Our analysis showed how adversarial examples took advantage of vulnerabilities of a LLM, leading to high misclassification rate. We evaluated the performance of LLMs on these adversarial scam messages and proposed strategies to improve their robustness. Chen-Wei Chang, Shailik Sarkar, Shutonu Mitra, Qi Zhang 0104, Hossein Salemi, Hemant Purohit, Fengxiu Zhang, Michin Hong, Jin-Hee Cho, Chang-Tien Lu |
IEEE Big Data | 6 |
| 2024 | ORIS: Online Active Learning Using Reinforcement Learning-based Inclusive Sampling for Robust Streaming Analytics SystemabstractEffective labeled data collection plays a critical role in developing and fine-tuning robust streaming analytics systems. However, continuously labeling documents to filter relevant information poses significant challenges like limited labeling budget or lack of high-quality labels. There is a need for efficient human-in-the-loop machine learning (HITL-ML) design to improve streaming analytics systems. One particular HITL-ML approach is online active learning, which involves iteratively selecting a small set of the most informative documents for labeling to enhance the ML model performance. The performance of such algorithms can get affected due to human errors in labeling. To address these challenges, we propose ORIS, a method to perform Online active learning using Reinforcement learning-based Inclusive Sampling of documents for labeling. ORIS aims to create a novel Deep Q-Network-based strategy to sample incoming documents that minimize human errors in labeling and enhance the ML model performance. We evaluate the ORIS method on emotion recognition tasks, and it outperforms traditional baselines in terms of both human labeling performance and the ML model performance. The code for this research is available at https://github.com/rpandey4/oris. Rahul Pandey, Ziwei Zhu 0001, Hemant Purohit |
IEEE Big Data | 3 |
| 2024 | SALSA: Salience-Based Switching Attack for Adversarial Perturbations in Fake News Detection Models
Chahat Raj, Anjishnu Mukherjee, Hemant Purohit, Antonios Anastasopoulos, Ziwei Zhu 0001 |
ECIR (5) | 3 |
| 2024 | Multilingual Serviceability Model for Detecting and Ranking Help Requests on Social Media during DisastersabstractSocial media users expect quick and high-quality responses from emergency services when seeking help. However, these organizations face difficulties in detecting and prioritizing critical requests due to the overwhelming amount of information on social media and their limited human resources to tackle it during mass emergencies or disaster events. The situation is exacerbated when users communicate in different or native languages, which can be expected during disasters. While recent studies have focused on characterizing and automatically detecting help requests on social media, they focused on non-behavioral features and monolingual data, primarily in English. Thus, a key gap exists in analyzing multilingual requests on social media for public services. In this paper, we introduce a knowledge distillation framework called MulTMR (Multiple Teachers Model for detecting and Ranking), which combines the power of both task-related and behavior-guided models as diverse teachers for training a student model to efficiently detect serviceable request messages across languages and regions on social media during natural disaster events. We demonstrate that the presented framework can enhance performance (with an AUC improvement of up to 10%) in various scenarios of multilingual test data. Our results, which were validated on real-world data collected in three languages during ten disasters across seven countries, indicate the use of behavior-guided teacher models in MulTMR can increase attention to relevant indicators of serviceability characteristics. The application of the MulTMR framework through a streaming data analytics tool could reduce the cognitive load on personnel within social media teams of emergency services. Further, its application could inform how to leverage human behavior characteristics in creating automated models for social media analytics to design systems in other public service domains beyond emergency management. Fedor Vitiugin, Hemant Purohit |
ICWSM | 2 |
| 2022 | Cross-Lingual Text Classification of Transliterated Hindi and MalayalamabstractTransliteration is very common on social media, but transliterated text is not adequately handled by modern neural models for various NLP tasks. In this work, we combine data augmentation approaches with a Teacher-Student training scheme to address this issue in a cross-lingual transfer setting for fine-tuning state-of-the-art pre-trained multilingual language models such as mBERT and XLM-R. We evaluate our method on transliterated Hindi and Malayalam, also introducing new datasets for benchmarking on real-world scenarios: one on sentiment classification in transliterated Malayalam, and another on crisis tweet classification in transliterated Hindi and Malayalam (related to the 2013 North India and 2018 Kerala floods). Our method yielded an average improvement of +5.6% on mBERT and +4.7% on XLM-R in F1 scores over their strong baselines.1 Jitin Krishnan, Antonios Anastasopoulos, Hemant Purohit, Huzefa Rangwala |
IEEE Big Data | 3 |
| 2022 | PROBER: A System for Real-time Propaganda Behavior Analytics on Social Media and Web Data StreamsabstractSocial media and online platforms provide a public space for many people to share opinions. Social media has numerous benefits to society; however, previous research has identified that individuals use social media for propagandizing purposes which can be detrimental to society, especially during humanitarian crises. Therefore, communities must look into this content to understand and effectively mitigate propaganda, especially when social media messages contain targeted hate or fake/disinformation. In this paper, we propose a human-centered system called PROBER, for propaganda behavior analytics in social and web data streams, which the relevant authorities, such as government institutions, could use for operational decision support and informing policy analysts for crisis management. Yasas Senarath, Antonios Anastasopoulos, Tonya Thornton, Hemant Purohit |
IEEE Big Data | 4 |
| 2021 | Practitioner-Centric Approach for Early Incident Detection Using Crowdsourced Data for Emergency ServicesabstractEmergency response is highly dependent on the time of incident reporting. Unfortunately, the traditional approach to receiving incident reports (e.g., calling 911 in the USA) has time delays. Crowdsourcing platforms such as Waze provide an opportunity for early identification of incidents. However, detecting incidents from crowdsourced data streams is difficult due to the challenges of noise and uncertainty associated with such data. Further, simply optimizing over detection accuracy can compromise spatial-temporal localization of the inference, thereby making such approaches infeasible for real-world deployment. This paper presents a novel problem formulation and solution approach for practitioner-centered incident detection using crowdsourced data by using emergency response management as a case-study. The proposed approach CROME (Crowdsourced Multi-objective Event Detection) quantifies the relationship between the performance metrics of incident classification (e.g., F1 score) and the requirements of model practitioners (e.g., 1 km. radius for incident detection). First, we show how crowdsourced reports, ground-truth historical data, and other relevant determinants such as traffic and weather can be used together in a Convolutional Neural Network (CNN) architecture for early detection of emergency incidents. Then, we use a Pareto optimization-based approach to optimize the output of the CNN in tandem with practitioner-centric parameters to balance detection accuracy and spatial-temporal localization. Finally, we demonstrate the applicability of this approach using crowdsourced data from Waze and traffic accident reports from Nashville, TN, USA. Our experiments demonstrate that the proposed approach outperforms existing approaches in incident detection while simultaneously optimizing the needs for real-world deployment and usability. Yasas Senarath, Ayan Mukhopadhyay, Sayyed Vazirizade, Hemant Purohit, Saideep Nannapaneni, Abhishek Dubey |
ICDM | 4 |
| 2020 | Unsupervised and Interpretable Domain Adaptation to Rapidly Filter Tweets for Emergency ServicesabstractDuring the onset of a natural or man-made crisis event, public often share relevant information for emergency services on social web platforms such as Twitter. However, filtering such relevant data in real-time at scale using social media mining is challenging due to the short noisy text, sparse availability of relevant data, and also, practical limitations in collecting large labeled data during an ongoing event. In this paper, we hypothesize that unsupervised domain adaptation through multi-task learning can be a useful framework to leverage data from past crisis events for training efficient information filtering models during the sudden onset of a new crisis. We present a novel method to classify relevant social posts during an ongoing crisis without seeing any new data from this event (fully unsupervised domain adaptation). Specifically, we construct a customized multi-task architecture with a multi-domain discriminator for crisis analytics: multi-task domain adversarial attention network (MT-DAAN). This model consists of dedicated attention layers for each task to provide model interpretability; critical for real-word applications. As deep networks struggle with sparse datasets, we show that this can be improved by sharing a base layer for multitask learning and domain adversarial training. The framework is validated with the public datasets of TREC incident streams that provide labeled Twitter posts (tweets) with relevant classes (Priority, Factoid, Sentiment) across 10 different crisis events such as floods and earthquakes. Evaluation of domain adaptation for crisis events is performed by choosing one target event as the test set and training on the rest. Our results show that the multi-task model outperformed its single-task counterpart. For the qualitative evaluation of interpretability, we show that the attention layer can be used as a guide to explain the model predictions and empower emergency services for exploring accountability of the model, by showcasing the words in a tweet that are deemed important in the classification process. Finally, we show a practical implication of our work by providing a use-case for the COVID-19 pandemic. Jitin Krishnan, Hemant Purohit, Huzefa Rangwala |
ASONAM | 2 |
| 2020 | Diversity-Based Generalization for Unsupervised Text Classification Under Domain Shift
Jitin Krishnan, Hemant Purohit, Huzefa Rangwala |
ECML/PKDD (2) | 2 |
| 2019 | Modeling human annotation errors to design bias-aware systems for social stream processingabstractHigh-quality human annotations are necessary to create effective machine learning systems for social media. Low-quality human annotations indirectly contribute to the creation of inaccurate or biased learning systems. We show that human annotation quality is dependent on the ordering of instances shown to annotators (referred as 'annotation schedule'), and can be improved by local changes in the instance ordering provided to the annotators, yielding a more accurate annotation of the data stream for efficient real-time social media analytics. Rahul Pandey, Carlos Castillo 0001, Hemant Purohit |
ASONAM | 3 |
| 2019 | Multi-stage Deep Classifier Cascades for Open World RecognitionabstractAt present, object recognition studies are mostly conducted in a closed lab setting with classes in test phase typically in training phase. However, real-world problem are far more challenging because: i)~new classes unseen in the training phase can appear when predicting; ii)~discriminative features need to evolve when new classes emerge in real time; and iii)~instances in new classes may not follow the "independent and identically distributed" (iid) assumption. Most existing work only aims to detect the unknown classes and is incapable of continuing to learn newer classes. Although a few methods consider both detecting and including new classes, all are based on the predefined handcrafted features that cannot evolve and are out-of-date for characterizing emerging classes. Thus, to address the above challenges, we propose a novel generic end-to-end framework consisting of a dynamic cascade of classifiers that incrementally learn their dynamic and inherent features. The proposed method injects dynamic elements into the system by detecting instances from unknown classes, while at the same time incrementally updating the model to include the new classes. The resulting cascade tree grows by adding a new leaf node classifier once a new class is detected, and the discriminative features are updated via an end-to-end learning strategy. Experiments on two real-world datasets demonstrate that our proposed method outperforms existing state-of-the-art methods. Xiaojie Guo 0002, Amir Alipour-Fanid, Lingfei Wu 0001, Hemant Purohit, Xiang Chen 0010, Kai Zeng 0001, Liang Zhao 0002 |
CIKM | 4 |
| 2018 | CitizenHelper-Adaptive: Expert-Augmented Streaming Analytics System for Emergency Services and Humanitarian OrganizationsabstractThere is an increasing amount of information posted on Web, especially on social media during real world events. Likewise, there is a vast amount of information and opinions posted about humanitarian issues on social media. Mining such data can provide timely knowledge to inform disaster resource allocation for who needs what and where as well as policies for humanitarian causes. However, information overload is a key challenge in leveraging this big data resource for organizations. We present an interactive user-feedback based streaming analytics system `CitizenHelper-Adaptive' to mine social media, news, and other public Web data streams for emergency services and humanitarian organizations. The system aims to collect, organize, and visualize the vast amounts of data across various user and content-based information attributes using the adaptive machine learning models, such as intent classification models to continuously identify requests for help or offers of help during disasters. This demonstration shows the first application of transfer-active learning methods for time-critical events, when there is an availability of abundant labeled data from past events but a scarcity of the sufficient labeled data for the ongoing event. The proposed system provides a user interface to solicit expert feedback on the predicted instances from pretrained models and actively learns to improve the models for efficient information processing and organization. Finally, the system regularly updates the predicted information categories in the visualization dashboard. We will demo CitizenHelper-Adaptive system for case studies in both mass emergency events and humanitarian related topics such as gender violence using datasets of more than 50 million Twitter messages and news streams collected between 2016 and 2018. Rahul Pandey, Hemant Purohit |
ASONAM | 2 |
| 2018 | Social-EOC: Serviceability Model to Rank Social Media Requests for Emergency Operation CentersabstractThe public expects a prompt response from emergency services to address requests for help posted on social media. However, the information overload of social media experienced by these organizations, coupled with their limited human resources, challenges them to timely identify and prioritize critical requests. This is particularly acute in crisis situations where any delay may have a severe impact on the effectiveness of the response. While social media has been extensively studied during crises, there is limited work on formally characterizing serviceable help requests and automatically prioritizing them for a timely response. In this paper, we present a formal model of serviceability called Social-EOC (Social Emergency Operations Center), which describes the elements of a serviceable message posted in social media that can be expressed as a request. We also describe a system for the discovery and ranking of highly serviceable requests, based on the proposed serviceability model. We validate the model for emergency services, by performing an evaluation based on real-world data from six crises, with ground truth provided by emergency management practitioners. Our experiments demonstrate that features based on the serviceability model improve the performance of discovering and ranking (nDCG up to 25%) service requests over different baselines. In the light of these experiments, the application of the serviceability model could reduce the cognitive load on emergency operation center personnel, in filtering and ranking public requests at scale. Hemant Purohit, Carlos Castillo 0001, Muhammad Imran 0002, Rahul Pandey |
ASONAM | 1 |
| 2018 | Distributional Semantics Approach to Detect Intent in Twitter Conversations on Sexual AssaultsabstractThe recent surge in women reporting sexual assault and harassment (e.g., #metoo campaign) has highlighted a longstanding societal crisis. This injustice is partly due to a culture of discrediting women who report such crimes and also, rape myths (e.g., 'women lie about rape'). Social web can facilitate the further proliferation of deceptive beliefs and culture of rape myths through intentional messaging by malicious actors. This multidisciplinary study investigates Twitter posts related to sexual assaults and rape myths for characterizing the types of malicious intent, which leads to the beliefs on discrediting women and rape myths. Specifically, we first propose a novel malicious intent typology for social media using the guidance of social construction theory from policy literature that includes Accusational, Validational, or Sensational intent categories. We then present and evaluate a malicious intent classification model for a Twitter post using semantic features of the intent senses learned with the help of convolutional neural networks. Lastly, we analyze a Twitter dataset of four months using the intent classification model to study narrative contexts in which malicious intents are expressed and discuss their implications for gender violence policy design. Rahul Pandey, Hemant Purohit, Bonnie Stabile, Aubrey Grant |
WI | 2 |
| 2018 | Ranking of Social Media Alerts with Workload Bounds in Emergency Operation CentersabstractExtensive research on social media usage during emergencies has shown its value to provide life-saving information, if a mechanism is in place to filter and prioritize messages. Existing ranking systems can provide a baseline for selecting which updates or alerts to push to emergency responders. However, prior research has not investigated in depth how many and how often should these updates be generated, considering a given bound on the workload for a user due to the limited budget of attention in this stressful work environment. This paper presents a novel problem and a model to quantify the relationship between the performance metrics of ranking systems (e.g., recall, NDCG) and the bounds on the user workload. We then synthesize an alert-based ranking system that enforces these bounds to avoid overwhelming end-users. We propose a Pareto optimal algorithm for ranking selection that adaptively determines the preference of top-k ranking and user workload over time. We demonstrate the applicability of this approach for Emergency Operation Centers (EOCs) by performing an evaluation based on real world data from six crisis events. We analyze the trade-off between recall and workload recommendation across periodic and realtime settings. Our experiments demonstrate that the proposed ranking selection approach can improve the efficiency of monitoring social media requests while optimizing the need for user attention. Hemant Purohit, Carlos Castillo 0001, Muhammad Imran 0002, Rahul Pandey |
WI | 1 |
| 2017 | CitizenHelper: A Streaming Analytics System to Mine Citizen and Web Data for Humanitarian Organizations
Prakruthi Karuna, Mohammad Rana, Hemant Purohit |
ICWSM | 3 |
| 2014 | On Understanding the Divergence of Online Social Group Discussion
Hemant Purohit, Yiye Ruan, David Fuhry, Srinivasan Parthasarathy 0001, Amit P. Sheth |
ICWSM | 1 |
| 2013 | Twitris v3: From Citizen Sensing to Analysis, Coordination and Action
Hemant Purohit, Amit P. Sheth |
ICWSM | 1 |
| 2012 | Finding Influential Authors in Brand-Page Communities
Hemant Purohit, Jitendra Ajmera, Sachindra Joshi, Ashish Verma 0001, Amit P. Sheth |
ICWSM | 1 |
| 2010 | A Qualitative Examination of Topical Tweet and Retweet Practices
Meena Nagarajan, Hemant Purohit, Amit P. Sheth |
ICWSM | 2 |