Sourav Kumar Dandapat

dblp:97/8395 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-2043-2356ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Computer networks · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploiting Joint Influence of Inter- and Intra-Clause Dependencies Towards Enhanced Emotion Cause Extraction
abstract
Emotion cause extraction (ECE) is vital in understanding the triggers of emotions portrayed in text, thereby enriching user interaction and relatability. Traditional methods such as rule-based and lexical matching struggle due to the paucity of lexical similarity between causes and emotions. While machine learning and deep learning have improved by modeling sequential dependencies and multilevel representations, they still fall short in handling long-distance dependencies and interclause interactions. Attention-based techniques have demonstrated efficacy in refining semantic context over longer text spans, but they frequently suffer from scalability issues for moderate-sized datasets. To tackle the aforementioned shortcomings, we propose a novel model incorporating inter and intraclause relationships, improving the identification of causes behind emotions. Our approach aims to capture fine-grained interactions within the clause (intra) using a dual-attention mechanism involving both cross- and self-attention on multisource features. This addresses scalability by incorporating numerical features as auxiliary information. To capture the dependencies between clauses (inter), we employ a contrastive learning approach to group relevant clauses together in feature space. This helps address long-distance dependency issues both within and between clauses. Through extensive experimentation on the RECCON benchmark dataset, we depict the effectiveness of our strategy in increasing the F1-score of the ECE task by$20{-}25$%. We also extend the evaluation of our model by depicting its application in the empathetic response generation task. The derived causes by our model, when used in response generation, enhance the empathy of a response.
Srishti Gupta 0001, Sourav Kumar Dandapat
IEEE Trans. Comput. Soc. Syst.2
2025 Beyond just saying it's false: explainable AI for multimodal misinformation detection
Saswata Roy, Manish Bhanu, Shalini Priya, Joydeep Chandra, Sourav Kumar Dandapat
Appl. Intell.5
2025 ATSumm: Auxiliary information enhanced approach for abstractive disaster tweet summarization with sparse training data
Piyush Kumar Garg, Roshni Chakraborty, Sourav Kumar Dandapat
Knowl. Based Syst.3
2025 When Voices Speak Louder: Leveraging Audio Signals in Emotion-Cause Extraction via Large Multilingual Multimodal Indian Dialogue Datasets
abstract
Multimodal emotion and cause recognition have progressed significantly since their inception. Existing multimodal English datasets and newer research extending to languages like Chinese and Polish have been developed to identify conversation emotion-cause pairs. However, these datasets often lack cultural relevance, particularly for Indian languages, which remain underrepresented in the field. This paper aims to bridge this gap by introducing new datasets in India's two most spoken languages: Hindi and Bengali. Our work provides essential resources for more culturally and linguistically diverse emotion and cause recognition tasks, contributing to the broader goal of enhancing detection in multilingual, multimodal Indian contexts.
Srishti Gupta 0001, Sourav Kumar Dandapat
IEEE Signal Process. Lett.3
2025 PORTRAIT: A Hybrid Approach to Create Extractive Ground-truth Summary for Disaster Event
abstract
Nowadays, X (formerly known as Twitter) is an important source of information and latest updates during ongoing events, such as disaster events. However, the huge number of tweets posted during a disaster makes identification of relevant information highly challenging. Therefore, a summary of the tweets can help the decision-makers to ensure efficient allocation of resources among the affected population. There exist several automated summarization approaches that can generate a summary given the tweets related to a disaster. Development of these automated summarization approaches require availability of ground-truth summary of the dataset for verification. However, the number of publicly available datasets along with the ground-truth summary for disaster events are still inadequate. To improve this situation, we need to create more ground-truth summaries. Existing approaches for ground-truth summary generation rely on the annotators’ wisdom and intuition. This process requires immense human effort and significant time. Moreover, the selection of the important tweets from the humongous set of input tweets often results in sub-optimal choice of tweets in the final summary. Therefore, to handle these challenges, we propose a hybrid approach (PORTRAIT) for ground-truth summary generation, where we partly automate the procedure to improve the quality of ground-truth summary and reduce human effort and time. We validate the effectiveness of PORTRAIT on nine disaster events through quantitative and qualitative analysis. We prepare and release the ground-truth summaries for nine disaster events, which consist of both natural and man-made disaster events belonging to five different continents.
Piyush Kumar Garg, Roshni Chakraborty, Sourav Kumar Dandapat
ACM Trans. Web3
2024 IKDSumm: Incorporating key-phrases into BERT for extractive disaster tweet summarization
Piyush Kumar Garg, Roshni Chakraborty, Srishti Gupta 0001, Sourav Kumar Dandapat
Comput. Speech Lang.4
2024 OntoDSumm: Ontology-Based Tweet Summarization for Disaster Events
abstract
The huge popularity of social media platforms, such as Twitter, attracts a large fraction of users to share real-time information and short situational messages during disasters. A summary of these tweets is required by the government organizations, agencies, and volunteers for efficient and quick disaster response. However, the huge influx of tweets makes it difficult to manually get a precise overview of ongoing events. To handle this challenge, several tweet summarization approaches have been proposed. In most of the existing literature, tweet summarization is broken into a two-step process where, in the first step, it categorizes tweets, and in the second step, it chooses representative tweets from each category. There are both supervised and unsupervised approaches found in the literature to solve the problem of first step. Supervised approaches require a huge amount of labeled data, which incurs cost as well as time. On the other hand, unsupervised approaches could not cluster tweet properly due to the overlapping keywords, vocabulary size, lack of understanding of semantic meaning, and so on, while, for the second step of summarization, existing approaches applied different ranking methods where those ranking methods are very generic, which fail to compute proper importance of a tweet with respect to a disaster. Both problems can be handled far better with proper domain knowledge. In this article, we exploited already existing domain knowledge by the means of ontology in both steps and proposed a novel disaster summarization method OntoDSumm. We evaluate this proposed method with six state-of-the-art methods using 12 disaster datasets. Evaluation results reveal that OntoDSumm outperforms the existing methods by approximately 2%–66% in terms of ROUGE-1 F1-score.
Piyush Kumar Garg, Roshni Chakraborty, Sourav Kumar Dandapat
IEEE Trans. Comput. Soc. Syst.3
2023 SEEC and CHASE: An emotion-cause pair-oriented approach and conversational dataset with heterogeneous emotions for empathetic response generation
Srishti Gupta 0001, Sourav Kumar Dandapat
Knowl. Based Syst.2
2022 Towards an orthogonality constraint-based feature partitioning approach to classify veracity and identify stance overlapping of rumors on twitter
Saswata Roy, Manish Bhanu, Sourav Kumar Dandapat, Joydeep Chandra
Expert Syst. Appl.3
2022 gDART: Improving rumor verification in social media with Discrete Attention Representations
Saswata Roy, Manish Bhanu, Shruti Saxena, Sourav Kumar Dandapat, Joydeep Chandra
Inf. Process. Manag.4
2022 Ripple: An approach to locate k nearest neighbours for location-based services
Pratima Biswas, Sourav Kumar Dandapat, Ashok Singh Sairam
Inf. Syst.2
2020 EnDeA: Ensemble based Decoupled Adversarial Learning for Identifying Infrastructure Damage during Disasters
abstract
Identifying tweets related to infrastructure damage during a crisis event is an important problem. However, the unavailability of labeled data during the early stages of a crisis event poses major challenge in training suitable models. Several domain adaptation strategies have been proposed for text classification that can be used to train models using available source data of previous crisis events and apply on a target data related to a current event. However, these approaches are insufficient to handle the distribution drift in the source and target data along with the class imbalance in the target data. In this paper we introduce an Ensemble learning approach with a Decoupled Adversarial (EnDeA) model to classify infrastructure damage tweets in a target tweet dataset. EnDeA is an ensemble of three different models two of which separately learn the event invariant and specific features of a target data from a set of source and target data. The third model which is an adversarial model helps to improve the prediction accuracy of both models. Unlike the existing approaches that also identify the domain invariant and specific properties of target data for sentiment classification, our method works for short texts and can better handle the distribution drift and class imbalance problem. We rigorously investigate the performance of the proposed approach using multiple public datasets and compare it with several state-of-the-art baselines. We discover that EnDeA outperforms these baselines with around 20% improvement in the 1 scores.
Shalini Priya, Apoorva Upadhyaya, Manish Bhanu, Sourav Kumar Dandapat, Joydeep Chandra
CIKM4
2020 TAQE: Tweet Retrieval-Based Infrastructure Damage Assessment During Disasters
abstract
Twitter is an active communication channel for the spreading of updated information in emergency situations. Retrieving specific information related to infrastructure damage offers the situational views to the concerned authorities, who can take necessary action to disburse help. However, such usages of Twitter demand significant accuracy of the retrieved information. Previous techniques on IR have not been able to capture the semantic variations satisfactorily in the tweets, due to low content quality and vocabulary gap, and consequently have failed to yield considerable performance. This has left ample scope for further improvement in this area of research. There are two major contributions of our work: 1) developing a relevant tweet retrieval framework that provides information about infrastructure damage and 2) assignment of a relative damage score to the affected regions so that the severity of the damage can be assessed. Our proposed technique involves a novel split-query-based mechanism with topic aligned query expansion (TAQE) to retrieve relevant tweets that are subsequently used for measuring the infrastructure damage across different locations. We report empirical results on multiple-crisis-related data sets to establish the efficacy of our approach to these events at different locations. Empirical validation of our proposed approach on manually annotated ground-truth data reveals considerably better performance metrics in terms of precision, recall, Bpref, and MAP over several state-of-the-art techniques.
Shalini Priya, Manish Bhanu, Sourav Kumar Dandapat, Kripabandhu Ghosh, Joydeep Chandra
IEEE Trans. Comput. Soc. Syst.3
2020 Understanding the Impact of Geographical Distance on Online Discussions
abstract
People in geographically close areas tend to show similarities in their interests such as sports, food, and festivals, which is often reflected in their discussion. Lifestyle and the way of communication among people living nearby play an essential role in this similarity. However, with the popularity of online social media platforms, communication with distant persons has become much easier and frequent. In one sense, social media platforms help in breaking the barrier of distance and bring people across geography closer. With this, a comprehensive study is required to understand whether geographical distance still has a significant impact on discussion or it has already faded and made us a global citizen. Moreover, if there is an impact of geographical distance on discussion, then it would be interesting to investigate whether this impact is uniform across different topics of discussion or not. This understanding will help in targeted marketing, advertisement, and policy-making. In this article, we analyze the geotagged tweet data collected for a period of around five months for three countries, USA, U.K., and India, each with diverse cultures and unique identities. We measure the impact of geographical distance at a finer granularity (within a country) in online discussions in terms of content similarity and other geographical parameters of topical tweets. This article reflects that there is a significant homogenization in online discussions with respect to geographical distance; however, this homogenization is not similar across all the topics of discussion and location. The impact varies depending on the topic of discussion and location.
Rimjhim, Nikhil Cheke, Joydeep Chandra, Sourav Kumar Dandapat
IEEE Trans. Comput. Soc. Syst.4
2019 Identifying infrastructure damage during earthquake using deep active learning
abstract
Twitter provides important information for emergency responders in the rescue process during disasters. However, tweets containing relevant information are sparse and are usually hidden in a vast set of noisy contents. This leads to inherent challenges in generating suitable training data that are required for neural network models. In this paper, we study the problem of retrieving the infrastructure damage information from tweets generated from different location during crisis using the model actively trained on past but similar events. We combine RNN and GRU based model coupled with active learning that gets trained on most uncertain samples and captures the latent features of different data distribution. It reduces the uses of around 90% less training data, thereby significantly reducing the manual annotation efforts. We use the model pre-trained using active learning based approach to retrieve the infrastructure damage tweets originated from different regions. We obtain a minimum of 18% gain on F1-measure and considerably on other metrics over recent state-of-the-art IR techniques.
Shalini Priya, Saharsh Singh, Sourav Kumar Dandapat, Kripabandhu Ghosh, Joydeep Chandra
ASONAM3
2019 Tweet Summarization of News Articles: An Objective Ordering-Based Perspective
abstract
Twitter has become an essential platform for the news media sources to disseminate news. The opinions expressed through Twitter can be mined by news media sources to obtain users' reactions centered around different news articles. A comprehensive summary of the users' reactions with respect to a news article can be crucial due to various reasons like: 1) understanding the sensitivity/importance of the news; 2) obtaining insights about the diverse opinions of the readers with respect to the news; and 3) understanding the key aspects that draw the interest of the readers. However, the selected summary tweets must fulfill multiple objectives, like relevance to the news article, diversity among the selected tweets, and should cover the entire spectrum of opinions expressed through the tweets. Existing methods primarily attempt to identify a set of relevant tweets from which the summary tweets are selected that maintains the diversity and coverage requirements. However, the noise and the nontemporal behavior of the article-specific tweets make the identification of such relevant tweets extremely difficult, resulting in poor summary quality. In this paper, through empirical investigations, we show that initially identifying the diverse opinions can lead to better identification of the relevant tweets, i.e., following a specific ordering of the objectives can lead to the improved summary. We, subsequently, propose a tweet summarization technique that follows such a specific ordering. Validation of our proposed approach for 800 news articles with 2.1 billion related tweets shows that the proposed approach produces 11.6%-34.8% improvement in summary quality as compared to existing state-of-the-art techniques.
Roshni Chakraborty, Maitry Bhavsar, Sourav Kumar Dandapat, Joydeep Chandra
IEEE Trans. Comput. Soc. Syst.3
2019 A Large-Scale Study of the Twitter Follower Network to Characterize the Spread of Prescription Drug Abuse Tweets
abstract
In this article, we perform a large-scale study of the Twitter follower network, involving around 0.42 million users who justify DA, to characterize the spreading of DA tweets across the network. Our observations reveal the existence of a very large giant component involving 99% of these users with dense local connectivity that facilitates the spreading of such messages. We further identify active cascades over the network and observe that the cascades of DA tweets get spread over a long distance through the engagement of several closely connected groups of users. Moreover, our observations also reveal a collective phenomenon, involving a large set of active fringe nodes (with a small number of follower and following) along with a small set of well-connected nonfringe nodes that work together toward such spread, thus potentially complicating the process of arresting such cascades. Furthermore, we discovered that the engagement of the users with respect to certain drugs, such as Vicodin, Percocet, and OxyContin, that were observed to be most mentioned in Twitter is instantaneous. On the other hand, for drugs, such as Lortab, that found lesser mentions, the engagement probability becomes high with increasing exposure to such tweets, thereby indicating that drug abusers engaged on Twitter remain vulnerable to adopting newer drugs, aggravating the problem further.
Ryan Sequeira, Avijit Gayen, Niloy Ganguly, Sourav Kumar Dandapat, Joydeep Chandra
IEEE Trans. Comput. Soc. Syst.4
2018 Forecasting Traffic Flow in Big Cities Using Modified Tucker Decomposition
Manish Bhanu, Shalini Priya, Sourav Kumar Dandapat, Joydeep Chandra, João Mendes-Moreira 0001
ADMA3
2018 Characterizing Infrastructure Damage After Earthquake: A Split-Query Based IR Approach
abstract
Retrieving relevant information from social media based on specific requirements has become a focus area for researchers. In this paper, we propose a framework for online retrieval of tweets providing information about possible infrastructure damages, caused due to earthquakes and use the same to determine a damage score for the possibly affected locations. Identifying such tweets would not only provide a holistic view of the affected areas but would also help in taking necessary relief actions. Existing works on this topic fail to effectively capture the semantic variation in the tweets, possibly due to poor content quality, thereby providing scopes for further improvement in the mechanisms involved. Our proposed technique relies on a novel split-query based mechanism along with a pseudo-relevance feedback approach to identify the relevant tweets. The pseudo-relevance feedback approach expands on an initial set of seed tweets obtained using a semi-automatic query generation mechanism that couples topic based clustering with human annotation. Empirical validation of our proposed method on a manually annotated ground truth data reveals a considerable improvement in precision, recall and mean average precision over several baseline methods.
Shalini Priya, Manish Bhanu, Sourav Kumar Dandapat, Kripabandhu Ghosh, Joydeep Chandra
ASONAM3
2017 A Network Based Stratification Approach for Summarizing Relevant Comment Tweets of News Articles
Roshni Chakraborty, Maitry Bhavsar, Sourav Kumar Dandapat, Joydeep Chandra
WISE (1)3
2015 ActivPass: Your Daily Activity is Your Password
abstract
This paper explores the feasibility of automatically extracting passwords from a user's daily activity logs, such as her Facebook activity, phone activity etc. As an example, a smartphone might ask the user: "Today morning from whom did you receive an SMS?" In this paper, we observe that infrequent activities (i.e., outliers) can be memorable and unpredictable. Building on this observation, we have developed an end to end system ActivPass and experimented with 70 users. With activity logs from Facebook, browsing history, call logs, and SMSs, the system achieves 95% success (authenticates legitimate users) and is compromised in 5.5% cases (authenticates impostors). While this level of security is obviously inadequate for serious authentication systems, certain practices such as password sharing can immediately be thwarted from the dynamic nature of passwords. With security improvements in the future, activity-based authentication could fill in for the inadequacies in today's password-based systems.
Sourav Kumar Dandapat, Swadhin Pradhan, Bivas Mitra, Romit Roy Choudhury, Niloy Ganguly
CHI1
2012 Framework for Collaborative Download in Wireless Mobile Environment
abstract
Proliferation of wireless technology and increasing demand of data result in traffic congestion. Congestion due to wireless Internet is increasing at an exponential rate and 50% of this traffic is due to video. However, as popularity of files in Internet follows power law and people with similar interest meet/interact more frequently, there is a good probability that many users in proximity would like to download similar files. In this paper we exploit this probability and propose a scheme for collaborative download where a number of users in proximity collaborate with each other for downloading a set of files. We have also developed a prototype in android platform, with basic features of collaborative download.
Sourav Kumar Dandapat, Ravi Niranjan, Niloy Ganguly
MDM1
2012 Distributed content storage for just-in-time streaming
abstract
We propose a content distribution strategy over municipal WiFi networks where Access Points (APs) collaboratively cache popular multimedia content, and disseminate them in a manner that each mobile device has the portion of the content just-in-time for playback. If successful, we envision that a child will be able to seamlessly watch a movie in a car, as her tablet downloads different parts of the movie over different WiFi APs at different times.
Sourav Kumar Dandapat, Sanyam Jain, Romit Roy Choudhury, Niloy Ganguly
SIGCOMM1
2012 Smart Association Control in Wireless Mobile Environment Using Max-Flow
abstract
WiFi clients must associate to a specific Access Point (AP) to communicate over the Internet. Current association methods are based on maximum Received Signal Strength Index (RSSI) implying that a client associates to the strongest AP around it. This is a simple scheme that has performed well in purely distributed settings. Modern wireless networks, however, are increasingly being connected by a wired backbone. The backbone allows for out-of-band communication among APs, opening up opportunities for improved protocol design. This paper takes advantage of this opportunity through a coordinated client association scheme where APs consider a global view of the network, and decide on the optimal client-AP association. We show that such an association outperforms RSSI based schemes in several scenarios, while remaining practical and scalable for wide-scale deployment. We also show that optimal association is a NP-Hard problem and our max-flow based heuristic is a promising solution.
Sourav Kumar Dandapat, Bivas Mitra, Romit Roy Choudhury, Niloy Ganguly
IEEE Trans. Netw. Serv. Manag.1
2010 Fair bandwidth allocation in wireless mobile environment using max-flow
abstract
Wireless clients must associate to a specific Access Point (AP) to communicate over the Internet. Current association methods are based on maximum Received Signal Strength Index (RSSI) implying that a client associates to the strongest AP around it. This is a simple scheme that has performed well in purely distributed settings. Modern wireless networks, however, are increasingly being connected by a wired backbone. The backbone allows for out-of-band communication among APs, opening up opportunities for improved protocol design. This paper takes advantage of this opportunity through a coordinated client association scheme where APs consider a global view of the network, and decides on the optimal client-AP association. We show that such an association outperforms RSSI based schemes in several scenarios, while remaining practical and scalable for wide-scale deployment. Although an early work in this direction, our basic analytical framework (based on a max-flow formulation) can be extended to sophisticated channel and traffic models. Our future work is focussed towards designing and evaluating these extensions.
Sourav Kumar Dandapat, Bivas Mitra, Niloy Ganguly, Romit Roy Choudhury
HiPC1
2010 Fair bandwidth allocation in wireless network using max-flow
abstract
This paper proposes a fair association scheme between clients and APs in WiFi network, exploiting the hybrid nature of the recent WLAN architecture. We show that such an association outperforms RSSI based schemes in several scenarios, while remaining practical and scalable for wide-scale deployment.
Sourav Kumar Dandapat, Bivas Mitra, Niloy Ganguly, Romit Roy Choudhury
SIGCOMM1