Roshni Chakraborty

dblp:171/2226 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-4476-403XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 CAPS: A Cross-Lingual Methodology for Detecting Misinformation in Estonian Health News
abstract
Health misinformation poses a significant public threat by eroding trust in scientific expertise and diminishing adherence to health guidelines, which collectively weaken community resilience to preventable diseases. For these reasons, detecting health misinformation is crucial to protect public health. However, manual detection requires substantial human effort and expertise, making it impractical at scale, particularly in low-resource settings where technological and linguistic resources are limited. Developing automated techniques for identifying false or misleading claims is therefore essential to ensure timely intervention. Advancing these automated detection methods depends on the development of robust datasets, as they enable more accurate modeling and adaptation for specific languages and contexts. To the best of the authors’ knowledge, no misinformation detection techniques or datasets have yet been developed specifically for the Estonian language within the health domain. Addressing this gap, the primary objective of this study is to develop a reliable system for generating ground truth labels for health misinformation in Estonian, thereby contributing to misinformation detection in low-resource settings. Leveraging pre-labeled datasets in English, the proposed Cross-lingual Alignment and Confident Prediction Sampling (CAPS) approach employs a hybrid two-phase methodology involving semantic similarity measurements, manual annotation, classification, and confidence sampling. This methodology enables the efficient generation of misinformation labels with minimal reliance on manual annotation, contributing a valuable resource for advancing misinformation detection in underrepresented languages. The resulting dataset of 8,795 annotated news articles represents a significant advancement in health misinformation detection for the Estonian language.
Li Tetsmann, Uku Kangur, Roshni Chakraborty, Rajesh Sharma 0002
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2025 Ensemble-Based Deep Learning Framework for Multi-Class Skin Lesion Diagnosis with Class Imbalance Mitigation
abstract
Accurate multi-class classification of skin lesions remains a challenging task due to the intrinsic complexity of dermoscopic images and the severe class imbalance present in medical datasets. This paper introduces an ensemble-based deep learning pipeline designed to achieve robust and generalizable multi-class skin lesion diagnosis. The proposed framework systematically addresses the primary challenges of class imbalance using an aggressive mix up augmentation. A five face structured pipeline encompassing data preparation, stratified crossvalidation, model training, ensemble inference, and performance evaluation is performed. The proposed ensemble integrates five heterogeneous deep architectures: Vision Transformer (ViT), Swin Transformer, ConvNeXt, EfficientNet, and DenseNet, each contributing complementary representational strengths. To enhance model robustness and fairness, the pipeline incorporates regularization and a hybrid loss combining focal Loss and label smoothing. Ensemble predictions are aggregated using a weighted soft voting strategy, ensuring stable and accurate classification outcomes across all lesion categories. Experimental evaluations demonstrate that the proposed system achieves state-of-the-art performance on the HAM10000 dataset, delivering high accuracy and balanced sensitivity across minority classes. The end-to-end design emphasizes reproducibility, scalability, and transparency, establishing a robust foundation for future research in automated dermatological diagnosis.
Ujjwal Jain, Santosh Prakash Chouhan, Roshni Chakraborty, Mahua Bhattacharya
BIBM3
2025 Stance Detection in Twitter Conversations Using Reply Support Classification
Parul Khandelwal, Preety Singh, Rajbir Kaur, Roshni Chakraborty
ICPRAM4
2025 Attention-based semi-supervised active learning for multi-class tweet classification
Parul Khandelwal, Preety Singh, Rajbir Kaur, Roshni Chakraborty
Knowl. Inf. Syst.4
2025 ATSumm: Auxiliary information enhanced approach for abstractive disaster tweet summarization with sparse training data
Piyush Kumar Garg, Roshni Chakraborty, Sourav Kumar Dandapat
Knowl. Based Syst.2
2025 PORTRAIT: A Hybrid Approach to Create Extractive Ground-truth Summary for Disaster Event
abstract
Nowadays, X (formerly known as Twitter) is an important source of information and latest updates during ongoing events, such as disaster events. However, the huge number of tweets posted during a disaster makes identification of relevant information highly challenging. Therefore, a summary of the tweets can help the decision-makers to ensure efficient allocation of resources among the affected population. There exist several automated summarization approaches that can generate a summary given the tweets related to a disaster. Development of these automated summarization approaches require availability of ground-truth summary of the dataset for verification. However, the number of publicly available datasets along with the ground-truth summary for disaster events are still inadequate. To improve this situation, we need to create more ground-truth summaries. Existing approaches for ground-truth summary generation rely on the annotators’ wisdom and intuition. This process requires immense human effort and significant time. Moreover, the selection of the important tweets from the humongous set of input tweets often results in sub-optimal choice of tweets in the final summary. Therefore, to handle these challenges, we propose a hybrid approach (PORTRAIT) for ground-truth summary generation, where we partly automate the procedure to improve the quality of ground-truth summary and reduce human effort and time. We validate the effectiveness of PORTRAIT on nine disaster events through quantitative and qualitative analysis. We prepare and release the ground-truth summaries for nine disaster events, which consist of both natural and man-made disaster events belonging to five different continents.
Piyush Kumar Garg, Roshni Chakraborty, Sourav Kumar Dandapat
ACM Trans. Web2
2024 IKDSumm: Incorporating key-phrases into BERT for extractive disaster tweet summarization
Piyush Kumar Garg, Roshni Chakraborty, Srishti Gupta 0001, Sourav Kumar Dandapat
Comput. Speech Lang.2
2024 OntoDSumm: Ontology-Based Tweet Summarization for Disaster Events
abstract
The huge popularity of social media platforms, such as Twitter, attracts a large fraction of users to share real-time information and short situational messages during disasters. A summary of these tweets is required by the government organizations, agencies, and volunteers for efficient and quick disaster response. However, the huge influx of tweets makes it difficult to manually get a precise overview of ongoing events. To handle this challenge, several tweet summarization approaches have been proposed. In most of the existing literature, tweet summarization is broken into a two-step process where, in the first step, it categorizes tweets, and in the second step, it chooses representative tweets from each category. There are both supervised and unsupervised approaches found in the literature to solve the problem of first step. Supervised approaches require a huge amount of labeled data, which incurs cost as well as time. On the other hand, unsupervised approaches could not cluster tweet properly due to the overlapping keywords, vocabulary size, lack of understanding of semantic meaning, and so on, while, for the second step of summarization, existing approaches applied different ranking methods where those ranking methods are very generic, which fail to compute proper importance of a tweet with respect to a disaster. Both problems can be handled far better with proper domain knowledge. In this article, we exploited already existing domain knowledge by the means of ontology in both steps and proposed a novel disaster summarization method OntoDSumm. We evaluate this proposed method with six state-of-the-art methods using 12 disaster datasets. Evaluation results reveal that OntoDSumm outperforms the existing methods by approximately 2%–66% in terms of ROUGE-1 F1-score.
Piyush Kumar Garg, Roshni Chakraborty, Sourav Kumar Dandapat
IEEE Trans. Comput. Soc. Syst.2
2023 SigGAN: Adversarial Model for Learning Signed Relationships in Networks
abstract
Signed link prediction in graphs is an important problem that has applications in diverse domains. It is a binary classification problem that predicts whether an edge between a pair of nodes is positive or negative. Existing approaches for link prediction in unsigned networks cannot be directly applied for signed link prediction due to their inherent differences. Furthermore, signed link prediction must consider the inherent characteristics of signed networks, such as structural balance theory. Recent signed link prediction approaches generate node representations using either generative models or discriminative models. Inspired by the recent success of Generative Adversarial Network (GAN) based models in several applications, we propose a GAN based model for signed networks, SigGAN. It considers the inherent characteristics of signed networks, such as integration of information from negative edges, high imbalance in number of positive and negative edges, and structural balance theory. Comparing the performance with state-of-the-art techniques on five real-world datasets validates the effectiveness of SigGAN.
Roshni Chakraborty, Ritwika Das, Joydeep Chandra
ACM Trans. Knowl. Discov. Data1
2023 Finding Representative Sampling Subsets in Sensor Graphs Using Time-series Similarities
abstract
With the increasing use of Internet-of-Things–enabled sensors, it is important to have effective methods to query the sensors. For example, in a dense network of battery-driven temperature sensors, it is often possible to query (sample) only a subset of the sensors at any given time, since the values of the non-sampled sensors can be estimated from the sampled values. If we can divide the set of sensors into disjoint so-calledrepresentative sampling subsets, in which each represents all the other sensors sufficiently well, then we can alternate between the sampling subsets and, thus, increase the battery life significantly of the sensor network. In this article, we formulate the problem of finding representative sampling subsets as a graph problem on a so-calledsensor graphwith the sensors as nodes. Our proposed solution,SubGraphSample, consists of two phases. In Phase-I, we create edges in thesimilarity graphbased on the similarities between the time-series of sensor values, analyzing six different techniques based on proven time-series similarity metrics. In Phase-II, we propose six different sampling techniques to find the maximum number ofrepresentative sampling subsets. Finally, we proposeAutoSubGraphSample, which auto-selects the best technique for Phase-I and Phase-II for a given dataset. Our extensive experimental evaluation shows thatAutoSubGraphSamplecan yield significant battery-life improvements within realistic error bounds.
Roshni Chakraborty, Josefine Kejser, Torben Bach Pedersen, Petar Popovski
ACM Trans. Sens. Networks1
2022 HCNA: Hyperbolic Contrastive Learning Framework for Self-Supervised Network Alignment
Shruti Saxena, Roshni Chakraborty, Joydeep Chandra
Inf. Process. Manag.2
2019 Tweet Summarization of News Articles: An Objective Ordering-Based Perspective
abstract
Twitter has become an essential platform for the news media sources to disseminate news. The opinions expressed through Twitter can be mined by news media sources to obtain users' reactions centered around different news articles. A comprehensive summary of the users' reactions with respect to a news article can be crucial due to various reasons like: 1) understanding the sensitivity/importance of the news; 2) obtaining insights about the diverse opinions of the readers with respect to the news; and 3) understanding the key aspects that draw the interest of the readers. However, the selected summary tweets must fulfill multiple objectives, like relevance to the news article, diversity among the selected tweets, and should cover the entire spectrum of opinions expressed through the tweets. Existing methods primarily attempt to identify a set of relevant tweets from which the summary tweets are selected that maintains the diversity and coverage requirements. However, the noise and the nontemporal behavior of the article-specific tweets make the identification of such relevant tweets extremely difficult, resulting in poor summary quality. In this paper, through empirical investigations, we show that initially identifying the diverse opinions can lead to better identification of the relevant tweets, i.e., following a specific ordering of the objectives can lead to the improved summary. We, subsequently, propose a tweet summarization technique that follows such a specific ordering. Validation of our proposed approach for 800 news articles with 2.1 billion related tweets shows that the proposed approach produces 11.6%-34.8% improvement in summary quality as compared to existing state-of-the-art techniques.
Roshni Chakraborty, Maitry Bhavsar, Sourav Kumar Dandapat, Joydeep Chandra
IEEE Trans. Comput. Soc. Syst.1
2017 A Network Based Stratification Approach for Summarizing Relevant Comment Tweets of News Articles
Roshni Chakraborty, Maitry Bhavsar, Sourav Kumar Dandapat, Joydeep Chandra
WISE (1)1
2015 Analyzing Link Dynamics in Scientific Collaboration Networks: A Social Yield Based Perspective
abstract
In this paper, we introduce social yield, a measure of collaboration success of the collaborating authors in a coauthorship network. We then attempt to empirically observe the link dynamics in collaboration networks induced by the social yield of the collaborations. Observation indicate that certain observed behavior like presence of large number of small sized communities and highly dynamic behavior of the links in collaboration networks can be explained based on the distribution of social yield of these collaborations. It is also observed that the distribution of social yield among the collaborations also affects the resilience of the collaboration networks to targeted link removal.
Arun Pandey, Roshni Chakraborty, Joydeep Chandra
ASONAM2