Sanasam Ranbir Singh

dblp:81/2295 · DBLP profile ↗
← Back
51ranked-venue papers
3as first author
33since 2021 · last 2026
0000-0003-0484-2144ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 1 first-author · 20 since 2021Databases, data management, data science and information retrieval · 17 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 4 since 2021Security and privacy · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Sentence-Level Back-Transliteration of Romanized Indian Languages: Performance Analysis and Challenges
Dhruvkumar Babubhai Kakadiya, Sanasam Ranbir Singh, Sukumar Nandi
LREC3
2026 Few-shot Prompting or Supervised Tuning? A Comparative Study of LLMs for Linguistically Distant Language Pairs in BDI
Deepen Naorem, Sanasam Ranbir Singh, Telem Joyson Singh, Priyankoo Sarmah
LREC2
2026 Insights from Romanized Manipuri Social Media Text: A Transliteration Corpus and Variation Analysis
Maisang Kamei Salice, Sanasam Ranbir Singh, Priyankoo Sarmah
LREC2
2026 AssamLegalTrans: A Parallel Corpus, Benchmark and Analysis for English-Assamese Machine Translation of Legal Judgments
Telem Joyson Singh, Hemanta Baruah, Sanasam Ranbir Singh, Anindita Talukdar, Nasrin Shahnaz, Okram Jimmy Singh, Priyankoo Sarmah, Pallav Kumar Dutta, Sukumar Nandi, Pranab Duara
LREC3
2026 Overcoming structural and similarity barriers: Graph neural network-based contextual reasoning for detecting incongruent and partially misleading news
Sujit Kumar, Sanasam Ranbir Singh
Neurocomputing3
2026 MVNIDS: A multiview-based network intrusion detection system
Sunit Kumar Nandi, Ritesh Ratti, Sanasam Ranbir Singh, Sukumar Nandi
J. Inf. Secur. Appl.3
2026 From Headlines and News to Truth: Leveraging LLMs With Cascade Classification and Prompt Tuning for Incongruent News Detection
abstract
Misleading and incongruent news on digital platforms is pivotal in spreading misinformation and disinformation, leading to widespread skepticism and misconception among audiences or content consumers. Existing studies employ similarity-based, summarization-based, and dual summarization methods to detect incongruent news. However, the incongruent news detection methods in existing studies exhibit the following key limitations. 1) Bag-of-words-based methods fail to capture the deep semantic relationships and complex negations between the news body sentences and the headline and between the news body and the headline. 2) Similarity-based methods struggle to detect incongruent news articles with concise headlines and large news bodies. 3) The existing methods in the literature are ineffective in partially incongruent news detection. 4) Methods in the literature are inadequate in handling topic inconsistency between sentences of the news body and topic dissimilarities between headlines and bodies for detecting incongruent articles. To address the aforementioned limitations of existing studies, this work fine-tunes bidirectional encoder representations from transformers (BERT) and robustly optimized BERT pretraining approach (RoBERTa) and proposescascade classificationandmasked language modeling with promptfor detecting misleading and incongruent news articles. We conduct our experiments on four publicly available benchmark datasets, and our results suggest that our proposed models outperform existing state-of-the-art methods in the literature and effectively detect incongruent and fake news of various forms, including misleading, satire, hoax, and propaganda. We also evaluate the robustness of our proposed methods in detecting incongruent and fake news by assessing their performance on a curated dataset of real-world misinformation and incongruent news instances. Our evaluations suggest that the proposed models are highly effective in detecting incongruent and fake news articles across diverse genres.
Sujit Kumar, Sanasam Ranbir Singh
IEEE Trans. Comput. Soc. Syst.2
2026 Prompt-based Masked Language Modeling for Numerical Reasoning
abstract
Headline generation, a crucial task in summarization, aims to summarize an entire article into a concise, single line. Despite the proficiency of sequence-to-sequence encoder-decoder models and transformer-based large language models (LLMs) in text generation and summarization, generating headlines that include numerals representing the numerals in the news body remains a significant challenge. Generating a numeral-aware headline requires the ability of models to solve numerical and mathematical reasoning capabilities to infer relationships between numerals in the news body. Given the challenges in numeral-aware headline generation and numerical reasoning over numerals in news bodies, this study conducts an empirical investigation of LLMs using various strategies, including Pretrained , Few-shot Prompting , and Chain-of-Thought (CoT) Prompting , for numeral-aware headline generation and numerical reasoning for headline generation. Building upon the insights gained from our empirical study on LLMs for numeral-aware headline generation and numerical reasoning, we propose two novel approaches: instruction tuning with LLMs for numeral-aware headline generation and prompt-based masked language modeling for numerical reasoning. We conducted our experiments on the NumHG dataset. We observed that our proposed method outperforms the Pretrained , Few-shot Prompting , and CoT Prompting setups of LLMs, as well as baseline models from the literature, on both numeral-aware headline generation and numerical reasoning tasks. Observations from the experimental results reveal that our proposed Prompt-based Masked Language Modeling significantly improves the performance of small and medium-sized language models on numerical reasoning tasks. We also study the robustness of our proposed models by evaluating their performance in fact-checking numerical claims and performing numerical reasoning for numeral-aware text summarization. Our findings suggest that the proposed Prompt-based Masked Language Modeling approach is also effective for numerical claim verification and numerical reasoning for numeral-aware text summarization.
Sujit Kumar, Tanveen, Sanasam Ranbir Singh
ACM Trans. Intell. Syst. Technol.5
2025 Deep Modality-Disentangled Prompt Tuning for Few-Shot Multimodal Sarcasm Detection
Soumyadeep Jana, Abhrajyoti Kundu, Sanasam Ranbir Singh
CIKM3
2025 SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification
abstract
Verifying scientific claims presents a significantly greater challenge than verifying political or news-related claims. Unlike the relatively broad audience for political claims, the users of scientific claim verification systems can vary widely, ranging from researchers testing specific hypotheses to everyday users seeking information on a medication. Additionally, the evidence for scientific claims is often highly complex, involving technical terminology and intricate domain-specific concepts that require specialized models for accurate verification. Despite considerable interest from the research community, there is a noticeable lack of large-scale scientific claim verification datasets to benchmark and train effective models. To bridge this gap, we introduce two large-scale datasets, SciClaimHunt and SciClaimHunt_Num, derived from scientific research papers. We propose several baseline models tailored for scientific claim verification to assess the effectiveness of these datasets. Additionally, we evaluate models trained on SciClaimHunt and SciClaimHunt_Num against existing scientific claim verification datasets to gauge their quality and reliability. Furthermore, we conduct human evaluations of the dataset’s claims and perform error analysis to assess the effectiveness of the proposed baseline models. Our findings indicate thatSciClaimHunt and SciClaimHunt_Num serve as highly reliable resources for training models in scientific claim verification.
Sujit Kumar, Anshul Sharma, Siddharth Hemant Khincha, Gargi Shroff, Sanasam Ranbir Singh, Rahul Mishra 0004
IJCNN5
2025 Chain-of-Morphemes Tuning: Injecting Morphology in LLM-based Machine Translation
abstract
Generative Large Language Models (LLMs) have recently achieved remarkable advancements in translation tasks. However, low-resource or unseen language words are underrepresented in the LLM vocabulary where words are frequently constructed from subwords that may not align with their morphological structures. In this study, we present a novel two-stage morphological knowledge injection strategy into LLMs to address the difficulties presented by low-resource, agglutinative languages focusing on English-Manipuri translation as a case study. The first stage involves extending the LLM vocabulary with morphemes and segmenting Manipuri words into fine-grained morphological units that preserve morpheme boundaries in the input representation. The second stage introduces a chain-of-morphemes (CoM) tuning approach that divides the translation process into two steps: first, generating word roots for semantic translation, and then applying morphological inflection. This methodology decouples semantic and morphological processing, improving translation quality. Evaluation on the WMT 2023 English-Manipuri dataset using mGPT and BLOOM models demonstrates that our approach outperforms existing baseline methods. Our findings emphasize the potential advantages of incorporating morphological knowledge in LLM-based machine translation for morphologically rich languages.
Telem Joyson Singh, Sanasam Ranbir Singh, Deepen Naorem, Priyankoo Sarmah
IJCNN2
2025 Fake News Detection using Hashtag Context
Sujit Kumar, Shifali Agrahari, Priyank Soni, Aayush Sachdeva, Sanasam Ranbir Singh
Pattern Recognit. Lett.5
2025 Distilling Knowledge in Machine Translation of Agglutinative Languages with Backward and Morphological Decoders
abstract
Agglutinative languages often have morphologically complex words (MCWs) composed of multiple morphemes arranged in a hierarchical structure, posing significant challenges in translation tasks. We present a novel Knowledge Distillation approach tailored for improving the translation of such languages. Our method involves an encoder, a forward decoder, and two auxiliary decoders: a backward decoder and a morphological decoder. The forward decoder generates target morphemes autoregressively and is augmented by distilling knowledge from the auxiliary decoders. The backward decoder incorporates future context, while the morphological decoder integrates target-side morphological information. We have also designed a reliability estimation method to selectively distill only the reliable knowledge from these auxiliary decoders. Our approach relies on morphological word segmentation. We show that the word segmentation method based on unsupervised morphology learning outperforms the commonly used Byte Pair Encoding method on highly agglutinative languages in translation tasks. Our experiments conducted on English-Tamil, English-Manipuri, and English-Marathi datasets show that our proposed approach achieves significant improvements over strong Transformer-based NMT baselines.
Telem Joyson Singh, Sanasam Ranbir Singh, Priyankoo Sarmah
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 AssameseBackTranslit: Back Transliteration of Romanized Assamese Social Media Text
abstract
This paper presents a novel back transliteration dataset capturing native language text originally composed in the Roman/Latin script, harvested from popular social media platforms, along with its corresponding representation in the native Assamese script. Assamese, categorized as a low-resource language within the Indo-Aryan language family, predominantly spoken in the north-east Indian state of Assam, faces a scarcity of linguistic resources. The dataset comprises a total of 60,312 Roman-native parallel transliterated sentences. This paper diverges from conventional forward transliteration datasets consisting mainly of named entities and technical terms, instead presenting a novel transliteration dataset cultivated from three prominent social media platforms, Facebook, Twitter(currently X), and YouTube, in the backward transliteration direction. The paper offers a comprehensive examination of ten state-of-the-art word-level transliteration models within the context of this dataset, encompassing transliteration evaluation benchmarks, extensive performance assessments, and a discussion of the unique challenges encountered during the processing of transliterated social media content. Our approach involves the initial use of two statistical transliteration models, followed by the training of two state-of-the-art neural network-based transliteration models, evaluation of three publicly available pre-trained models, and ultimately fine-tuning one existing state-of-the-art multilingual transliteration model along with two pre-trained large language models using the collected datasets. Notably, the Neural Transformer model outperforms all other baseline transliteration models, achieving the lowest Word Error Rate (WER) and Character Error Rate (CER), and the highest BLEU (up to 4 gram) score of 55.05, 19.44, and 69.15, respectively.
Hemanta Baruah, Sanasam Ranbir Singh, Priyankoo Sarmah
LREC/COLING2
2024 Improving linear orthogonal mapping based cross-lingual representation using ridge regression and graph centrality
Deepen Naorem, Sanasam Ranbir Singh, Priyankoo Sarmah
Comput. Speech Lang.2
2024 Chart classification: a survey and benchmarking of different state-of-the-art methods
Jennil Thiyam, Sanasam Ranbir Singh, Prabin Kumar Bora
Int. J. Document Anal. Recognit.2
2024 Sentiment analysis of tweets using text and graph multi-views learning
abstract
Abstract With the surge of deep learning framework, various studies have attempted to address the challenges of sentiment analysis of tweets (data sparsity, under-specificity, noise, and multilingual content) through text and network-based representation learning approaches. However, limited studies on combining the benefits of textual and structural (graph) representations for sentiment analysis of tweets have been carried out. This study proposes a multi-view learning framework ( end-to-end and ensemble-based ) that leverages both text-based and graph-based representation learning approaches to enrich the tweet representation for sentiment classification. The efficacy of the proposed framework is evaluated over three datasets using suitable baseline counterparts. From various experimental studies, it is observed that combining both textual and structural views can achieve better performance of sentiment classification tasks than its counterparts.
Loitongbam Gyanendro Singh, Sanasam Ranbir Singh
Knowl. Inf. Syst.2
2024 Transliteration Characteristics in Romanized Assamese Language Social Media Text and Machine Transliteration
abstract
This article aims to understand different transliteration behaviors of Romanized Assamese text on social media. Assamese, a language that belongs to the Indo-Aryan language family, is also among the 22 scheduled languages in India. With the increasing popularity of social media in India and also the common use of the English Qwerty keyboard, Indian users on social media express themselves in their native languages, but using the Roman/Latin script. Unlike some other popular South Asian languages (say Pinyin for Chinese), Indian languages do not have a common standard romanization convention for writing on social media platforms. Assamese and English are two very different orthographical languages. Thus, considering both orthographic and phonemic characteristics of the language, this study tries to explain how Assamese vowels, vowel diacritics, and consonants are represented in Roman transliterated form. From a dataset of romanized Assamese social media texts collected from three popular social media sites: (Facebook, YouTube, and X (formerly known as Twitter)), 1 we have manually labeled them with their native Assamese script. A comparison analysis is also carried out between the transliterated Assamese social media texts with six different Assamese romanization schemes that reflect how Assamese users on social media do not adhere to any fixed romanization scheme. We have built three separate character-level transliteration models from our dataset. One using a traditional phrase-based statistical machine transliteration model, (1) PBSMT model and two separate neural transliteration models, (2) BiLSTM neural seq2seq model with attention, and (3) Neural transformer model. A thorough error analysis has been performed on the transliteration result obtained from the three state-of-the-art models mentioned above. This may help to build a more robust machine transliteration system for the Assamese social media domain in the future. Finally, an attention analysis experiment is also carried out with the help of attention weight scores taken from the character-level BiLSTM neural seq2seq transliteration model built from our dataset.
Hemanta Baruah, Sanasam Ranbir Singh, Priyankoo Sarmah
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 Gated Recursive and Sequential Deep Hierarchical Encoding for Detecting Incongruent News Articles
abstract
With the increase in misinformation across digital platforms, incongruent news detection is becoming an important research problem. Earlier, researchers have exploited various feature engineering approaches and deep learning models with embedding to capture incongruity between news headlines and their respective bodies. Studies have broadly considered different combinations of bag-of-words-based features, sequential encoding, hierarchical encoding, headline-guided attention-based encoding, and so on, of the text in headlines and bodies. In this article, we focus on addressing two important limitations observed with hierarchical encoding and headline-guided attention-based encoding methods. The existing hierarchical encoding-based studies limit the hierarchical structure of the body of a news article to paragraph level only, undermining the importance of incorporating long-term dependence from word level to sentence, paragraph, and body. Furthermore, the existing headline-guided attention-based encoding focuses on contextually similar contents in the body of the headline, undermining the importance of incorporating contextually dissimilar contents. Motivated by the above observations, this article proposes a gated recursive and sequential deep hierarchical encoding (GraSHE) method for detecting incongruent news articles by extending the hierarchical structure of the news body from the body to the word level and incorporating incongruity weight. From various experimental setups over three publicly available benchmarks datasets, the experimental results indicate that the proposed model outperforms baseline models with bag-of-word-based features, sequential, hierarchical, and headline-guided attention-based encoding methods. To further validate the performance of the proposed model, we conduct several ablation studies. The following key observations can be made from the ablation study: 1) models with hierarchical encoding outperform models with nonhierarchical encoding; 2) recursive encoding of sentences boosts the performance of models as compared with sequential encoding of sentences within paragraphs; and 3) incongruent news article detection is domain-dependent. Incorporating explicit features further boosts the performance of proposed model and also decreases the domain dependence of models.
Sujit Kumar, Sanasam Ranbir Singh
IEEE Trans. Comput. Soc. Syst.3
2024 A Headline-Centric Graph-Based Dual Context Matching Approach for Incongruent News Detection
abstract
The prevalence of incongruent news has demonstrated its significant role in propagating fake news, which catalyzes the dissemination of both misinformation and disinformation. Consequently, detecting incongruent news articles is an important research problem to counter early spreading of misinformation. In the literature, researchers have explored various bag-of-word-based features, news body-centric and news headline-centric encoding methods for incongruent news article detection. However, headline-centric and body-centric approaches in the literature fail to detect partially incongruent articles efficiently. Motivated by the above limitations, this study proposes graph-based dual context matching (GDCM), which first represents headlines and news bodies as abigramnetwork to capture contextual relations between words and document structure. For every word in the headline, GDCM extracts dual contexts (positive and negative) from thebigramnetwork representing news body and estimates similarity between dual contexts and the headline for incongruent news detection. We conduct extensive experiments on three publicly available benchmark datasets and compare its performance with 16 baseline models. Our experimental results suggest that the proposed model outperforms existing state-of-the-art models and efficiently detects partially incongruent news. We further validate the performance of the proposed model through several ablation studies. The following key observations can be made from the ablation studies: 1) extracting dualbigramcontext of words in the headline from different segments of news body and then estimating the similarity between dualbigramcontexts from news body and the headline helps in incongruent news detection and also helps in detecting partial incongruent news efficiently; and 2) representing news headlines and bodies in the form of a network based onbigramcontext helps to capture better nonlinear and contextual relationships between headline and body.
Sujit Kumar, Sanasam Ranbir Singh
IEEE Trans. Comput. Soc. Syst.3
2023 Assamese Back Transliteration - An Empirical Study Over Canonical and Non-canonical Datasets
Hemanta Baruah, Sanasam Ranbir Singh, Priyankoo Sarmah
PACLIC2
2023 Subwords to Word Back Composition for Morphologically Rich Languages in Neural Machine Translation
Telem Joyson Singh, Sanasam Ranbir Singh, Priyankoo Sarmah
PACLIC2
2023 Protocol Aware Unsupervised Network Intrusion Detection System
abstract
In recent years the number of attacks on computer networks has increased exponentially due to the easy availability of sophisticated tools and attack techniques. These attacks are possible due to existing vulnerabilities in networking protocols. Most of the machine learning based intrusion detection systems proposed earlier, to mitigate these attacks, consider training a model for the group of attacks, which doesn’t consider protocol-specific properties into account and is biased toward attacks where most of the data is available. In this paper, we propose protocol aware unsupervised method based on an autoencoder-based learning approach to detect the attack in network flows by training the model using only normal traffic and using reconstruction error as the parameter to classify the attack event. Our proposed method is based on building protocol aware model by combining individual protocol-specific encoders and learning the protocol channel importance using attention mechanism. We perform various experiments on different recent datasets like CICDDoS2019, and CICIDS2018, and experimental results show that the proposed protocol aware model performs better than the non-protocol aware method.
Ritesh Ratti, Sanasam Ranbir Singh, Sukumar Nandi
TrustCom2
2023 Network based Intrusion Detection using Time aware LSTM Autoencoder
abstract
With the advancement of Internet technologies Cyber attacks have become a significant risk to overall security, therefore, intelligent security systems are required to strengthen the network security against these threats. Machine learning has played a pivotal role in the detection and mitigation of these attacks over the years. However, to identify the zero-day attacks and incorporate frequently changing attack scenarios, techniques need to be developed that can work with minimally labeled data. In this paper, we propose Time aware LSTM Autoencoder-based learning approach to detect the attack in network flows by training the model using only normal traffic and using reconstruction error as the parameter to classify the attack event. We perform the experiments on different recent datasets like CICDDoS2019, & CICIDS2018 and experimental results exhibit that the proposed model overall provides better classification metrics.
Ritesh Ratti, Sanasam Ranbir Singh, Sukumar Nandi
TrustCom2
2023 Effect of attention and triplet loss on chart classification: a study on noisy charts and confusing chart pairs
Jennil Thiyam, Sanasam Ranbir Singh, Prabin Kumar Bora
J. Intell. Inf. Syst.2
2023 Integrated document segmentation and region identification: textual, equation and graphical
Jennil Thiyam, Sanasam Ranbir Singh, Prabin Kumar Bora
Multim. Syst.2
2023 Hashtag-Based Tweet Expansion for Improved Topic Modeling
abstract
Topic modeling on tweets is known to experience under-specificity and data sparsity due to its character limitation. In earlier studies, researchers attempted to address this problem by either 1) tweet aggregation, where related tweets are combined into a single document or 2) tweet expansion with related text from external sources. The first approach faces the problem of losing the topic distribution in individual tweets. While finding a relevant text from the external source for a random tweet in the second approach is challenging for various reasons like differences in writing styles, multilingual content, and informal text. In contrast to adding context from external resources or combining related tweets into a pool, this study uses the internal vocabulary (hashtags) to counter under-specificity and sparsity in tweets. Earlier studies have indicated hashtags to be an important feature for representing the underlying context present in the tweet. Sequential models like Bi-directional Long Short Term Memory (BiLSTM) and Convolution Neural Network (CNN) over distributed representation of words have shown promising results in capturing semantic relationships between words of a tweet in the past. Motivated by the above, this article proposes a unified framework of hashtag-based tweet expansion exploiting text-based and network-based representation learning methods such as BiLSTM, BERT, and Graph Convolution Network (GCN). The hashtag-based expanded tweets using the proposed framework have significantly improved topic modeling performance compared to un-expanded (raw) tweets and hashtag-pooling-based approaches over two real-world tweet datasets of different nature. Furthermore, this article also studies the significance of hashtags in topic modeling performance by experimenting with different combination of word types such as hashtags, keywords, and user mentions.
Loitongbam Gyanendro Singh, Sanasam Ranbir Singh
IEEE Trans. Comput. Soc. Syst.3
2022 Revisiting Link Prediction on Heterogeneous Graphs with a Multi-view Perspective
abstract
In this work, we present a novel approach for link prediction on heterogeneous networks – networks that accommodate multiple types of nodes as well as multiple types of relations among them. Specifically, we propose a multi-view network representation learning framework to incorporate structural intuitions from the underlying graph and enrich the relational representations for link prediction. The method relies on the metapath view, the community view, and the subgraph view between a source and target node pair whose linkage is to be predicted. Furthermore, our proposed model leverages a relation-aware attention mechanism to aggregate the candidate contexts in a principled way. Empirically, we demonstrate that the proposed architecture outperforms state-of-the-art transductive and inductive methods in link prediction by a significant margin. A detailed ablation study and attention weight visualizations suggest that the chosen views are complementary and useful to predict links robustly.
Anasua Mitra, Priyesh Vijayan, Sanasam Ranbir Singh, Diganta Goswami, Srinivasan Parthasarathy 0001, Balaraman Ravindran
ICDM3
2022 SwitchNet: Learning to switch for word-level language identification in code-mixed social media text
abstract
Abstract Word-level language identification is an essential prerequisite for extracting useful information from code-mixed social media content. Previous studies in word-level language identification show two important observations. First, the local context is an important indicator of the language of a word when a word is valid in multiple languages. Second, considering the word in isolation from its context leads to more effective language classification when a word is borrowed or embedded into sentences of other languages. In this paper, we propose a framework for language identification that makes use of a dynamic switching mechanism for effective language classification of both words that are borrowed or embedded from other languages as well as words that are valid in multiple languages. For a given input, the proposed switching mechanism makes a dynamic decision to bias its prediction either towards the prediction obtained by the contextual information or that obtained by the word in isolation. In contrast to existing studies that rely upon large amounts of annotated data for robust performance in a multilingual environment, the proposed approach uses minimal annotated resources and no external resources, making it easily extendible to newer languages. Evaluation over a corpus of transliterated Facebook comments shows that the proposed approach outperforms its baseline counterparts: classification based on the contextual information, classification based on the word in isolation, as well as an ensemble of the two classifiers.
Neelakshi Sarma, Sanasam Ranbir Singh, Diganta Goswami
Nat. Lang. Eng.2
2022 Synonymy Expansion Using Link Prediction Methods: A Case Study of Assamese WordNet
abstract
WordNets built for low-resource languages, such as Assamese, often use the expansion methodology. This may result in missing lexical entries and missing synonymy relations. As the Assamese WordNet is also built using the expansion method, using the Hindi WordNet, it also has missing synonymy relations. As WordNets can be visualized as a network of unique words connected by synonymy relations, link prediction in complex network analysis is an effective way of predicting missing relations in a network. Hence, to predict the missing synonyms in the Assamese WordNet, link prediction methods were used in the current work that proved effective. It is also observed that for discovering missing relations in the Assamese WordNet, simple local proximity-based methods might be more effective as compared to global and complex supervised models using network embedding. Further, it is noticed that though a set of retrieved words are not synonyms per se, they are semantically related to the target word and may be categorized as semantic cohorts.
Bornali Phukan, Akash Anil, Sanasam Ranbir Singh, Priyankoo Sarmah
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2021 Challenges in chart image classification: a comparative study of different deep learning methods
abstract
Charts are commonly used forms of visualizing scientific observations from research findings or commercial trends. They provide an abstraction of the underlying information in a more understandable way. Over time, different forms of charts are developed. With the increase in the number of scientific documents present on the internet with different types of charts, automatic chart classification is becoming an important task for various applications. There have been several studies on chart classification with methods ranging from traditional machine learning approaches like SVM, KNN, and HMM to recent deep learning models like VGG, ResNet, and Xception. However, inconsistencies in experimental results are evident. This paper evaluates nine of the recently proposed deep learning-based models on three datasets (one curated and annotated by authors, and two publicly available), and systematically studies their performances over various setups to understand the reason for observing inconsistent results.
Jennil Thiyam, Sanasam Ranbir Singh, Prabin Kumar Bora
DocEng2
2021 Semi-Supervised Deep Learning for Multiplex Networks
abstract
Multiplex networks are complex graph structures in which a set of entities are connected to each other via multiple types of relations, each relation representing a distinct layer. Such graphs are used to investigate many complex biological, social, and technological systems. In this work, we present a novel semi-supervised approach for structure-aware representation learning on multiplex networks. Our approach relies on maximizing the mutual information between local node-wise patch representations and label correlated structure-aware global graph representations to model the nodes and cluster structures jointly. Specifically, it leverages a novel cluster-aware, node-contextualized global graph summary generation strategy for effective joint-modeling of node and cluster representations across the layers of a multiplex network. Empirically, we demonstrate that the proposed architecture outperforms state-of-the-art methods in a range of tasks: classification, clustering, visualization, and similarity search on seven real-world multiplex networks for various experiment settings.
Anasua Mitra, Priyesh Vijayan, Sanasam Ranbir Singh, Diganta Goswami, Srinivasan Parthasarathy 0001, Balaraman Ravindran
KDD3
2021 Empirical study of sentiment analysis tools and techniques on societal topics
Loitongbam Gyanendro Singh, Sanasam Ranbir Singh
J. Intell. Inf. Syst.2
2020 Sentiment Analysis of Tweets using Heterogeneous Multi-layer Network Representation and Embedding
abstract
Sentiment classification on tweets often needs to deal with the problems of under-specificity, noise, and multilingual content.This study proposes a heterogeneous multi-layer networkbased representation of tweets to generate multiple representations of a tweet and address the above issues.The generated representations are further ensembled and classified using a neural-based early fusion approach.Further, we propose a centrality aware random-walk for node embedding and tweet representations suitable for the multi-layer network.From various experimental analysis, it is evident that the proposed method can address the problem of under-specificity, noisy text, and multilingual content present in a tweet and provides better classification performance than the textbased counterparts.Further, the proposed centrality aware based random walk provides better representations than unbiased and other biased counterparts.
Loitongbam Gyanendro Singh, Anasua Mitra, Sanasam Ranbir Singh
EMNLP (1)3
2020 Predicting emotion dynamics sequence on Twitter via deep learning approach
abstract
Exploring the mechanism about users' emotion dynamics towards social events and further predicting their future emotions have attracted great attention to the researchers. Despite the concreteness of the online expressions in written form, it remains unpredictable which kinds of emotions will be expressed in individual messages of Twitter users influenced by his/her friends. To investigate this, we perform an investigation on observing emotions unfolding in a consecutive sequence of tweets for a particular user based on his/her past history. In this paper, we propose an Emotion-based User Sequential Influence Model (E-USIM) on given a set of tweets related with some events (identified by the usage of a hashtag), determines how those sentiments will be distributed on behalf of a person within a conversation. We then apply the developed model to predict users' future emotions by combing of personal and interpersonal influence.
Debashis Naskar, Eva Onaindia, Miguel Rebollo, Sanasam Ranbir Singh
MoMM4
2020 SHE: Sentiment Hashtag Embedding Through Multitask Learning
abstract
Recent studies have shown the importance of utilizing hashtags for sentiment analysis task on social media data. However, as the hashtag generation process is less restrictive, it throws several challenges, such as hashtag normalization, topic modeling, and semantic similarity. Recently, researchers have tried to address the above-mentioned challenges through representation learning. However, most of the studies on hashtag embedding try to capture the semantic distribution of hashtags and often fail to capture the sentiment polarity. Furthermore, generating a task-specific hashtag embedding can distort its semantic representation, which is undesirable for sentiment representation of hashtag. Therefore, this article proposes a semisupervised sentiment hashtag embedding (SHE) model, which is capable of preserving both semantic as well as sentiment distribution of the hashtags. In particular, SHE leverages a multitask learning approach using an autoencoder and a convolutional neural network-based classifier. To assess the efficacy of hashtag embedding, we compare the performance of SHE against suitable baselines for three different tasks, namely, hashtag sentiment classification, tweet sentiment classification, and retrieval of semantically similar hashtags. It is evident from various experimental results that SHE outperforms the majority of the baselines with significant margins.
Loitongbam Gyanendro Singh, Akash Anil, Sanasam Ranbir Singh
IEEE Trans. Comput. Soc. Syst.3
2020 Emotion Dynamics of Public Opinions on Twitter
abstract
Recently, social media has been considered the fastest medium for information broadcasting and sharing. Considering the wide range of applications such as viral marketing, political campaigns, social advertisement, and so on, influencing characteristics of users or tweets have attracted several researchers. It is observed from various studies that influential messages or users create a high impact on a social ecosystem. In this study, we assume that public opinion on a social issue on Twitter carries a certain degree of emotion, and there is an emotion flow underneath the Twitter network. In this article, we investigate social dynamics of emotion present in users’ opinions and attempt to understand (i) changing characteristics of users’ emotions toward a social issue over time, (ii) influence of public emotions on individuals’ emotions, (iii) cause of changing opinion by social factors, and so on. We study users’ emotion dynamics over a collection of 17.65M tweets with 69.36K users and observe 63% of the users are likely to change their emotional state against the topic into their subsequent tweets. Tweets were coming from the member community shows higher influencing capability than the other community sources. It is also observed that retweets influence users more than hashtags, mentions, and replies.
Debashis Naskar, Sanasam Ranbir Singh, Sukumar Nandi, Eva Onaindia
ACM Trans. Inf. Syst.2
2019 Influence of social conversational features on language identification in highly multilingual online conversations
Neelakshi Sarma, Sanasam Ranbir Singh, Diganta Goswami
Inf. Process. Manag.2
2018 Exploiting reciprocity toward link prediction
Niladri Sett, Devesh, Sanasam Ranbir Singh, Sukumar Nandi
Knowl. Inf. Syst.3
2018 Mining heterogeneous terrorist attack network using personalized PageRank
abstract
Majority of the studies on counter-terrorism using social network analysis consider homogeneous networks built over either terrorists or terrorist organizations. However, terrorist attacks are often defined using heterogeneous attributes such as location, time, target type, organization, etc. This paper constructs a heterogeneous terrorist attack network considering heterogeneous attributes and captures heterogeneous influences propagated from different attributes using Personalized PageRank. Personalized PageRank is a flexible model capable of propagating supervised information while traversing over a network using different personalized parameters. This study investigates effects of various parametric setups to study influence of various factors such as news media discussions, historical activities, temporal behavior etc., on one of the important counter-terrorism problems; prediction of future attacks of a terrorist organization. From various experimental observations, it is evident that news media discussion and network’s temporal behavior have a positive influence on future activities of a terrorist organization. Further, this paper investigates responses of various node proximity based link prediction methods on predicting future relationships between a terrorist organization with other attributes (such as country, city and target types). Majority of the studies on link prediction using node proximity ignore node importance. However, in a heterogeneous environment, nodes from different classes may have different importance. This paper proposes new variants of four proximity based link prediction methods, namely, Adamic Adar, Jaccard Coefficient, Resource Allocation, and Common Neighbor, which have the capability to incorporate node importance. With suitable experiments, we show that the proposed variants of link predictors are more accurate at predicting relations than their state of art counterparts.
Akash Anil, Sanasam Ranbir Singh, Ranjan Sarmah
Web Intell.2
2018 Temporal link prediction in multi-relational network
Niladri Sett, Saptarshi Basu, Sukumar Nandi, Sanasam Ranbir Singh
World Wide Web4
2016 Automatic Syllabification for Manipuri language
abstract
Development of hand crafted rule for syllabifying words of a language is an expensive task. This paper proposes several data-driven methods for automatic syllabification of words written in Manipuri language. Manipuri is one of the scheduled Indian languages. First, we propose a language-independent rule-based approach formulated using entropy based phonotactic segmentation. Second, we project the syllabification problem as a sequence labeling problem and investigate its effect using various sequence labeling approaches. Third, we combine the effect of sequence labeling and rule-based method and investigate the performance of the hybrid approach. From various experimental observations, it is evident that the proposed methods outperform the baseline rule-based method. The entropy based phonotactic segmentation provides a word accuracy of 96%, CRF (sequence labeling approach) provides 97% and hybrid approach provides 98% word accuracy.
Loitongbam Gyanendro Singh, Lenin Laitonjam, Sanasam Ranbir Singh
COLING3
2016 Personalised PageRank as a Method of Exploiting Heterogeneous Network for Counter Terrorism and Homeland Security
abstract
Majority of the social network analysis studies for counter-terrorism and homeland security consider homogeneous network. However, a terrorist activity (attack) is often defined by several attributes such as terrorist organisation, time, place, attack type etc. To capture inherent dependency between the attributes, we need to adopt a network which is capable of capturing the dependency between the attributes. In this paper, we define a heterogeneous network to represent a collection of terrorist activities. Further, we propose personalised PageRank (PPR) as a method capable of performing various analytical operations over heterogeneous network just by changing model parameters without changing the underlying model. Using global terrorist data (GTD), behavioural network, and news discussion network, we show various applications of PPR for counter-terrorism over heterogeneous network just by changing the model parameter. In addition we propose heterogeneous version of four local proximity based link prediction methods, namely, Common Neighbour, Adamic-Adar, Jaccard Coefficient, and Resource Allocation.
Akash Anil, Sanasam Ranbir Singh, Ranjan Sarmah
WI2
2016 A Time Aware Method for Predicting Dull Nodes and Links in Evolving Networks for Data Cleaning
abstract
Existing studies on evolution of social network largely focus on addition of new nodes and links in the network. However, as network evolves, existing relationships degrade and break down, and some nodes go to hibernation or decide not to participate in any kind of activities in the network where it belongs. Such nodes and links, which we refer as "dull", may affect analysis and prediction tasks in networks. This paper formally defines the problem of predicting dull nodes and links at an early stage, and proposes a novel time aware method to solve it. Pruning of such nodes and links is framed as "network data cleaning" task. As the definitions of dull node and link are non-trivial and subjective, a novel scheme to label such nodes and links is also proposed here. Experimental results on two real network datasets demonstrate that the proposed method accurately predicts potential dull nodes and links. This paper further experimentally validates the need for data cleaning by investigating its effect on the well-known "link prediction" problem.
Niladri Sett, Subhrendu Chattopadhyay, Sanasam Ranbir Singh, Sukumar Nandi
WI3
2016 Influence of edge weight on node proximity based link prediction methods: An empirical analysis
Niladri Sett, Sanasam Ranbir Singh, Sukumar Nandi
Neurocomputing2
2014 Modeling evolution of a social network using temporalgraph kernels
abstract
Majority of the studies on modeling the evolution of a social network using spectral graph kernels do not consider temporal effects while estimating the kernel parameters. As a result, such kernels fail to capture structural properties of the evolution over the time. In this paper, we propose temporal spectral graph kernels of four popular graph kernels namely path counting, triangle closing, exponential and neumann. Their responses in predicting future growth of the network have been investigated in detail, using two large datasets namely Facebook and DBLP. It is evident from various experimental setups that the proposed temporal spectral graph kernels outperform all of their non-temporal counterparts in predicting future growth of the networks.
Akash Anil, Niladri Sett, Sanasam Ranbir Singh
SIGIR3
2010 Inference Based Query Expansion Using User's Real Time Implicit Feedback
Sanasam Ranbir Singh, Hema A. Murthy, Timothy A. Gonsalves
IC3K1
2008 Determining user's interest in real time
abstract
Most of the search engine optimization techniques attempt to predict users interest by learning from the past information collected from different sources. But, a user's current interest often depends on many factors which are not captured in the past information. In this paper, we attempt to identify user's current interest in real time from the information provided by the user in the current query session. By identifying user's interest in real time, the engine could adapt differently to different users in real time. Experimental verification indicates that our approach is encouraging for short queries
Sanasam Ranbir Singh, Hema A. Murthy, Timothy A. Gonsalves
WWW1
2007 Estimating the Rate of Web Page Updates
Sanasam Ranbir Singh
IJCAI1
2006 An algorithm for discovering the frequent closed itemsets in a large database
abstract
Previous research revealed that the problem of discovering a complete set of frequent itemsets from a large database can be reduced to the problem of discovering the frequent closed itemsets, and this process results in a much smaller set of itemsets without information loss. This article is based on the observation that the set of all itemsets can be grouped into non-overlapping clusters such that each cluster is identified by a unique closed tidset. It is also found that there is only one closed itemset in each cluster and it is the superset of all itemsets with the same support. Therefore, the problem of discovering closed itemsets can be further considered as the problem of clustering the set of itemsets and then identifying each cluster by a unique closed tidset. This article presents CloseMiner, a new algorithm for discovering all frequent closed itemsets by grouping the set of itemsets into non-overlapping clusters. Experimental evaluation based on a number of real and synthetic databases has proved that CloseMiner outperforms the existing systems APRIORI and CHARM.
Ningthoujam Gourakishwar Singh, Sanasam Ranbir Singh, Anjana Kakoti Mahanta, Bhanu Prasad 0001
J. Exp. Theor. Artif. Intell.2
2005 CloseMiner: Discovering Frequent Closed Itemsets Using Frequent Closed Tidsets
abstract
Complete set of itemsets can be grouped into non-overlapping clusters identified by closed tidsets. Each cluster has only one closed itemset and is the superset of all itemsets with the same support. Number of closed itemsets is identical to the number of clusters. Therefore, the problem of discovering closed itemsets can be considered as the problem of clustering the complete set of itemsets by closed tidsets. In this paper, we present CloseMiner, a new algorithm for discovering all frequent closed itemsets by grouping the complete set of itemsets into non-overlapping clusters identified by closed tidsets. An extensive experimental evaluation on a number of real and synthetic databases shows that CloseMiner outperforms Apriori and CHARM.
Ningthoujam Gourakishwar Singh, Sanasam Ranbir Singh, Anjana Kakoti Mahanta
ICDM2