Sukomal Pal

dblp:06/3884 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0001-8743-9830ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 10 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 MUSIA: Multilingual Story Illustration Corpus for Cross-Cultural Alignment and Generation
Krishna Tewari, Supriya Chanda, Nirmit Patil, Sukomal Pal
LREC4
2025 HiDePCC: A Novel Dual-Pronged Untargeted Attack on Federated Recommendation via Gradient Perturbation and Cluster Crafting
Yamini Jha, Krishna Tewari, Sukomal Pal
RecSys3
2025 Advancing Hindi Text Summarization: Named Entity Recognition and Content Augmentation Strategies
abstract
We explore advancements in Hindi text summarization, a critical area in natural language processing that aids in managing information overload. Despite a growing corpus of Hindi data, there’s a significant gap in practical summarization tools due to intricate linguistic features and limited resources compared to English. Previous works focused on extractive methods, but recent shifts towards abstractive approaches promise more natural and coherent summaries by understanding and paraphrasing content. Our research introduces novel methodologies, Named Entity Aware-Abstractive Text Summarization (NEA-ATS) and Query-Driven Content Augmentation for Summarization (QDCAS), aimed at enhancing the accuracy and richness of Hindi summaries. NEA-ATS integrates Named Entity Recognition to prioritize crucial information, improving language model attention to critical details but occasionally disrupting context. While NEA-ATS shows some improvements, it occasionally disrupts the text’s context, leading to only marginal gains in summary quality. Meanwhile, QDCAS addresses extrinsic hallucinations—common in state-of-the-art models—by augmenting source documents with relevant content through focused web crawling—a technique to selectively gather topic-specific web pages—broadening contextual understanding and refining outputs. Empirical results demonstrate the effectiveness of QDCAS, showing marginal improvements in ROUGE and BERTScores over traditional language models. This work advances Hindi text summarization and explores content-rich strategies, potentially expanding to other languages and domains.
Saumay Gupta, Sukomal Pal
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 Sentiment analysis on Hindi tweets during COVID-19 pandemic
abstract
Abstract A gap among the people has been created due to a lack of social interactions. The physical void has led to an increase in online interaction among users on social media platforms. Sentiment analysis of such interactions can help us analyze the general public psychology during the pandemic. However, the lack of data in non‐English and low‐resource languages like ‘Hindi’ makes it difficult to study it among native and non‐English speaking masses. Here, we create a small collection of ‘Hindi’ tweets on COVID‐19 during the pandemic containing 10,011 tweets for sentiment analysis, which is named as sentiment analysis for Hindi (SAFH). In this article, we describe the process of collecting, creating, annotating the corpus, and sentiment classification. The claims have been verified using different word embedding with a deep learning classifier through the proposed model. The achieved accuracy of the proposed model yields up to a permissible rate of 90.9%.
Anita Saroj, Akash Thakur, Sukomal Pal
Comput. Intell.3
2024 Water chicken swarm optimization-based deep segmental neural network for spoken term detection using bayesian filtering
Sushil Venkatesh Kulkarni, Sukomal Pal
Multim. Tools Appl.2
2024 Ensemble-based domain adaptation on social media posts for irony detection
Anita Saroj, Sukomal Pal
Multim. Tools Appl.2
2024 Diversified recommendation using implicit content node embedding in heterogeneous information network
Naina Yadav, Sukomal Pal, Anil Kumar Singh 0001
Multim. Tools Appl.2
2023 Building a text retrieval system for the Sanskrit language: Exploring indexing, stemming, and searching issues
Siba Sankar Sahu, Sukomal Pal
Comput. Speech Lang.2
2023 A Study on Corpus-based Stopword Lists in Indian Language IR
abstract
We explore and evaluate the effect of different stopword lists (non-corpus-based and corpus-based) in the information retrieval (IR) tasks with different Indian languages such as Bengali, Marathi, Gujarati, Hindi, and English. The issue was investigated from three viewpoints. Is there any performance difference between non-corpus-based and corpus-based stopword removal in chosen Indian languages? Can corpus-based stopword lists improve performance in Indian languages IR? If yes, to what extent? Among the different corpus-based stopword lists, which stopword list provides the best IR performance? Does the length of a corpus-based stopword list affect the retrieval performance in Indian languages? If yes, to what extent? It was observed that a corpus-based stopword list provides better retrieval performance than a non-corpus-based stopword list in different Indian languages. Among the different corpus-based stopword lists generated and experimented with, Zipf’s law-based stopword list (idf-based one) provides the best retrieval performance in various Indian languages. The aggregation1-based stopword list provides better retrieval than the aggregation2-based list in Indian languages, but in English, the aggregation2-based stopword list performs better than the aggregation1-based list. The best performing idf-based stopword list improves MAP score by 5.43% in Bengali, 1.91% in Marathi, 5.4% in Gujarati, 1.5% in Hindi, and 2.12% in English, respectively, over their baseline counterparts. The probabilistic retrieval models (BM25 and TF-IDF) perform best in different Indian languages. A smaller length of corpus-based stopword lists performs better than a larger length of non-corpus-based stopword lists for all the Indian languages considered. The proposed schemes demonstrate that a stopword list can be heuristically generated in a language-independent statistical method and effectively used for IR tasks with performance comparable, to or even better than non-corpus-based stopword lists.
Siba Sankar Sahu, Sukomal Pal
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2022 Improved self-attentive Musical Instrument Digital Interface content-based music recommendation system
abstract
Abstract Automatic music recommendation is an open research problem that has seen much work in recent years. A common and successful music recommendation approach is collaborative filtering, which has worked well in this domain. One major drawback of this method is that it suffers from a cold‐start problem, and it requires a lot of user‐personalized information. It is an ineffective mechanism for recommending new and unpopular songs as well as for new users. In this article, we report a hybrid methodology that uses the song's content information. We use MIDI (Musical Instrument Digital Interface) content data, a compressed version of an audio song that contains digital information about a song and is machine‐readable. We describe a model called MSA‐SRec (MIDI Based Self Attentive Sequential Music Recommendation), a latent factor‐based self‐attentive deep learning model that uses a substantial amount of sequential information as content information of the song for recommendation generation. We use MIDI data of a song that is under‐explored content information for music recommendation. We show that using MIDI as content data with user and item latent vector produces reasonable recommendations. We also demonstrate that using MIDI over other music metadata performs better with various state‐of‐the‐art models of recommendation systems.
Naina Yadav, Anil Kumar Singh 0001, Sukomal Pal
Comput. Intell.3
2022 A fusion of variants of sentence scoring methods and collaborative word rankings for document summarization
abstract
Abstract Document summarization is an important task in natural language processing that helps deal with the problem of information overload occurring due to the existence of redundant content. Summary generation with highly relevant contents and maximum coverage is particularly challenging which can only be achieved when redundancy is minimized. This article introduces a novel approach for automatic text summarization based on sentence scoring and collaborative ranking to produce summaries with minimal redundancy and improved overall performance of summarization. The proposed model is a fusion of weighted and unweighted features‐based sentence scoring methods. To learn optimal weights of text features, it has been modelled as an optimization problem. Moreover, the proposed model exploits the strength of collaborative ranking to generate the summary of a given document. Three similarity factors (proximity, significance and singularity)‐based models have been employed to find the similarity between weighted and unweighted sentence scores. The results of the comparison experiment demonstrate that the proposed (PS + Jac) method generates a closer summary to the reference summary with minimal redundant contents. On average, the proposed (PS + Jac) method generates the summaries with 61% accurate contents with greater improved rates up to 40%. The statistical testing also confirms that the performance improvement is significant at a 5% level of significance.
Pradeepika Verma, Anshul Verma, Sukomal Pal
Expert Syst. J. Knowl. Eng.3
2021 A proactive decision support system for reviewer recommendation in academia
Tribikram Pradhan, Suchit Sahoo, Sukomal Pal
Expert Syst. Appl.4
2021 A deep neural architecture based meta-review generation and final decision prediction of a scholarly article
Tribikram Pradhan, Chaitanya Bhatia, Sukomal Pal
Neurocomputing4
2021 CLAVER: An integrated framework of convolutional layer, bidirectional LSTM with attention mechanism based scholarly venue recommendation
Tribikram Pradhan, Sukomal Pal
Inf. Sci.3
2021 Query Expansion for Transliterated Text Retrieval
abstract
With Web 2.0, there has been exponential growth in the number of Web users and the volume of Web content. Most of these users are not only consumers of the information but also generators of it. People express themselves here in colloquial languages, but using Roman script (transliteration). These texts are mostly informal and casual, and therefore seldom follow grammar rules. Also, there does not exist any prescribed set of spelling rules in transliterated text. This freedom leads to large-scale spelling variations, which is a major challenge in mixed script information processing. This article studies different existing phonetic algorithms to handle the issue of spelling variation, points out the limitations of them, and proposes a novel phonetic encoding approach with two different flavors in the light of Hindi transliteration. Experiments performed over Hindi song lyrics retrieval in mixed script domain with three different retrieval models show that proposed approaches outperform the existing techniques in a majority of the cases (sometimes statistically significantly) for a number of metrics like nDCG@1, nDCG@5, nDCG@10, MAP, MRR, and Recall.
Dinesh Kumar Prabhakar, Sukomal Pal, Chiranjeev Kumar
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2020 MetaGen: An academic Meta-review Generation system
abstract
Peer reviews form an essential part of scientific communications. Research papers and proposals are reviewed by several peers before they are finally accepted or rejected. The procedure followed requires experts to review the research work. Then the area/program chair/ editor writes a meta-review summarizing the review comments and taking a call based on the reviewers' decisions. In this paper, we present MetaGen, a novel meta-review generation system which takes the peer reviews as input and produces an assistive meta-review. This meta-review generation can help the area/program chair writing a meta-review and taking the final decision on the paper/proposal. Thus it can also help to speed up the review process for conference/journals where a large number of submissions need to be handled within a stipulated time. Our approach first generates an extractive draft and then uses fine-tuned UniLM (Unified Langauge Model) for predicting the acceptance decision and making the final meta-review in an abstractive manner. To the best of our knowledge, this is the first work in the direction of meta-review generation. Evaluation based on ROUGE score shows promising results and comparison with few state-of-the-art summarizers demonstrates the effectiveness of the system.
Chaitanya Bhatia, Tribikram Pradhan, Sukomal Pal
SIGIR3
2020 A hybrid personalized scholarly venue recommender system integrating social network analysis and contextual similarity
abstract
Rapidly developing academic venues throw a challenge to researchers in identifying the most appropriate ones that are in-line with their scholarly interests and of high relevance. Even a high-quality paper is sometimes rejected due to a mismatch between the area of the paper, and the scope of the journal attempted to. Recommending appropriate academic venues can, therefore, enable researchers to identify and take part in relevant conferences and to publish in impactful journals. Although a researcher may know a few leading high-profile venues for her specific field of interest, a venue recommender system becomes particularly helpful when one explores a new field or when more options are needed. We propose DISCOVER: A Diversified yet Integrated Social network analysis and COntextual similarity-based scholarly VEnue Recommender system. Our work provides an integrated framework incorporating social network analysis, including centrality measure calculation, citation and co-citation analysis, topic modeling based contextual similarity, and key-route identification based main path analysis of a bibliographic citation network. The paper also addresses cold start issues for a new researcher and a new venue along with a considerable reduction in data sparsity, computational costs, diversity, and stability problems. Experiments based on the Microsoft Academic Graph (MAG) dataset show that the proposed DISCOVER outperforms state-of-the-art recommendation techniques using standard metrics of [email protected], [email protected], accuracy, MRR, F−measuremacro, diversity, stability, and average venue quality.
Tribikram Pradhan, Sukomal Pal
Future Gener. Comput. Syst.2
2020 HASVRec: A modularized Hierarchical Attention-based Scholarly Venue Recommender system
Tribikram Pradhan, Sukomal Pal
Knowl. Based Syst.3
2020 CNAVER: A Content and Network-based Academic VEnue Recommender system
Tribikram Pradhan, Sukomal Pal
Knowl. Based Syst.2
2020 A multi-level fusion based decision support system for academic collaborator recommendation
Tribikram Pradhan, Sukomal Pal
Knowl. Based Syst.2
2019 Diversity in Recommendation System: A Cluster Based Approach
Naina Yadav, Rajesh Kumar Mundotiya, Anil Kumar Singh 0001, Sukomal Pal
HIS4
2019 A Comparative Analysis on Hindi and English Extractive Text Summarization
abstract
Text summarization is the process of transfiguring a large documental information into a clear and concise form. In this article, we present a detailed comparative study of various extractive methods for automatic text summarization on Hindi and English text datasets of news articles. We consider 13 different summarization techniques, namely, TextRank, LexRank, Luhn, LSA, Edmundson, ChunkRank, TGraph, UniRank, NN-ED, NN-SE, FE-SE, SummaRuNNer, and MMR-SE, and we evaluate their performance using various performance metrics, such as precision, recall,F1, cohesion, non-redundancy, readability, and significance. A thorough analysis is done in eight different parts that exhibits the strengths and limitations of these methods, effect of performance over the summary length, impact of language of a document, and other factors as well. A standard summary evaluation tool (ROUGE) and extensive programmatic evaluation using Python 3.5 in Anaconda environment are used to evaluate their outcome.
Pradeepika Verma, Sukomal Pal, Hari Om
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2016 Recent developments in social spam detection and combating techniques: A survey
Manajit Chakraborty, Sukomal Pal, Rahul Pramanik, C. Ravindranath Chowdary
Inf. Process. Manag.2
2011 Evaluation effort, reliability and reusability in XML retrieval
abstract
The Initiative for the Evaluation of XML retrieval (INEX) provides a TREC-like platform for evaluating content-oriented XML retrieval systems. Since 2007, INEX has been using a set of precision-recall based metrics for its ad hoc tasks. The authors investigate the reliability and robustness of these focused retrieval measures, and of the INEX pooling method. They explore four specific questions: How reliable are the metrics when assessments are incomplete, or when query sets are small? What is the minimum pool/query-set size that can be used to reliably evaluate systems? Can the INEX collections be used to fairly evaluate “new” systems that did not participate in the pooling process? And, for a fixed amount of assessment effort, would this effort be better spent in thoroughly judging a few queries, or in judging many queries relatively superficially? The authors' findings validate properties of precision-recall-based metrics observed in document retrieval settings. Early precision measures are found to be more error-prone and less stable under incomplete judgments and small topic-set sizes. They also find that system rankings remain largely unaffected even when assessment effort is substantially (but systematically) reduced, and confirm that the INEX collections remain usable when evaluating nonparticipating systems. Finally, they observe that for a fixed amount of effort, judging shallow pools for many queries is better than judging deep pools for a smaller set of queries. However, when judging only a random sample of a pool, it is better to completely judge fewer topics than to partially judge many topics. This result confirms the effectiveness of pooling methods.
Sukomal Pal, Mandar Mitra, Jaap Kamps
J. Assoc. Inf. Sci. Technol.1
2010 The FIRE 2008 Evaluation Exercise
abstract
The aim of the Forum for Information Retrieval Evaluation (FIRE) is to create an evaluation framework in the spirit of TREC (Text REtrieval Conference), CLEF (Cross-Language Evaluation Forum), and NTCIR (NII Test Collection for IR Systems), for Indian language Information Retrieval. The first evaluation exercise conducted by FIRE was completed in 2008. This article describes the test collections used at FIRE 2008, summarizes the approaches adopted by various participants, discusses the limitations of the datasets, and outlines the tasks planned for the next iteration of FIRE.
Prasenjit Majumder, Mandar Mitra, Dipasree Pal, Ayan Bandyopadhyay, Samaresh Maiti, Sukomal Pal, Deboshree Modak, Sucharita Sanyal
ACM Trans. Asian Lang. Inf. Process.6
2008 Text collections for FIRE
abstract
The aim of the Forum for Information Retrieval Evaluation (FIRE) is to create a Cranfield-like evaluation framework in the spirit of TREC, CLEF and NTCIR, for Indian Language Information Retrieval. For the first year, six Indian languages have been selected: Bengali, Hindi, Marathi, Punjabi, Tamil, and Telugu. This poster describes the tasks as well as the document and topic collections that are to be used at the FIRE workshop.
Prasenjit Majumder, Mandar Mitra, Dipasree Pal, Ayan Bandyopadhyay, Samaresh Maiti, Sukanya Mitra, Aparajita Sen, Sukomal Pal
SIGIR8