Sivaji Bandyopadhyay

dblp:04/2897 · DBLP profile ↗
← Back
72ranked-venue papers
3as first author
21since 2021 · last 2024
0000-0003-2607-1774ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 11 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 An empirical study of a novel multimodal dataset for low-resource machine translation
Loitongbam Sanayai Meetei, Thoudam Doren Singh, Sivaji Bandyopadhyay
Knowl. Inf. Syst.3
2024 Text summary evaluation based on interpretable semantic textual similarity
Goutam Majumder, Vikrant Rajput, Partha Pakray, Sivaji Bandyopadhyay, Benoît Favre
Multim. Tools Appl.4
2024 Exploiting multiple correlated modalities can enhance low-resource machine translation quality
Loitongbam Sanayai Meetei, Thoudam Doren Singh, Sivaji Bandyopadhyay
Multim. Tools Appl.3
2024 Consensus-Based Machine Translation for Code-Mixed Texts
abstract
Multilingualism in India is widespread due to its long history of foreign acquaintances. This leads to the presence of an audience familiar with conversing using more than one language. Additionally, due to the social media boom, the usage of multiple languages to communicate has become extensive. Hence, the need for a translation system that can serve the novice and monolingual user is the need of the hour. Such translation systems can be developed by methods such as statistical machine translation and neural machine translation, where each approach has its advantages as well as disadvantages. In addition, the parallel corpus needed to build a translation system, with code-mixed data, is not readily available. In the present work, we present two translation frameworks that can leverage the individual advantages of these pre-existing approaches by building an ensemble model that takes a consensus of the final outputs of the preceding approaches and generates the target output. The developed models were used for translating English-Bengali code-mixed data (written in Roman script) into their equivalent monolingual Bengali instances. A code-mixed to monolingual parallel corpus was also developed to train the preceding systems. Empirical results show improved BLEU and TER scores of 17.23 and 53.18 and 19.12 and 51.29, respectively, for the developed frameworks.
Sainik Kumar Mahata, Dipankar Das 0001, Sivaji Bandyopadhyay
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2023 English-Assamese neural machine translation using prior alignment and pre-trained language model
Sahinur Rahman Laskar, Bishwaraj Paul, Pankaj Dadure, Riyanka Manna, Partha Pakray, Sivaji Bandyopadhyay
Comput. Speech Lang.6
2023 An extract-then-abstract based method to generate disaster-news headlines using a DNN extractor followed by a transformer abstractor
Sumanta Banerjee, Shyamapada Mukherjee, Sivaji Bandyopadhyay, Partha Pakray
Inf. Process. Manag.3
2023 Correction to: Attention based video captioning framework for Hindi
Alok Singh 0007, Thoudam Doren Singh, Sivaji Bandyopadhyay
Multim. Syst.3
2023 Detection of Diabetic Retinopathy using Convolutional Neural Networks for Feature Extraction and Classification (DRFEC)
Dolly Das, Saroj K. Biswas 0001, Sivaji Bandyopadhyay
Multim. Tools Appl.3
2022 Rule extraction from decision tree: Transparent expert system of rules
abstract
Abstract A system which is transparent and has less decision rules is an efficient, user‐convincing system and moreover convenient and manageable to fields like banking, business, and medical. Decision Tree (DT) is a data mining technique which is transparent and produces a set of production rules for decision‐making. However sometimes it creates some unnecessary and redundant rules which diminish its comprehensibility. Thus a system named Transparent Expert System of Rules (TESR) is proposed in this paper to efficiently improve comprehensibility of the DT by reducing the number of rules drastically without compromising accuracy. The proposed system adopts a Sequential Hill Climbing method with a flexible heuristic function to prune the insignificant rules from decision rules generated by DT. Finally, the proposed TESR system produces a transparent and comprehensible rule set for a decision. The proposed TESR performance is evaluated using 10 datasets and is compared with simple DT (ID3, C4.5, and Classification and Regression Trees) and also two of the existing transparent systems with respect to comprehensibility, accuracy, precision, recall, and F‐measures.
Arpita Nath Boruah, Saroj K. Biswas 0001, Sivaji Bandyopadhyay
Concurr. Comput. Pract. Exp.3
2022 A formula embedding approach for semantic similarity and relatedness between formulas
abstract
Summary The mathematical formula is one of the most vital components in a scientific document, which can explicitly describe various complex concepts and ideas. In addition to numerical calculations, they are also used to clarify definitions and disambiguate explanations transcribed in natural language. Nevertheless, the formulas have a noteworthy impact in the scientific documents, the existing information retrieval systems have limited access to scientific documents based on formulas‐based queries. To accomplish this, in this research, we have studied and implemented the formula embedding approach, which encodes the formula into the embedded vector. For encoding the formula, we have used pretrained sentence bidirectional encoder representations from transformers model. The proposed embedding model takes the latex formula as input and generates an upshot as a fixed dimensional embedding representation. In addition to this, the Siamese network is used to reform the semantic meaning of the formulas. Furthermore, the embedding of the formulas and the queried formula are compared, and cosine similarity is estimated. The performance of the suggested methodology is verified using a math stack exchange corpus of ARQMath 2020, and obtained results have shown a remarkable contribution in the task of formula retrieval.
Pankaj Dadure, Partha Pakray, Sivaji Bandyopadhyay
Concurr. Comput. Pract. Exp.3
2022 A multi-task learning based approach for efficient breast cancer detection and classification
abstract
Abstract Automatic segmentation and classification of breast tumours in ultrasound images using deep learning approaches can help early detect breast cancer. Such predictive modelling can potentially significantly improve the survival chances of the involved patients. Most of the typical deep convolutional neural network (CNN) based approaches consider segmentation and classification tasks separately. But this loses important supervisory information to help achieve better model training. This work proposes the integrated learning of both of these tasks in an end‐to‐end manner, using a multi‐task learning based approach. More specifically, a convolutional encoder‐decoder based architecture is coupled with a residual CNN for performing segmentation and classification together. The level‐wise feature maps from both the encoder and decoder parts of the segmentation network are utilized for classification in the proposed approach. From experimental analysis on a publicly available breast ultrasound image (BUSI) dataset, it has been observed that the proposed approach can achieve impressive performances, both with respect to tumour segmentation and classification. A mean test set AUC of 0.97 and a mean dice score of 0.74 is achieved, establishing a new state‐of‐the‐art performance on the BUSI dataset. From the impressive experimental observations, it can be concluded that learning to perform both segmentation and classification simultaneously can have a very high positive impact on the overall quality of the predictive model. Such observations suggest that the proposed approach can be beneficial in providing real‐time decision support to the involved diagnostic radiologists, which can help improve the survival chances of the corresponding patients.
Arnab Kumar Mishra, Pinki Roy, Sivaji Bandyopadhyay, Sujit Kumar Das
Expert Syst. J. Knowl. Eng.3
2022 Attention based video captioning framework for Hindi
Alok Singh 0007, Thoudam Doren Singh, Sivaji Bandyopadhyay
Multim. Syst.3
2022 Perspective of AI system for COVID-19 detection using chest images: a review
Dolly Das, Saroj K. Biswas 0001, Sivaji Bandyopadhyay
Multim. Tools Appl.3
2022 A critical review on diagnosis of diabetic retinopathy using machine learning and deep learning
Dolly Das, Saroj K. Biswas 0001, Sivaji Bandyopadhyay
Multim. Tools Appl.3
2022 Feature fusion based machine learning pipeline to improve breast cancer prediction
Arnab Kumar Mishra, Pinki Roy, Sivaji Bandyopadhyay, Sujit Kumar Das
Multim. Tools Appl.3
2022 V2T: video to text framework using a novel automatic shot boundary detection algorithm
Alok Singh 0007, Thoudam Doren Singh, Sivaji Bandyopadhyay
Multim. Tools Appl.3
2022 Simplification of English and Bengali Sentences for Improving Quality of Machine Translation
Sainik Kumar Mahata, Avishek Garain, Dipankar Das 0001, Sivaji Bandyopadhyay
Neural Process. Lett.4
2021 Breast ultrasound tumour classification: A Machine Learning - Radiomics based approach
abstract
Abstract Prediction of breast tumour malignancy using ultrasound imaging, is an important step for early detection of breast cancer. An efficient prediction system can be a great help to improve the survival chances of the involved patients. In this work, a machine learning (ML)—radiomics based classification pipeline is proposed, to perform this predictive modelling task, in a much more efficient manner. Multiple different types of image features of the region of interests are considered in this work, followed by a recursive feature elimination based feature selection step. Furthermore, a synthetic minority oversampling technique based step is also included in the pipeline, to deal with the class imbalance problem, that is often present in medical imaging datasets. Various ML models are considered in the subsequent model training phase, on a publicly available breast ultrasound image dataset (BUSI). From experimental analysis it has been observed that, shape, texture and histogram oriented gradients related features are the most informative, with respect to the predictive modelling task. Furthermore, it was observed that ensemble learners such as random forest, gradient boosting and AdaBoost classifiers are able to achieve significant results with respect to multiple evaluation metrics. The proposed approach achieved the state‐of‐the‐art accuracy, area under the curve, F1‐score and Mathews correlation coefficient values of 0.974, 0.97, 0.94 and 0.959, respectively, on the BUSI dataset. Such kind of impressive results suggest that the proposed approach can have a very high practical utility, in real medical diagnostic settings.
Arnab Kumar Mishra, Pinki Roy, Sivaji Bandyopadhyay, Sujit Kumar Das
Expert Syst. J. Knowl. Eng.3
2021 Investigating the roles of sentiment in machine translation
Sainik Kumar Mahata, Dipankar Das 0001, Sivaji Bandyopadhyay
Mach. Transl.3
2021 Predictive approaches for the UNIX command line: curating and exploiting domain knowledge in semantics deficit data
Thoudam Doren Singh, Abdullah Faiz Ur Rahman Khilji, Divyansha, Apoorva Vikram Singh, Surmila Thokchom, Sivaji Bandyopadhyay
Multim. Tools Appl.6
2021 An encoder-decoder based framework for hindi image caption generation
Alok Singh 0007, Thoudam Doren Singh, Sivaji Bandyopadhyay
Multim. Tools Appl.3
2018 Says Who? Deep Learning Models for Joint Speech Recognition, Segmentation and Diarization
abstract
The field of speech recognition has seen tremendous advances in the recent past owing to the development of powerful deep learning architectures. However, the closely related fields of speech segmentation and di-arization are still primarily dominated by sophisticated variants of hierarchical clustering algorithms. We propose a powerful adaptation of the state-of-the-art Speech Recognition models for these tasks and demonstrate the effectiveness of our techniques on standard datasets. Our architectures are a combination of Bidirectional Long Short Term Memory (LSTM) Networks, Convolutional Networks, and Fully Connected Networks, trained by Gradient Descent to minimize the Cross Entropy and the Connectionist Temporal Classification (CTC) losses. We adapt the Libri Speech corpus for the task of segmentation and diarization. We obtained comparable results with respect to state-of-the-art in both tasks.
Amitrajit Sarkar, Surajit Dasgupta, Sudip Kumar Naskar, Sivaji Bandyopadhyay
ICASSP4
2018 WME 3.0: An Enhanced and Validated Lexicon of Medical Concepts
abstract
Information extraction in the medical domain is laborious and time-consuming due to the insufficient number of domainspecific lexicons and lack of involvement of domain experts such as doctors and medical practitioners.Thus, in the present work, we are motivated to design a new lexicon, WME 3.0 (WordNet of Medical Events), which contains over 10,000 medical concepts along with their part of speech, gloss (descriptive explanations), polarity score, sentiment, similar sentiment words, category, affinity score and gravity score features.In addition, the manual annotators help to validate the overall as well as individual category level of medical concepts of WME 3.0 using Cohen's Kappa agreement metric.The agreement score indicates almost correct identification of medical concepts and their assigned features in WME 3.0.
Anupam Mondal, Dipankar Das 0001, Erik Cambria, Sivaji Bandyopadhyay
GWC4
2018 Multimodal mood classification of Hindi and Western songs
Braja Gopal Patra, Dipankar Das 0001, Sivaji Bandyopadhyay
J. Intell. Inf. Syst.3
2017 Textual Entailment Using Machine Translation Evaluation Metrics
Tanik Saikh, Sudip Kumar Naskar, Asif Ekbal, Sivaji Bandyopadhyay
CICLing (1)4
2017 Labeling data and developing supervised framework for hindi music mood analysis
Braja Gopal Patra, Dipankar Das 0001, Sivaji Bandyopadhyay
J. Intell. Inf. Syst.3
2016 A Multilevel Approach to Sentiment Analysis of Figurative Language in Twitter
Braja Gopal Patra, Soumadeep Mazumdar, Dipankar Das 0001, Paolo Rosso, Sivaji Bandyopadhyay
CICLing (2)5
2016 Multimodal Mood Classification - A Case Study of Differences in Hindi and Western Songs
abstract
Music information retrieval has emerged as a mainstream research area in the past two decades. Experiments on music mood classification have been performed mainly on Western music based on audio, lyrics and a combination of both. Unfortunately, due to the scarcity of digitalized resources, Indian music fares poorly in music mood retrieval research. In this paper, we identified the mood taxonomy and prepared multimodal mood annotated datasets for Hindi and Western songs. We identified important audio and lyric features using correlation based feature selection technique. Finally, we developed mood classification systems using Support Vector Machines and Feed Forward Neural Networks based on the features collected from audio, lyrics, and a combination of both. The best performing multimodal systems achieved F-measures of 75.1 and 83.5 for classifying the moods of the Hindi and Western songs respectively using Feed Forward Neural Networks. A comparative analysis indicates that the selected features work well for mood classification of the Western songs and produces better results as compared to the mood classification systems for Hindi songs.
Braja Gopal Patra, Dipankar Das 0001, Sivaji Bandyopadhyay
COLING3
2016 Statistical Natural Language Generation from Tabular Non-textual Data
Joy Mahapatra, Sudip Kumar Naskar, Sivaji Bandyopadhyay
INLG3
2016 WME: Sense, Polarity and Affinity based Concept Resource for Medical Events
abstract
In order to overcome the lack of medical corpora, we have developed a WordNet for Medical Events (WME) for identifying medical terms and their sense related information using a seed list.The initial WME resource contains 1654 medical terms or concepts.In the present research, we have reported the enhancement of WME with 6415 number of medical concepts along with their conceptual features viz.Parts-of-Speech (POS), gloss, semantics, polarity, sense and affinity.Several polarity lexicons viz.SentiWordNet, SenticNet, Bing Liu's subjectivity list and Taboda's adjective list were introduced with WordNet synonyms and hyponyms for expansion.The semantics feature guided us to build a semantic co-reference relation based network between the related medical concepts.These features help to prepare a medical concept network for better sense relation based visualization.Finally, we evaluated with respect to Adaptive Lesk Algorithm and conducted an agreement analysis for validating the expanded WME resource.
Anupam Mondal, Dipankar Das 0001, Erik Cambria, Sivaji Bandyopadhyay
GWC4
2015 Identifying Temporal Information and Tracking Sentiment in Cancer Patients' Interviews
Braja Gopal Patra, Nilabjya Ghosh, Dipankar Das 0001, Sivaji Bandyopadhyay
CICLing (2)4
2015 Textual Entailment Using Different Similarity Metrics
Tanik Saikh, Sudip Kumar Naskar, Chandan Giri, Sivaji Bandyopadhyay
CICLing (1)4
2014 Cross Lingual Snippet Generation Using Snippet Translation System
Pintu Lohar, Pinaki Bhaskar, Santanu Pal, Sivaji Bandyopadhyay
CICLing (2)4
2014 Word Alignment-Based Reordering of Source Chunks in PB-SMT
Santanu Pal, Sudip Kumar Naskar, Sivaji Bandyopadhyay
LREC3
2013 An Empirical Study of Combing Multiple Models in Bengali Question Classification
Somnath Banerjee 0001, Sivaji Bandyopadhyay
IJCNLP2
2013 Construction of Emotional Lexicon Using Potts Model
Braja Gopal Patra, Hiroya Takamura, Dipankar Das 0001, Manabu Okumura, Sivaji Bandyopadhyay
IJCNLP5
2013 MWE Alignment in Phrase Based Statistical Machine Translation
Santanu Pal, Sudip Kumar Naskar, Sivaji Bandyopadhyay
MTSummit3
2012 The 5W Structure for Sentiment Summarization-Visualization-Tracking
Amitava Das 0001, Sivaji Bandyopadhyay, Björn Gambäck
CICLing (1)2
2012 Roles of Event Actors and Sentiment Holders in Identifying Event-Sentiment Association
Anup Kumar Kolya, Dipankar Das 0001, Asif Ekbal, Sivaji Bandyopadhyay
CICLing (1)4
2012 Will the Identification of Reduplicated Multiword Expression (RMWE) Improve the Performance of SVM Based Manipuri POS Tagging?
Aribam Umananda Sharma, Laishram Martina Devi, Nepoleon Keisam, Khangengbam Dilip Singh, Sivaji Bandyopadhyay
CICLing (1)6
2012 Emotion Tracking on Blogs - A Case Study for Bengali
Dipankar Das 0001, Sagnik Roy, Sivaji Bandyopadhyay
IEA/AIE3
2012 A Classifier Based Approach to Emotion Lexicon Construction
Dipankar Das 0001, Soujanya Poria, Sivaji Bandyopadhyay
NLDB3
2011 Emotions on Bengali Blog Texts: Role of Holder and Topic
abstract
The paper presents an approach to identify the emotions of the bloggers on different topics provided in the Bengali blog documents. The rule based identification of emotion holders and topics along with their corresponding emotional expressions forms the baseline system. A Support Vector Machine (SVM) based supervised framework is also employed to identify the three components from the blog sentences and it outperforms the baseline system. As the topic of a document is not always conveyed at the sentence level, the similarity between a document's overall topic and sentential topic is measured through semantic clustering approach. Two different approaches are adopted to identify the many to many relationships among the holders and topics on Ekman's six emotions. One is based from the perspectives of the holders and other is with respect to topics. The two way evaluation of Ekman's six emotions achieves precision, recall and F-Score of 65.02%, 76.23% and 70.18% for 10 bloggers and 71.02%, 78.47% and 74.55% for 8 different topics on 512 test sentences respectively.
Dipankar Das 0001, Sivaji Bandyopadhyay
ASONAM2
2011 Temporal Analysis of Sentiment Events - A Visual Realization and Tracking
Dipankar Das 0001, Anup Kumar Kolya, Asif Ekbal, Sivaji Bandyopadhyay
CICLing (1)4
2011 Identification of Reduplicated Multiword Expressions Using CRF
Dhiraj Laishram, Naorem Bikramjit Singh, Ngariyanbam Mayekleima Chanu, Sivaji Bandyopadhyay
CICLing (1)5
2011 Answer Validation Using Textual Entailment
Partha Pakray, Alexander F. Gelbukh, Sivaji Bandyopadhyay
CICLing (2)3
2011 Integration of Reduplicated Multiword Expressions and Named Entities in a Phrase Based Statistical Machine Translation System
Thoudam Doren Singh, Sivaji Bandyopadhyay
IJCNLP2
2011 Handling Multiword Expressions in Phrase-Based Statistical Machine Translation
Santanu Pal, Tanmoy Chakraborty 0002, Sivaji Bandyopadhyay
MTSummit3
2010 Emotion Holder for Emotional Verbs - The Role of Subject and Syntax
Dipankar Das 0001, Sivaji Bandyopadhyay
CICLing2
2010 A Syntactic Textual Entailment System Based on Dependency Parser
Partha Pakray, Alexander F. Gelbukh, Sivaji Bandyopadhyay
CICLing3
2010 JU_CSE_GREC10: Named Entity Generation at GREC 2010
Amitava Das 0001, Tanik Saikh, Tapabrata Mondal, Sivaji Bandyopadhyay
INLG4
2010 A Query Focused Multi Document Automatic Summarization
Pinaki Bhaskar, Sivaji Bandyopadhyay
PACLIC2
2010 Identifying Emotional Expressions, Intensities and Sentence Level Emotion Tags Using a Supervised Framework
Dipankar Das 0001, Sivaji Bandyopadhyay
PACLIC2
2010 Finding Emotion Holder from Bengali Blog Texts---An Unsupervised Syntactic Approach
Dipankar Das 0001, Sivaji Bandyopadhyay
PACLIC2
2010 Towards the Global SentiWordNet
Amitava Das 0001, Sivaji Bandyopadhyay
PACLIC2
2010 A Supervised Machine Learning Approach for Event-Event Relation Identification
Anup Kumar Kolya, Asif Ekbal, Sivaji Bandyopadhyay
PACLIC3
2009 Voted Approach for Part of Speech Tagging in Bengali
Asif Ekbal, Mohammed Hasanuzzaman, Sivaji Bandyopadhyay
PACLIC3
2009 Named Entity Recognition for Manipuri Using Support Vector Machine
Thoudam Doren Singh, Asif Ekbal, Sivaji Bandyopadhyay
PACLIC4
2008 Invited Talk: Multilingual Named Entity Recognition
Sivaji Bandyopadhyay
IJCNLP1
2008 Bengali, Hindi and Telugu to English Ad-hoc Bilingual Task
Sivaji Bandyopadhyay, Tapabrata Mondal, Sudip Kumar Naskar, Asif Ekbal, Rejwanul Haque, Srinivasa Rao Godhavarthy
IJCNLP1
2008 Bengali Named Entity Recognition Using Support Vector Machine
Asif Ekbal, Sivaji Bandyopadhyay
IJCNLP2
2008 Named Entity Recognition in Bengali: A Conditional Random Field Approach
Asif Ekbal, Rejwanul Haque, Sivaji Bandyopadhyay
IJCNLP3
2008 Language Independent Named Entity Recognition in Indian Languages
Asif Ekbal, Rejwanul Haque, Amitava Das 0001, Venkateswarlu Poka, Sivaji Bandyopadhyay
IJCNLP5
2008 Generation of Referring Expression Using Prefix Tree Structure
Sibabrata Paladhi, Sivaji Bandyopadhyay
IJCNLP2
2008 A Document Graph Based Query Focused Multi-Document Summarizer
Sibabrata Paladhi, Sivaji Bandyopadhyay
IJCNLP2
2008 Design of a Rule-based Stemmer for Natural Language Text in Bengali
Sandipan Sarkar, Sivaji Bandyopadhyay
IJCNLP2
2008 Morphology Driven Manipuri POS Tagger
Thoudam Doren Singh, Sivaji Bandyopadhyay
IJCNLP2
2008 JU-PTBSGRE: GRE Using Prefix Tree Based Structure
Sibabrata Paladhi, Sivaji Bandyopadhyay
INLG2
2008 Multi-Engine Approach for Named Entity Recognition in Bengali
Asif Ekbal, Sivaji Bandyopadhyay
PACLIC2
2006 A Modified Joint Source-Channel Model for Transliteration
Asif Ekbal, Sudip Kumar Naskar, Sivaji Bandyopadhyay
ACL3
2005 Generating Headline Summary from a Document Set
Kamal Sarkar, Sivaji Bandyopadhyay
CICLing2
2000 Detection and Correction of Phonetic Errors with a New Orthographic Dictionary
Sivaji Bandyopadhyay
PACLIC1