Santosh Kumar Mishra

dblp:183/0480 · DBLP profile ↗
← Back
19ranked-venue papers
13as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 11 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Span-level target and harmfulness detection in code-mixed hinglish using cross-lingual and multilingual transformers
Ashutosh Upadhyay, Eishaan Khatri, Harish Pratap Singh, Santosh Kumar Mishra, Vikash Kumar Rai
Expert Syst. Appl.4
2025 Multimodal Multi-task deep learning framework for classification of sentiment, emotion, humor, sarcasm and toxicity from speech
Harish Pratap Singh, Puneet Prashar, Gaddam Sai Bharath Chandra Reddy, Santosh Kumar Mishra
Knowl. Based Syst.4
2025 Hi-MetaCap: Configuring Object Relational Transformer in Meta-Learning Environment for Image Captioning in Hindi
abstract
This article proposes a meta-learning-based, few-shot image captioning framework based on an ensemble of object-relational transformer models and a self-distillation strategy. Unlike traditional approaches, the proposed framework can train models using non-paired images and captions. In each iteration, multiple base models are trained on distinct data samples, forming an ensemble that generates pseudo-captions with confidence-based weighting. An efficient pseudo-feature generation method based on gradient descent enables learning from non-paired captions. These pseudo-features and pseudo-captions are then used to train the base models in subsequent iterations. The obtained results show that using just 1% of the paired image-caption training data significantly improves performance and generates meaningful captions. The proposed framework demonstrates the potential to reduce reliance on large, paired datasets while maintaining high-quality image captioning performance. 1
Santosh Kumar Mishra, Soham Chakraborty 0006, Sriparna Saha 0001, Pushpak Bhattacharyya
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2025 Continual Learning for Image Captioning in Hindi
abstract
Continual learning, alternatively referred to as incremental learning or lifelong learning, represents a learning paradigm enabling an agent to acquire new information without compromising its retention of previously learned knowledge. Continual learning research has resulted in a number of ways to prevent catastrophic forgetting in deep neural networks. Apparently, continuous learning of recurrent models has never been used to solve the image captioning problems in the Hindi language. This research presents a novel approach to employ continual learning in the context of generating image captions in Hindi. Given its widespread usage in South Asia and its official status in India, Hindi ranks as the third most spoken language worldwide. We look at the continual learning of Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) based models for image captioning in a systematic way. In continuous image captioning problems, we incorporate an attention-based method that explicitly addresses the transient aspect of vocabulary. Using MS-COCO datasets, we employ our methods to solve the incremental image captioning problem. Our findings show that the proposed method can learn five captioning tasks in consecutive order without forgetting the one it has already learned.
Santosh Kumar Mishra, Sriparna Saha 0001, Pushpak Bhattacharyya
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2023 GAGPT-2: A Geometric Attention-based GPT-2 Framework for Image Captioning in Hindi
abstract
Image captioning frameworks usually employ an encoder-decoder paradigm, with the encoder receiving abstract image feature vectors as input and the decoder for language modeling. Nowadays, most prominent architectures employ features from region proposals derived from object detection modules. In this work, we propose a novel architecture for image captioning. We employ the object detection module integrated with transformer architecture as an encoder and GPT-2 (Generative Pre-trained Transformer) as a decoder. The encoder utilizes the information of the spatial relationships among detected objects. We introduce a unique methodology for image caption generation in Hindi, which is widely spoken in South Asia and India and is the world’s third most spoken language as well as India’s official language. In terms of BLEU scores, the proposed approach’s performance is comparable to those of other baselines, and the results illustrate that the proposed approach outperforms the other baselines. The efficacy of the proposed approach in generating correct captions is further determined by human assessment in terms of adequacy and fluency.
Santosh Kumar Mishra, Soham Chakraborty 0006, Sriparna Saha 0001, Pushpak Bhattacharyya
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2023 An Object Localization-based Dense Image Captioning Framework in Hindi
abstract
Dense image captioning is a task that requires generating localized captions in natural language for multiple regions of an image. This task leverages its functionalities from both computer vision for recognizing regions in an image and natural language processing for generating captions. Numerous works have been carried out on dense image captioning for resource-rich languages like English; however, resource-poor languages like Hindi are not explored. Hindi is one of India’s official languages and is the third most spoken language in the world. This article proposes a dense image captioning model to describe different segments of an image by generating more than one caption in the Hindi language. For localized image recognition and language modeling, we employ Faster R-CNN and Long Short-Term Memory (LSTM), respectively. Apart from this, we conduct various experiments using gated recurrent units (GRUs) and attention mechanism. By manually translating the well-known Visual Genome dataset from English to Hindi, a dataset has been created for dense image captioning in Hindi. The experiments conducted on the newly constructed Hindi dense image captioning dataset illustrate the efficacy of the proposed method over the state-of-the-art methods.
Santosh Kumar Mishra, Sriparna Saha 0001, Pushpak Bhattacharyya
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2023 Dynamic Convolution-based Encoder-Decoder Framework for Image Captioning in Hindi
abstract
In sequence-to-sequence modeling tasks, such as image captioning, machine translation, and visual question answering, encoder-decoder architectures are state of the art. An encoder, convolutional neural network (CNN) encodes input images into fixed dimensional vector representation in the image captioning task, whereas a decoder, a recurrent neural network, performs language modeling and generates the target descriptions. Recent CNNs use the same operation over every pixel; however, all the image pixels are not equally important. To address this, the proposed method uses a dynamic convolution-based encoder for image encoding or feature extraction, Long-Short-Term-Memory as a decoder for language modeling, and X-Linear attention to make the system robust. Encoders, attentions, and decoders are important aspects of the image captioning task; therefore, we experiment with various encoders, decoders, and attention mechanisms. Most of the works for image captioning have been carried out for the English language in the existing literature. We propose a novel approach for caption generation from images in Hindi. Hindi, widely spoken in South Asia and India, is the fourth most-spoken language globally; it is India’s official language. The proposed method utilizes dynamic convolution operation on the encoder side to obtain a better image encoding quality. The Hindi image captioning dataset is manually created by translating the popular MSCOCO dataset from English to Hindi. In terms of BLEU scores, the performance of the proposed method is compared with other baselines, and the results obtained show that the proposed method outperforms different baselines. Manual human assessment in terms of adequacy and fluency of the captions generated further determines the efficacy of the proposed method in generating good-quality captions.
Santosh Kumar Mishra, Sushant Sinha, Sriparna Saha 0001, Pushpak Bhattacharyya
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2022 Bug report summarization using multi-view multi-objective optimization framework
abstract
Existing bug text reports are widely used by software engineers to assist them in understanding important components of individual defects and adjustments done to resolve the fault. However, bug reports are typically long and require significant effort to comprehend. Summarization of bug reports appears to be beneficial in this regard, covering essential and diverse information. We frame the Bug report summarization problem as a clustering-based optimization problem and solve it using a multi-view multi-objective optimization framework. To represent the bug reports, semantic and syntactic representations, which are regarded as separate views, are taken into account. Several cluster quality measures computed on partitionings obtained using distinct views are optimized simultaneously using a multi-objective optimization-based approach known as archived multi-objective simulated annealing. To determine the consensus between the partitionings generated using different views, an agreement index is computed, which is also optimized simultaneously along with other cluster quality measures. The proposed methodology automatically determines the number of clusters. The experiments are carried out using the two benchmark datasets (SDS and ADS) and evaluated using the well-known ROUGE, Precision, Recall, and F-measure evaluation metrics. The obtained results show that the proposed methodology outperforms state-of-the-art methods.
Santosh Kumar Mishra, Harshavardhan Kundarapu, Sayantan Mitra, Sriparna Saha 0001, Pushpak Bhattacharyya
GECCO1
2022 Investigations in Psychological Stress Detection from Social Media Text using Deep Architectures
abstract
Psychological stress is a feeling of mental stress and pressure. Within this context we try to understand the tweets which express psychological stress. Psychological stress detection is a complex task majorly because the stressed community contains mostly introverts. With the increase in people using social-media sites, the information related to mental health of the people extracted from their posts are increasing day-by-day. We use this information to design an AI enabled framework for automatic stress detection. However these data are noisy and complex, therefore deep learning based models are utilized for automatic extraction of features rather than manual extraction of features. Several deep learning based architectures including Multichannel CNN, CNN, GRU, Capsule network and BERT model are explored for solving this task of detecting tweets having mentions about mental stress. Experimental results on a standard Twitter dataset reveal that Multichannel CNN attains the best performance with accuracy of 97.5%, precision, recall and f-score values of 96.8%, 97.5% and 97.2%, respectively.
Bishal Shaw, Sriparna Saha 0001, Santosh Kumar Mishra, Angshuman Ghosh
ICPR3
2022 A Deep Learning based Framework for Image Paragraph Generation in Hindi
Santosh Kumar Mishra, Sushant Sinha, Sriparna Saha 0001, Pushpak Bhattacharyya
PACLIC1
2022 Scientific document summarization in multi-objective clustering framework
Santosh Kumar Mishra, Naveen Saini, Sriparna Saha 0001, Pushpak Bhattacharyya
Appl. Intell.1
2022 Efficient Channel Attention Based Encoder-Decoder Approach for Image Captioning in Hindi
abstract
Image captioning refers to the process of generating a textual description that describes objects and activities present in a given image. It connects two fields of artificial intelligence, computer vision, and natural language processing. Computer vision and natural language processing deal with image understanding and language modeling, respectively. In the existing literature, most of the works have been carried out for image captioning in the English language. This article presents a novel method for image captioning in the Hindi language using encoder–decoder based deep learning architecture with efficient channel attention. The key contribution of this work is the deployment of an efficient channel attention mechanism with bahdanau attention and a gated recurrent unit for developing an image captioning model in the Hindi language. Color images usually consist of three channels, namely red, green, and blue. The channel attention mechanism focuses on an image’s important channel while performing the convolution, which is basically to assign higher importance to specific channels over others. The channel attention mechanism has been shown to have great potential for improving the efficiency of deep convolution neural networks (CNNs). The proposed encoder–decoder architecture utilizes the recently introduced ECA-NET CNN to integrate the channel attention mechanism. Hindi is the fourth most spoken language globally, widely spoken in India and South Asia; it is India’s official language. By translating the well-known MSCOCO dataset from English to Hindi, a dataset for image captioning in Hindi is manually created. The efficiency of the proposed method is compared with other baselines in terms of Bilingual Evaluation Understudy (BLEU) scores, and the results obtained illustrate that the method proposed outperforms other baselines. The proposed method has attained improvements of 0.59%, 2.51%, 4.38%, and 3.30% in terms of BLEU-1, BLEU-2, BLEU-3, and BLEU-4 scores, respectively, with respect to the state-of-the-art. Qualities of the generated captions are further assessed manually in terms of adequacy and fluency to illustrate the proposed method’s efficacy.
Santosh Kumar Mishra, Gaurav Rai, Sriparna Saha 0001, Pushpak Bhattacharyya
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2022 On Multimodal Microblog Summarization
abstract
Microblog summarization systems are gaining importance during natural disasters. A lot of tweets are posted along with multimedia content during the occurrence of any natural disaster event. Extracting relevant information/summary from these tweets is important for the smooth functioning of the rescue operation. Moreover, because of the limited size of the tweets, in many cases, tweets are associated with images. The current work is the first of its kind where both the image and the tweet text are utilized simultaneously to generate a summary from microblog data generated during a disaster event. Different aspects, such as syntactic similarity, the maximum length of the tweets, retweet score, and antiredundancy, are considered as objective functions and those are simultaneously optimized using a metaheuristic population-based evolutionary strategy to select a good set of tweets to form a good quality summary. In order to extract information from images, a dense captioning model is utilized and the dense captions are further utilized for calculating the antiredundancy measure. We employed word mover distance to capture the semantic similarity between two tweets. Due to the unavailability of the dataset for multimodal microblog summarization tasks in a disaster-event scenario, datasets are created and made openly available to the community. The obtained summarization results are evaluated using the well-known ROUGE measure.
Naveen Saini, Sriparna Saha 0001, Pushpak Bhattacharyya, Shubhankar Mrinal, Santosh Kumar Mishra
IEEE Trans. Comput. Soc. Syst.5
2021 Let's Summarize Scientific Documents! A Clustering-Based Approach via Citation Context
Santosh Kumar Mishra, Naveen Saini, Sriparna Saha 0001, Pushpak Bhattacharyya
NLDB1
2021 Dense Image Captioning in Hindi
abstract
Much work has been done on dense image captioning in the English language. In this paper, we propose a Encoder-Decoder architecture for the first time, to the best of our knowledge, to generate Hindi dense image captions. We leverage the proven encoder-decoder architecture and provide a Hindi data-set along with results of dense captioning on the data-set. We also propose a novel post-processing technique to further improve the quality of captions by using multiple discriminators (language discriminator, length penalty and dissimilarity score). Hindi is the fourth most spoken language in the world. It is the official language of India and to the best of our knowledge, this is the first attempt at dense image captioning in the Hindi language. The data-set is manually created by translating Visual Genome captions from English to Hindi under manual supervision. Both qualitative (in terms of adequacy and fluency) and quantitative (in terms of BLEU scores) analysis of the generated captions reveal the effectiveness of our proposed approach compared to several strong baselines.
Karanjit Gill, Sriparna Saha 0001, Santosh Kumar Mishra
SMC3
2021 An Information Multiplexed Encoder-Decoder Network for Image Captioning in Hindi
abstract
Image captioning is a multi-modal problem linking computer vision and natural language processing, which combines image analysis and text generation challenges. In the literature, most of the image captioning works have been accomplished in the English language only. This paper proposes a new approach for image captioning in the Hindi language using deep learning-based encoder-decoder architecture. Hindi, widely spoken in India and South Asia, is the fourth most spoken language globally; it is India’s official language. In recent years, significant advancement has been made in image captioning, utilizing encoder-decoder architectures based on convolutional neural networks (CNNs) and recurrent neural networks (RNNs). Encoder CNN extracts features from input images, whereas decoder RNN performs language modeling. The proposed encoder-decoder architecture utilizes information multiplexing in the encoder CNN to achieve a performance gain in feature extraction. Extensive experimentation is carried out on the benchmark MSCOCO Hindi dataset, and significant improvements in BLEU score are reported compared to the baselines. Manual human evaluation in terms of adequacy and fluency of the generated captions further establishes the proposed method’s efficacy in generating good quality captions.
Santosh Kumar Mishra, Mahesh Babu Peethala, Sriparna Saha 0001, Pushpak Bhattacharyya
SMC1
2021 MEABRS: A Multi-objective Evolutionary Framework for Software Bug Report Summarization
abstract
Software developers frequently use existing bug text reports to help them in understanding the different aspects of the specific defects and changes made to resolve that bug. But, bug reports are usually lengthy in nature and require considerable effort in understanding. In this direction, summarization of bug reports seems to be useful, covering relevant and diversified information. In the current article, we investigate the use of Multi-objective Evolutionary Algorithm (MEA) for BRS and thus, we name our approach as MEABRS. For MEA, we utilize the search capability of multi-objective bi-nary differential evolution (MOBDE) where we simultaneously optimize different aspects of BRS including diversity among sentences, sentence relevance using term weighting scheme, and length of the sentence. A keyword-based objective function is also incorporated in our optimization process to improve the quality of a bug report’s summary and in order to do so, rapid automatic keyword extraction (RAKE) toolkit is utilized. The generated summaries are evaluated with the available gold summaries corresponding to the two benchmark datasets (ADS and SDS) in terms of precision, recall, F-measure, and ROUGE measure. Results obtained demonstrate the efficacy of MEABRS with an average improvement (over both datasets) of 4.5% in terms of F1-Measure. Further, results are also validated using a statistical significance t-test.
Anuj Shastri, Naveen Saini, Sriparna Saha 0001, Santosh Kumar Mishra
SMC4
2021 Design of Fractional Calculus based differentiator for edge detection in color images
Santosh Kumar Mishra, Koushlendra Kumar Singh, Richa Dixit, Manish Kumar Bajpai
Multim. Tools Appl.1
2021 A Hindi Image Caption Generation Framework Using Deep Learning
abstract
Image captioning is the process of generating a textual description of an image that aims to describe the salient parts of the given image. It is an important problem, as it involves computer vision and natural language processing, where computer vision is used for understanding images, and natural language processing is used for language modeling. A lot of works have been done for image captioning for the English language. In this article, we have developed a model for image captioning in the Hindi language. Hindi is the official language of India, and it is the fourth most spoken language in the world, spoken in India and South Asia. To the best of our knowledge, this is the first attempt to generate image captions in the Hindi language. A dataset is manually created by translating well known MSCOCO dataset from English to Hindi. Finally, different types of attention-based architectures are developed for image captioning in the Hindi language. These attention mechanisms are new for the Hindi language, as those have never been used for the Hindi language. The obtained results of the proposed model are compared with several baselines in terms of BLEU scores, and the results show that our model performs better than others. Manual evaluation of the obtained captions in terms of adequacy and fluency also reveals the effectiveness of our proposed approach. Availability of resources : The codes of the article are available at https://github.com/santosh1821cs03/Image_Captioning_Hindi_Language ; The dataset will be made available: http://www.iitp.ac.in/∼ai-nlp-ml/resources.html .
Santosh Kumar Mishra, Rijul Dhir, Sriparna Saha 0001, Pushpak Bhattacharyya
ACM Trans. Asian Low Resour. Lang. Inf. Process.1