Amit Kumar Jaiswal 0001

dblp:76/9560-1 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0001-8848-7041ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Multimodal RAG Enhanced Visual Description
abstract
Textual descriptions for multimodal inputs entail recurrent refinement of queries to produce relevant output images. Despite efforts to address challenges such as scaling model size and data volume, the cost associated with pre-training and fine-tuning remains substantial. However, pre-trained large multimodal models (LMMs) encounter a modality gap, characterised by a misalignment between textual and visual representations within a common embedding space. Although fine-tuning can potentially mitigate this gap, it is typically expensive and impractical due to the requirement for extensive domain-driven data. To overcome this challenge, we propose a lightweight training-free approach utilising Retrieval-Augmented Generation (RAG) to extend across the modality using a linear mapping, which can be computed efficiently. Our reproducible code can be found in https://github.com/amitkumarj441/mRAG-gim. During inference, this mapping is applied to images embedded by an LMM enabling retrieval of closest textual descriptions from the training set. These textual descriptions, in conjunction with an instruction, cater as an input prompt for the language model to generate new textual descriptions. In addition, we introduce an iterative technique for distilling the mapping by generating synthetic descriptions via the language model facilitating optimisation for standard utilised image description measures. Experimental results on two benchmark multimodal datasets demonstrate significant improvements.
Amit Kumar Jaiswal 0001, Haiming Liu 0002, Ingo Frommholz
CIKM1
2025 DHOW '25: 2nd International Workshop on Diffusion of Harmful Content on Online Web
abstract
With the advancement of digital technologies and gadgets, online content has become easily accessible. At the same time, harmful content also spread widely. There are different harmful content types present on various platforms in multiple languages. The topic of harmful content is broad and covers multiple research directions. Users of platforms are affected by all of them. In research, the different forms are mostly analysed separately, e.g. misinformation, cyber-bullying and hate speech. Most research has been conducted for only one platform, for a monolingual situation or on a particular issue. Counter-measures like blocking are down-ranking can make harmful content spreaders to switch platforms and languages to continuously reach a user base. Harmful content does not only appear on social media but also on news media. Spreader share harmful content in posts, news articles, comments and hyperlinks. There is a great need to study harmful content across platforms, languages, and topics. We plan to bring the research on harmful content under one umbrella such that different approaches and novel methods can be shared. The workshop will also cover the currently ongoing issues of war and elections. We propose the workshop, DHOW: Diffusion of Harmful Content on Online Web, which brings together the research on different topics of harmful content. We expect to discuss innovative research work and future research directions. The proposed workshop is the next iteration of DHOW 2024. https://dhow-workshop.github.io previously organized at ACM WebSci 2024 in Stuttgart, Germany.
Amit Kumar Jaiswal 0001, Thomas Mandl 0001, Gautam Kishore Shahi, Durgesh Nandini, Haiming Liu 0002
ACM Multimedia1
2025 Multimodal medical image fusion algorithm in the era of big data
abstract
Abstract In image-based medical decision-making, different modalities of medical images of a given organ of a patient are captured. Each of these images will represent a modality that will render the examined organ differently, leading to different observations of a given phenomenon (such as stroke). The accurate analysis of each of these modalities promotes the detection of more appropriate medical decisions. Multimodal medical imaging is a research field that consists in the development of robust algorithms that can enable the fusion of image information acquired by different sets of modalities. In this paper, a novel multimodal medical image fusion algorithm is proposed for a wide range of medical diagnostic problems. It is based on the application of a boundary measured pulse-coupled neural network fusion strategy and an energy attribute fusion strategy in a non-subsampled shearlet transform domain. Our algorithm was validated in dataset with modalities of several diseases, namely glioma, Alzheimer’s, and metastatic bronchogenic carcinoma, which contain more than 100 image pairs. Qualitative and quantitative evaluation verifies that the proposed algorithm outperforms most of the current algorithms, providing important ideas for medical diagnosis.
Prayag Tiwari, Hari Mohan Pandey, Catarina Moreira, Amit Kumar Jaiswal 0001
Neural Comput. Appl.5
2024 FakeClaim: A Multiple Platform-Driven Dataset for Identification of Fake News on 2023 Israel-Hamas War
Gautam Kishore Shahi, Amit Kumar Jaiswal 0001, Thomas Mandl 0001
ECIR (5)2
2023 Lightweight Adaptation of Neural Language Models via Subspace Embedding
abstract
Traditional neural word embeddings are usually dependent on a richer diversity of vocabulary. However, the language models recline to cover major vocabularies via the word embedding parameters, in particular, for multilingual language models that generally cover a significant part of their overall learning parameters. In this work, we present a new compact embedding structure to reduce the memory footprint of the pre-trained language models with a sacrifice of up to 4% absolute accuracy. The embeddings vectors reconstruction follows a set of subspace embeddings and an assignment procedure via the contextual relationship among tokens from pre-trained language models. The subspace embedding structure1 calibrates to masked language models, to evaluate our compact embedding structure on similarity and textual entailment tasks, sentence and paraphrase tasks. Our experimental evaluation shows that the subspace embeddings achieve compression rates beyond 99.8% in comparison with the original embeddings for the language models on XNLI and GLUE benchmark suites.
Amit Kumar Jaiswal 0001, Haiming Liu 0002
CIKM1
2023 A Model-Agnostic Framework for Recommendation via Interest-aware Item Embeddings
abstract
Item representation holds significant importance in recommendation systems, which encompasses domains such as news, retail, and videos. Retrieval and ranking models utilise item representation to capture the user-item relationship based on user behaviours. While existing representation learning methods primarily focus on optimising item-based mechanisms, such as attention and sequential modelling. However, these methods lack a modelling mechanism to directly reflect user interests within the learned item representations. Consequently, these methods may be less effective in capturing user interests indirectly. To address this challenge, we propose a novel Interest-aware Capsule network (IaCN) recommendation model, a model-agnostic framework that directly learns interest-oriented item representations. IaCN serves as an auxiliary task, enabling the joint learning of both item-based and interest-based representations. This framework adopts existing recommendation models without requiring substantial redesign. We evaluate the proposed approach on benchmark datasets, exploring various scenarios involving different deep neural networks, behaviour sequence lengths, and joint learning ratios of interest-oriented item representations. Experimental results demonstrate significant performance enhancements across diverse recommendation models, validating the effectiveness of our approach.
Amit Kumar Jaiswal 0001
RecSys1
2023 Predicting users' behavior using mouse movement information: an information foraging theory perspective
Amit Kumar Jaiswal 0001, Prayag Tiwari, M. Shamim Hossain
Neural Comput. Appl.1
2023 Securing Blockchain Transactions Using Quantum Teleportation and Quantum Digital Signature
Sheetal Singh, Nikhil Kumar Rajput, Vipin Kumar Rathi, Hari Mohan Pandey, Amit Kumar Jaiswal 0001, Prayag Tiwari
Neural Process. Lett.5
2022 SANTM: Efficient Self-attention-driven Network for Text Matching
abstract
Self-attention mechanisms have recently been embraced for a broad range of text-matching applications. Self-attention model takes only one sentence as an input with no extra information, i.e., one can utilize the final hidden state or pooling. However, text-matching problems can be interpreted either in symmetrical or asymmetrical scopes. For instance, paraphrase detection is an asymmetrical task, while textual entailment classification and question-answer matching are considered asymmetrical tasks. In this article, we leverage attractive properties of self-attention mechanism and proposes an attention-based network that incorporates three key components for inter-sequence attention: global pointwise features, preceding attentive features, and contextual features while updating the rest of the components. Our model follows evaluation on two benchmark datasets cover tasks of textual entailment and question-answer matching. The proposed efficient Self-attention-driven Network for Text Matching outperforms the state of the art on the Stanford Natural Language Inference and WikiQA datasets with much fewer parameters.
Prayag Tiwari, Amit Kumar Jaiswal 0001, Sahil Garg, Ilsun You
ACM Trans. Internet Techn.2
2021 Entity-aware capsule network for multi-class classification of big data: A deep learning approach
Amit Kumar Jaiswal 0001, Prayag Tiwari, Sahil Garg, M. Shamim Hossain
Future Gener. Comput. Syst.1
2020 Utilising Information Foraging Theory for User Interaction with Image Query Auto-Completion
Amit Kumar Jaiswal 0001, Haiming Liu 0002, Ingo Frommholz
ECIR (1)1
2020 Evidence of power-law behavior in cognitive IoT applications
Sujit Bebortta, Dilip Senapati, Nikhil Kumar Rajput, Vipin Kumar Rathi, Hari Mohan Pandey, Amit Kumar Jaiswal 0001, Jia Qian, Prayag Tiwari
Neural Comput. Appl.7