VLDB 2026 Research / reviewers in the wild / expert
Kishor Datta Gupta
dblp:231/4821
· DBLP profile ↗
14ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-6867-2653ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | POSTER: TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge DevicesabstractAs autonomous AI agents increasingly rely on Large Language Models (LLMs) for decision making; the necessity for local, privacy preserving execution on edge devices has become paramount. This paper investigates the effectiveness of Small Language Models (SLMs) with equal or fewer than 4 billion parameters for agentic tasks, specifically function and API calling. Using the Berkeley Function Calling Leaderboard (BFCL), we provide a comprehensive evaluation of models including xLAM-2, Qwen3, and TinyLlama. We introduce a Direct Preference Optimization (DPO) pipeline that leverages AgentBank (ALFRED) trajectories to align SLMs toward expert level tool use behavior. Our experimental results demonstrate a clear performance hierarchy where models in the 1-3B range significantly outperform sub 1B variants, achieving up to 65.74% overall accuracy and 55.62% in multi turn scenarios. Furthermore, we validate deployment feasibility through 4-bit quantization and hardware simulation, showing that these models can operate within the 8GB VRAM constraints of an Nvidia Jetson Nano while maintaining a functional throughput of ∼ 2.6 tokens/sec. Mohd Ariful Haque, Fahad Rahman, Kishor Datta Gupta, Khalil Shujaee, Roy George |
CF | 3 |
| 2025 | Physical Fuzzy Rule Based Unsupervised News Article ClusteringabstractThis study introduces the Fuzzy-Rule-Enhanced Autoencoder (FREA), a novel hybrid approach for interpretable text clustering that addresses the limitations of traditional methods in comprehending nuanced linkages and incorporating domain-specific knowledge. FREA integrates fuzzy physical rules generated by Large Language Models (LLMs) that derive weighted related terms for provisional news labels within an autoencoder framework. This approach enhances interpretability and flexibility by including these principles into the feature learning process, aligning them with Word2Vec and RoBERTa embeddings to ensure contextually relevant and semantically robust representations. The fuzzy rule layer adaptively adjusts feature learning based on the similarity of inputs to predefined linguistic rules, allowing the model to incorporate domain-specific constraints in a completely trainable manner. FREA demonstrates its effectiveness and adaptability through evaluations on benchmark datasets utilizing both hard and soft clustering methods, along with different quantitative evaluation metrics, primarily focusing on news articles as the primary test case. This study amalgamates the benefits of fuzzy rules and deep learning, offering a flexible, interpretable, and human-aligned methodology for unsupervised clustering problems, therefore effectively linking rule-based reasoning with neural networks. Marufa Kamal, Kishor Datta Gupta, Masrura Tasnim, Mohd Ariful Haque, Roy George |
COMPSAC | 2 |
| 2024 | Utilizing Structural Metrics from Knowledge Graphs to Enhance the Robustness Quantification of Large Language Models (Extended Abstract)abstractThe goal of this study is to determine whether large language models (LLMs) like CodeLlama, Mistral, and Vicuna can be used to build knowledge graphs (KGs) from textual data. We create class descriptions for well-known KGs such as DBpedia, YAGO, and Google Knowledge Graph, from which we extract RDF triples and enhance these graphs using different preprocessing methods. Six structural quality measures are used in the study to compare the constructed and existing KGs. Our results demonstrate how important LLMs are to improving KG construction and provide insightful information for KG construction researchers. Moreover, an in-depth analysis of popular open-source LLM models enables researchers to identify the most efficient model for various tasks, ensuring optimal performance in specific applications. Mohd Ariful Haque, Marufa Kamal, Roy George, Kishor Datta Gupta |
DSAA | 4 |
| 2023 | Case Study-Based Approach of Quantum Machine Learning in Cybersecurity: Quantum Support Vector Machine for Malware Classification and ProtectionabstractQuantum machine learning (QML) is an emerging field of research that leverages quantum computing to improve the classical machine learning approach to solve complex real-world problems. QML has the potential to address cybersecurity-related challenges. Considering the novelty and complex architecture of QML, resources are not yet explicitly available that can pave cybersecurity learners to instill efficient knowledge of this emerging technology. In this research, we design and develop QML-based ten learning modules covering various cybersecurity topics by adopting student centering case-study based learning approach. We apply one subtopic of QML on a cybersecurity topic comprised of pre-lab, lab, and post-lab activities towards providing learners with hands-on QML experiences in solving real-world security problems. In order to engage and motivate students in a learning environment that encourages all students to learn, pre-lab offers a brief introduction to both the QML subtopic and cybersecurity problem. In this paper, we utilize quantum support vector machine (QSVM) for malware classification and protection where we use open source Pennylane QML framework on the drebin215 dataset. We demonstrate our QSVM model and achieve an accuracy of 95% in malware classification and protection. We will develop all the modules and introduce them to the cybersecurity community in the coming days. Mst. Shapna Akter, Hossain Shahriar, Sheikh Iqbal Ahamed, Kishor Datta Gupta, Muhammad Rahman 0006, Atef Mohamed, Mohammad Ashiqur Rahman, Akond Ashfaque Ur Rahman, Fan Wu 0013 |
COMPSAC | 4 |
| 2022 | A Safer Approach to Build Recommendation Systems on Unidentifiable Data
Kishor Datta Gupta, Akib Sadmanee, Nafiz Sadman |
ICAART (3) | 1 |
| 2022 | HeteroGenius: An Improvised 'Intelligence' in Heterogeneous Graph TransformersabstractHeterogeneous graphs can capture pragmatic relations between entities (or nodes) better than homogeneous graphs. Heterogeneous graphs are crucial in search and classification problems and can correlate with social network graphs. However, this increases complexity and demands a clear understanding of the relationships, and rankings of the network. The ranking can take the form of various scoring-based systems, or finding the importance of the relations between two entities. In this research, we practice the use of incorporating a meta-edge definition, ‘Force’, between nodes to embed meaning to its dimensionality. We add this ‘Force’ to the Heterogeneous Graph Transformer, which we term as ‘HeteroGenius’, and experimentally demonstrate that the addition increases the overall accuracy by 2%. Nafiz Sadman, Akib Sadmanee, Kishor Datta Gupta, Roy George |
ICMLA | 3 |
| 2022 | Identifying Anomalous Flight Trajectories by leveraging ensembled outlier detection frameworkabstractIncreased traffic density with a greater degree of increased automation in aviation is expected within the next decade. Therefore, airspace capacity will become more congested and result in increasing challenges for detecting conflicts between aerial vehicles. Furthermore, because these vehicles rely on surrounding vehicles following a planned path, it is essential to identify flights not following a planned direction. In this paper, we utilize an ensemble of the existing outlier detection approaches for identifying the anomalous flight trajectories. In the initial step, flight trajectories are preprocessed to extract and process vital features, with the next step of having twenty different outlier detection algorithms assembled to classify trajectories. Throughout our extensive experiments and comparison studies, promising results are shown including the effectiveness of different anomaly detection algorithms and how utilizing feature engineering can improve the results of these outlier detection methods. Mikol Forney, Xuyang Yan, Kishor Datta Gupta, Mahmoud Nabil 0001, Abdollah Homaifar |
IJCNN | 3 |
| 2022 | A Data-driven Approach for Travel Time Prediction and AnalysisabstractRealtime estimation of travel time is a key traffic parameter for designing and planning for transportation systems, particularly when providing mobility-on-demand (MOD) services. However, the analysis and prediction of travel time can be delayed significantly due to the complexity and huge computational requirements of microsimulation models. Thus, as an alternative solution, we propose a data-driven approach for the efficient and reliable prediction of travel time. Our approach takes advantage of the strengths of SVM and ARIMA for fully capturing the traffic patterns in the traffic data. We introduce a new parameter $\kappa$ into the SVM-ARIMA model to adjust the weight of the ARIMA component, which significantly improves the performance. We validate the performance of the proposed approach using data generated from a microsimulation platform. Our experimental results and comparisons with the existing ML-based methods demonstrates the efficacy of the proposed data-driven approach. Benjamin Lartey, Lydia Zeleke, Xuyang Yan, Kishor Datta Gupta, Abdollah Homaifar, Ali Karimoddini |
SMC | 4 |
| 2022 | Interpretable Convolutional Learning Classifier System (C-LCS) for higher dimensional datasetsabstractThe purpose of this paper is to devise an interpretable hybrid classification model for Convolutional Neural Networks (CNN) and a Learning Classifier System (LCS). The presented hybrid system integrates the fundamental attributes from both types of these classifiers. In the proposed hybrid model CNN works as an automatic feature extractor, and LCS works to provide interpretable rule-based classification results. Although LCS has limitations working on higher dimensional datasets, we resolve this limitation by using CNN as a feature extractor. The other concept of the non-interpretability of CNN is addressed by using the LCS rule. Furthermore, our experiment with higher dimensional datasets like CIFAR-10 and Fashion-MNIST shows that extended LCS provides comparable performance to the standard neural network model while also providing interpretable results. We named this extended LCS method Convolutional Learning Classifier Cystem (C-LCS). Jelani Owens, Kishor Datta Gupta, Xuyang Yan, Lydia Asrat Zeleke, Abdollah Homaifar |
SMC | 2 |
| 2022 | A clustering-based active learning method to query informative and representative samples
Xuyang Yan, Shabnam Nazmi, Biniam Gebru, Mohd Anwar, Abdollah Homaifar, Mrinmoy Sarkar, Kishor Datta Gupta |
Appl. Intell. | 7 |
| 2021 | Using Negative Detectors for Identifying Adversarial Data Manipulation in Machine LearningabstractWith the increased popularity of Machine Learning (ML) in real-world applications, adversarial attacks are emerging to subvert the ML-based decision support systems. It appears that the existing adversarial defenses are ineffective against adaptive attacks since these are highly depend on knowledge of prior attacks and the ML model architecture. To alleviate the challenges, We propose a negative filtering strategy that does not require any adversarial knowledge and can work independent of ML models. This filtering strategy relies on salient features of clean (training) data and employs a complementary approach to cover possible attack surface in an application. Our empirical experiments with different data sets demonstrate that the negative filters could effectively detect wide-range of adversarial inputs and update itself to protect against adaptive attacks. Kishor Datta Gupta, Dipankar Dasgupta |
IJCNN | 1 |
| 2020 | An Empirical Study on Algorithmic BiasabstractIn all goal-oriented selection activities, an existence of certain level of bias is unavoidable and may be desired for efficient artificial intelligence based decision support systems. However, a fair independent comparison of all eligible entities is essential to alleviate explicit bias in competitive marketplace. For example, searching online for a good or service, it is expected that the underlying algorithm will provide fair results by searching all available entities in the category mentioned. However, a biased search can make a narrow or collaborative query, ignoring competitive outcomes, resulting customers in costing more or getting lower quality products or services for the money they spend. This paper describes algorithmic bias in different contexts with examples and scenarios, best practices to detect bias, and two case studies to identify algorithmic bias. Sajib Sen, Dipankar Dasgupta, Kishor Datta Gupta |
COMPSAC | 3 |
| 2020 | CIDMP: Completely Interpretable Detection of Malaria Parasite in Red Blood Cells using Lower-dimensional Feature SpaceabstractPredicting if red blood cells (RBC) are infected with the malaria parasite is an important problem in Pathology. Recently, supervised machine learning approaches have been used for this problem, and they have had reasonable success. In particular, state-of-the-art methods such as Convolutional Neural Networks automatically extract increasingly complex feature hierarchies from the image pixels. While such generalized automatic feature extraction methods have significantly reduced the burden of feature engineering in many domains, for niche tasks such as the one we consider in this paper, they result in two major problems. First, they use a very large number of features (that may or may not be relevant) and therefore training such models is computationally expensive. Further, more importantly, the large feature-space makes it very hard to interpret which features are truly important for predictions. Thus, a criticism of such methods is that learning algorithms pose as opaque blackboxes to its users, in this case medical experts. The recommendation of such algorithms can be understood easily, but the reason for their recommendation is not clear. This is the problem of non-interpretability of the model, and the best-performing algorithms are usually the least interpretable. To address these issues, in this paper, we propose an approach to extract a very small number of aggregated features that are easy to interpret and compute, and empirically show that we obtain high prediction accuracy even with a significantly reduced feature-space. Anik Khan, Kishor Datta Gupta, Deepak Venugopal, Nirman Kumar |
IJCNN | 2 |
| 2019 | A Hybrid POW-POS Implementation Against 51 percent Attack in Cryptocurrency SystemabstractThe success and popularity of Bitcoin mainly focuses the underlying blockchain technology which is totally immutable distributed ledger, highly secured by its P2P network consensus named Proof of Work (PoW). One of the worst threats to a Proof-of-Work based cryptocurrency is 51% attack. If one or more dishonest network peer gains more than 50% of resource such as processing power, then they will become the majority decision maker in the network. It is already proved that mixing of two or more existing protocol that is called hybrid protocol can make the network enough resistive to this attack. The recent implementations of hybrid protocols have other limitations and problems that they are facing and striving to resolve. But their main weakness is in distribution of block mining reward to the investors. From the perspective of an investor, an investor invests his hard-earned money in a cryptocurrency for making proper profit from his investment. The main source of this profit is the block reward which is generated and given to the miner on successful mining of a block. So, to ensure this profit is given to proper user on proper time interval, the consistency of block generation time interval is a vital factor. The voting system, ticket system etc. are not time controlled and over all block reward generation interval will not show a uniform distribution of profit. Another big issue is diversifying the peers by creating special committee and groups of validators the concept of P2P network is violated. In this paper we will describe a step by step process to implement a Hybrid PoW-PoS based consensus protocol. In our proposed system, the PoW mining process is only used to regulate the block generation time. The actual block generation is done by the same user with PoS consensus mechanism. There is no voting or validating committee. The entire network will validate each block. This is the major difference with other discussed system. The system will not only be able to tackle the 51% attack, it provides a uniform distribution of mining reward to the stake holders and investors by maintaining a precise block generation interval with difficulty adjustment in PoW mining and probability calculation for stake holders according to their matured staking balance. We will not only show how to make the system non-vulnerable to this attack but also describe in detail about how to validate the transactions and blocks in different stage of creating the block chain. Kishor Datta Gupta, Abdur Rahman Khan Jehad, Subash Poudyal, Mohammad Nurul Huda, M. A. Parvez Mahmud |
CloudCom | 1 |