Arif Ahmed 0002

dblp:154/3197-2 · also Arif Ahmed Sekh, Sk. Arif Ahmed · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-0706-2565ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Exploring gallbladder cancer prognosis using machine learning and explainable AI
abstract
Abstract Gallbladder cancer (GBC) is globally rare but prevalent in India. GBC is often diagnosed at an advanced stage, leading to a poor outcome. The identification of key prognostic factors and accurate survival prediction are crucial to optimizing treatment strategies. Retrospective data from 698 patients with gallbladder cancer treated at Tata Medical Center, Kolkata were obtained from a published dataset. Different machine learning models were used to predict survival outcomes. The performance of the model was evaluated, and explainable AI techniques were used to interpret the model’s output and identify significant prognostic factors. The analysis identified elevated liver enzymes and bilirubin levels as significant prognostic factors. Advanced age has also been shown to be correlated with terminal disease outcomes. When comparing different modeling systems, the stacking model obtained perfect matrix scores. These results confirmed the resilience of the stacking model for survival prediction. Other models also showed high prediction performance. Explainable AI techniques provide details of the relative importance of these prognostic factors. Compared to prevailing models of survival prediction for gallbladder cancer, this one represents a major improvement, correlating a rigorously validated stacking ensemble with comprehensive explainability, as demonstrated by SHAP, PDP, ICE, and LIME. This results in almost perfect discrimination (AUROC = 0.9949) and a significantly boosted interpretability of prognostic factors. The partial dependence plots present the effects of liver and glycemic markers on survival prognosis. SHAP values illustrated that the final stage of the tumor and the absence of surgery are the top negative prognostic indicators. The explainability of SHAP also demonstrated the importance of liver enzymes, bilirubin, age, and albumin levels in predicting survival. Hence, explainable machine learning models help predict more accurately and also support an understanding of the main factors contributing to survival. If they are ever integrated into clinical environments, such models can help personalize treatment strategies and improve care outcomes. All codes and datasets used in this study are available at https://github.com/devnarayan87/GBC_Cancer_Data .
Shantam Srivastava, Ankita Dutta, Debashree Guha, Debnarayan Khatua, Dilip K. Prasad, Arif Ahmed 0002
Discov. Comput.6
2025 Haphazard Inputs as Images in Online Learning
abstract
The field of varying feature space in online learning settings, also known as haphazard inputs, is very prominent nowadays due to its applicability in various fields. However, the current solutions to haphazard inputs are model-dependent and cannot benefit from the existing advanced deep-learning methods, which necessitate inputs of fixed dimensions. Therefore, we propose to transform the varying feature space in an online learning setting to a fixed-dimension image representation on the fly. This simple yet novel approach is model-agnostic, allowing any vision-based models to be applicable for haphazard inputs, as demonstrated using ResNet and ViT. The image representation handles the inconsistent input data seamlessly, making our proposed approach scalable and robust. We show the efficacy of our method on four publicly available datasets. The code is available at https://github.com/Rohit102497/HaphazardInputsAsImages.
Aryan Dessai, Arif Ahmed 0002, Krishna Agarwal, Alexander Horsch, Dilip K. Prasad
IJCNN3
2025 Graph-based hostile content detection in Hindi language
abstract
Abstract Organizations and governments are struggling to handle the hostile content on social media sites ( $$Facebook^{TM}$$ , $$Twitter^{TM}$$ , etc.). While extensive research exists for English-language content, regional languages like Hindi lack robust tools and datasets for effective moderation. This study proposes a scalable AI-based framework for detecting hostile posts in Hindi, the most widely spoken language in the Indian subcontinent and the third most spoken globally. We employ both binary (coarse-grained) and multi-class, multi-label (fine-grained) classification using contextual and semantic features. Our approach integrates various BERT-based embeddings with Relational Graph Convolutional Networks (R-GCN), forming a hybrid BRGCN architecture trained on the Constraint 2021 Hindi dataset. To enhance performance, we implement a hard voting-based ensemble classifier. The proposed model achieves superior F1-scores compared to existing baselines: 0.98 for coarse-grained classification and 0.84, 0.61, 0.49, and 0.64 for the fine-grained categories of Fake, Hate, Defamation, and Offensive, respectively. Code and data will be made publicly available in https://github.com/mani-design/B-RGCN .
Angana Chakraborty, Subhankar Joardar, Dilip K. Prasad, Arif Ahmed 0002
Discov. Comput.4
2025 BangleFIR: bridging the gap in fashion image retrieval with a novel dataset of bangles
Sk Maidul Islam, Subhankar Joardar, Arif Ahmed 0002
Multim. Tools Appl.3
2024 Blend & Predict: Domain-Adaptable Few-Shot Learning for Microscopy Imaging
abstract
Accurate classification of microscopy images is critical for the analysis of biological samples. The availability of large-scale labeled datasets has contributed to recent progress in training large, deep classification models in the medical imaging domain, but methods that cater to a variety of microscopy modalities across a range of biological samples and length scales are scarce. A key reason is that curating labeled data for microscopy images is costly and needs tedious and timeconsuming effort of AI and domain experts. We propose a novel few-shot learning technique, specifically “Blend & Predict” that uses small labeled datasets for training and infers unlabeled datasets. We evaluated the performance and generalizability of our approach using three medical image datasets, each with a different microscopy modality and addressing a different biomedical question on different samples. We achieved results comparable to state-of-the-art models like GoogleNet, VGG16, RestNet50 that used large datasets for training.
Ayush Somani, Arif Ahmed 0002, Krishna Agarwal, Dilip K. Prasad
ICIP3
2024 Customizable and Programmable Deep Learning
Ratnabali Pal, Samarjit Kar, Arif Ahmed 0002
ICPR (1)3
2024 Attention Seekers U-Net with Mamba for Sub-cellular Segmentation
Pratik Sinha, Arif Ahmed 0002
ICPR (12)2
2024 Biomedical term extraction using fuzzy association
Bidyut Das, Mukta Majumder, Santanu Phadikar, Arif Ahmed 0002
Soft Comput.4
2024 Ensemble Classifier for Hindi Hostile Content Detection
abstract
Detection of hostile content from social media posts (Facebook, Twitter, etc.) is a demanding task in the field of Natural Language Processing. The increase of hostile content in different electronic media has opened up new challenges in language understanding. It becomes more difficult in regional languages. AI-based solutions are required to identify hostile content on a large scale. Although a satisfactory amount of research has been carried out in the English language, finding hostile content in regional languages is still under development due to the unavailability of suitable datasets and tools. In terms of the number of speakers, Hindi ranks third in the world and first on the Indian subcontinent. The objective of this article is to design a hostile content detection system in Hindi using coarse-grained (binary) classification and fine-grained (multi-class, multi-label) classification. We note that different baseline learning methods with different pre-trained language models perform differently. Using the Constraint 2021 Hindi Dataset, this research proposes a Bidirectional Encoder Representations from Transformers–(BERT) based contextual embedding technique with a concatenation of emoji2vec embeddings to classify social media posts in Hindi Devanagari script as hostile or non-hostile. Additionally, for the fine-grained tasks where hostile posts are sub-categorized as defamation, fake, hate, and offensive, we develop an ensemble classifier varying different learning methods and embedding models. With an F1-Score of 0.9721, it is found that our proposed Indic-BERT+emoji model outperforms the baseline model and other existing models for the coarse-grained task. We have also observed that our proposed ensemble method provides better results than the existing models and the baseline model for the fine-grained tasks with F1-Scores of 0.43, 0.82, 0.58, and 0.62 for the defamation, fake, hate, and offensive classes, respectively. The code and the data are available at https://github.com/skarifahmed/hostile .
Angana Chakraborty, Subhankar Joardar, Arif Ahmed 0002
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2023 DSSN: dual shallow Siamese network for fashion image retrieval
Sk Maidul Islam, Subhankar Joardar, Arif Ahmed 0002
Multim. Tools Appl.3
2023 Classification of Ayurveda constitution types: a deep learning approach
Debnarayan Khatua, Arif Ahmed 0002, Rintu Kutum, Mitali Mukherji, Bhavana Prasher, Samarjit Kar
Soft Comput.2
2022 Person re-identification in indoor videos by information fusion using Graph Convolutional Networks
Komal Soni, Debi Prosad Dogra, Arif Ahmed 0002, Samarjit Kar, Heeseung Choi, Ig-Jae Kim
Expert Syst. Appl.3
2022 Emotionally charged text classification with deep learning and sentiment semantic
abstract
Abstract Text classification is one of the widely used phenomena in different natural language processing tasks. State-of-the-art text classifiers use the vector space model for extracting features. Recent progress in deep models, recurrent neural networks those preserve the positional relationship among words achieve a higher accuracy. To push text classification accuracy even higher, multi-dimensional document representation, such as vector sequences or matrices combined with document sentiment, should be explored. In this paper, we show that documents can be represented as a sequence of vectors carrying semantic meaning and classified using a recurrent neural network that recognizes long-range relationships. We show that in this representation, additional sentiment vectors can be easily attached as a fully connected layer to the word vectors to further improve classification accuracy. On the UCI sentiment labelled dataset, using the sequence of vectors alone achieved an accuracy of 85.6%, which is better than 80.7% from ridge regression classifier—the best among the classical technique we tested. Additional sentiment information further increases accuracy to 86.3%. On our suicide notes dataset, the best classical technique—the Naíve Bayes Bernoulli classifier, achieves accuracy of 71.3%, while our classifier, incorporating semantic and sentiment information, exceeds that at 75% accuracy.
Jeow Li Huan, Arif Ahmed 0002, Hiok Chai Quek, Dilip K. Prasad
Neural Comput. Appl.2
2021 Multiple-choice question generation with auto-generated distractors for computer-assisted educational assessment
Bidyut Das, Mukta Majumder, Santanu Phadikar, Arif Ahmed 0002
Multim. Tools Appl.4
2021 Can deep learning solve a preschool image understanding problem?
Bidyut Das, Arif Ahmed 0002, Mukta Majumder, Santanu Phadikar
Neural Comput. Appl.2
2021 RS-HeRR: a rough set-based Hebbian rule reduction neuro-fuzzy system
abstract
Abstract Interpretabilty is one of the desired characteristics in various classification task. Rule-based system and fuzzy logic can be used for interpretation in classification. The main drawback of rule-based system is that it may contain large complex rules for classification and sometimes it becomes very difficult in interpretation. Rule reduction is also difficult for various reasons. Removing important rules may effect in classification accuracy. This paper proposes a hybrid fuzzy-rough set approach named RS-HeRR for the generation of effective, interpretable and compact rule set. It combines a powerful rule generation and reduction fuzzy system, called Hebbian-based rule reduction algorithm (HeRR) and a novel rough-set-based attribute selection algorithm for rule reduction. The proposed hybridization leverages upon rule reduction through reduction in partial dependency as well as improvement in system performance to significantly reduce the problem of redundancy in HeRR, even while providing similar or better accuracy. RS-HeRR demonstrates these characteristics repeatedly over four diverse practical classification problems, such as diabetes identification, urban water treatment monitoring, sonar target classification, and detection of ovarian cancer. It also demonstrates excellent performance for highly biased datasets. In addition, it competes very well with established non-fuzzy classifiers and outperforms state-of-the-art methods that use rough sets for rule reduction in fuzzy systems.
Feng Liu 0043, Arif Ahmed 0002, Hiok Chai Quek, Geok See Ng, Dilip K. Prasad
Neural Comput. Appl.2
2020 Learning Nanoscale Motion Patterns of Vesicles in Living Cells
abstract
Detecting and analyzing nanoscale motion patterns of vesicles, smaller than the microscope resolution (~250 nm), inside living biological cells is a challenging problem. State-of-the-art CV approaches based on detection, tracking, optical flow or deep learning perform poorly for this problem. We propose an integrative approach, built upon physics based simulations, nanoscopy algorithms, and shallow residual attention network to make it possible for the first time to analysis sub-resolution motion patterns in vesicles that may also be of sub-resolution diameter. Our results show state-of-the-art performance, 89% validation accuracy on simulated dataset and 82% testing accuracy on an experimental dataset of living heart muscle cells imaged under three different pathological conditions. We demonstrate automated analysis of the motion states and changed in them for over 9000 vesicles. Such analysis will enable large scale biological studies of vesicle transport and interaction in living cells in the future.
Arif Ahmed 0002, Ida Sundvor Opstad, Åsa Birna Birgisdottir, Truls Myrmel, Balpreet Singh Ahluwalia, Krishna Agarwal, Dilip K. Prasad
CVPR1
2020 Person Re-identification in Videos by Analyzing Spatio-temporal Tubes
abstract
Abstract Typical person re-identification frameworks search for k best matches in a gallery of images that are often collected in varying conditions. The gallery usually contains image sequences for video re-identification applications. However, such a process is time consuming as video re-identification involves carrying out the matching process multiple times. In this paper, we propose a new method that extracts spatio-temporal frame sequences or tubes of moving persons and performs the re-identification in quick time. Initially, we apply a binary classifier to remove noisy images from the input query tube. In the next step, we use a key-pose detection-based query minimization technique. Finally, a hierarchical re-identification framework is proposed and used to rank the output tubes. Experiments with publicly available video re-identification datasets reveal that our framework is better than existing methods. It ranks the tubes with an average increase in the CMC accuracy of 6-8% across multiple datasets. Also, our method significantly reduces the number of false positives. A new video re-identification dataset, named Tube-based Re-identification Video Dataset (TRiViD), has been prepared with an aim to help the re-identification research community.
Arif Ahmed 0002, Debi Prosad Dogra, Heeseung Choi, Seungho Chae, Ig-Jae Kim
Multim. Tools Appl.1
2020 Can we automate diagrammatic reasoning?
abstract
Diagrammatic reasoning (DR) problems are well known. However, solving DR problems represented in 4 × 1 Raven’s Progressive Matrix (RPM) form using computer vision and pattern recognition has not yet been tried. Emergence of deep learning techniques aided by advanced computing can be exploited to solve such DR problems. In this paper, we propose a new learning framework by combining LSTM and Convolutional LSTM to solve 4 × 1 DR problems. Initially, the elementary geometrical shapes in such problems are detected using a typical CNN-based detector. Next, relations of various shapes are analyzed and a high-level feature set is produced and processed in the LSTM framework. A new 4 × 1 DR dataset has been prepared and made available to the research community. We believe, it will be helpful in advancing this research further. We have compared our method with some of the existing frameworks that can be used for solving RPM-guided DR problems. We have recorded 18–20% increase in the average prediction accuracy as compared to the prior frameworks when applied to RPM-guided DR problems. We believe the CV research community will be interested to carry out similar research, particularly to investigate the feasibility of solving other types of known DR problems.
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001, Dilip K. Prasad
Pattern Recognit.1
2020 Video trajectory analysis using unsupervised clustering and multi-criteria ranking
abstract
Abstract Surveillance camera usage has increased significantly for visual surveillance. Manual analysis of large video data recorded by cameras may not be feasible on a larger scale. In various applications, deep learning-guided supervised systems are used to track and identify unusual patterns. However, such systems depend on learning which may not be possible. Unsupervised methods relay on suitable features and demand cluster analysis by experts. In this paper, we propose an unsupervised trajectory clustering method referred to as t-Cluster. Our proposed method prepares indexes of object trajectories by fusing high-level interpretable features such as origin, destination, path, and deviation. Next, the clusters are fused using multi-criteria decision making and trajectories are ranked accordingly. The method is able to place abnormal patterns on the top of the list. We have evaluated our algorithm and compared it against competent baseline trajectory clustering methods applied to videos taken from publicly available benchmark datasets. We have obtained higher clustering accuracies on public datasets with significantly lesser computation overhead.
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001
Soft Comput.1
2020 Query-Based Video Synopsis for Intelligent Traffic Monitoring Applications
abstract
Synopsis of a long-duration video has many applications in intelligent transportation systems. It can help to monitor traffic with lesser manpower. However, generating meaningful synopsis of a long-duration video recording can be challenging. Often summarized outputs include redundant contents or activities that may not be helpful to the observer. Moving object trajectories are possible sources of information that can be used to generate the synopsis of long-duration videos. The synopsis generation faces challenges due to object tracking, grouping of the trajectories with respect to activity type, object category, and contextual information, and generating smooth synopsis according to a query. In this paper, we propose a method to generate meaningful and smooth synopsis of long-duration videos according to the users' query. We have tracked moving objects and adopted deep learning to classify the objects into known categories (e.g., car, bike, and pedestrians). We then identify regions in the surveillance scene with the help of unsupervised clustering. Each tube (spatiotemporal object trajectory) is represented by the source and the destination. In the final stage, we take a query from the user and generate the synopsis video by smoothly blending the appropriate tubes over the background frame through energy minimization. The proposed method has been evaluated on two publicly available datasets and our own surveillance datasets. We have compared the method with popular state-of-the-art techniques. The experiments reveal that the proposed method is superior to the existing techniques and it produces visually seamless video synopsis.
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Renuka Patnaik, Seung-Cheol Lee, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim
IEEE Trans. Intell. Transp. Syst.1
2019 Fingertip detection and tracking for recognition of air-writing in videos
Sohom Mukherjee, Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001
Expert Syst. Appl.2
2019 Trajectory-Based Surveillance Analysis: A Survey
abstract
Due to the advancement of camera hardware and machine learning techniques, video object tracking for surveillance has received noticeable attention from the computer vision research community. Object tracking and trajectory modeling have important applications in surveillance video analysis. For example, trajectory clustering, summarization or synopsis generation, and detection of anomalous or abnormal events in videos are mainly being exploited by the research community. However, barring one research work (which is almost a decade old), there is no recent review that emphasizes the use of video object trajectories, particularly in the perspective of visual surveillance. This paper presents a survey of trajectory-based surveillance applications with a focus on clustering, anomaly detection, summarization, and synopsis generation. The methods reviewed in this paper broadly summarize the abovementioned applications. The main purpose of this survey is to summarize the state-of-the-art video object trajectory analysis techniques used in the indoor and outdoor surveillance.
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Surveillance scene representation and trajectory abnormality detection using aggregation of multiple concepts
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001
Expert Syst. Appl.1
2018 Unsupervised classification of erroneous video object trajectories
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001
Soft Comput.1
2017 Localization of region of interest in surveillance scene
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Byung-Gyu Kim, Paul R. Hill, Harish Bhaskar
Multim. Tools Appl.1
2016 Smart video summarization using mealy machine-based trajectory modelling for surveillance applications
Debi Prosad Dogra, Arif Ahmed 0002, Harish Bhaskar
Multim. Tools Appl.2