Dilip K. Prasad

dblp:93/8841 · also Dilip Kumar Prasad · DBLP profile ↗
← Back
38ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0002-3693-6973ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Exploring gallbladder cancer prognosis using machine learning and explainable AI
abstract
Abstract Gallbladder cancer (GBC) is globally rare but prevalent in India. GBC is often diagnosed at an advanced stage, leading to a poor outcome. The identification of key prognostic factors and accurate survival prediction are crucial to optimizing treatment strategies. Retrospective data from 698 patients with gallbladder cancer treated at Tata Medical Center, Kolkata were obtained from a published dataset. Different machine learning models were used to predict survival outcomes. The performance of the model was evaluated, and explainable AI techniques were used to interpret the model’s output and identify significant prognostic factors. The analysis identified elevated liver enzymes and bilirubin levels as significant prognostic factors. Advanced age has also been shown to be correlated with terminal disease outcomes. When comparing different modeling systems, the stacking model obtained perfect matrix scores. These results confirmed the resilience of the stacking model for survival prediction. Other models also showed high prediction performance. Explainable AI techniques provide details of the relative importance of these prognostic factors. Compared to prevailing models of survival prediction for gallbladder cancer, this one represents a major improvement, correlating a rigorously validated stacking ensemble with comprehensive explainability, as demonstrated by SHAP, PDP, ICE, and LIME. This results in almost perfect discrimination (AUROC = 0.9949) and a significantly boosted interpretability of prognostic factors. The partial dependence plots present the effects of liver and glycemic markers on survival prognosis. SHAP values illustrated that the final stage of the tumor and the absence of surgery are the top negative prognostic indicators. The explainability of SHAP also demonstrated the importance of liver enzymes, bilirubin, age, and albumin levels in predicting survival. Hence, explainable machine learning models help predict more accurately and also support an understanding of the main factors contributing to survival. If they are ever integrated into clinical environments, such models can help personalize treatment strategies and improve care outcomes. All codes and datasets used in this study are available at https://github.com/devnarayan87/GBC_Cancer_Data .
Shantam Srivastava, Ankita Dutta, Debashree Guha, Debnarayan Khatua, Dilip K. Prasad, Arif Ahmed 0002
Discov. Comput.5
2025 Hedging Is Not All You Need: A Simple Baseline for Online Learning Under Haphazard Inputs
abstract
Handling haphazard streaming data, such as data from edge devices, presents a challenging problem. Over time, the incoming data becomes inconsistent, with missing, faulty, or new inputs reappearing. Therefore, it requires models that are reliable. Recent methods to solve this problem depend on a hedging-based solution and require specialized elements like auxiliary dropouts, forked architectures, and intricate network design. We observed that hedging can be reduced to a special case of weighted residual connection; this motivated us to approximate it with plain self-attention. In this work, we propose HapNet, a simple baseline that is scalable, does not require online backpropagation, and is adaptable to varying input types. All present methods are restricted to scaling with a fixed window; however, we introduce a more complex problem of scaling with a variable window where the data becomes positionally uncorrelated, and cannot be addressed by present methods. We demonstrate that a variant of the proposed approach can work even for this complex scenario. We extensively evaluated the proposed approach on five benchmarks and found competitive performance. https://bit.ly/Online-Learning-Transformers.
Himanshu Buckchash, Momojit Biswas, Dilip K. Prasad
ICASSP4
2025 Haphazard Inputs as Images in Online Learning
abstract
The field of varying feature space in online learning settings, also known as haphazard inputs, is very prominent nowadays due to its applicability in various fields. However, the current solutions to haphazard inputs are model-dependent and cannot benefit from the existing advanced deep-learning methods, which necessitate inputs of fixed dimensions. Therefore, we propose to transform the varying feature space in an online learning setting to a fixed-dimension image representation on the fly. This simple yet novel approach is model-agnostic, allowing any vision-based models to be applicable for haphazard inputs, as demonstrated using ResNet and ViT. The image representation handles the inconsistent input data seamlessly, making our proposed approach scalable and robust. We show the efficacy of our method on four publicly available datasets. The code is available at https://github.com/Rohit102497/HaphazardInputsAsImages.
Aryan Dessai, Arif Ahmed 0002, Krishna Agarwal, Alexander Horsch, Dilip K. Prasad
IJCNN6
2025 Compact representation for memory-efficient storage of images using genetic algorithm-guided key pixel selection
abstract
In the past few years, we have observed rapid growth in digital content. Even in the biological domain, the arrival of microscopic and nanoscopic images and videos captured for biological investigations increases the need for space to store them. Hence, storing these data in a storage-efficient manner is a pressing need. In this work, we have introduced a compact image representation technique with an eye on preserving the shape that can shrink the memory requirement to store. The compact image representation is different from image compression since it does not include any encoding mechanism. Rather, the idea is that this mechanism stores the positions of key pixels, and when required, the original image can be regenerated. The genetic algorithm is used to select key pixels, while the Gaussian kernel performs the reconstruction task with the help of the positions of the selected key pixels. The model is tested on four different datasets. The proposed technique shrinks the memory requirement by 87% to 98% while evaluated using the bit reduction rate. However, the reconstructed images’ quality is a bit low when evaluated using metrics like structural similarity index (ranges between 0.81 to 0.94), or root means squared error (ranges between 0.06 to 0.08). To investigate the impact of quality reduction in reconstructed images in real-life applications, we performed image classification using reconstructed samples and found 0.13% to 2.30% classification accuracy reduction compared to when classification is done using original samples. The proposed model’s performance is comparable to state-of-the-art’s similar solutions.
Samir Malakar, Nirwan Banerjee, Dilip K. Prasad
Eng. Appl. Artif. Intell.3
2025 Graph-based hostile content detection in Hindi language
abstract
Abstract Organizations and governments are struggling to handle the hostile content on social media sites ( $$Facebook^{TM}$$ , $$Twitter^{TM}$$ , etc.). While extensive research exists for English-language content, regional languages like Hindi lack robust tools and datasets for effective moderation. This study proposes a scalable AI-based framework for detecting hostile posts in Hindi, the most widely spoken language in the Indian subcontinent and the third most spoken globally. We employ both binary (coarse-grained) and multi-class, multi-label (fine-grained) classification using contextual and semantic features. Our approach integrates various BERT-based embeddings with Relational Graph Convolutional Networks (R-GCN), forming a hybrid BRGCN architecture trained on the Constraint 2021 Hindi dataset. To enhance performance, we implement a hard voting-based ensemble classifier. The proposed model achieves superior F1-scores compared to existing baselines: 0.98 for coarse-grained classification and 0.84, 0.61, 0.49, and 0.64 for the fine-grained categories of Fake, Hate, Defamation, and Offensive, respectively. Code and data will be made publicly available in https://github.com/mani-design/B-RGCN .
Angana Chakraborty, Subhankar Joardar, Dilip K. Prasad, Arif Ahmed 0002
Discov. Comput.3
2024 Blend & Predict: Domain-Adaptable Few-Shot Learning for Microscopy Imaging
abstract
Accurate classification of microscopy images is critical for the analysis of biological samples. The availability of large-scale labeled datasets has contributed to recent progress in training large, deep classification models in the medical imaging domain, but methods that cater to a variety of microscopy modalities across a range of biological samples and length scales are scarce. A key reason is that curating labeled data for microscopy images is costly and needs tedious and timeconsuming effort of AI and domain experts. We propose a novel few-shot learning technique, specifically “Blend & Predict” that uses small labeled datasets for training and infers unlabeled datasets. We evaluated the performance and generalizability of our approach using three medical image datasets, each with a different microscopy modality and addressing a different biomedical question on different samples. We achieved results comparable to state-of-the-art models like GoogleNet, VGG16, RestNet50 that used large datasets for training.
Ayush Somani, Arif Ahmed 0002, Krishna Agarwal, Dilip K. Prasad
ICIP5
2024 Partial Binarization of Neural Networks for Budget-Aware Efficient Learning
abstract
Binarization is a powerful compression technique for neural networks, significantly reducing FLOPs, but often results in a significant drop in model performance. To address this issue, partial binarization techniques have been developed, but a systematic approach to mixing binary and full-precision parameters in a single network is still lacking. In this paper, we propose a controlled approach to partial binarization, creating a budgeted binary neural network (B2NN) with our MixBin strategy. This method optimizes the mixing of binary and full-precision components, allowing for explicit selection of the fraction of the network to remain binary. Our experiments show that B2NNs created using MixBin outperform those from random or iterative searches and state-of-the-art layer selection methods by up to 3% on the ImageNet-1K dataset. We also show that B2NNs outperform the structured pruning baseline by approximately 23% at the extreme FLOP budget of 15%, and perform well in object tracking, with up to a 12.4% relative improvement over other baselines. Additionally, we demonstrate that B2NNs developed by MixBin can be transferred across datasets, with some cases showing improved performance over directly applying MixBin on the downstream data.1
Udbhav Bamba, Neeraj Anand, Saksham Aggarwal, Dilip K. Prasad
WACV4
2024 pNNCLR: Stochastic pseudo neighborhoods for contrastive learning based unsupervised representation learning problems
abstract
Nearest neighbor (NN) sampling provides more semantic variations than predefined transformations for self-supervised learning (SSL) based image recognition problems. However, its performance is restricted by the quality of the support set, which holds positive samples for the contrastive loss. In this work, we show that the quality of the support set plays a crucial role in any nearest neighbor based method for SSL. We then provide a refined baseline (pNNCLR) to the nearest neighbor based SSL approach (NNCLR). To this end, we introduce pseudo nearest neighbors (pNN) to control the quality of the support set, wherein, rather than sampling the nearest neighbors, we sample in the vicinity of hard nearest neighbors by varying the magnitude of the resultant vector and employing a stochastic sampling strategy to improve the performance. Additionally, to stabilize the effects of uncertainty in NN-based learning, we employ a smooth-weight-update approach for training the proposed network. Evaluation of the proposed method on multiple public image recognition and medical image recognition datasets shows that it performs up to 8 percent better than the baseline nearest neighbor method, and is comparable to other previously proposed SSL methods. The code is available at https://github.com/mb16biswas/pnnclr.
Momojit Biswas, Himanshu Buckchash, Dilip K. Prasad
Neurocomputing3
2023 Mabnet: Master Assistant Buddy Network With Hybrid Learning for Image Retrieval
abstract
Image retrieval has garnered a growing interest in recent times. The current approaches are either supervised or self-supervised. These methods do not exploit the benefits of hybrid learning using both supervision and self-supervision. We present a novel Master Assistant Buddy Network (MAB-Net) for image retrieval which incorporates both the learning mechanisms. MABNet consists of master and assistant block, both learning independently through supervision and collectively via self-supervision. The master guides the assistant by providing its knowledge base as a reference for self-supervision and the assistant reports its knowledge back to the master by weight transfer. We perform extensive experiments on the public datasets with and without post-processing.
Gyanendra Das, Saksham Aggarwal, Alexander Horsch, Dilip K. Prasad
ICASSP5
2023 On Designing Light-Weight Object Trackers Through Network Pruning: Use CNNS or Transformers?
abstract
Object trackers deployed on low-power devices need to be light-weight, however, most of the current state-of-the-art (SOTA) methods rely on using compute-heavy backbones built using CNNs or Transformers. Large sizes of such models do not allow their deployment in low-power conditions and designing compressed variants of large tracking models is of great importance. This paper demonstrates how highly compressed light-weight object trackers can be designed using neural architectural pruning of large CNN and Transformer based trackers. Further, a comparative study on architectural choices best suited to design lightweight trackers is provided. A comparison between SOTA trackers using CNNs, Transformers as well as the combination of the two is presented to study their stability at various compression ratios. Finally results for extreme pruning scenarios going as low as 1% in some cases are shown to study the limits of network pruning in object tracking. This work provides deeper insights into designing highly efficient trackers from existing SOTA methods.1
Saksham Aggarwal, Taneesh Gupta, Pawan Kumar Sahu, Arnav Chavan, Rishabh Tiwari, Dilip K. Prasad
ICASSP6
2022 Mapping Functional Changes in the Embryonic Heart of Atlantic Salmon Post Viral Infection Using AI Technique
abstract
The problem of analysing videos of beating embryonic salmon hearts is addressed here with the aim of classifying infected and healthy embryonic salmon hearts. Determination of infection in a very early stage with simple means can reduce the cost of Salmon farming greatly and can boost the economy of countries dependent on Atlantic salmon fish farming. We converted data-dense videos into physically relevant single variable time-signal, namely area of heart as a function of time. This helped us derive simple physical attributes like spectrograms to describe the heart function in a more informative way. We used these spectrogram images with ensemble of deep learning algorithms to show successful classification of infected and healthy salmon hearts, with an average accuracy of 82%. We are unaware of such studies where simple but powerful physically relevant features are used for achieving this accuracy for videos with unconventional dynamics, such as beating heart.
A. Malachowski, Jaya Kumari Swain, R. A. Dalmo, Alexander Horsch, Dilip K. Prasad
ICIP6
2022 Auxiliary Network: Scalable and Agile Online Learning for Dynamic System with Inconsistently Available Inputs
Krishna Agarwal, Alexander Horsch, Dilip K. Prasad
ICONIP (1)4
2022 Learning-based Ellipse Detection for Robotic Grasps of Cylinders and Ellipsoids
abstract
In our daily life, there are many objects represented by cylindrical shapes and ellipsoids. The tops of these objects are formed by elliptic shape primitives. Thus, it is available for a robot to manipulate these objects by ellipse detection. In this work, we propose a novel approach to generating ground truth for training the model based on domain randomization. Using synthetic data generated in this manner, we build an end-to-end deep neural network with a detection backbone and then, combine multiple branches archived from the backbone for sharing the multiple-scale features; further, after employing active rotation filters, the features pass through the region proposal net to form the prediction branches of the box, orientation regression, and object classification; finally, these branches are fused to do ellipse detection, allowing robotic manipulations of cylinders and ellipsoids. To demonstrate the capabilities of the proposed detector, we show the comparison results with the state-of-the-art detector on synthetic and public datasets. The proposed model for ellipse detection and data generation pipeline based on domain randomization in a simulation are evaluated by a series of robotic manipulations implemented in real application scenarios. The results illustrate a high success rate on real-world grasp attempts despite having only been trained on a synthetic dataset. (A video of some robotic experiments is available on YouTube: https://youtu.be/Ueg1XSI2S98).
Huixu Dong, Jiadong Zhou, Dilip K. Prasad, I-Ming Chen 0001
ICRA4
2022 Emotionally charged text classification with deep learning and sentiment semantic
abstract
Abstract Text classification is one of the widely used phenomena in different natural language processing tasks. State-of-the-art text classifiers use the vector space model for extracting features. Recent progress in deep models, recurrent neural networks those preserve the positional relationship among words achieve a higher accuracy. To push text classification accuracy even higher, multi-dimensional document representation, such as vector sequences or matrices combined with document sentiment, should be explored. In this paper, we show that documents can be represented as a sequence of vectors carrying semantic meaning and classified using a recurrent neural network that recognizes long-range relationships. We show that in this representation, additional sentiment vectors can be easily attached as a fully connected layer to the word vectors to further improve classification accuracy. On the UCI sentiment labelled dataset, using the sequence of vectors alone achieved an accuracy of 85.6%, which is better than 80.7% from ridge regression classifier—the best among the classical technique we tested. Additional sentiment information further increases accuracy to 86.3%. On our suicide notes dataset, the best classical technique—the Naíve Bayes Bernoulli classifier, achieves accuracy of 71.3%, while our classifier, incorporating semantic and sentiment information, exceeds that at 75% accuracy.
Jeow Li Huan, Arif Ahmed 0002, Hiok Chai Quek, Dilip K. Prasad
Neural Comput. Appl.4
2021 Motivation detection using EEG signal analysis by residual-in-residual convolutional neural network
abstract
While we know that motivated students learn better than non-motivated students but detecting motivation is challenging. Here we present a game-based motivation detection approach from the EEG signals. We take an original approach of using EEG-based brain computer interface to assess if motivation state is manifest in physiological EEG signals as well, and what are suitable conditions in order to achieve the goal? To the best of our knowledge, detection of motivation level from brain signals is proposed for the first time in this paper. In order to resolve the central obstacle of small EEG datasets containing deep features, we propose a novel and unique ‘residual-in-residual architecture of convolutional neural network (RRCNN)’ that is capable of reducing the problem of over-fitting on small datasets and vanishing gradient. Having accomplished this, several aspects of using EEG signals for motivation detection are considered, including channel selection and accuracy obtained using alpha or beta waves of EEG signals. We also include a detailed validation of the different aspects of our methodology, including detailed comparison with other works as relevant. Our approach achieves 89% accuracy in using EEG signals to detect motivation state while learning, where alpha wave signals of frontal asymmetry channels are employed. A more robust (less sensitive to learning conditions) 88% accuracy is achieved using beta waves signals of frontal asymmetry channels. The results clearly indicate the potential of detecting motivation states using EEG signals, provided suitable methodologies such as proposed in this paper, are employed.
Soham Chattopadhyay, Laila Zary, Hiok Chai Quek, Dilip K. Prasad
Expert Syst. Appl.4
2021 RS-HeRR: a rough set-based Hebbian rule reduction neuro-fuzzy system
abstract
Abstract Interpretabilty is one of the desired characteristics in various classification task. Rule-based system and fuzzy logic can be used for interpretation in classification. The main drawback of rule-based system is that it may contain large complex rules for classification and sometimes it becomes very difficult in interpretation. Rule reduction is also difficult for various reasons. Removing important rules may effect in classification accuracy. This paper proposes a hybrid fuzzy-rough set approach named RS-HeRR for the generation of effective, interpretable and compact rule set. It combines a powerful rule generation and reduction fuzzy system, called Hebbian-based rule reduction algorithm (HeRR) and a novel rough-set-based attribute selection algorithm for rule reduction. The proposed hybridization leverages upon rule reduction through reduction in partial dependency as well as improvement in system performance to significantly reduce the problem of redundancy in HeRR, even while providing similar or better accuracy. RS-HeRR demonstrates these characteristics repeatedly over four diverse practical classification problems, such as diabetes identification, urban water treatment monitoring, sonar target classification, and detection of ovarian cancer. It also demonstrates excellent performance for highly biased datasets. In addition, it competes very well with established non-fuzzy classifiers and outperforms state-of-the-art methods that use rough sets for rule reduction in fuzzy systems.
Feng Liu 0043, Arif Ahmed 0002, Hiok Chai Quek, Geok See Ng, Dilip K. Prasad
Neural Comput. Appl.5
2021 Object Pose Estimation via Pruned Hough Forest With Combined Split Schemes for Robotic Grasp
abstract
Robotic grasp in complex open-world scenarios requires an effective and generalizable perception. Estimating object’s pose is needed in a variety of practical grasping scenarios. Here we present a novel approach of pose estimation of textureless and textured objects. The algorithm utilizes a single RGB-D image to exploit depth invariant, oriented point pair feature as well as local contextual sensitivity in cluttered environments. To enhance the performance of the voting process and improve learning efficiency, we employ a global pruning algorithm that reduces the risk of overfitting and simplifies the structure of decision trees after compensating for the complementary information among multiple trees by optimizing a designed global objective function. Finally, we also refine the pose obtained from the above stage. The proposed approach of estimating 6-D (degree of freedom) poses of textured and textureless objects is evaluated on publicly available data sets against the recent works under various conditions. It illustrates that our framework is superior to these recent works. Further, we perform extensive qualitative experiments of robotic grasp to illustrate the proposed approach can be applied to practical scenarios.Note to Practitioners—This article is motivated by the problem of the pose estimation of textured and textureless objects in clutter environments. It is difficult for conventional works to address the issue of estimating textured or textureless objects’ poses in such scenarios. We considered that a novel system should be able to obtain the 6-D poses of objects. Therefore, we investigate the combined use of multiple split functions with different characteristics. Learning the model based on Hough forests always cost much computational resource; therefore, we construct a novel pruned Hough forest for solving this issue. Through the comparison and robotic grasp verifications, the behavior of our system can be used in practical applications. In future, we will deploy the proposed system in robotic assembling tasks.
Huixu Dong, Dilip K. Prasad, I-Ming Chen 0001
IEEE Trans Autom. Sci. Eng.2
2020 Temporal Motionless Analysis of Video using CNN in MPSoC
abstract
This paper proposes a novel human-inspired methodology called IRON-MAN (Integrated RatiONal prediction and Motionless ANalysis of videos) on mobile multi-processor systems-on-chips (MPSoCs). The methodology integrates analysis of the previous image frames of the video to represent the analysis of the current frame in order to perform Temporal Motionless Analysis of the Video (TMAV). This is the first work on TMAV using Convolutional Neural Network (CNN) for scene prediction in MPSoCs. Experimental results show that our methodology outperforms state-of-the-art. We also introduce a metric named, Energy Consumption per Training Image (ECTI) to assess the suitability of using a CNN model in mobile MPSoCs with a focus on energy consumption of the device.
Somdip Dey, Amit Kumar Singh 0002, Dilip K. Prasad, Klaus D. McDonald-Maier
ASAP3
2020 Learning Nanoscale Motion Patterns of Vesicles in Living Cells
abstract
Detecting and analyzing nanoscale motion patterns of vesicles, smaller than the microscope resolution (~250 nm), inside living biological cells is a challenging problem. State-of-the-art CV approaches based on detection, tracking, optical flow or deep learning perform poorly for this problem. We propose an integrative approach, built upon physics based simulations, nanoscopy algorithms, and shallow residual attention network to make it possible for the first time to analysis sub-resolution motion patterns in vesicles that may also be of sub-resolution diameter. Our results show state-of-the-art performance, 89% validation accuracy on simulated dataset and 82% testing accuracy on an experimental dataset of living heart muscle cells imaged under three different pathological conditions. We demonstrate automated analysis of the motion states and changed in them for over 9000 vesicles. Such analysis will enable large scale biological studies of vesicle transport and interaction in living cells in the future.
Arif Ahmed 0002, Ida Sundvor Opstad, Åsa Birna Birgisdottir, Truls Myrmel, Balpreet Singh Ahluwalia, Krishna Agarwal, Dilip K. Prasad
CVPR7
2020 GEMM-eMFIS (FRI/E): A Novel General Episodic Memory Mechanism For Fuzzy Neural Networks
abstract
In fields such as finance, medicine, engineering, and science, making real-time predictions during transient periods characterized by sudden and large changes is a hard challenge for machine learning. Humans keep memory of these transient events, abstractly learn the most relevant rules and reuse them when similar events occur, which stems from episodic memory that allows storage and recall of similar events. This paper proposes a novel online general episodic memory mechanism (GEMM) and demonstrates its integration into the Neuro-Fuzzy system (NFS) architecture called evolving Mamdani Fuzzy Inference System (eMFIS) with Fuzzy Rule Interpolation and Extrapolation (FRI/E). Our proposition, called GEMM-eMFIS(FRI/E), learns from past events by storing and retrieving them from an episodic memory cache during event-driven transient behavior, thereby boosting performance while using a few rules only. GEMM-eMFIS(FRI/E) further has several in-built mechanisms that enable it to learn effectively from continuous stream of online data. They include associative-dissociative learning theory to keep its rule base updated, 2-stage incremental clustering; (2SIC) to determine cluster width, interpolation and extrapolation of rules to deal with concept shifts and drifts in the time-variant data, and rule pruning and merging to keep the rule base compact. GEMM-eMFIS (FRI/E) is benchmarked against other NFS' on various time-variant datasets such as stock index prices and rainfall runoff with 3%-5% improvement during transient period and shows strong forecasting performances with 4%-5% more interpretability with lesser rules.
Sheng Wei Pang, Hiok Chai Quek, Dilip K. Prasad
IJCNN3
2020 IT2-GSETSK: An evolving interval Type-II TSK fuzzy neural system for online modeling of noisy data
Mohammad Ashrafi, Dilip K. Prasad, Hiok Chai Quek
Neurocomputing2
2020 Can we automate diagrammatic reasoning?
abstract
Diagrammatic reasoning (DR) problems are well known. However, solving DR problems represented in 4 × 1 Raven’s Progressive Matrix (RPM) form using computer vision and pattern recognition has not yet been tried. Emergence of deep learning techniques aided by advanced computing can be exploited to solve such DR problems. In this paper, we propose a new learning framework by combining LSTM and Convolutional LSTM to solve 4 × 1 DR problems. Initially, the elementary geometrical shapes in such problems are detected using a typical CNN-based detector. Next, relations of various shapes are analyzed and a high-level feature set is produced and processed in the LSTM framework. A new 4 × 1 DR dataset has been prepared and made available to the research community. We believe, it will be helpful in advancing this research further. We have compared our method with some of the existing frameworks that can be used for solving RPM-guided DR problems. We have recorded 18–20% increase in the average prediction accuracy as compared to the prior frameworks when applied to RPM-guided DR problems. We believe the CV research community will be interested to carry out similar research, particularly to investigate the feasibility of solving other types of known DR problems.
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001, Dilip K. Prasad
Pattern Recognit.5
2020 Are Object Detection Assessment Criteria Ready for Maritime Computer Vision?
abstract
Maritime vessels equipped with visible and infrared cameras can complement other conventional sensors for object detection. However, application of computer vision techniques in maritime domain received attention only recently. The maritime environment offers its own unique requirements and challenges. Assessment of the quality of detections is a fundamental need in computer vision. However, the conventional assessment metrics suitable for usual object detection are deficient in the maritime setting. Thus, a large body of related work in computer vision appears inapplicable to the maritime setting at the first sight. We discuss the problem of defining assessment metrics suitable for maritime computer vision. We consider new bottom edge proximity metrics as assessment metrics for maritime computer vision. These metrics indicate that existing computer vision approaches are indeed promising for maritime computer vision and can play a foundational role in the emerging field of maritime computer vision.
Dilip K. Prasad, Huixu Dong, Deepu Rajan, Hiok Chai Quek
IEEE Trans. Intell. Transp. Syst.1
2019 Object Detection in a Maritime Environment: Performance Evaluation of Background Subtraction Methods
abstract
This paper provides a benchmark of the performance of 23 classical and state-of-the-art background subtraction (BS) algorithms on visible range and near infrared range videos in the Singapore Maritime dataset. Importantly, our study indicates the limitations of the conventional performance evaluation criteria for maritime vision and proposes new performance evaluation criteria that is better suited to this problem. This paper provides insight into the specific challenges of BS in maritime vision. We identify four open challenges that plague BS methods in maritime scenario. These include spurious dynamics of water, wakes, ghost effect, and multiple detections. Poor recall and extremely poor precision of all the 23 methods, which have been otherwise successful for other challenging BS situations, allude to the need for new BS methods custom designed for maritime vision.
Dilip K. Prasad, Chandrashekar Krishna Prasath, Deepu Rajan, Lily Rachmawati, Eshan Rajabally, Hiok Chai Quek
IEEE Trans. Intell. Transp. Syst.1
2019 Real-Time Robotic Manipulation of Cylindrical Objects in Dynamic Scenarios Through Elliptic Shape Primitives
abstract
Robotic manipulation employs the object detection in images to create a scene awareness and locate an object's pose. In dynamic scenarios, fast multiobject detection and tracking are crucial. Many objects commonly found in household and industrial environments are represented by cylindrical shapes. Thus, it is available for robots to manipulate them through the real-time detection of elliptic shape primitives formed by the circular tops of these objects. We devise an efficient algorithm of the detection of elliptic shape primitives, which in turn enables robust and real-time robotic manipulations of such objects. The proposed algorithm incorporates the information of elliptic edge curvature, splits complex curves into arcs, classifies the arcs into different quadrants of a candidate elliptic shape, determines the quality of arc selection for ellipse fitting, and then retrieves the corresponding elliptic shape primitive. Our algorithm provides either faster or more accurate ellipse detection results than the current state-of-the-art methods, irrespective of challenging scenarios such as occluded or overlapping ellipses. This is verified by performance comparison with six state-of-the-art elliptic shape detection algorithms on four public image datasets. The algorithm has been integrated on robots to demonstrate the ability to carry out accurate robotic manipulations (tracking, grasping, and stacking) of cylindrical objects in real time. We show that the robotic manipulator, empowered by the elliptic shape primitive algorithm, performs well in complex manipulation experiments as well as dynamic scenarios.
Huixu Dong, Ehsan Asadi, Guangbin Sun, Dilip K. Prasad, I-Ming Chen 0001
IEEE Trans. Robotics4
2018 Efficient Pose Estimation from Single RGB-D Image via Hough Forest with Auto-Context
abstract
We propose a high efficient learning approach to estimating 6D (Degree of Freedom) pose of the textured or texture-less objects for grasping purposes in a cluttered environment where the objects might be partially occluded. The method comprises three main steps. Given a single RGB-D image, we first deploy appropriate features and the random forest to deduce the object class probability and cast votes for the 6D pose in Hough space by joint regression and classification framework, adopting reservoir sampling and summarizing the pose distribution by clustering. Next, we integrate the auto-context into cascaded Hough forests to improve the efficiency of learning. Extensive experiments on various public datasets and robotic grasps indicate that our method presents some improvements over the state-of-art and reveals the capability for estimating poses in practical applications efficiently.
Huixu Dong, Dilip K. Prasad, Qilong Yuan, Jiadong Zhou, Ehsan Asadi, I-Ming Chen 0001
IROS2
2018 PIE-RSPOP: A brain-inspired pseudo-incremental ensemble rough set pseudo-outer product fuzzy neural network
Aparna Ramesh Iyer, Dilip K. Prasad, Hiok Chai Quek
Expert Syst. Appl.2
2018 Sentiment analysis using EEG activities for suicidology
Dilip K. Prasad, Shen-Hsing Annabel Chen, Hiok Chai Quek
Expert Syst. Appl.1
2018 Accurate detection of ellipses with false detection control at video rates using a gradient analysis
Huixu Dong, Dilip K. Prasad, I-Ming Chen 0001
Pattern Recognit.2
2017 Robust ellipse detection via arc segmentation and classification
abstract
In this paper, we propose a novel ellipse detection algorithm for synthetic and real images. Existing ellipse detection methods are too slow when used with limited hardware resources. The proposed method demonstrates the capability of detecting ellipses with an excellent accuracy at an acceptable speed level in three public datasets. The excellent performance is attributed to the novel combination of classification of arcs into different quadrants of a candidate ellipse, edge curvature and convexity-concavity analysis, and an elliptic geometry constraint.
Huixu Dong, I-Ming Chen 0001, Dilip K. Prasad
ICIP3
2017 Video Processing From Electro-Optical Sensors for Object Detection and Tracking in a Maritime Environment: A Survey
abstract
We present a survey on maritime object detection and tracking approaches, which are essential for the development of a navigational system for autonomous ships. The electro-optical (EO) sensor considered here is a video camera that operates in the visible or the infrared spectra, which conventionally complements radar and sonar for situational awareness at sea and has demonstrated its effectiveness over the last few years. This paper provides a comprehensive overview of various approaches of video processing for object detection and tracking in the maritime environment. We follow an approach-based taxonomy wherein the advantages and limitations of each approach are compared. The object detection system consists of the following modules: horizon detection, static background subtraction, and foreground segmentation. Each of these has been studied extensively in maritime situations and has been shown to be challenging due to the presence of background motion especially due to waves and wakes. The key processes involved in object tracking include video frame registration, dynamic background subtraction, and the object tracking algorithm itself. The challenges for robust tracking arise due to camera motion, dynamic background, and low contrast of tracked object, possibly due to environmental degradation. The survey also discusses multisensor approaches and commercial maritime systems that use EO sensors. The survey also highlights methods from computer vision research, which hold promise to perform well in maritime EO data processing. Performance of several maritime and computer vision techniques is evaluated on Singapore Maritime Dataset.
Dilip K. Prasad, Deepu Rajan, Lily Rachmawati, Eshan Rajabally, Hiok Chai Quek
IEEE Trans. Intell. Transp. Syst.1
2014 Raw-to-Raw: Mapping between Image Sensor Color Responses
abstract
Camera images saved in raw format are being adopted in computer vision tasks since raw values represent minimally processed sensor responses. Camera manufacturers, however, have yet to adopt a standard for raw images and current raw-rgb values are device specific due to different sensors spectral sensitivities. This results in significantly different raw images for the same scene captured with different cameras. This paper focuses on estimating a mapping that can convert a raw image of an arbitrary scene and illumination from one camera's raw space to another. To this end, we examine various mapping strategies including linear and non-linear transformations applied both in a global and illumination-specific manner. We show that illumination-specific mappings give the best result, however, at the expense of requiring a large number of transformations. To address this issue, we introduce an illumination-independent mapping approach that uses white-balancing to assist in reducing the number of required transformations. We show that this approach achieves state-of-the-art results on a range of consumer cameras and images of arbitrary scenes and illuminations.
Nguyen Ho Man Rang, Dilip K. Prasad, Michael S. Brown
CVPR2
2014 Training-Based Spectral Reconstruction from a Single RGB Image
Nguyen Ho Man Rang, Dilip K. Prasad, Michael S. Brown
ECCV (7)2
2014 DEB: Definite Error Bounded Tangent Estimator for Digital Curves
abstract
We propose a simple and fast method for tangent estimation of digital curves. This geometric-based method uses a small local region for tangent estimation and has a definite upper bound error for continuous as well as digital conics, i.e., circles, ellipses, parabolas, and hyperbolas. Explicit expressions of the upper bounds for continuous and digitized curves are derived, which can also be applied to nonconic curves. Our approach is benchmarked against 72 contemporary tangent estimation methods and demonstrates good performance for conic, nonconic, and noisy curves. In addition, we demonstrate a good multigrid and isotropic performance and low computational complexity of O(1) and better performance than most methods in terms of maximum and average errors in tangent computation for a large variety of digital curves.
Dilip K. Prasad, Maylor K. H. Leung, Hiok Chai Quek, Michael S. Brown
IEEE Trans. Image Process.1
2013 ElliFit: An unconstrained, non-iterative, least squares based geometric Ellipse Fitting method
Dilip K. Prasad, Maylor K. H. Leung, Hiok Chai Quek
Pattern Recognit.1
2012 A novel framework for making dominant point detection methods non-parametric
abstract
Most dominant point detection methods require heuristically chosen control parameters. One of the commonly used control parameter is maximum deviation. This paper uses a theoretical bound of the maximum deviation of pixels obtained by digitization of a line segment for constructing a general framework to make most dominant point detection methods non-parametric. The derived analytical bound of the maximum deviation can be used as a natural bench mark for the line fitting algorithms and thus dominant point detection methods can be made parameter-independent and non-heuristic. Most methods can easily incorporate the bound. This is demonstrated using three categorically different dominant point detection methods. Such non-parametric approach retains the characteristics of the digital curve while providing good fitting performance and compression ratio for all the three methods using a variety of digital, non-digital, and noisy curves.
Dilip K. Prasad, Maylor K. H. Leung, Hiok Chai Quek, Siu-Yeung Cho
Image Vis. Comput.1
2012 Edge curvature and convexity based ellipse detection method
Dilip K. Prasad, Maylor K. H. Leung, Siu-Yeung Cho
Pattern Recognit.1
2010 Reliability / precision uncertainity in shape fitting problems
abstract
The precision/uncertainty duality has been long known in the context of Hough transform, where a shape in an image cannot be fit precisely using the Hough transform without compromising the certainty or reliability of the fitting. This paper mathematically shows that such duality also exists while using the least squares based method. This paper also proposes a method to quantify the reliability of a fit. Further, based on the proposed measure of reliability, an optimization scheme to strike a balance between the precision and reliability is suggested. Though the mathematical formulations deal with only straight line, considering it as the simplest and basic geometric primitive, it is argued that such duality exists for any shape fitting and applies to any shape fitting method.
Dilip K. Prasad, Maylor K. H. Leung
ICIP1