Tushar Sandhan

dblp:140/7939 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
13since 2021 · last 2026
0009-0004-1138-4756ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021
YearPublicationVenuePosition
2026 NeuralPathLite: Fast and Robust Diffusion-Based Path Planning for Autonomous Navigation
Pratham Gupta, Tushar Sandhan
IEEE Trans Autom. Sci. Eng.3
2025 Tagsim: Topic-Informed Attention Guided Similarity Metric for Image Caption Comparison
abstract
Existing image caption evaluation metrics, such as BLEU, ROUGE, and CIDEr primarily rely on high-level similarities like n-gram matching. Here, we propose TAGSim, a novel metric that automatically incorporates topics or concepts as the caption’s topic contains the main summary of the image. TAGSim creates topic-weighted latent representations, thereby acquiring semantics by integrating both con-text and topic information. It leverages a regression model to create an attention-based novel similarity metric, while internally building caption representations. TAGSim moves be-yond lexical overlap by focusing on meaningful semantic relationships, better aligning captions with the core image topic. It handles paraphrasing and diverse expressions, ensuring a more nuanced and reliable evaluation across languages and styles. It outperforms the strong similarity baseline by an average 0.72 points (SCI evaluation index) across 3 datasets. It is language agnostic, empirically established that it is a pseudometric, and correlates well with standard caption evaluation metrics. Our code and datasets will be publicly available.
Vipul Chanchlani, Vishal Himmatsinghka, Ayush Himmatsinghka, Jivnesh Sandhan, Tushar Sandhan
ICIP5
2025 Daafnet: Domain Adaptive Augmented Feature Network for Biosignal based Emotion Recognition
abstract
Scalp-Recorded biosignals, like EEG, are very effective tools in capturing neurodynamics due to their high temporal resolution, providing useful insights on complex cognitive processes such as emotion. However, due to the low spatial resolution of EEG along with its low signal power to noise ratio and highly complex nature of the brain, the task of emotion recognition tends to be challenging. Furthermore, inter-subject and inter-trial variability in EEG studies interferes with the deep learning models’ ability to generalize. To overcome these challenges, here, we propose DAAFNet, a Domain Adaptive Augmented Feature Network that utilizes quantitative EEG features as multi-channel images and EEG time series data to perform emotion classification. The benchmark datasets, DEAP and SEED, are utilized to perform extensive experiments to demonstrate the performance of our method in comparison with the state-of-the-art methods. We present an extensive analysis pipeline for personalized emotion classification. For reproducible research, complete implementation code is made available at: https://github.com/Mythili98/DAAFNet.
Swathi Pratapa, Tushar Sandhan
ICIP2
2025 ENADL: Towards Performance Improvement of IoT Networks Using Deep Learning-Based Node Fault Prediction
abstract
The Internet of Things (IoT) has grown explosively with wireless technology integration. Several IoT applications require high data throughput, low data transmission latency, and high data gathering reliability. Since, the IoT network (IoTN) is generally dynamic and utilizes a multi-hop data transmission scheme for such applications, the throughput, latency, and network lifetime tend to degrade as the hops increase. Moreover, IoT devices (IoD) are low-cost, less computationally capable, and battery-limited, further impacting performance. A faulty IoD worsens network lifetime and throughput. Predicting faulty nodes and re-routing data can significantly enhance performance. This work proposes a node fault prediction framework to enhance data routing in dynamic IoTN, maximizing throughput and lifetime. The network is represented as a graph in which the IoD are the nodes. Then a novel deep learning model is proposed utilizing various node and edge features to predict the faulty IoDs. Particularly, the proposed edge and node features-accumulation deep learning (ENADL) method exploits features, such as Euclidean distance between nodes, residual energy level of nodes, and type and number of messages passed between edges to predict the forthcoming faulty IoD. Thereafter, data routing is performed over the updated network topology. Furthermore, to improve the network lifetime, the node's degree and betweenness centrality measures-based energy allocation method is also proposed. Finally, numerical results on simulated and real-field testbeds demonstrate the ENADL method.s effectiveness in predicting faulty nodes and re-routing data packets. This results in maximized network throughput and lifetime as compared to several existing methods.
Shraddha Tripathi, Faheem Nizar, Om Jee Pandey, Tushar Sandhan, Rajesh M. Hegde
IEEE Trans. Reliab.4
2024 swCNN: A Small World Convolutional Neural Network for Efficient Training
Shubham Dwivedi, Tushar Sandhan, Om Jee Pandey, Rajesh M. Hegde
ICPR (8)2
2024 Severity of Flood Damage Estimation from Aerial Scenery
Tarakeswara Rao Landa, Tushar Sandhan
ICPR (18)2
2024 Ensembling YOLO and ViT for Plant Disease Detection
Debojyoti Misra, Suryansh Goel, Tushar Sandhan
ICPR (21)3
2024 Robust Leaf Detection using Shape Priors within Smaller Datasets
Debojyoti Misra, Tushar Sandhan
ICPR (15)2
2024 Bandwise Attention in CycleGAN for Fructose Estimation from Hyperspectral Images
Divyani Tyagi, Tushar Sandhan
ICPR (23)2
2024 Scalable and Time-Efficient Bin-Picking for Unknown Objects in Dense Clutter
abstract
The task of fully automated picking of novel bin objects that are placed in a densely cluttered pile poses a significant challenge. It becomes even more challenging if the objects are of various shapes, sizes, colors, and textures. Generally, grasp planning for a given scene begins with sampling several grasp poses, which are then evaluated to determine the final grasp pose for the robot action. With the increment in clutter level in the bin, the fraction of graspable locations becomes smaller. Hence, for a scalable grasp pose planning strategy, the pose sampling method should be intelligent enough to find suitable grasp regions in a time-efficient manner, irrespective of the amount of clutter and the workspace size. In this paper, we present a scalable robotic bin-picking method (SE-RoB) that performs equally well amidst the increasing level of clutter in the scene. In a real-world challenging bin-picking setup, our proposed method has shown significantly better performance ($13\%$improvement) compared to the state-of-the-art methods in terms of grasp success rate in a time-efficient manner (6 Hz inference speed).Note to Practitioners—Picking objects from a cluttered pile can be challenging, especially when dealing with objects of different shapes, sizes, colors, and textures. We have created a new method called SE-RoB that can help robots pick up unknown objects from a pile reliably and in a time-efficient manner. In our experiments, we tested this method in real-world situations where the pile of objects was densely cluttered and had many different types of objects. SE-RoB worked much better compared to the state-of-the-art methods. Additionally, this method works at around 6 Hz inference speed, which means it is suitable for real-life scenarios such as warehouse automation. We believe that SE-RoB can be a valuable tool for practitioners who are looking to automate their bin-picking processes using robots.
Prem Raj, Laxmidhar Behera, Tushar Sandhan
IEEE Trans Autom. Sci. Eng.3
2023 Active Perception System for Enhanced Visual Signal Recovery Using Deep Reinforcement Learning
abstract
Deep neural networks have demonstrated excellent object detection and segmentation performance from RGB data. However, these models can only recognize and predict segmentation masks with great accuracy when RGB data have sufficient information about the objects of interest. In this paper, we suggest an intelligent, active perception system that can adjust its 3D position to improve signal acquisition. The segmentation score of cluttered scene is improved a lot due to this proposed system, which can also enhance grasp pose detection for the robotic manipulator. The ResNet-50 backbone of the proposed perception system is initialized using pre-trained weights to extract a latent state from an RGB image of the cluttered scene. A Reinforcement Learning (RL) agent uses these retrieved states to reposition the visual perception system for enhancement of the underlying computer vision tasks such as segmentation of the cluttered scene. Our trained RL agent can anticipate the better position of the visual perception system, which ensures enhanced signal recovery. The effectiveness of the proposed approach is tested in a pybullet simulation environment.
Gaurav Chaudhary, Laxmidhar Behera, Tushar Sandhan
ICASSP3
2023 Graph Based Semantic Ensemble of Riemannian Neural Structured Learning for BCI-EEG Signal Classification
abstract
Machine Learning (ML) classifiers have been made more robust in recent years by leveraging the graph structure between the inputs using Neural Structured Learning (NSL). However, researchers have not taken full advantage of it for the Brain-Computer Interface (BCI) classification tasks. While the traditional NSL faces the issues of a very minimized use of graph structural properties and optimal similarity metric, in this paper, we propose a Node Impact Multi Metric Threshold NSL (NI-MT-NSL) to overcome these issues. For the first time, the node-influence properties from graph theory are incorporated to alter the way different EEG samples influence the training, while an ensemble of semantic graphs is used in the NSL module to capture different semantic relations between the EEG trial data. The proposed model is assessed on the standard BCI IV 2a dataset. On comparing its test accuracies with the traditional Riemannian classifiers and the baseline NSL, we have found improved accuracies over all subjects. We have found a tremendous improvement in classification, with a mean gain of 7% for the subjects having very poor accuracy even with state-of-the-art methods.
Vinay Gupta, Laxmidhar Behera, Tushar Sandhan
ICASSP3
2022 A Novel Multi-Task Learning Approach for Context-Sensitive Compound Type Identification in Sanskrit
abstract
The phenomenon of compounding is ubiquitous in Sanskrit. It serves for achieving brevity in expressing thoughts, while simultaneously enriching the lexical and structural formation of the language. In this work, we focus on the Sanskrit Compound Type Identification (SaCTI) task, where we consider the problem of identifying semantic relations between the components of a compound word. Earlier approaches solely rely on the lexical information obtained from the components and ignore the most crucial contextual and syntactic information useful for SaCTI. However, the SaCTI task is challenging primarily due to the implicitly encoded context-sensitive semantic relation between the compound components. Thus, we propose a novel multi-task learning architecture which incorporates the contextual information and enriches the complementary syntactic information using morphological tagging and dependency parsing as two auxiliary tasks. Experiments on the benchmark datasets for SaCTI show 6.1 points (Accuracy) and 7.7 points (F1-score) absolute gain compared to the state-of-the-art system. Further, our multi-lingual experiments demonstrate the efficacy of the proposed architecture in English and Marathi languages.
Jivnesh Sandhan, Hrishikesh Terdalkar, Tushar Sandhan, Suvendu Samanta, Laxmidhar Behera, Pawan Goyal 0002
COLING4
2020 Separating Particulate Matter From a Single Microscopic Image
abstract
Particulate matter (PM) is the blend of various solid and liquid particles suspended in atmosphere. These submicron particles are imperceptible for usual hand-held camera photography, but become a great obstacle in microscopic imaging. PM removal from a single microscopic image is a highly ill-posed and one of the challenging image denoising problems. In this work, we thoroughly analyze the physical properties of PM, microscope and their inevitable interaction; and propose an optimization scheme, which removes the PM from a high-resolution microscopic image within a few seconds. Experiments on real world microscopic images show that the proposed method significantly outperforms other competitive image denoising methods. It preserves the comprehensive microscopic foreground details while clearly separating the PM from a single monochromatic or color image.
Tushar Sandhan, Jin Young Choi 0002
CVPR1
2017 Anti-Glare: Tightly Constrained Optimization for Eyeglass Reflection Removal
abstract
Absence of a clear eye visibility not only degrades the aesthetic value of an entire face image but also creates difficulties in many computer vision tasks. Even mild reflections produce the undesired superpositions of visual information, whose decomposition into the background and reflection layers using a single image is a highly ill-posed problem. In this work, we enforce the tight constraints derived by thoroughly analysing the properties of an eyeglass reflection. In addition, our strategy regularizes gradients of the reflection layer to be highly sparse and proposes the facial symmetry prior via formulating a non-convex optimization scheme, which removes the reflections within a few iterations. Experiments on frontal face image inputs demonstrate the high quality reflection removal results and improvement of the iris detection rate.
Tushar Sandhan, Jin Young Choi 0002
CVPR1
2017 Simultaneous Detection and Removal of High Altitude Clouds from an Image
abstract
Interestingly, shape of the high-altitude clouds serves as a beacon for weather forecasting, so its detection is of vital importance. Besides these clouds often cause hindrance in an endeavor of satellites to inspect our world. Even thin clouds produce the undesired superposition of visual information, whose decomposition into the clear background and cloudy layer using a single satellite image is a highly ill-posed problem. In this work, we derive sophisticated image priors by thoroughly analyzing the properties of high-altitude clouds and geological images; and formulate a non-convex optimization scheme, which simultaneously detects and removes the clouds within a few seconds. Experimental results on real world RGB images demonstrate that the proposed method outperforms the other competitive methods by retaining the comprehensive background details and producing the precise shape of the cloudy layer.
Tushar Sandhan, Jin Young Choi 0002
ICCV1
2017 Audio Classification Using Class-Specific Learned Descriptors
Sukanya Sonowal, Tushar Sandhan, In Kyu Choi, Nam Soo Kim
INTERSPEECH2
2014 Multi-task learning with over-sampled time-series representation of a trajectory for traffic motion pattern recognition
abstract
This paper proposes an efficient feature sampling and multi-task learning scheme for traffic scene analysis, where all classifiers are trained simultaneously by exploiting the correlations among different motion patterns. We make feature descriptors by high dimensional embedding of the time series data for traffic pattern representation. They preserve detailed spatio-temporal information of the underlying event. Pattern specific details are extracted from raw trajectories and embedded into feature descriptors, which ensures their great discriminability. Training data scarcity problem is tackled through amplification of the patterns hidden in raw trajectory via strategic oversampling and employment of joint feature selection procedure while training the models. Experimental results on 4 surveillance datasets, show great improvement in the motion pattern recognition performance, importance of joint feature selection and fast incremental learning ability of the proposed framework.
Tushar Sandhan, Young Joon Yoo, Hanjoo Yoo, Sangdoo Yun, Moonsub Byeon
AVSS1
2014 Frequencygrams and multi-feature joint sparse representation for action and gesture recognition
abstract
Features play a vital role in human action recognition (HAR), as they encapsulate the underlying dynamics of the action. We propose the features (frequencygrams) based on frequency domain analysis of histograms of the motion and its spatiotemporal gradient (rate of change in motion flow). Feature extraction is quite simple and can be performed in real time using sparse or interest point motion flow. They are resilient to delayed initiated actions, scale variation, moving background, sudden illumination changes (high frequency noise) and avoid the overload of person detection and tracking. Being robust to camera motions, they also provide a natural, compact and discriminative representation for reciprocating motions by preserving comprehensive temporal information of the action sequences. As other global features also bear some action semantics, we fuse all these features together in a systematic way to improve the overall HAR performance, by employing the joint sparse representation with group sparsity regularization. The extensive experimental results, on three benchmark action datasets and one gesture recognition dataset, show the effectiveness and generality of the proposed method.
Tushar Sandhan, Jin Young Choi 0002
ICIP1
2014 Handling Imbalanced Datasets by Partially Guided Hybrid Sampling for Pattern Recognition
abstract
Occurrence of high imbalance in real-world domains is a direct result of rarity of interesting events, which results in skewed datasets. Without dataset rebalancing, the learning algorithm will encounter extremely low minority class samples therefore it gets biased towards the majority class in the classification tasks. Hence properly handling the imbalanced dataset is a crucial issue in the pattern recognition domain. We have employed bootstrapping by simultaneous oversampling of the minority class and under sampling of the majority class to build the ensemble of classifiers. Oversampling is partially guided by the extracted hidden patterns from minority class, which prevents its over-generalization and amplify subtle vital patterns. The proposed framework is evaluated on four highly imbalanced datasets with employing a series of classifiers like, support vector machine, logistic regression, nearest neighbor and Gaussian process classifier. Experimental results showed that the pattern classification performance for various tasks improves after rebalancing datasets using the proposed framework.
Tushar Sandhan, Jin Young Choi 0002
ICPR1
2014 View invariant action recognition using generalized 4D features
Sun Jung Kim, Soo Wan Kim, Tushar Sandhan, Jin Young Choi 0002
Pattern Recognit. Lett.3
2013 Towards simultaneous clustering and motif-modeling for a large number of protein family
abstract
In this paper, we propose a novel clustering and motif modeling framework for analyzing large number of protein family using k-mer. Our approach of using k-mers utilizes both occurring frequency and position information of k-mers that essential for classification yet not fully used in previous methods. We found that the structure has close relationship between motif of protein family and hence well describe important biological features or motifs of each protein family. The classification/clustering procedure are executed in incremental manner which was difficult for previous algorithms and is modeled by using bipartite model. Furthermore, the method can be efficiently implemented using parallel computing and hash. Experimental results using the entire COG family database shows that our model can model a large number of protein families without sacrificing accuracy. In addition, the classification structure, path of the graph for protein sequences, explains characteristic subsequences or motif of each family quite well. Thus the proposed method has the potential to model both protein families and motifs, even for a large number of families.
Young Joon Yoo, Tushar Sandhan, Jin Young Choi 0002, Sun Kim
BIBM2
2013 Abstracted radon profiles for fingerprint recognition
abstract
Conventional minutiae-based fingerprint recognition approaches consider only local characteristics and their accuracy dramatically decreases as the number of available minutiae decreases. We propose new features based on Abstracted Radon Profile (ARP). Proposed method uses global properties of an image and it does not necessitate any heavy preprocessing as in classical methods. By using independent gradual patching via proposed multilayer architecture, local characteristics of an image are also preserved. ARP features have an advantage of being robust to zero mean additive noise. For sparse signal representation, dictionary is constructed from the ARP features of the training samples. Recognition is done by ℓ1-minimization with quadratic constraints, so this framework can handle dense noise by exploiting the fact that these errors are often sparse. Experimental results in assessing recognition performance demonstrate the proposed approach outperforms the conventional approaches in correlation and distance based comparisons. Computational time comparison result shows the proposed feature is more efficient than brute-force method of image alignment and promising for handling other pattern recognition problems as well.
Tushar Sandhan, Hyung Jin Chang, Jin Young Choi 0002
ICIP1