Nilanjan Ray

dblp:19/6409 · DBLP profile ↗
← Back
81ranked-venue papers
18as first author
13since 2021 · last 2025
0000-0002-7588-5400ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 53 · 13 first-author · 6 since 2021Artificial intelligence and machine learning · 24 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Learning Instance-Specific Parameters of Black-Box Models Using Differentiable Surrogates
abstract
Tuning parameters of a non-differentiable or black-box compute is challenging. Existing methods rely mostly on random sampling or grid sampling from the parameter space. Further, with all the current methods, it is not possible to supply any input specific parameters to the black-box. To the best of our knowledge, for the first time, we are able to learn input-specific parameters for a black box in this work. As a test application, we choose a popular image de-noising method BM3D as our black-box compute. Then, we use a differentiable surrogate model (a neural network) to approximate the black-box behaviour. Next, another neural network is used in an end-to-end fashion to learn input instance-specific parameters for the black-box. Moti-vated by prior advances in surrogate-based optimization, we applied our method to the Smartphone Image Denoising Dataset (SIDD) and the Color Berkeley Segmentation Dataset (CBSD68) for image denoising. The results are compelling, demonstrating a significant increase in PSNR and a notable improvement in SSIM nearing 0.93. Experimental results underscore the effectiveness of our approach in achieving substantial improvements in both model performance and optimization efficiency. For code and implementation details, please refer to our GitHub repository: https://github.com/arnisha-k/instance-specific-param.
Arnisha Khondaker, Nilanjan Ray
WACV2
2025 Unpaired document image denoising for OCR using BiLSTM enhanced CycleGAN
Katyani Singh, Ganesh Tata, Eric Van Oeveren, Nilanjan Ray
Int. J. Document Anal. Recognit.4
2024 ShadowSense: Unsupervised Domain Adaptation and Feature Fusion for Shadow-Agnostic Tree Crown Detection from RGB-Thermal Drone Imagery
abstract
Accurate detection of individual tree crowns from remote sensing data poses a significant challenge due to the dense nature of forest canopy and the presence of diverse environmental variations, e.g., overlapping canopies, occlusions, and varying lighting conditions. Additionally, the lack of data for training robust models adds another limitation in effectively studying complex forest conditions. This paper presents a novel method for detecting shadowed tree crowns and provides a challenging dataset comprising roughly 50k paired RGB-thermal images to facilitate future research for illumination-invariant detection. The proposed method (ShadowSense) is entirely self-supervised, leveraging domain adversarial training without source domain annotations for feature extraction and foreground feature alignment for feature pyramid networks to adapt domain-invariant representations by focusing on visible foreground regions, respectively. It then fuses complementary information of both modalities to effectively improve upon the predictions of an RGB-trained detector and boost the overall accuracy. Extensive experiments demonstrate the superiority of the proposed method over both the baseline RGB-trained detector and state-of-the-art techniques that rely on unsupervised domain adaptation or early image fusion. Our code and data are available: https://github.com/rudrakshkapil/ShadowSense.
Rudraksh Kapil, Seyed Mojtaba Marvasti-Zadeh, Nadir Erbilgin, Nilanjan Ray
WACV4
2024 Training-Based Model Refinement and Representation Disagreement for Semi-Supervised Object Detection
abstract
Semi-supervised object detection (SSOD) aims to improve the performance and generalization of existing object detectors by utilizing limited labeled data and extensive unlabeled data. Despite many advances, recent SSOD methods are still challenged by inadequate model refinement using the classical exponential moving average (EMA) strategy, the consensus of Teacher-Student models in the latter stages of training (i.e., losing their distinctiveness), and noisy/misleading pseudo-labels. This paper proposes a novel training-based model refinement (TMR) stage and a simple yet effective representation disagreement (RD) strategy to address the limitations of classical EMA and the consensus problem. The TMR stage of Teacher-Student models optimizes the lightweight scaling operation to refine the model’s weights and prevent overfitting or forgetting learned patterns from unlabeled data. Meanwhile, the RD strategy helps keep these models diverged to encourage the student model to explore additional patterns in unlabeled data. Our approach can be integrated into established SSOD methods and is empirically validated using two baseline methods, with and without cascade regression, to generate more reliable pseudo-labels. Extensive experiments demonstrate the superior performance of our approach over state-of-the-art SSOD methods. Specifically, the proposed approach outperforms the baseline Unbiased-Teacher-v2 (& Unbiased-Teacher-v1) method by an average mAP margin of 2.23, 2.1, and 3.36 (& 2.07, 1.9 and 3.27) on COCO-standard, COCO-additional, and Pascal VOC datasets, respectively.
Seyed Mojtaba Marvasti-Zadeh, Nilanjan Ray, Nadir Erbilgin
WACV2
2023 A Closer Look at Weak Supervision's Limitations in WSI Recurrence Score Prediction
abstract
Histological examination remains the gold standard for breast cancer diagnosis, prognosis assessment and treatment guidance. Commercial molecular signature test, ONCOTYPEDX®is routinely used for luminal breast cancers to predict the probabilities of response to chemotherapy and disease recurrence. We attempted to predict RS using digital pathology and Weakly Supervised (WS) attention-based models like CLAM (Clustering-constrained Attention Multiple Instance Learning) [1] and TransMIL (Transformer based Correlated Multiple Instance Learning) [2] on our in-house dataset. In tissue samples, the malignant component is haphazardly admixed with the nonmalignant component in variable proportions. This represents a challenge for the WS attention-based models to identify high-valued diagnostic/prognostic areas within whole slide images (WSIs). To address this, we propose an interactive approach with a human in the middle (supervised) by creating a user-friendly Graphical User Interface (GUI) that allows a pathologist to provide feedback to heatmaps generated by any WS attention-based model. We incorporate the feedback from the GUI as expected scores and penalize its difference with the attention scores in the successive training process (current scores). We observe an improvement in RS prediction after the pathologist’s feedback - a 5% rise in AUC and 4% in accuracy for CLAM and a 4.5% increase in AUC and 3% in accuracy for TransMIL. We analyze the generated heatmaps and notice an improvement in cosine similarity between the expected scores and attention scores before and after the feedback - 5% and 10% increase for CLAM and TransMIL, respectively. The implementation of the proposed approach and the dataset is available for download1. Our adaptive, interactive feedback system harmonizes attention scores with expert intuition and instills higher confidence in the system’s predictions. This study establishes a potent synergy between AI and pathologists.
Namitha Guruprasad, Penny J. Barnes, Amir Akbarnejad, Gilbert Bigras, Nilanjan Ray
BIBM5
2023 GPEX, A Framework For Interpreting Artificial Neural Networks
abstract
The analogy between Gaussian processes (GPs) and deep artificial neural networks (ANNs) has received a lot of interest, and has shown promise to unbox the blackbox of deep ANNs. Existing theoretical works put strict assumptions on the ANN (e.g. requiring all intermediate layers to be wide, or using specific activation functions). Accommodating those theoretical assumptions is hard in recent deep architectures, and those theoretical conditions need refinement as new deep architectures emerge. In this paper we derive an evidence lower-bound that encourages the GP's posterior to match the ANN's output without any requirement on the ANN. Using our method we find out that on 5 datasets, only a subset of those theoretical assumptions are sufficient. Indeed, in our experiments we used a normal ResNet-18 or feed-forward backbone with a single wide layer in the end. One limitation of training GPs is the lack of scalability with respect to the number of inducing points. We use novel computational techniques that allow us to train GPs with hundreds of thousands of inducing points and with GPU acceleration. As shown in our experiments, doing so has been essential to get a close match between the GPs and the ANNs on 5 datasets. We implement our method as a publicly available tool called GPEX: https://github.com/amirakbarnejad/gpex. On 5 datasets (4 image datasets, and 1 biological dataset) and ANNs with 2 types of functionality (classifier or attention-mechanism) we were able to find GPs whose outputs closely match those of the corresponding ANNs. After matching the GPs to the ANNs, we used the GPs' kernel functions to explain the ANNs' decisions. We provide more than 200 explanations (around 30 in the paper and the rest in the supplementary) which are highly interpretable by humans and show the ability of the obtained GPs to unbox the ANNs' decisions.
Amir Akbarnejad, Gilbert Bigras, Nilanjan Ray
NeurIPS3
2023 Crown-CAM: Interpretable Visual Explanations for Tree Crown Detection in Aerial Images
abstract
Visual explanation of “black-box” models allows researchers inexplainable artificial intelligence(XAI) to interpret the model’s decisions in a human-understandable manner. In this paper, we proposeinterpretable class activation mapping for tree crown detection(Crown-CAM) that overcomes inaccurate localization & computational complexity of previous methods while generating reliable visual explanations for the challenging and dynamic problem of tree crown detection in aerial images. It consists of an unsupervised selection of activation maps, computation of local score maps, and non-contextual background suppression to efficiently provide fine-grain localization of tree crowns in scenarios with dense forest trees or scenes without tree crowns. Additionally, twoIntersection over Union(IoU)-based metrics are introduced to effectively quantify both the accuracy and inaccuracy of generated explanations with respect to regions with or even without tree crowns in the image. Empirical evaluations demonstrate that the proposed Crown-CAM outperforms the Score-CAM, Augmented Score-CAM, and Eigen-CAM methods by an average IoU margin of 8.7, 5.3, and 21.7 (and 3.3, 9.8, and 16.5) respectively in improving the accuracy (and decreasing inaccuracy) of visual explanations on the challenging NEON tree crown dataset.
Seyed Mojtaba Marvasti-Zadeh, Devin Goodsman, Nilanjan Ray, Nadir Erbilgin
IEEE Geosci. Remote. Sens. Lett.3
2022 Dynamic Background Subtraction by Generative Neural Networks
abstract
Background subtraction is a significant task in computer vision and an essential step for many real world applications. One of the challenges for background subtraction methods is dynamic background, which constitutes stochastic movements in some parts of the background. In this paper, we have proposed a new background subtraction method, called DBSGen, which uses two generative neural networks, one for dynamic motion removal and another for background generation. At the end, the foreground moving objects are obtained by a pixel-wise distance threshold based on a dynamic entropy map. DBSGen is an end-to-end, unsupervised optimization method with a near real-time frame rate. The performance of the method is evaluated over dynamic background sequences and it outperforms most of state-of-the-art unsupervised methods. Our code is publicly available at https://github.com/FatemeBahri/DBSGen.
Fateme Bahri, Nilanjan Ray
AVSS2
2022 Deep Learning Based Parametrization of Diffeomorphic Image Registration for the Application of Cardiac Image Segmentation
abstract
Cardiac segmentation from magnetic resonance imaging (MRI) is one of the essential tasks to analyze the anatomy and function of the heart for the assessment and diagnosis of cardiac diseases. However, manual annotation is difficult and time consuming. This study proposes a novel end-to-end supervised cardiac MRI segmentation framework based on a diffeomorphic deformable registration that can segment the left ventricle from 2D and 3D images or volumes. In order to represent the actual cardiac deformation, the methodology parameterizes the transformation using radial and rotational components, computed using a deep learning approach The method was evaluated over three different data sets and showed significant improvements compared to exacting learning and non-learning based methods in terms of the Dice score and Hausdorff distance metrics.
Ameneh Sheikhjafari, Deepa Krishnaswamy, Michelle Noga, Nilanjan Ray, Kumaradevan Punithakumar
BIBM4
2022 Fire Together Wire Together: A Dynamic Pruning Approach with Self-Supervised Mask Prediction
abstract
Dynamic model pruning is a recent direction that allows for the inference of a different sub-network for each input sample during deployment. However, current dynamic methods rely on learning a continuous channel gating through regularization by inducing sparsity loss. This formulation introduces complexity in balancing different losses (e.g task loss, regularization loss). In addition, regularization based methods lack transparent tradeoff hyper- parameter selection to realize a computational budget. Our contribution is two-fold: 1) decoupled task and pruning losses. 2) Simple hyperparameter selection that enables FLOPs reduction estimation before training. Inspired by the Hebbian theory in Neuroscience: “neurons that fire together wire together”, we propose to predict a mask to process k filters in a layer based on the activation of its previous layer. We pose the problem as a self-supervised binary classification problem. Each mask predictor module is trained to predict if the log-likelihood for each filter in the current layer belongs to the top-k activated filters. The value k is dynamically estimated for each input based on a novel criterion using the mass of heatmaps. We show experiments on several neural architectures, such as VGG, ResNet and MobileNet on CIFAR and ImageNet datasets. On CIFAR, we reach similar accuracy to SOTA methods with 15% and 24% higher FLOPs reduction. Similarly in ImageNet, we achieve lower drop in accuracy with up to 13% improvement in FLOPs reduction.
Sara Elkerdawy, Mostafa Elhoushi, Hong Zhang 0013, Nilanjan Ray
CVPR4
2022 Towards Positive Jacobian: Learn to Postprocess for Diffeomorphic Image Registration with Matrix Exponential
abstract
We present a postprocessing layer for deformable image registration to make a registration field more diffeomorphic by encouraging Jacobians of the transformation to be positive. Diffeomorphic image registration is important for medical imaging studies because of the properties like invertibility, smoothness of the transformation, and topology preservation/non-folding of the grid. Violation of these properties can lead to destruction of the neighbourhood and the connectivity of anatomical structures during image registration. Most of the recent deep learning methods do not explicitly address this folding problem and try to solve it with a smoothness regularization on the registration field. In this paper, we propose a differentiable layer, which takes any registration field as its input, computes exponential of the Jacobian matrices of the input and reconstructs a new registration field from the exponentiated Jacobian matrices using Poisson reconstruction. Our proposed Poisson reconstruction loss enforces positive Jacobians for the final registration field. Thus, our method acts as a post-processing layer without any learnable parameters of its own and can be placed at the end of any deep learning pipeline to form an end-to-end learnable framework. We show the effectiveness of our proposed method for a popular deep learning registration method Voxelmorph and evaluate it with a dataset containing 3D brain MRI scans. Our results show that our post-processing can effectively decrease the number of non-positive Jacobians by a significant amount without any noticeable deterioration of the registration accuracy, thus making the registration field more diffeomorphic. Our code is available online at https://github.com/Soumyadeep-Pal/Diffeomorphic-Image-Registration-Postprocess
Soumyadeep Pal, Matthew Tennant, Nilanjan Ray
ICPR3
2022 A training-free recursive multiresolution framework for diffeomorphic deformable image registration
Ameneh Sheikhjafari, Michelle Noga, Kumaradevan Punithakumar, Nilanjan Ray
Appl. Intell.4
2021 Unknown-Box Approximation to Improve Optical Character Recognition Performance
Ayantha Randika, Nilanjan Ray, Allegra Latimer
ICDAR (1)2
2020 To Filter Prune, or to Layer Prune, That Is the Question
Sara Elkerdawy, Mostafa Elhoushi, Abhineet Singh, Hong Zhang 0013, Nilanjan Ray
ACCV (3)5
2020 Multiview 3-D Echocardiography Image Fusion with Mutual Information Neural Estimation
abstract
Multiview three-dimensional echocardiography (M3DE) fuses volumetric datasets acquired from complementary acoustic windows to expand field-of-view and allow for visualization of the entire heart. This is of great importance for cardiac chamber quantification. The M3DE also allows for image quality improvement through fusion of single views on overlapping regions. However, shape variations and increase in noise stemming from the nature of ultrasound physics make fusion a challenging task. This study proposes a novel machine learning-based fusion method to combine ultrasound views that are spatially apart, namely, apical and parasternal. Our method jointly uses: 1) an autoencoder framework to generate the fused image; and 2) a mutual information neural estimation network to maximize the mutual information between source and fused images. The experimental evaluations show promising results and the fused image generated by the proposed method improves the signal-to-noise ratio by up to 18.23 dB and the contrast-to-noise ratio by up to 21.76 dB compared to the state-of-art approaches.
Juiwen Ting, Kumaradevan Punithakumar, Nilanjan Ray
BIBM3
2020 One-Shot Layer-Wise Accuracy Approximation For Layer Pruning
abstract
Recent advances in neural networks pruning have made it possible to remove a large number of filters without any perceptible drop in accuracy. However, the gain in speed depends on the number of filters per layer. In this paper, we propose a one-shot layer-wise proxy classifier to estimate layer importance that in turn allows us to prune a whole layer. In contrast to existing filter pruning methods which attempt to reduce the layer width of a dense model, our method reduces its depth and can thus guarantee inference speed up. In our proposed method, we first go through the training data once to construct proxy classifiers for each layer using imprinting. Next, we prune layers with smallest accuracy difference from their preceding layer till a latency budget is achieved. Finally, we fine-tune the newly pruned model to improve accuracy. Experimental results showed 43.70% latency reduction with 1.27% accuracy increase on CIFAR100 for the pruned VGG19. Further, we achieved 16% and 25% latency reduction with 0.58% increase and 0.01% decrease in accuracy respectively on ImageNet for ResNet-50. The major advantage of our proposed method is that these latency reductions cannot be achieved with existing filter pruning methods as they are bounded by the original model's depth. Code is available at https://github.com/selkerdawy/one-shot-layer-pruning.
Sara Elkerdawy, Mostafa Elhoushi, Abhineet Singh, Hong Zhang 0013, Nilanjan Ray
ICIP5
2020 Foreground-focused domain adaption for object detection
abstract
Object detectors suffer from accuracy loss caused by domain shift from a source to a target domain. Unsupervised domain adaptation (UDA) approaches mitigate this loss by training with unlabeled target domain images. A popular processing pipeline applies adversarial training that aligns the distributions of the features from the two domains. We advocate that aligning the full image level features is not ideal for UDA object detection due to the presence of varied background areas during inference. Thus, we propose a novel foreground-focused domain adaptation (FFDA) framework which mines the loss of the domain discriminators to concentrate on the backpropagation of foreground loss. We obtain mining masks by collecting target predictions and source labels to outline foreground regions, and apply the masks to image and instance level domain discriminators to allow backpropagation only on the mined regions. By reinforcing this foreground-focused adaptation throughout multiple layers in the detector model, we gain a significant accuracy boost on the target domain prediction. Compared to previous methods, our method reaches the new state-of-the-art accuracy on adapting Cityscape to Foggy Cityscape dataset and demonstrates competitive accuracy on other datasets that include various scenarios for autonomous driving applications.
Nilanjan Ray
ICPR2
2020 A Flexible Method for Performance Evaluation of Robot Localization
abstract
An important research issue in mobile robotics is performance assessment of robot SLAM algorithms in terms of their localization accuracy. Typically, SLAM algorithms are evaluated with the help of benchmark datasets or expensive equipment such as motion capture. Benchmark datasets however, are environment-specific, and use of motion capture constrains spatial coverage and affordability. In this paper, we present a novel method for SLAM performance evaluation, which only uses distinctive markers (such as AR tags), randomly placed in the robot navigation environment at arbitrary locations, and observes these markers with a camera onboard of the robot. Formulated as a generative latent optimization (GLO) problem, our method uses the local robot-to-marker poses to evaluate the global robot pose estimates by a SLAM algorithm and therefore its performance. Through extensive experiments on two robots, three localization/SLAM algorithms and both LiDAR and RGB-D sensors, we demonstrate the feasibility and accuracy of our proposed method.
Sean Scheideman, Nilanjan Ray, Hong Zhang 0013
ICRA2
2020 Animal Detection in Man-made Environments
abstract
Automatic detection of animals that have strayed into human inhabited areas has important security and road safety applications. This paper attempts to solve this problem using deep learning techniques from a variety of computer vision fields including object detection, segmentation, tracking and edge detection. Several interesting insights into transfer learning are elicited while adapting models trained on benchmark datasets for real world deployment. Empirical evidence is presented to demonstrate the inability of detectors to generalize from training images of animals in their natural habitats to deployment scenarios of man-made environments. A solution is also proposed using semi-automated synthetic data generation for domain specific training. Code and data used in the experiments are made available to facilitate further work in this domain.
Abhineet Singh, Marcin Pietrasik, Gabriell Natha, Nehla Ghouaiel, Ken Brizel, Nilanjan Ray
WACV6
2020 River Ice Segmentation With Deep Learning
abstract
This article deals with the problem of computing surface concentrations for two types of river ice from digital images acquired during freeze-up. It presents the results of attempting to solve this problem using several state-of-the-art semantic segmentation methods based on deep convolutional neural networks (CNNs). This task presents two main challenges-very limited availability of labeled training data and presence of noisy labels due to the great difficulty of visually distinguishing between the two types of ice, even for human experts. The results are used to analyze the extent to which some of the best deep learning methods currently in existence can handle these challenges. The code and data used in the experiments are made publicly available to facilitate further work in this domain.
Abhineet Singh, Hayden Kalke, Mark Loewen, Nilanjan Ray
IEEE Trans. Geosci. Remote. Sens.4
2019 Lightweight Monocular Depth Estimation Model by Joint End-to-End Filter Pruning
abstract
Convolutional neural networks (CNNs) have emerged as the state-of-the-art in multiple vision tasks including depth estimation. However, memory and computing power requirements remain as challenges to be tackled in these models. Monocular depth estimation has significant use in robotics and virtual reality that requires deployment on low-end devices. Training a small model from scratch results in a significant drop in accuracy and it does not benefit from pre-trained large models. Motivated by the literature of model pruning, we propose a lightweight monocular depth model obtained from a large trained model. This is achieved by removing the least important features with a novel joint end-to-end filter pruning. We propose to learn a binary mask for each filter to decide whether to drop the filter or not. These masks are trained jointly to exploit relations between filters at different layers as well as redundancy within the same layer. We show that we can achieve around 5x compression rate with small drop in accuracy on the KITTI driving dataset. We also show that masking can improve accuracy over the baseline with fewer parameters, even without enforcing compression loss.
Sara Elkerdawy, Hong Zhang 0013, Nilanjan Ray
ICIP3
2019 Fast Large-Scale Spectral Clustering via Explicit Feature Mapping
abstract
We propose an efficient spectral clustering method for large-scale data. The main idea in our method consists of employing random Fourier features to explicitly represent data in kernel space. The complexity of spectral clustering thus is shown lower than existing Nyström approximations on largescale data. With m training points from a total of n data points, Nyström method requires O(nmd + m3+ nm2) operations, where d is the input dimension. In contrast, our proposed method requires O(nDd + D3+ n'D2), where n' is the number of data points needed until convergence and D is the kernel mapped dimension. In large-scale datasets where n ≪ n hold true, our explicitly mapping method can significantly speed up eigenvector approximation and benefit prediction speed in spectral clustering. For instance, on MNIST (60000 data points), the proposed method is similar in clustering accuracy to Nyström methods while its speed is twice as fast as Nyström.
Li He 0002, Nilanjan Ray, Yisheng Guan, Hong Zhang 0013
IEEE Trans. Cybern.2
2019 Training Convolutional Neural Networks and Compressed Sensing End-to-End for Microscopy Cell Detection
abstract
Automated cell detection and localization from microscopy images are significant tasks in biomedical research and clinical practice. In this paper, we design a new cell detection and localization algorithm that combines deep convolutional neural network (CNN) and compressed sensing (CS) or sparse coding (SC) for end-to-end training. We also derive, for the first time, a backpropagation rule, which is applicable to train any algorithm that implements a sparse code recovery layer. The key innovation behind our algorithm is that the cell detection task is structured as a point object detection task in computer vision, where the cell centers (i.e., point objects) occupy only a tiny fraction of the total number of pixels in an image. Thus, we can apply compressed sensing (or equivalently SC) to compactly represent a variable number of cells in a projected space. Subsequently, CNN regresses this compressed vector from the input microscopy image. The SC/CS recovery algorithm (L1optimization) can then recover sparse cell locations from the output of CNN. We train this entire processing pipeline end-to-end and demonstrate that end-to-end training improves accuracy over a training paradigm that treats CNN and CS-recovery layers separately. We have validated our algorithm on five benchmark datasets with excellent results.
Gilbert Bigras, Judith Hugh, Nilanjan Ray
IEEE Trans. Medical Imaging4
2018 Using Accelerometric and Gyroscopic Data to Improve Blood Pressure Prediction from Pulse Transit Time Using Recurrent Neural Network
abstract
We propose a method for estimating blood pressure (BP) non-invasively from electrocardiogram (ECG) and photoplethysmogram (PPG) signals. This method has potential to be used as a continuous form of BP estimation. Along with these signals, to our knowledge, for the first time in the BP measurement studies, we included accelerometric and gyroscopic signals from a wearable device to compensate for motion during continuous BP prediction. Our prediction model is a long-short-term-memory (LSTM) architecture of a recurrent neural network (RNN), which accommodates the multiscale temporal dependency between the sequential raw signal values and the corresponding systolic and diastolic BP values. We performed a study with 50 healthy volunteers. The mean difference ± standard deviation (SD) of the RNN-based approach were 0.02±4.8 for SBP and 1.5±3.7 for DBP in seated position & 2.6±6.0 for SBP and 2.7±4.5 for DBP while walking. These values meet current validation standard requirements for measurement accuracy. Our experiments also demonstrate that the proposed RNN-based approach outperformed the classical linear regression model for BP prediction.
Shrimanti Ghosh, Ankur Banerjee, Nilanjan Ray, Peter W. Wood, Pierre Boulanger, Raj Padwal
ICASSP3
2018 Object Classification With Joint Projection and Low-Rank Dictionary Learning
abstract
For an object classification system, the most critical obstacles toward real-world applications are often caused by large intra-class variability, arising from different lightings, occlusion, and corruption, in limited sample sets. Most methods in the literature would fail when the training samples are heavily occluded, corrupted or have significant illumination or viewpoint variations. Besides, most of the existing methods and especially deep learning-based methods, need large training sets to achieve a satisfactory recognition performance. Although using the pre-trained network on a generic large-scale data set and fine-tune it to the small-sized target data set is a widely used technique, this would not help when the content of base and target data sets are very different. To address these issues simultaneously, we propose a joint projection and low-rank dictionary learning method using dual graph constraints. Specifically, a structured class-specific dictionary is learned in the low-dimensional space, and the discrimination is further improved by imposing a graph constraint on the coding coefficients, that maximizes the intra-class compactness and inter-class separability. We enforce structural incoherence and low-rank constraints on sub-dictionaries to reduce the redundancy among them, and also make them robust to variations and outliers. To preserve the intrinsic structure of data, we introduce a supervised neighborhood graph into the framework to make the proposed method robust to small-sized and high-dimensional data sets. Experimental results on several benchmark data sets verify the superior performance of our method for object classification of small-sized data sets, which include a considerable amount of different kinds of variation, and may have high-dimensional feature vectors.
Homa Foroughi, Nilanjan Ray, Hong Zhang 0013
IEEE Trans. Image Process.2
2017 Face recognition using multi-modal low-rank dictionary learning
abstract
Face recognition has been widely studied due to its importance in different applications; however, most of the proposed methods fail when face images are occluded or captured under illumination and pose variations. Recently several low-rank dictionary learning methods have been proposed and achieved promising results for noisy observations. While these methods are mostly developed for single-modality scenarios, recent studies demonstrated the advantages of feature fusion from multiple inputs. We propose a multi-modal structured low-rank dictionary learning method for robust face recognition, using raw pixels of face images and their illumination invariant representation. The proposed method learns robust and discriminative representations from contaminated face images, even if there are few training samples with large intra-class variations. Extensive experiments on different datasets validate the superior performance and robustness of our method to severe illumination variations and occlusion.
Homa Foroughi, Moein Shakeri, Nilanjan Ray, Hong Zhang 0013
ICIP3
2017 Automated 3D muscle segmentation from MRI data using convolutional neural network
abstract
In this paper, we propose an automated segmentation algorithm for human leg muscles from 3D MRI data using a deep convolutional neural network (CNN). Using a generalized cylinder model, a 3D human leg muscle is represented by two smooth 2D images. The CNN predicts these two images from raw 3D voxels. For our base CNN, we use a pre-trained AlexNet that is coupled with a principle components analysis (PCA)-head. The AlexNet predicts a compressed vector, which is then back-projected by the PCA-head into two 2D images representing a 3D leg muscle boundary. This structured-output CNN architecture is fine-tuned in an end-to-end fashion. Our proposed CNN outperforms the conventional model-based approach, the active appearance model (AAM) image segmentation algorithm. The average Dice score between the ground truth segmentation and the obtained segmentation image is 0.85 using our CNN model, whereas the AAM yields a Dice score of 0.60.
Shrimanti Ghosh, Pierre Boulanger, Scott T. Acton, Silvia S. Blemker, Nilanjan Ray
ICIP5
2017 A two-stage minimum spanning tree (MST) based clustering algorithm for 2D deformable registration of time sequenced images
abstract
Significant cardiac and respiratory motion of the living subject, occasional spells of defocus, drifts in the field of view, and long image sequences make the registration of in-vivo microscopy image sequences used in atherosclerosis study an onerous task. In this study we developed and implemented a novel Minimum Spanning Tree (MST)-based clustering method for image sequence registration that first constructs a minimum spanning tree for the input image sequence. The spanning tree re-orders the images in such a way where poor quality images appear at the end of the sequence. Then the spanning tree is clustered into several groups based on the similarity of the images. Subsequently deformable registration is conducted locally within the group with respect to the local anchor image selected automatically from the images in the group. After that coarse registration is performed to find the global anchor and then a deformable registration is performed globally to incorporate larger drift and distortion. Two-stage deformable registration incrementally incorporates larger drifts and distortions present in the longer sequence. Our algorithm involves very few tuning parameters, the optimal value of these parameters can be easily learned from data. Our method outperforms other methods on microscopy image sequences of mouse arteries.
Baidya Nath Saha, Nilanjan Ray, Sara McArdle, Klaus Ley
ICIP2
2017 Convolutional gated recurrent networks for video segmentation
abstract
Semantic segmentation has recently witnessed major progress, but most of the previous work focused on improving single image segmentation. In this paper, we introduce a novel approach to implicitly utilize temporal data in videos for online segmentation. This design receives a sequence of consecutive video frames and outputs the segmentation of the last frame. Convolutional gated recurrent networks are used for the recurrent part to preserve spatial connectivities in the image. This architecture is tested for both binary and semantic video segmentation tasks. Experiments are conducted on the recent benchmarks in SegTrack V2, Davis, Camvid, and Synthia. Using recurrent fully convolutional networks improved the baseline network performance in all of our experiments. Namely, 5% and 3% improvement of F-measure in SegTrack2 and Davis respectively, 5.7% and 1.6% improvement in mean IoU in Synthia and Camvid. Thus, RFCN networks can be seen as a method to improve any baseline segmentation network by embedding them into a recurrent module that utilizes temporal data.
Mennatullah Siam, Sepehr Valipour, Martin Jägersand, Nilanjan Ray
ICIP4
2017 A novel framework to integrate convolutional neural network with compressed sensing for cell detection
abstract
The ability to detect certain types of cells in a microscopy image is important for a wide range of clinical applications. Cells often present huge variations in density and appearance, and often occupy only a small portion of an image. Consequently, general object detection methods in computer vision do not meet accuracy requirements: false / missed detections prevail. In this paper, we apply convolutional neural network (CNN) to regress a fixed length vector from a microscopy image. Then, L1minimization / compressed sensing (CS) recovers a variable number of cell locations from this fixed-length predicted vector. Our contribution in this work is combining CS with CNN to solve cell detection problem in a regression framework, which needs to handle a variable number of cells. Our method relies on the observation that the number of pixels indicating cell centroid locations is a tiny fraction of the total image size. Thus, we utilize this sparsity property by CS. The proposed method is evaluated with several state-of-the-art approaches on public cell datasets and it obtains superior or comparable performances.
Nilanjan Ray, Judith Hugh, Gilbert Bigras
ICIP2
2017 Selecting the Optimal Sequence for Deformable Registration of Microscopy Image Sequences Using Two-Stage MST-based Clustering Algorithm
Baidya Nath Saha, Nilanjan Ray, Sara McArdle, Klaus Ley
MICCAI (1)2
2017 Recurrent Fully Convolutional Networks for Video Segmentation
abstract
Image segmentation is an important step in most visual tasks. While convolutional neural networks have shown to perform well on single image segmentation, to our knowledge, no study has been done on leveraging recurrent gated architectures for video segmentation. Accordingly, we propose and implement a novel method for online segmentation of video sequences that incorporates temporal data. The network is built from a fully convolutional network and a recurrent unit that works on a sliding window over the temporal data. We use convolutional gated recurrent unit that preserves the spatial information and reduces the parameters learned. Our method has the advantage that it can work in an online fashion instead of operating over the whole input batch of video frames. The network is tested on video segmentation benchmarks in Segtrack V2 and Davis. It proved to have 5% improvement in Segtrack and 3% improvement in Davis in F-measure over a plain fully convolutional network.
Sepehr Valipour, Mennatullah Siam, Martin Jägersand, Nilanjan Ray
WACV4
2017 Error bound of Nyström-approximated NCut eigenvectors and its application to training size selection
Li He 0002, Nilanjan Ray, Hong Zhang 0013
Neurocomputing2
2017 Deep deformable registration: Enhancing accuracy by fully convolutional neural net
Sayan Ghosal, Nilanjan Ray
Pattern Recognit. Lett.2
2016 MISTICA: Minimum Spanning Tree-Based Coarse Image Alignment for Microscopy Image Sequences
abstract
Registration of an in vivo microscopy image sequence is necessary in many significant studies, including studies of atherosclerosis in large arteries and the heart. Significant cardiac and respiratory motion of the living subject, occasional spells of focal plane changes, drift in the field of view, and long image sequences are the principal roadblocks. The first step in such a registration process is the removal of translational and rotational motion. Next, a deformable registration can be performed. The focus of our study here is to remove the translation and/or rigid body motion that we refer to here as coarse alignment. The existing techniques for coarse alignment are unable to accommodate long sequences often consisting of periods of poor quality images (as quantified by a suitable perceptual measure). Many existing methods require the user to select an anchor image to which other images are registered. We propose a novel method for coarse image sequence alignment based on minimum weighted spanning trees (MISTICA) that overcomes these difficulties. The principal idea behind MISTICA is to reorder the images in shorter sequences, to demote nonconforming or poor quality images in the registration process, and to mitigate the error propagation. The anchor image is selected automatically making MISTICA completely automated. MISTICA is computationally efficient. It has a single tuning parameter that determines graph width, which can also be eliminated by the way of additional computation. MISTICA outperforms existing alignment methods when applied to microscopy image sequences of mouse arteries.
Nilanjan Ray, Sara McArdle, Klaus Ley, Scott T. Acton
IEEE J. Biomed. Health Informatics1
2015 Joint Feature Selection with Low-rank Dictionary Learning
Homa Foroughi, Moein Shakeri, Nilanjan Ray, Hong Zhang 0013
BMVC3
2015 Robust people counting using sparse representation and random projection
Homa Foroughi, Nilanjan Ray, Hong Zhang 0013
Pattern Recognit.2
2015 Unique people count from monocular videos
Satarupa Mukherjee, Stephanie Gil, Nilanjan Ray
Vis. Comput.3
2014 People counting with image retrieval using compressed sensing
abstract
The estimation of the number of people present in an image has many applications such as intelligent transportation, urban planning and crowd surveillance. Rather than conventional counting by detection or regression/machine-learning methods, we propose an image retrieval approach, which uses an image descriptor to estimate the people count. We review the performance of several image descriptors. In addition, we propose a straightforward global image descriptor for image retrieval based on compressed sensing theory. Extensive evaluations on existing crowd analysis benchmark datasets demonstrate the effectiveness of our image retrieval-based approach compared to state-of-the-art regression-based people counting methods.
Homa Foroughi, Nilanjan Ray, Hong Zhang 0013
ICASSP2
2014 Registering sequences of in vivo microscopy images for cell tracking using dynamic programming and minimum spanning trees
abstract
Registration of in vivo microscopy image sequences is important for tracking of cells. Registering a long sequence of in vivo microscopy images is particularly challenging for several reasons, which include motion artifacts created by the cardiac cycle and breathing movements of the living subject, occasional defocussing, illumination change, and noise in image acquisition. To accommodate these variations, we sample time points redundantly during microscopic image acquisition. Second, we use dynamic programming to select image frames with tolerable motion and eliminate those with large motion. Third, we employ a novel method based on the minimum spanning tree algorithm to register the selected image frames. Testing on actual in vivo image sequences reveals that our approach excels over three existing registration methods in terms of structural image similarity of the registered images.
Sara McArdle, Scott T. Acton, Klaus Ley, Nilanjan Ray
ICIP4
2014 A robust convergence index filter for breast cancer cell segmentation
abstract
COnvergence INdex (COIN) filter, a successful tool for cell localization, evaluates the degree of convergence of the gradient vectors within the neighborhood (region of support) toward a pixel of interest. All previous efforts were to increase the adaptability of the region of support to make the COIN filter robust and accurate. However, improving the quality of the image gradient map was ignored, which results in poor performance of the members of the COIN family in noisy settings. We propose a new Robust Convergence Index (RCI) filter that tailors the COIN filter in a noisy environment by (a) spreading the gradient vectors within non-homogeneous object regions by convolving an Aggregated Edge Probability Map (AEPM) with an edge preserving gradient vector kernel, and (b) increasing the convergence of the gradient vectors through the integration of the sine and cosine distribution as well as the magnitude of the gradient vectors. AEPM is computed through the consensus of the responses of a number of edge detectors over a wide range of scales, which lessens the effects of clutter by enforcing higher weights to the actual edges, and a non-parametric Kernel Density Estimation (KDE) is used to compute the edge probability map. Experimental results demonstrate that it obtains state-of-the-art performance.
Baidya Nath Saha, Amritpal Saini, Nilanjan Ray, Russell Greiner, Judith Hugh, Mauro Tambasco
ICIP3
2014 VFCCV snake: A novel active contour model combining edge and regional information
abstract
Active contour models have been widely used for image segmentation. Among leading models of active contour is vector-field convolution (VFC), a parametric active contour that improves the popular gradient vector flow (GVF) model. However VFC is still sensitive to noise and can be easily trapped in cluttered regions of an image because it only considers edge information. Based on the geometric active contour model proposed by Chan and Vese, this paper introduces a novel active contour model that incorporates region information in VFC in order to take advantage of edge and regional information. This new model, which we refer to as VFCCV snake, is implemented in the parametric active contour framework, and has control on topology especially in noisy images and images with boundary gaps. Experimental results on both synthetic and real images show superior performance of our VFCCV snake to state-of-the-art leading active contour methods.
Jiuyu Sun, Nilanjan Ray, Hong Zhang 0013
ICIP2
2013 AR-Boost: Reducing Overfitting by a Robust Data-Driven Regularization Strategy
Baidya Nath Saha, Gautam Kunapuli, Nilanjan Ray, Joseph A. Maldjian, Sriraam Natarajan
ECML/PKDD (3)3
2013 Adaptive shape prior in graph cut image segmentation
Hong Zhang 0013, Nilanjan Ray
Pattern Recognit.3
2012 Spectrogram based features selection using multiple kernel learning for speech/music discrimination
abstract
This paper presents a multiple kernel learning (MKL) approach to speech/music discrimination (SMD). The time-frequency representation (spectrogram) implemented by short-time Fourier transform (STFT) of audio segment is decomposed by wavelet packet transform into different subband levels. The subbands, which contain rich texture information, are used as features for this discrimination problem. MKL technique is used to select the optimal subbands to discriminate the audio signals. The proposed MKL based algorithm is applied for SMD of a standard dataset. The experimental results show that the proposed technique yields noticeable improvements in classification accuracy and tolerance toward different noise types compared to the existing methods.
Sharmin Nilufar, Nilanjan Ray, Md. Khademul Islam Molla, Keikichi Hirose
ICASSP2
2012 Wavelet subband-based steam detection by multiple kernel learning
abstract
Wavelet transform coefficients have been shown as significant features for detecting steam and smoke. Wavelet transform is multi-resolution in nature; moreover, at each resolution, wavelet transform coefficients form a high dimensional feature set. In this paper we handle both these issues in a multiple kernel learning (MKL) framework. First, high dimensionality is handled by using a kernel function that measures similarity between two sets of wavelet coefficients at the same resolution. Next, we consider a convex combination of these kernel functions that correspond to all the available resolutions of the wavelet transform. The proposed MKL uses an L1norm linear support vector machine (SVM) for sparse learning of the convex combination. Then, this mixture kernel function is used in an L2norm nonlinear SVM for binary classification- image with steam or without steam. Our method yields encouraging results and outperforms other competing methods.
Sharmin Nilufar, Nilanjan Ray, Hong Zhang 0013
ICIP2
2012 Seeing through clutter: Snake computation with dynamic programming for particle segmentation
Nilanjan Ray, Scott T. Acton, Hong Zhang 0013
ICPR1
2012 Shape based appearance model for kernel tracking
Zhijie Wang 0003, Mohamed Ben Salah, Hong Zhang 0013, Nilanjan Ray
Image Vis. Comput.4
2012 Clump splitting via bottleneck detection and shape classification
Hong Zhang 0013, Nilanjan Ray
Pattern Recognit.3
2012 Shape based local thresholding for binarization of document images
Jichuan Shi, Nilanjan Ray, Hong Zhang 0013
Pattern Recognit. Lett.2
2012 Object Detection With DoG Scale-Space: A Multiple Kernel Learning Approach
abstract
Difference of Gaussians (DoG) scale-space for an image is a significant way to generate features for object detection and classification. While applying DoG scale-space features for object detection/classification, we face two inevitable issues: dealing with high dimensional data and selecting/weighting of proper scales. The scale selection process is mostly ad-hoc to date. In this paper, we propose a multiple kernel learning (MKL) method for both DoG scale selection/weighting and dealing with high dimensional scale-space data. We design a novel shift invariant kernel function for DoG scale-space. To select only the useful scales in the DoG scale-space, a novel framework of MKL is also proposed. We utilize a 1-norm support vector machine (SVM) in the MKL optimization problem for sparse weighting of scales from DoG scale-space. The optimized data-dependent kernel accommodates only a few scales that are most discriminatory according to the large margin principle. With a 2-norm SVM this learned kernel is applied to a challenging detection problem in oil sand mining: to detect large lumps in oil sand videos. We tested our method on several challenging oil sand data sets. Our method yields encouraging results on these difficult-to-process images and compares favorably against other popular multiple kernel methods.
Sharmin Nilufar, Nilanjan Ray, Hong Zhang 0013
IEEE Trans. Image Process.2
2011 Median filter with absolute value norm spatial regularization
abstract
We provide a novel formulation for computing median filter with spatial regularization as minimizing a cost function composed of absolute value norms. We turn this cost minimization into an equivalent linear programming (LP) and solve its dual LP as a minimum cost flow (MCF) problem. The MCF is solved over a graph constructed for an input image, and the primal LP solution is retrieved as the filtered image. For solving the MCF, we utilize an efficient network simplex algorithm. Numerical results show that the proposed median filter with a spatial regularization term outperforms median filters and a decision theoretic filter for impulse noise removal.
Nilanjan Ray
ICASSP1
2011 Anovel framework for automatic passenger counting
abstract
We propose a novel framework for counting passengers in a railway station. The framework has three components: people detection, tracking and validation. We detect every person using Hough circle when he or she enters the field of view. The person is then tracked using optical flow until (s)he leaves the field of view. Finally, the tracker generated trajectory is validated through a spatiotemporal background subtraction technique. The number of valid trajectories provides passenger count. Each of the three components of the proposed framework has been compared with competitive methods on three datasets of varying crowd densities. Extensive experiments have been conducted on the datasets having top views of the passengers. Experimental results demonstrate that the proposed algorithmic framework performs well both on dense and sparse crowds and it can successfully detect and track persons with different hair colors, hoodies, caps, long winter jackets, bags and so on. The proposed algorithm shows promising results also for people moving in different directions. The proposed framework can detect up to 30% more accurately and 20% more precisely than other competitive methods.
Satarupa Mukherjee, Baidya Nath Saha, Iqbal Jamal, Richard Leclerc, Nilanjan Ray
ICIP5
2011 Clump splitting via bottleneck detection
abstract
Under-segmentation of an image with multiple objects is a common problem in image segmentation algorithms. This paper presents a novel approach for the splitting of clumps formed by multiple objects due to under-segmentation. The algorithm includes two steps: finding a pair of points for clump splitting, and joining the pair of selected points. In the first step, a pair of points for splitting is detected using a bottleneck rule, under the assumption that the desired objects have roughly convex shape. In the second step, the selected pair of splitting points is joined by finding the optimal splitting line between them, based on minimizing an image energy. The performance of this method is evaluated using images from various applications. Experimental results show that the proposed approach has several advantages over existing splitting methods in identifying points for splitting as well as finding an accurate split line.
Hong Zhang 0013, Nilanjan Ray
ICIP3
2011 Computation of Fluid and Particle Motion From a Time-Sequenced Image Pair: A Global Outlier Identification Approach
abstract
Fluid motion estimation from time-sequenced images is a significant image analysis task. Its application is widespread in experimental fluidics research and many related areas like biomedical engineering and atmospheric sciences. In this paper, we present a novel flow computation framework to estimate the flow velocity vectors from two consecutive image frames. In an energy minimization-based flow computation, we propose a novel data fidelity term, which: 1) can accommodate various measures, such as cross-correlation or sum of absolute or squared differences of pixel intensities between image patches; 2) has a global mechanism to control the adverse effect of outliers arising out of motion discontinuities, proximity of image borders; and 3) can go hand-in-hand with various spatial smoothness terms. Further, the proposed data term and related regularization schemes are both applicable to dense and sparse flow vector estimations. We validate these claims by numerical experiments on benchmark flow data sets.
Nilanjan Ray
IEEE Trans. Image Process.1
2010 Automating Snakes for Multiple Objects Detection
Baidya Nath Saha, Nilanjan Ray, Hong Zhang 0013
ACCV (3)2
2010 A concave cost formulation for parametric curve fitting: Detection of leukocytes from intravital microscopy images
abstract
We formulate parametric curve fitting as a concave cost minimization problem. Our formulation is general encompassing any parametric curve where parameters can be free or constrained. The proposed concave cost opens the future possibility of applying several available concave programming algorithms in curve fitting. In this paper, we propose a fast local minimization of the concave cost and utilize it to a specific application- automated detection of ellipse-shaped leukocytes (white blood cells) from microscopy images. We illustrate that our solution can cope well with outliers compared to other competitive methods of ellipse fitting and leukocyte detection.
Nilanjan Ray
ICIP1
2010 Cell Tracking in Video Microscopy Using Bipartite Graph Matching
abstract
Automated visual tracking of cells from video microscopy has many important biomedical applications. In this paper, we model the problem of cell tracking over pairs of video microscopy image frames as a minimum weight matching problem in bipartite graphs. The bipartite matching essentially establishes one-to-one correspondences between the cells in different frames. A key advantage of using bipartite matching is the inherent scalability, which arises from its polynomial time-complexity. We propose two different tracking methods based on bipartite graph matching and properties of Gaussian distributions. In both the methods, i) the centers of the cells appearing in two frames are treated as vertices of a bipartite graph and ii) the weight matrix contains information about distance between the cells (in two frames) and cell velocity. In the first method, we identify fast-moving cells based on distance and filter them out using Gaussian distributions before the matching is applied. In the second method, we remove false matches using Gaussian distributions after the bipartite graph matching is employed. Experimental results indicate that both the methods are promising while the second method has higher accuracy.
Ananda S. Chowdhury, Rohit Chatterjee, Mayukh Ghosh, Nilanjan Ray
ICPR4
2010 Automatic segmentation of spinal cord mri using symmetric boundary tracing
abstract
We develop an adaptive active contour tracing algorithm for extraction of spinal cord from MRI that is fully automatic, unlike existing approaches that need manually chosen seeds. We can accurately extract the target spinal cord and construct the volume of interest to provide visual guidance for strategic rehabilitation surgery planning.
Dipti Prasad Mukherjee, Irene Cheng 0001, Nilanjan Ray, Vivian Mushahwar, R. Marc Lebel, Anup Basu
IEEE Trans. Inf. Technol. Biomed.3
2009 Optimum kernel function design from scale space features for object detection
abstract
Scale-space representation of an image is a significant way to generate features for classification. However, for a specific classification task, the entire scale-space may not be useful; only a part of it is typically effective. Toward this end, we design a data dependent classification kernel function, which is a weighted mixture of kernels defined on individual scales. In order to choose the optimum weights in the mixture kernel function (MKF), we propose an optimization criterion that leads to the minimization of Raleigh quotient in the positive orthant. This optimization is in general a difficult, non-convex, quadratically constrained quadratic programming. Utilizing a property of ratio of functions, we reduce the aforementioned optimization into a novel binary search, which is essentially a series of quadratic programming. As an application we choose a significant detection problem in oil sands mining called large lump detection from videos. Employing support vector classifier with our MKF yields encouraging results on these difficult-to-process images and compares favorably against the kernel alignment method as well as Fisher criterion adopted in.
Sharmin Nilufar, Nilanjan Ray, Hong Zhang 0013
ICIP2
2009 Solidity based local threshold for oil sand image segmentation
abstract
A novel local threshold algorithm for images with poor illumination and complex texture surface is presented in this paper. This algorithm improves segmentation quality by selecting local thresholds according to object level information incorporating prior knowledge, specifically the solidity features. Local thresholds are searched by maximizing the probability of solidity, and fragments with lower segmentation quality are filtered by the stability of solidity. Since thresholding results are produced with object level information, our algorithm is robust in dealing with images of poor quality. Experiments on oil sand images show the proposed algorithm has superior performance to existing local threshold approaches in terms of segmentation quality.
Jichuan Shi, Hong Zhang 0013, Nilanjan Ray
ICIP3
2009 Tracking of multiple interacting objects using a novel prediction model
abstract
Tracking multiple interacting objects is an interesting and difficult task in computer vision. Two common problems in this field are a single object with multiple tracks and a single track with multiple objects. Most of the existing algorithms address the first problem but not the second one. In this paper, to solve the second problem we propose a new algorithm with a novel prediction model, which exploits the idea of penalizing outliers in statistics. The experiments show that our proposed algorithm is more robust than the existing algorithms in tackling both the aforementioned problems.
Zhijie Wang 0003, Hong Zhang 0013, Nilanjan Ray
ICIP3
2009 Image thresholding by variational minimax optimization
Baidya Nath Saha, Nilanjan Ray
Pattern Recognit.2
2009 Snake Validation: A PCA-Based Outlier Detection Method
abstract
We utilize outlier detection by principal component analysis (PCA) as an effective step to automate snakes/active contours for object detection. The principle of our approach is straightforward: we allow snakes to evolve on a given image and classify them into desired object and non-object classes. To perform the classification, an annular image band around a snake is formed. The annular band is considered as a pattern image for PCA. Extensive experiments have been carried out on oil-sand and leukocyte images and the performance of the proposed method has been compared with two other automatic initialization and two gradient-based outlier detection techniques. Results show that the proposed algorithm improves the performance of automatic initialization techniques and validates snakes more accurately than other outlier detection methods, even when considerable object localization error is present.
Baidya Nath Saha, Nilanjan Ray, Hong Zhang 0013
IEEE Signal Process. Lett.2
2008 Computing oil sand particle size distribution by snake-PCA algorithm
abstract
An important measure in various stages of oil sand mining is particle size distribution (PSD) of oil sand particles. Currently PSD is found by time consuming manual inspection. An effective automation of PSD computation can play a significant role in improving the mining process. Toward this goal we propose an algorithm (snake-PCA) to detect oil sands from conveyor belt images, which pose considerable challenges to automated analysis. The novelty in snake-PCA is as follows. First, snake-PCA evolves a number of snakes based on a novel variation of gradient vector flow requiring only a point as initialization. Oil sand is then detected by applying a threshold on PCA reconstruction error of a novel pattern image formed on each evolved snake. We show the discriminative property of the proposed pattern image here. Also, our detection experiments with snake-PCA produce a PSD matching well with a manually found PSD.
Baidya Nath Saha, Nilanjan Ray, Hong Zhang 0013
ICASSP2
2008 Oil sand image segmentation using the inclusion filter
abstract
Oil sands may constitute two thirds of the world's oil reserves. To efficiently harvest this important resource, image analysis is required to quantify production related performance in terms of particle size distribution. We utilize connected filters to simplify the oil sand images and to generate a robust segmentation. Specifically, a self-dual operator called the inclusion filter is applied to the difficult segmentation problem. The inclusion filter removes minor interior regions and clutter based on the connected component relationships defined by the adjacency forest. We show that the use of the inclusion filter significantly improves the edge fidelity and the insensitivity to initialization for the oil sand application.
Nilanjan Ray, Baidya Nath Saha, Scott T. Acton
ICIP1
2008 Graph-cut optimization of the ratio of functions and its application to image segmentation
abstract
Optimizing the ratio of two functions of binary variables is a common task in many image analysis applications. In general, such a ratio is not amenable to graph-cut based optimization. In this paper, we show that if the numerator and the denominator of a ratio are individually graph-representable functions, then their ratio can be optimized via graph-cut based technique. As an example of such a ratio function we choose Yezzi et al.’s energy function [2], minimization of which produces a binary labeling of an image. Through examples, we illustrate the advantage of working with graph-cut-based optimization for the aforementioned ratio in finding a global solution as opposed to the local solutions found by level set methods proposed in [2].
Maddy Hui Wang, Nilanjan Ray, Hong Zhang 0013
ICIP2
2007 Edge Sensitive Variational Image Thresholding
abstract
In this paper we propose a locally adaptive image threshold technique via variational energy minimization. The novelty of the proposed method is that from an image it automatically computes the weights on the data fidelity and the regularization terms in the energy functional, unlike many other previously proposed variational formulations that require manual input of these weights by laborious trial and error. To achieve the automatic setting of the weighting parameters we propose a non-linear convex combination of the data fidelity and the regularization terms in the energy functional and seek the optimum threshold surface via minimax principle. Our choice of the novel energy functional allows fast computation of the unique minimax solution. As a specific segmentation application, the proposed technique shows promising results when applied to find lung boundary from MR imagery. Illustrative examples are also provided where the proposed method is observed to retain texture information better than other competing methods.
Nilanjan Ray, Baidya Nath Saha
ICIP (6)1
2005 Spatiotemporal Segmentation for Validation of Rolling Leukocyte Tracking Data
abstract
Processing of bulk microscopy video data requires automated tracking of rolling leukocytes in the hundreds to compute a rolling velocity distribution, which is an indispensable descriptor in inflammation research and anti/pro-inflammatory drug testing. However, for any automated tracking method to be successful, an automated validation process must exist to accept or reject the output of tracking. In this paper, we propose an automated validation technique that first generates a spatiotemporal image from the cell locations output by a tracking method; then, it segments the spatiotemporal image to detect the presence or absence of a leukocyte by employing an edge-response filter followed by an active contour method. The proposed direction sensitive edge-response filter, the maximum absolute average directional derivative (MAADD), computes the magnitude of the mean directional derivative over an oriented line segment and chooses the maximum of all such values within a range of orientations of the line segment. Our validation experiments show that the proposed method is successful in 93% of the trials using manual tracking, in 83% using correlation tracking and in 84% using active contour tracking method.
Nilanjan Ray, Scott T. Acton
ICASSP (2)1
2005 Tracking multiple cells by correspondence resolution in a sequential Bayesian framework
abstract
We propose a multi-target tracking (MTT) algorithm in a sequential Bayesian framework that computes cell velocities from video microscopy. Unlike the traditional tracking methods, our formulation does not involve the estimation of target states; instead, we estimate one-to-one target correspondences by way of a sequential Markov chain Monte Carlo (MCMC) algorithm. The proposed probabilistic framework also automatically accounts for a variable number of targets. We have tested the proposed tracking algorithm on two different in vitro and one in vivo microscopy experiments. The three experiments show that the method holds promise in terms of low false positive and false negative rates as well as low rates of correspondence error.
Nilanjan Ray, Gang Dong, Scott T. Acton
ICIP (1)1
2005 Inclusion filters: a class of self-dual connected operators
abstract
In this paper, we define a connected operator that either fills or retains the holes of the connected sets depending on application-specific criteria that are increasing in the set theoretic sense. We refer to this class of connected operators as inclusion filters, which are shown to be increasing, idempotent, and self dual (gray-level inversion invariance). We demonstrate self duality for 8-adjacency on a discrete Cartesian grid. Inclusion filters are defined first for binary-valued images, and then the definition is extended to grayscale imagery. It is also shown that inclusion filters are levelings, a larger class of connected operators. Several important applications of inclusion filters are demonstrated-automatic segmentation of the lung cavities from magnetic resonance imagery, user interactive shape delineation in content-based image retrieval, registration of intravital microscopic video sequences, and detection and tracking of cells from these sequences. The numerical performance measures on 100-cell tracking experiments show that the use of inclusion filter improves the total number of frames successfully tracked by five times and provides a threefold reduction in the overall position error.
Nilanjan Ray, Scott T. Acton
IEEE Trans. Image Process.1
2005 Intravital leukocyte detection using the gradient inverse coefficient of variation
abstract
The problem of identifying and counting rolling leukocytes within intravital microscopy is of both theoretical and practical interest. Currently, methods exist for tracking rolling leukocytes in vivo, but these methods rely on manual detection of the cells. In this paper we propose a technique for accurately detecting rolling leukocytes based on Bayesian classification. The classification depends on a feature score, the gradient inverse coefficient of variation (GICOV), which serves to discriminate rolling leukocytes from a cluttered environment. The leukocyte detection process consists of three sequential steps: the first step utilizes an ellipse matching algorithm to coarsely identify the leukocytes by finding the ellipses with a locally maximal GICOV. In the second step, starting from each of the ellipses found in the first step, a B-spline snake is evolved to refine the leukocytes boundaries by maximizing the associated GICOV score. The third and final step retains only the extracted contours that have a GICOV score above the analytically determined threshold. Experimental results using 327 rolling leukocytes were compared to those of human experts and currently used methods. The proposed GICOV method achieves 78.6% leukocyte detection accuracy with 13.1% false alarm rate.
Gang Dong, Nilanjan Ray, Scott T. Acton
IEEE Trans. Medical Imaging2
2004 Level set analysis for leukocyte detection and tracking
abstract
We propose a cell detection and tracking solution using image-level sets computed via threshold decomposition. In contrast to existing methods where manual initialization is required to track individual cells, the proposed approach can automatically identify and track multiple cells by exploiting the shape and intensity characteristics of the cells. The capture of the cell boundary is considered as an evolution of a closed curve that maximizes image gradient along the curve enclosing a homogeneous region. An energy functional dependent upon the gradient magnitude along the cell boundary, the region homogeneity within the cell boundary and the spatial overlap of the detected cells is minimized using a variational approach. For tracking between frames, this energy functional is modified considering the spatial and shape consistency of a cell as it moves in the video sequence. The integrated energy functional complements shape-based segmentation with a spatial consistency based tracking technique. We demonstrate that an acceptable, expedient solution of the energy functional is possible through a search of the image-level lines: boundaries of connected components within the level sets obtained by threshold decomposition. The level set analysis can also capture multiple cells in a single frame rather than iteratively computing a single active contour for each individual cell. Results of cell detection using the energy functional approach and the level set approach are presented along with the associated processing time. Results of successful tracking of rolling leukocytes from a number of digital video sequences are reported and compared with the results from a correlation tracking scheme.
Dipti Prasad Mukherjee, Nilanjan Ray, Scott T. Acton
IEEE Trans. Image Process.2
2004 Motion gradient vector flow: an external force for tracking rolling leukocytes with shape and size constrained active contours
abstract
Recording rolling leukocyte velocities from intravital microscopic video imagery is a critical task in inflammation research and drug validation. Since manual tracking is excessively time consuming, an automated method is desired. This paper illustrates an active contour based automated tracking method, where we propose a novel external force to guide the active contour that takes the hemodynamic flow direction into account. The construction of the proposed force field, referred to as motion gradient vector flow (MGVF), is accomplished by minimizing an energy functional involving the motion direction, and the image gradient magnitude. The tracking experiments demonstrate that MGVF can be used to track both slow- and fast-rolling leukocytes, thus extending the capture range of previously designed cell tracking techniques.
Nilanjan Ray, Scott T. Acton
IEEE Trans. Medical Imaging1
2003 Self-dual inclusion filters for grayscale imagery
abstract
Using the structure of an adjacency-tree for binary-valued images, we define inclusion filters, a class of connected operators. Inclusion filters modify the image by filling or retaining the holes of the connected components of foreground and those of the background of a binary image depending on application-specific criteria, which are increasing in the set theoretic sense. We demonstrate a straightforward method to achieve self-duality (gray level inversion invariance) of inclusion filters on the discrete Cartesian domain by considering only 8-adjacency. Inclusion filters are extended to the grayscale images by the threshold decomposition principle. As an application, inclusion filters are shown to improve the performance of snake-based tracking of leukocytes observed in intravital microscopic video imagery. In this set of experiments the mean position error is reduced by a factor of 2.5 using the inclusion filter.
Nilanjan Ray, Scott T. Acton
ICIP (1)1
2003 Merging Parametric Active Contours Within Homogeneous Image Regions for MRI-Based Lung Segmentation
abstract
Inhaled hyperpolarized helium-3 (3He) gas is a new magnetic resonance (MR) contrast agent that is being used to study lung functionality. To evaluate the total lung ventilation from the hyperpolarized 3He MR images, it is necessary to segment the lung cavities. This is difficult to accomplish using only the hyperpolarized 3He MR images, so traditional proton (1H) MR images are frequently obtained concurrent with the hyperpolarized 3He MR examination. Segmentation of the lung cavities from traditional proton (1H) MRI is a necessary first step in the analysis of hyperpolarized 3He MR images. In this paper, we develop an active contour model that provides a smooth boundary and accurately captures the high curvature features of the lung cavities from the 1H MR images. This segmentation method is the first parametric active contour model that facilitates straightforward merging of multiple contours. The proposed method of merging computes an external force field that is based on the solution of partial differential equations with boundary condition defined by the initial positions of the evolving contours. A theoretical connection with fluid flow in porous media and the proposed force field is established. Then by using the properties of fluid flow we prove that the proposed method indeed achieves merging and the contours stop at the object boundary as well. Experimental results involving merging in synthetic images are provided. The segmentation technique has been employed in lung 1H MR imaging for segmenting the total lung air space. This technology plays a key role in computing the functional air space from MR images that use hyperpolarized 3He gas as a contrast agent.
Nilanjan Ray, Scott T. Acton, Talissa A. Altes, Eduard E. de Lange, James R. Brookeman
IEEE Trans. Medical Imaging1
2002 Tracking fast-rolling leukocytes in vivo with active contours
abstract
We propose and demonstrate an active contour technique to track fast-rolling leukocytes observed in vivo from video microscopy. A rolling leukocyte is an activated white blood cell that interacts with the vessel wall (the endothelium) in the inflammatory process. Tracking is enhanced here to accommodate fast-moving cells. To tackle the task of tracking wherein only low temporal resolution is possible, we have introduced an energy-minimizing framework and obtained a partial differential equation (PDE) based active contour evolution technique. The proposed PDEs are shown to be an initialization-insensitive version of the gradient vector flow (GVF) proposed by Xu and Prince (1998). We modify the GVF-PDEs by adding a Dirichlet type boundary condition (BC) based on the initial position of the active contour and the direction of cell movement. Using actual intravital experiments, we compare the performance of the proposed active contour tracker with the Dirichlet BC, the active contour tracker without the BC, the correlation tracker and the centroid tracker. The comparative results provide evidence of the advantages of the proposed method in terms of increased number of frames successfully tracked and reduced localization error.
Nilanjan Ray, Scott T. Acton
ICIP (3)1
2002 Tracking Leukocytes In Vivo with Shape and Size Constrained Active Contours
abstract
Inflammatory disease is initiated by leukocytes (white blood cells) rolling along the inner surface lining of small blood vessels called postcapillary venules. Studying the number and velocity of rolling leukocytes is essential to understanding and successfully treating inflammatory diseases. Potential inhibitors of leukocyte recruitment can be screened by leukocyte rolling assays and successful inhibitors validated by intravital microscopy. In this paper, we present an active contour or snake-based technique to automatically track the movement of the leukocytes. The novelty of the proposed method lies in the energy functional that constrains the shape and size of the active contour. This paper introduces a significant enhancement over existing gradient-based snakes in the form of a modified gradient vector flow. Using the gradient vector flow, we can track leukocytes rolling at high speeds that are not amenable to tracking with the existing edge-based techniques. We also propose a new energy-based implicit sampling method of the points on the active contour that replaces the computationally expensive explicit method. To enhance the performance of this shape and size constrained snake model, we have coupled it with Kalman filter so that during coasting (when the leukocytes are completely occluded or obscured), the tracker may infer the location of the center of the leukocyte. Finally, we have compared the performance of the proposed snake tracker with that of the correlation and centroid-based trackers. The proposed snake tracker results in superior performance measures, such as reduced error in locating the leukocyte under tracking and improvements in the percentage of frames successfully tracked. For screening and drug validation, the tracker shows promise as an automated data collection tool.
Nilanjan Ray, Scott T. Acton, Klaus Ley
IEEE Trans. Medical Imaging1
2001 MRI ventilation analysis by merging parametric active contours
abstract
A novel technique that combines MR imaging of hyperpolarized helium gas and conventional MR imaging facilitates, high resolution imaging of lung functionality for the first time. We put forth a segmentation method for measuring the total lung air space and a classification approach to computing the functional air space. For segmentation, we introduce a parametric active contour that allows automated merging of multiple contours. The active contour technique uses gradient vector flow modified and strengthened by a boundary condition that inhibits contour crossover. The active contour approach is computationally inexpensive and is independent of initial contour placement. For classification of the functional lung air space in the helium images, a fuzzy c-means technique is applied. The classification results, in conjunction with the segmentation, allow the analysis of ventilation. The resultant biomedical image analysis tool can used in determining the efficacy of certain respiratory treatments.
Nilanjan Ray, Scott T. Acton, Talissa A. Altes, Eduard E. de Lange
ICIP (2)1
2001 Active contour segmentation guided by AM-FM dominant component analysis
abstract
For the first time, we explore the application of active contours in the modulation domain by computing snakes on image modulations. As we demonstrate in the examples, such snakes are able to utilize information inherent in the dominant image modulations to acquire and track visually and semantically meaningful structures within the image. We use nonlinear AM-FM image representations to capture regions that are homogeneous in intensity and in texture. A geometric snake approach utilizing a fuzzy classifier is then applied to the image modulations. The combination of AM-FM analysis and the active contour evolution produces an efficacious image partition. As a preliminary demonstration of this novel approach, we apply the modulation domain snakes to the classical texture segmentation problem.
Nilanjan Ray, Joseph P. Havlicek, Scott T. Acton, Marios S. Pattichis
ICIP (1)1
2001 A fast and flexible multiresolution snake with a definite termination criterion
Nilanjan Ray, Bhabatosh Chanda, Jyotirmay Das
Pattern Recognit.1