Amit Sethi

dblp:32/455 · DBLP profile ↗
← Back
50ranked-venue papers
2as first author
27since 2021 · last 2026
0000-0002-8634-1804ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 13 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Weight Entropy-Maximised Evidential Metamodel for Uncertainty Quantification (Student Abstract)
abstract
Reliable uncertainty quantification (UQ) is crucial for deploying deep learning models in safety-critical domains. Existing UQ methods often either rely on multi-pass inference, which increases computational cost, or restrict expressiveness by using only final-layer embeddings. In this work, we propose a lightweight evidential meta-model that leverages multi-layer feature fusion from a pretrained backbone, capturing both low-level features and high-level semantics to better estimate uncertainty. To further enhance epistemic fidelity, we integrate maximum weight-entropy (Max-WEnt) regularization, which encourages hypothesis diversity without altering the base network or adding test-time overhead. Experiments across two benchmark settings, medical (BACH, HAM10000, BreakHIS, DIV2K) and natural (ImageNet, SVHN, Fashion-MNIST, ImageNet-C) datasets, demonstrate consistent improvements in AUROC of out-of-distribution detection compared to prior post-hoc UQ methods. Our findings show that combining multi-layer evidential modeling with Max-WEnt provides a robust, efficient, and practical framework for trustworthy AI in high-stakes applications. The meta-model adds only ~0.8M parameters and trains in under four hours on a single 48GB GPU, making it practical for real-world deployment.
Gouranga Bala, Abhimanyu Chauhan, Amit Sethi
AAAI3
2026 UniVarFL: Uniformity and Variance Regularized Federated Learning for Heterogeneous Data (Student Abstract)
abstract
Federated Learning (FL) often suffers from severe performance degradation when faced with non-IID data, largely due to local classifier bias. Traditional remedies such as global model regularization or layer freezing either incur high computational costs or struggle to adapt to feature shifts. In this work, we propose UniVarFL, a novel FL framework that emulates IID-like training dynamics directly at the client level, eliminating the need for global model dependency. UniVarFL leverages two complementary regularization strategies during local training: Classifier Variance Regularization, which aligns class-wise probability distributions with those expected under IID conditions, effectively mitigating local classifier bias; and Hyperspherical Uniformity Regularization, which encourages a uniform distribution of feature representations across the hypersphere, thereby enhancing the model’s ability to generalize under diverse data distributions. Extensive experiments on multiple benchmark datasets demonstrate that UniVarFL outperforms existing methods in accuracy, highlighting its potential as a highly scalable and efficient solution for real-world FL deployments, especially in resource-constrained settings.
Sunny Gupta, Nikita Jangid, Amit Sethi
AAAI3
2026 Federated Cross-Modal Style-Aware Prompt Generation (Student Abstract)
abstract
Existing federated prompt learning methods for vision-language models like CLIP rely solely on text-based prompts and final-layer visual features, missing crucial multiscale visual details and client-specific style variations. This limits generalization across non-IID distributions and novel classes. We introduce FedCSAP (Federated Cross-Modal Style-Aware Prompt Generation), which harnesses multiscale features from CLIP's vision encoder alongside domain-aware style statistics from client data. By fusing these visual representations with textual context, FedCSAP generates adaptive, context-aware prompts that enhance robustness across seen and unseen classes. Our privacy-preserving approach operates through local training and global aggregation, effectively handling heterogeneous client distributions. Experiments on multiple image classification datasets demonstrate that FedCSAP significantly outperforms existing federated prompt learning methods in both accuracy and generalization.
Suraj Prasad, Navyansh Mahla, Sunny Gupta, Amit Sethi
AAAI4
2026 Network Inversion for Uncertainty-Aware Out-of-Distribution Detection (Student Abstract)
abstract
Out-of-distribution (OOD) detection and uncertainty estimation (UE) are critical components for building safe machine learning systems. In this work, we propose a novel framework that combines network inversion with classifier training to simultaneously address both OOD detection and uncertainty estimation. We extend a standard n-class classifier by adding an (n+1)-th "garbage" class to capture outliers, initially populated with random Gaussian noise. After each training epoch, we use network inversion to reconstruct inputs for all classes; incoherent reconstructions are assigned to the garbage class for retraining. This iterative cycle of training, inversion, and exclusion continues until inverted samples resemble in-distribution data and uncertainty drops, indicating learned decision boundaries and cleaner class manifolds. At inference, the model detects OOD inputs by classifying them as garbage, with confidence scores estimating uncertainty. Unlike prior methods, this approach requires no external OOD data or post-hoc calibration, providing a simple, unified solution for robust classification, OOD detection, and uncertainty estimation.
Pirzada Suhail, Rehna Afroz, Gouranga Bala, Amit Sethi
AAAI4
2026 Shortcut Learning Susceptibility in Vision Classifiers (Student Abstract)
abstract
Shortcut learning, where machine learning models exploit spurious correlations in data instead of capturing meaningful features, poses a significant challenge to building generalizable models. Vision classifiers based on Convolutional Neural Networks (CNNs), Multi-Layer Perceptrons (MLPs), and Vision Transformers (ViTs) leverage distinct architectural principles to process spatial and structural information, making them differently susceptible to shortcut learning. In this study, we systematically evaluate these architectures by introducing deliberate shortcuts into the dataset that are correlated with class labels both positionally and via intensity, creating a controlled setup to assess whether models rely on these artificial cues or learn actual distinguishing features. We perform both quantitative evaluation by training on the shortcut-modified dataset and testing on two different test sets—one containing the same shortcuts and another without them—to determine the extent of reliance on shortcuts. Additionally, qualitative evaluation is performed using network inversion-based reconstruction techniques to analyze what the models internalize in their weights, aiming to reconstruct the training data as perceived by the classifiers. Further, we evaluate susceptibility to shortcut learning across different learning rates. Our analysis reveals that CNNs at lower learning rates tend to be more reserved against entirely picking up shortcut features, while ViTs, particularly those without positional encodings, almost entirely ignore the distinctive image features in the presence of shortcuts.
Pirzada Suhail, Vrinda Goel, Amit Sethi
AAAI3
2025 WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency (Student Abstract)
abstract
Recent advancements in single image super-resolution have been predominantly driven by token mixers and transformer architectures. WaveMixSR utilized the WaveMix architecture, employing a two-dimensional discrete wavelet transform for spatial token mixing, achieving superior performance in super-resolution tasks with remarkable resource efficiency. In this work, we present an enhanced version of the WaveMixSR architecture by (1) replacing the traditional transpose convolution layer with a pixel shuffle operation and (2) implementing a multistage design for higher resolution tasks (4x). Our experiments demonstrate that our enhanced model -- WaveMixSR-V2 -- outperforms other architectures in multiple super-resolution tasks, achieving state-of-the-art for the BSD100 dataset, while also consuming fewer resources and exhibiting higher parameter efficiency and throughput.
Pranav Jeevan, Neeraj Nixon, Amit Sethi
AAAI3
2025 Network Inversion of Convolutional Neural Nets (Student Abstract)
abstract
Neural networks have emerged as powerful tools across various applications, yet their decision-making process often remains opaque, leading to them being perceived as "black boxes." This opacity raises concerns about their interpretability and reliability, especially in safety-critical scenarios. Network inversion techniques offer a solution by allowing us to peek inside these black boxes, revealing the features and patterns learned by the networks behind their decision-making processes and thereby provide valuable insights into how neural networks arrive at their conclusions, making them more interpretable and trustworthy. This paper presents a simple yet effective approach to network inversion using a meticulously conditioned generator that learns the data distribution in the input space of the trained neural network, enabling the reconstruction of inputs that would most likely lead to the desired outputs. To capture the diversity in the input space for a given output, instead of simply revealing the conditioning labels to the generator, we encode the conditioning label information into vectors and intermediate matrices and further minimize the cosine similarity between features of the generated images.
Pirzada Suhail, Amit Sethi
AAAI2
2025 Softmax-Weighted Pseudo-Label Refinement for Enhancing Robustness Against Label Noise
abstract
Deep neural networks often require large-scale, accurately labeled datasets to perform well, but in practice the labels are frequently corrupted by noise in medical imaging, especially instance-dependent noise. In this work, we propose a novel framework to address instance-dependent label noise by integrating three key components: (i) self- supervised pretraining using SimCLR to learn robust, noise-agnostic feature representations; (ii) an iterative pseudo-label refinement strategy employing a stage-wise consensus mechanism to progressively correct mislabeled samples; and (iii) a softmax-weighted crossentropy loss that dynamically down-weighs uncertain predictions. We validate our approach on benchmark datasets such as CIFAR10 and CIFAR-100 corrupted with synthetic noise at 20 %, 30 % and 50 % levels, demonstrating significant improvements over state-of-the-art methods. We further validated our method on Chest X-rays medical imaging datasets.
Gouranga Bala, Amit Sethi
BIBE3
2024 Clustered Patch Embeddings for Permutation-Invariant Classification of Whole Slide Images
abstract
Whole Slide Imaging (WSI) is a cornerstone of digital pathology, offering detailed insights critical for diagnosis and research. Yet, the gigapixel size of WSIs imposes significant computational challenges, limiting their practical utility. Our novel approach addresses these challenges by leveraging various encoders for intelligent data reduction and employing a different classification model to ensure robust, permutationinvariant representations of WSIs. A key innovation of our method is the ability to distill the complex information of an entire WSI into a single vector, effectively capturing the essential features needed for accurate analysis. This approach significantly enhances the computational efficiency of WSI analysis, enabling more accurate pathological assessments without the need for extensive computational resources. This breakthrough equips us with the capability to effectively address the challenges posed by large image resolutions in whole-slide imaging, paving the way for more scalable and effective utilization of WSIs in medical diagnostics and research, marking a significant advancement in the field.
Ravi Kant Gupta, Shounak Das, Amit Sethi
BIBE3
2024 Efficient Whole Slide Image Classification Through Fisher Vector Representation
abstract
The advancement of digital pathology, particularly through computational analysis of whole slide images (WSI), is poised to significantly enhance diagnostic precision and efficiency. However, the large size and complexity of WSIs make it difficult to analyze and classify them using computers. This study introduces a novel method for WSI classification by automating the identification and examination of the most informative patches, thus eliminating the need to process the entire slide. Our method involves two-stages: firstly, it extracts only a few patches from the WSIs based on their pathological significance; and secondly, it employs Fisher vectors (FVs) for representing features extracted from these patches, which is known for its robustness in capturing fine-grained details. This approach not only accentuates key pathological features within the WSI representation but also significantly reduces computational overhead, thus making the process more efficient and scalable. We have rigorously evaluated the proposed method across multiple datasets to benchmark its performance against comprehensive WSI analysis and contemporary weakly-supervised learning methodologies. The empirical results indicate that our focused analysis of select patches, combined with Fisher vector representation, not only aligns with, but at times surpasses, the classification accuracy of standard practices. Moreover, this strategy notably diminishes computational load and resource expenditure, thereby establishing an efficient and precise framework for WSI analysis in the realm of digital pathology.
Ravi Kant Gupta, Dadi Dharani, Shambhavi Shanker, Amit Sethi
BIBE4
2024 Classification and Morphological Analysis of Dlbcl Subtypes in H&E-Stained Slides
abstract
We address the challenge of automated classification of diffuse large B-cell lymphoma (DLBCL) into its two primary subtypes: activated B-cell-like (ABC) and germinal center B-cell-like (GCB). Accurate classification between these subtypes is essential for determining the appropriate therapeutic strategy, given their distinct molecular profiles and treatment responses. Our proposed deep learning model demonstrates robust performance, achieving an average area under the curve (AUC) of$(87.4 \pm 5.7) \%$during cross-validation. It shows a high positive predictive value (PPV), highlighting its potential for clinical application, such as triaging for molecular testing. To gain biological insights, we performed an analysis of morphological features of ABC and GCB subtypes. We segmented cell nuclei using a pre-trained deep neural network and compared the statistics of geometric and color features for ABC and GCB. We found that the distributions of these features were not very different for the two subtypes, which suggests that the visual differences between them are more subtle. These results underscore the potential of our method to assist in more precise subtype classification and can contribute to improved treatment management and outcomes for patients of DLBCL.
Ravi Kant Gupta, Mohit Jindal, Garima Jain, Epari Sridhar, Subhash Yadav, Hasmukh Jain, Tanuja Shet, Uma Sakhdeo, Manju Sengar, Lingaraj Nayak, Bhausaheb Bagal, Umesh Apkare, Amit Sethi
BIBE13
2024 HER2 and Fish Status Prediction in Breast Biopsy H&E-Stained Images Using Deep Learning
abstract
The current standard for detecting human epidermal growth factor receptor 2 (HER2) status in breast cancer patients relies on HER2 expression identified through immunohistochemistry (IHC) or amplification identified through fluorescence in situ hybridization (FISH). However, hematoxylin and eosin (H&E) tumor stains are more widely available, and accurately predicting HER2 status using H&E could reduce costs and expedite treatment selection. Deep Learning algorithms for H&E have shown effectiveness in predicting various cancer features and clinical outcomes, including moderate success in HER2 status prediction. In this work, we employed a customized weak supervision classification technique combined with MoCov2 contrastive learning for self-supervised feature extraction training to predict HER2 status. We trained our pipeline on 182 publicly available H&E whole slide images (WSIs) from The Cancer Genome Atlas (TCGA), for which annotations by the pathology team at Yale School of Medicine are publicly available. Our pipeline achieved an Area Under the Curve (AUC) of 0.85$\pm {0. 0 2}$across four different test folds. Additionally, we tested our model on 44 H&E slides from the TCGA-BRCA dataset, which had an HER2 score of$2+$and included corresponding HER2 status and FISH test results. These cases are considered equivocal for IHC, requiring an expensive FISH test on their IHC slides for disambiguation. Our pipeline demonstrated an AUC of 0.81 on these challenging H & E slides. Reducing the need for FISH test can have significant implications in cancer treatment equity for underserved populations.
Ardhendu Sekhar, Vrinda Goel, Garima Jain, Abhijeet Patil, Ravi Kant Gupta, Tripti Bameta, Swapnil Rane, Amit Sethi
BIBE8
2024 FLeNS: Federated Learning with Enhanced Nesterov-Newton Sketch
abstract
Federated learning faces a critical challenge in balancing communication efficiency with rapid convergence, especially for second-order methods. While Newton-type algorithms achieve linear convergence in communication rounds, transmitting full Hessian matrices is often impractical due to quadratic complexity. We introduce Federated Learning with Enhanced Nesterov-Newton Sketch (FLeNS), a novel method that harnesses both the acceleration capabilities of Nesterov’s method and the dimensionality reduction benefits of Hessian sketching. FLeNS approximates the centralized Newton’s method without relying on the exact Hessian, significantly reducing communication overhead. By combining Nesterov’s acceleration with adaptive Hessian sketching, FLeNS preserves crucial second-order information while preserving the rapid convergence characteristics. Our theoretical analysis, grounded in statistical learning, demonstrates that FLeNS achieves super-linear convergence rates in communication rounds - a notable advancement in federated optimization. We provide rigorous convergence guarantees and characterize tradeoffs between acceleration, sketch size, and convergence speed. Extensive empirical evaluation validates our theoretical findings, showcasing FLeNS’s state-of-the-art performance with reduced communication requirements, particularly in privacy-sensitive and edge-computing scenarios. The code is available at https://github.com/sunnyinAI/FLeNS
Sunny Gupta, Mohit Jindal, Pankhi Kashyap, Pranav Jeevan, Amit Sethi
IEEE Big Data5
2024 IFSENet: Harnessing Sparse Iterations for Interactive Few-Shot Segmentation Excellence
Shreyas Chandgothia, Ardhendu Sekhar, Amit Sethi
ICPR (4)3
2024 IDAL: Improved Domain Adaptive Learning for Natural Images Dataset
Ravi Kant Gupta, Shounak Das, Amit Sethi
ICPR (1)3
2024 Adversarial Transport Terms for Unsupervised Domain Adaptation
Chirag P, Mukta Wagle, Ravi Kant Gupta, Pranav Jeevan, Amit Sethi
ICPR (2)5
2024 WaveMixSR: Resource-efficient Neural Network for Image Super-resolution
abstract
Image super-resolution research recently has been dominated by transformer models which need higher computational resources than CNNs due to the quadratic complexity of self-attention. We propose a new neural network – WaveMixSR – for image super-resolution based on the WaveMix architecture which uses a 2D-discrete wavelet transform for spatial token-mixing. Unlike transformer-based models, WaveMixSR does not unroll the image as a sequence of pixels/patches. It uses the inductive bias of convolutions along with the lossless token-mixing property of wavelet transform to achieve higher performance while requiring fewer resources and training data. We compare the performance of our network with other state-of-the-art methods for image super-resolution. Our experiments show that WaveMixSR achieves competitive performance in all datasets and reaches state-of-the-art performance in the BSD100 dataset on multiple super-resolution tasks. Our model is able to achieve this performance using less training data and computational resources while maintaining high parameter efficiency compared to current state-of-the-art models.
Pranav Jeevan, Akella Srinidhi, Pasunuri Prathiba, Amit Sethi
WACV4
2024 Reverse Knowledge Distillation: Training a Large Model using a Small One for Retinal Image Matching on Limited Data
abstract
Retinal image matching (RIM) plays a crucial role in monitoring disease progression and treatment response as retina is the only tissue where blood vessels can be directly observed. However, datasets with matched keypoints between temporally separated pairs of images are not available in abundance to train transformer-based models. Firstly, we release keypoint annotations for retinal images from multiple datasets to aid further research on RIM. Secondly, we propose a novel approach based on reverse knowledge distillation to train large models with limited data while preventing overfitting. We propose architectural modifications to a CNN-based semi-supervised method called SuperRetina [22] that helps improve its results on a publicly available dataset. We train a computationally heavier model based on a vision transformer encoder, utilizing the lighter CNN-based model. This approach, which we call reverse knowledge distillation (RKD), further improves the matching results even though it contrasts with the conventional knowledge distillation where lighter models are trained based on heavier ones is the norm. Further, we show that our technique generalizes to other domains, such as facial landmark matching.
Sahar Almahfouz Nasser, Nihar Gupte, Amit Sethi
WACV3
2024 The ACROBAT 2022 challenge: Automatic registration of breast cancer tissue
abstract
The alignment of tissue between histopathological whole-slide-images (WSI) is crucial for research and clinical applications. Advances in computing, deep learning, and availability of large WSI datasets have revolutionised WSI analysis. Therefore, the current state-of-the-art in WSI registration is unclear. To address this, we conducted the ACROBAT challenge, based on the largest WSI registration dataset to date, including 4,212 WSIs from 1,152 breast cancer patients. The challenge objective was to align WSIs of tissue that was stained with routine diagnostic immunohistochemistry to its H&E-stained counterpart. We compare the performance of eight WSI registration algorithms, including an investigation of the impact of different WSI properties and clinical covariates. We find that conceptually distinct WSI registration methods can lead to highly accurate registration performances and identify covariates that impact performances across methods. These results provide a comparison of the performance of current WSI registration methods and guide researchers in selecting and developing methods.
Philippe Weitz, Masi Valkonen, Leslie Solorzano, Circe Carr, Kimmo Kartasalo, Constance Boissin, Sonja Koivukoski, Aino Kuusela, Dusan Rasic, Yanbo Feng, Sandra Kristiane Sinius Pouplier, Kajsa Ledesma Eriksson, Stephanie Robertson, Christian Marzahl, Chandler Gatenbee, Alexander R. A. Anderson, Marek Wodzinski, Artur Jurgas, Niccolò Marini, Manfredo Atzori, Henning Müller, Daniel Budelmann, Nick Weiss, Stefan Heldmann, Johannes Lotz 0002, Jelmer M. Wolterink, Bruno De Santi, Abhijeet Patil, Amit Sethi, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Mahtab Farrokh, Neeraj Kumar 0002, Russell Greiner, Leena Latonen, Anne-Vibeke Laenkholm, Johan Hartman, Pekka Ruusuvuori, Mattias Rantalainen
Medical Image Anal.30
2023 Robust Semi-Supervised Learning for Histopathology Images Through Self-Supervision Guided Out-of-Distribution Scoring
abstract
Semi-supervised learning (semi-SL) offers a promising solution for challenging scenarios in medical image analysis where acquiring sufficient labeled data is difficult. However, practical applications often violate the assumption that the unlabeled data distribution matches the labeled samples, especially in medical imaging. This leads to out-of-distribution (OOD) samples that hinder algorithm efficiency. Filtering methods commonly used for outlier samples may not be suitable for diverse anatomical structures and rare morphologies in medical images. To address these challenges in digital histology images, we propose a novel pipeline. Our approach utilizes self-supervised learning to estimate an OOD score for each unlabeled data point, enabling calibration of the subsequent semi-SL framework. By modulating sample selection based on the outlier score, we prioritize samples aligned with the labeled distribution during the semi-SL stage. Our framework is compatible with any semi-SL approach and has been validated on two digital pathology datasets. Extensive experiments on a colorectal histology dataset and a TCGA-BRCA whole slide image dataset demonstrate the superiority of our method compared to popular semi-SL techniques, particularly the widely used Mixmatch framework. Our approach enhances medical image analysis performance and shows promise in addressing challenges related to open-set supervised learning.
Nikhil Cherian Kurian, S. Varsha, Abhijit Patil, Shashikant Khade, Amit Sethi
BIBE5
2022 Deep Multi-Scale U-Net Architecture and Label-Noise Robust Training Strategies for Histopathological Image Segmentation
abstract
Although the U-Net architecture has been extensively used for the segmentation of medical images, we address two of its shortcomings in this work. Firstly, the accuracy of vanilla U-Net degrades when the target regions for segmentation exhibit significant variations in shape and size. Even though the U-Net already possesses some capability to analyze features at various scales, we propose to explicitly add multi-scale feature maps in each convolutional module of the U-Net encoder to improve the segmentation of histology images. Secondly, the accuracy of a U-Net model also suffers when the annotations for supervised learning are noisy or incomplete. This can happen due to the inherent difficulty for a human expert to identify and delineate all instances of specific pathology very precisely and accurately. We address this challenge by introducing auxiliary confidence maps that emphasize less on the boundaries of the given target regions. Further, we utilize the bootstrapping properties of the deep network to address the missing annotation problem intelligently. In our experiments on a private dataset of breast cancer lymph nodes, where the primary task was to segment germinal centers and sinus histiocytosis, we observed substantial improvement over baselines based on the two proposed schemes.
Nikhil Cherian Kurian, Amit Lehan, Gregory Verghese, Nimish Dharamshi, Swati Meena, Cheryl Gillet, Swapnil Rane, Anita Grigoriadis, Amit Sethi
BIBE11
2022 Shallow Water Bathymetry Survey Using an Autonomous Surface Vehicle
abstract
Accurate and cost effective mapping of water bodies has an enormous significance for environmental understanding and navigation. However, the quantity and quality of information we acquire from such environmental features is limited by various factors, including cost, time, security, and the capabilities of existing data collection techniques. Measurement of water depth is an important part of such mapping, particularly in shallow locations that could provide navigational risk or have important ecological functions. Erosion and deposition at these locations, for example, due to storms and erosion, can cause rapid changes that require repeated measurements. In this paper, we describe a low-cost, resilient, unmanned autonomous surface vehicle for bathymetry data collection using side-scan sonar. We discuss the adaptation of equipment and sensors for the collection of navigation, control, and bathymetry data and also give an overview of the vehicle setup. This autonomous surface vehicle has been used to collect bathymetry from the Powai Lake in Mumbai, India.
Bibin Wilson, Amit Sethi
IGARSS3
2022 Resource-efficient Hybrid X-formers for Vision
abstract
Although transformers have become the neural architectures of choice for natural language processing, they require orders of magnitude more training data, GPU memory, and computations in order to compete with convolutional neural networks for computer vision. The attention mechanism of transformers scales quadratically with the length of the input sequence, and unrolled images have long sequence lengths. Plus, transformers lack an inductive bias that is appropriate for images. We tested three modifications to vision transformer (ViT) architectures that address these shortcomings. Firstly, we alleviate the quadratic bottleneck by using linear attention mechanisms, called X-formers (such that, X ∈{Performer, Linformer, Nyströmformer}), thereby creating Vision X-formers (ViXs). This resulted in up to a seven times reduction in the GPU memory requirement. We also compared their performance with FNet and multi-layer perceptron mixers, which further reduced the GPU memory requirement. Secondly, we introduced an inductive prior for images by replacing the initial linear embedding layer by convolutional layers in ViX, which significantly increased classification accuracy without increasing the model size. Thirdly, we replaced the learnable 1D position embeddings in ViT with Rotary Position Embedding (RoPE), which increases the classification accuracy for the same model size. We believe that incorporating such changes can democratize transformers by making them accessible to those with limited data and computing resources.
Pranav Jeevan, Amit Sethi
WACV2
2022 Appraisal of Resistivity Inversion Models With Convolutional Variational Encoder-Decoder Network
abstract
Recovering the actual subsurface electrical resistivity properties from the electrical resistivity tomography data is challenging because the inverse problem is nonlinear and ill-posed. This paper proposes a Variational Encoder-Decoder (VED) based network to obtain resistivity model, which maps the apparent resistivity data (input) to true resistivity data (output). Since deep learning (DL) models are highly dependent on training sets and providing a meaningful geological resistivity model is complex, we have first developed an algorithm to construct many realistic resistivity synthetic models. Our algorithm automatically constructs an apparent resistivity pseudo-section from these resistivity models. We further computed the resistivity from two different neural architectures for comparison – UNet, and attention UNet with and without input depth encoding apparent data. In the end, we have compared our deep learning results with traditional inversion and borewell data on apparent resistivity datasets collected for aquifer mapping in the hard rock terrain of the West Medinipur district of West Bengal, India. A detailed qualitative and quantitative evaluation reveals that our VED approach is the most effective for the inversion compared to other approaches considered.
Bibin Wilson, Amit Sethi
IEEE Trans. Geosci. Remote. Sens.3
2022 Author's Reply to "MoNuSAC2020: A Multi-Organ Nuclei Segmentation and Classification Challenge"
abstract
We had released MoNuSAC2020 as one of the largest publicly available, manually annotated, curated, multi-class, and multi-instance medical image segmentation datasets. Based on this dataset, we had organized a challenge at the International Symposium on Biomedical Imaging (ISBI) 2020. Along with the challenge participants, we had published an article summarizing the results and findings of the challenge (Verma et al., 2021). Foucart et al. (2022) in their "Analysis of the MoNuSAC 2020 challenge evaluation and results: metric implementation errors" have pointed ways in which the computation of the segmentation performance metric for the challenge can be corrected or improved. After a careful examination of their analysis, we have found a small bug in our code and an erroneous column-header swap in one of our result tables. Here, we present our response to their analysis, and issue an errata. After fixing the bug the challenge rankings remain largely unaffected. On the other hand, two of Foucart et al.'s other suggestions are good for future consideration, but it is not clear that those should be immediately implemented. We thank Foucart et al. for their detailed analysis to help us fix the two errors.
Ruchika Verma, Neeraj Kumar 0002, Abhijeet Patil, Nikhil Cherian Kurian, Swapnil Rane, Amit Sethi
IEEE Trans. Medical Imaging6
2021 PAL : Pretext-based Active Learning
Shubhang Bhatnagar, Sachin Goyal, Darshan Tank, Amit Sethi
BMVC4
2021 MoNuSAC2020: A Multi-Organ Nuclei Segmentation and Classification Challenge
abstract
Detecting various types of cells in and around the tumor matrix holds a special significance in characterizing the tumor micro-environment for cancer prognostication and research. Automating the tasks of detecting, segmenting, and classifying nuclei can free up the pathologists' time for higher value tasks and reduce errors due to fatigue and subjectivity. To encourage the computer vision research community to develop and test algorithms for these tasks, we prepared a large and diverse dataset of nucleus boundary annotations and class labels. The dataset has over 46,000 nuclei from 37 hospitals, 71 patients, four organs, and four nucleus types. We also organized a challenge around this dataset as a satellite event at the International Symposium on Biomedical Imaging (ISBI) in April 2020. The challenge saw a wide participation from across the world, and the top methods were able to match inter-human concordance for the challenge metric. In this paper, we summarize the dataset and the key findings of the challenge, including the commonalities and differences between the methods developed by various participants. We have released the MoNuSAC2020 dataset to the public.
Ruchika Verma, Neeraj Kumar 0002, Abhijeet Patil, Nikhil Cherian Kurian, Swapnil Rane, Simon Graham, Quoc Dang Vu, Mieke Zwager, Shan E Ahmed Raza, Nasir M. Rajpoot, Xiyi Wu, Huai Chen, Lisheng Wang, Hyun Jung, G. Thomas Brown, Shuolin Liu, Seyed Alireza Fatemi Jahromi, Aliasghar Khani, Ehsan Montahaei, Mahdieh Soleymani Baghshah, Hamid Behroozi, Pavel Semkin, Alexandr Rassadin, Prasad Dutande, Romil Lodaya, Ujjwal Baid, Bhakti Baheti, Sanjay N. Talbar, Amirreza Mahbod, Rupert Ecker, Isabella Ellinger, Bin Dong 0006, Zhengyu Xu, Yuehan Yao, Ming Feng, Kele Xu, Hasib Zunair, A. Ben Hamza, Steven M. Smiley, Tang-Kai Yin, Qi-Rui Fang, Shikhar Srivastava 0001, Dwarikanath Mahapatra, Lubomira Trnavska, Hanyun Zhang, Priya Lakshmi Narayanan, Justin Law, Yinyin Yuan, Abhiroop Tejomay, Aditya Mitkari, Dinesh Koka, Vikas Ramachandra, Lata Kini, Amit Sethi
IEEE Trans. Medical Imaging58
2020 Visualization for Histopathology Images using Graph Convolutional Neural Networks
abstract
With an increase in the use of deep learning for computer-aided diagnosis in medical images, the criticism of the black-box nature of the deep learning models is also on the rise. The medical community prefers interpretable models for its due diligence and advancing the understanding of disease and treatment mechanisms. For instance, in histology, while cells and their spatial relationships manifest in rich detail, it is difficult to modify convolutional neural networks to point out the relevant visual features. We adopt an approach to model the histology of a cancer tissue as a graph of its constituent nuclei. We analyze this graph using two novel graph convolutional network frameworks- one based on node occlusion, and another based on attention mechanism- for disease classification and visualization. The proposed methods highlight the relative contribution of each cell nucleus in the disease diagnosis. As proofs of concept, our frameworks not only distinguish accurately between IDC and DCIS breast cancers as well as Gleason 3 and 4 prostate cancers, but they also highlight important visual details, such as boundaries of tumor nests in DCIS and those of glands in Gleason 3.
Mookund Sureka, Abhijeet Patil, Deepak Anand, Amit Sethi
BIBE4
2020 Satellite-Derived Bathymetry Using Deep Convolutional Neural Network
abstract
Our goal is to develop technique for assessing bathymetry maps, which show the topography of the floors of water-bodies, using satellite or aerial imagery. The advent of deep neural networks has enabled the use of new techniques in analysing and creating depth maps from high resolution satellite images. In this paper we report a pilot study in exploring the potential use of a deep learning architecture by framing bathymetry problem as a pixel-wise classification task. We took the Sentinel-2 bands as the satellite image and independent depth measurements from Humminbird™data. The data for the supervised training was carefully prepared and processed to facilitate the use of a powerful deep learning segmentation models. The efficiency of the model was quantified by the average F1-score (Dice score) on the held out dataset on quantized depth bands.
Bibin Wilson, Nikhil Cherian Kurian, Amit Sethi
IGARSS4
2020 A Multi-Organ Nucleus Segmentation Challenge
abstract
Generalized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics.
Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi
IEEE Trans. Medical Imaging87
2019 Pixel-wise Segmentation of Right Ventricle of Heart
abstract
One of the first steps in the diagnosis of most cardiac diseases, such as pulmonary hypertension, coronary heart disease is the segmentation of ventricles from cardiac magnetic resonance (MRI) images. Manual segmentation of the right ventricle requires diligence and time, while its automated segmentation is challenging due to shape variations and ill-defined borders. We propose a deep learning based method for the accurate segmentation of right ventricle, which does not require post-processing and yet it achieves the state-of-the-art performance of 0.86 Dice coefficient and 6.73 mm Hausdorff distance on RVSC-MICCAI 2012 dataset. We use a novel adaptive cost function to counter extreme class-imbalance in the dataset. We present a comprehensive comparative study of loss functions, architectures, and ensembling techniques to build a principled approach for biomedical segmentation tasks.
Yaman Dang, Deepak Anand, Amit Sethi
TENCON3
2019 Improving Histopathology Classification using Learnable Preprocessing
abstract
A deep learning classifier trained on a source dataset often performs poorly on a target dataset, even for the same classification task, due to the differences in the distributions of the two datasets. In histopathology, the problem of dataset bias is even more severe due to the differences in specific tissue preparation and imaging set ups across patient cohorts. With the objective of improving the generalization across datasets, we propose a set of learnable preprocessing operations - an approach that has not been extensively explored - for a supervised deep learning framework that can be trained separately or together with the rest of the neural network. Through preprocessing, the data from a target domain is transformed before being fed to a classification module trained on the source domain to increase the overlap of the former's distribution with that of the latter. Through an extensive set of experiments on histopathology and face datasets, we show the particular and general utility of the proposed preprocessing operations for domain adaptation and compare it to previous approaches.
Viraf Patrawala, Nikhil Cherian Kurian, Amit Sethi
TENCON3
2019 Hyperspectral Tissue Image Segmentation Using Semi-Supervised NMF and Hierarchical Clustering
abstract
Hyperspectral imaging (HSI) of tissue samples in the mid-infrared (mid-IR) range provides spectro-chemical and tissue structure information at sub-cellular spatial resolution. Disease states can be directly assessed by analyzing the mid-IR spectra of different cell types (e.g., epithelial cells) and sub-cellular components (e.g., nuclei), provided that we can accurately classify the pixels belonging to these components. The challenge is to extract information from hundreds of noisy mid-IR bands at each pixel, where each band is not very informative in itself, making annotations of unstained tissue HSI images particularly tricky. Because the tissue structure is not necessarily identical between the two sections, only a few regions in unstained HSI image can be annotated with high confidence, even when serial (or adjacent) hematoxylin and eosin stained section is used as a visual guide. In order to completely use both labeled and unlabeled pixels in training images, we have developed an HSI pixel classification method that uses semi-supervised learning for both spectral dimension reduction and hierarchical pixel clustering. Compared to the supervised classifiers, the proposed method was able to account for the vast differences in the spectra of sub-cellular components of the same cell type and to achieve an F1 score of 71.18% on twofold cross-validation across 20 tissue images. To generate further interest in this promising modality, we have released our source code and also showed that disease classification is straightforward after HSI image segmentation.
Neeraj Kumar 0002, Phani Krishna Uppala, Karthik Duddu, Hari Sreedhar, Vishal K. Varma, Grace Guzman, Michael J. Walsh 0005, Amit Sethi
IEEE Trans. Medical Imaging8
2018 Super Resolution by Comprehensively Exploiting Dependencies of Wavelet Coefficients
abstract
We propose an algorithm for single image super resolution (SR) using wavelet decomposition and machine learning. Wavelets have been used for SR before due to their ability to capture scale invariant properties of natural images. However previous techniques used only a subset of relationships that exist between multiscale wavelet coefficients. We present a first-of-its-kind analysis of the wavelet properties relevant to SR that leads to insights for a novel SR algorithm. In particular we discovered that to estimate a desired finer scale detail coefficient it is not enough to use only its parent detail coefficient at the coarser scale as was done by the previous techniques. The estimation can be improved a lot by using certain additional coarser level detail coefficients and even finer scale approximation coefficients whose relative locations are suggested by our analysis. Additionally the previous wavelet-based techniques used generative learning frameworks for SR. However we show that SR is a type of problem on which discriminative frameworks excel. These improvements allowed our technique to far surpass the reconstruction accuracy of the previous wavelet-based SR algorithms on a large set of images. Additionally our algorithm also surpassed other state-of-the-art SR algorithms that are not based on wavelets in reconstruction quality and training and testing speeds. We thus reestablish the utility of the wavelets for SR.
Neeraj Kumar 0002, Amit Sethi
IEEE Trans. Multim.2
2017 Action recognition using spatio-temporal differential motion
abstract
This paper presents human action recognition using spatio-temporal differential motion maps. The concept of differential motion in space and time helps in overcoming several challenges in action recognition such as camera motion and multiple actions in the same scene. Spatially differential motion in a frame is represented using divergence of optical flow. Divergence map of each frame in a video is projected onto three orthogonal Cartesian planes. A map of spatio-temporal differential motion is formed by accumulation of the absolute differences between projected maps of pairs of consecutive frames through an entire video sequence for each projection. A feature vector is formed from these three spatiotemporal maps of differential motion which represents the action performed in the video. Classification of action was done by using l2-regularized collaborative representation with a distance-weighted Tikhonov matrix. We tested on two popular datasets, KTH and UCF11, and got better performance than state-of-the-art methods. A comparison of differential motion and optical flow (any motion with respect to the camera) was also done to show that differential motion gives better feature representation than simply using optical flow.
Gaurav Kumar Yadav, Amit Sethi
ICIP2
2017 Convolutional neural networks for wavelet domain super resolution
Neeraj Kumar 0002, Ruchika Verma, Amit Sethi
Pattern Recognit. Lett.3
2017 A Dataset and a Technique for Generalized Nuclear Segmentation for Computational Pathology
abstract
Nuclear segmentation in digital microscopic tissue images can enable extraction of high-quality features for nuclear morphometrics and other analysis in computational pathology. Conventional image processing techniques, such as Otsu thresholding and watershed segmentation, do not work effectively on challenging cases, such as chromatin-sparse and crowded nuclei. In contrast, machine learning-based segmentation can generalize across various nuclear appearances. However, training machine learning algorithms requires data sets of images, in which a vast number of nuclei have been annotated. Publicly accessible and annotated data sets, along with widely agreed upon metrics to compare techniques, have catalyzed tremendous innovation and progress on other image classification problems, particularly in object recognition. Inspired by their success, we introduce a large publicly accessible data set of hematoxylin and eosin (H&E)-stained tissue images with more than 21000 painstakingly annotated nuclear boundaries, whose quality was validated by a medical doctor. Because our data set is taken from multiple hospitals and includes a diversity of nuclear appearances from several patients, disease states, and organs, techniques trained on it are likely to generalize well and work right out-of-the-box on other H&E-stained images. We also propose a new metric to evaluate nuclear segmentation results that penalizes object- and pixel-level errors in a unified manner, unlike previous metrics that penalize only one type of error. We also propose a segmentation technique based on deep learning that lays a special emphasis on identifying the nuclear boundaries, including those between the touching or overlapping nuclei, and works well on a diverse set of test images.
Neeraj Kumar 0002, Ruchika Verma, Sanuj Sharma, Surabhi Bhargava, Abhishek Vahadane, Amit Sethi
IEEE Trans. Medical Imaging6
2016 Action recognition using interest points capturing differential motion information
abstract
Human action recognition has been a challenging task in computer vision because of intra-class variability. State-of-the-art methods have shown good performance for constrained videos but have failed to achieve good results for complex scenes. Reasons for their failing include treating spatial and temporal dimensions without distinction as well as not capturing temporal information in video representation. To address these problems we propose principled changes to an action recognition framework that is based on video interest points (IP) detection with capturing differential motion as the central theme. First, we propose to detect points with high curl of optical flow, which captures relative motion boundaries in a frame. We track these points to form dense trajectories. Second, we discard points on the trajectories that do not represent change in motion of the same object, yielding temporally localized IPs. Third, we propose a video representation based on spatio-temporal arrangement of IPs with respect to their neighboring IPs. The proposed approach yields a compact and information-dense representation without using any local descriptor around the detected IPs. It significantly outperforms state-of-the-art methods on UCF youtube dataset, which has complex action classes, as well as on KTH dataset, which has simple action classes.
Gaurav Kumar Yadav, Prakhar Shukla, Amit Sethi
ICASSP3
2016 Detecting multiple sub-types of breast cancer in a single patient
abstract
Determining the molecular or genomic sub-type of a cancer of a particular organ is important for prognosis and treatment planning. While clinical tests determine the dominant subtype of cancer in a patient, it is believed that some cancer treatments targeting the dominant sub-types ultimately prove ineffective because they ignore the existence of additional sub-types in the same patient. We present a method to detect co-existence of two common sub-types of breast cancer - HER2 and basal-like - in whole slide images of H&E-stained tumors in the same patient. Our main contribution is in formulating this problem in precision medicine and in preparation of its training data. The detectors tested to be highly accurate in classifying patches taken from test cases with well-separated molecular profiles, confirming that image classification was effective in detecting the two sub-types. Interestingly, test patients whose molecular profiles were not very clearly separated appeared to include heterogeneous patches - some HER2 and other basal-like - indicating existence of a secondary sub-type within a single patient. The detection of a secondary sub-type in a single patient may impact the prognosis and treatment of the cancer.
Ruchika Verma, Neeraj Kumar 0002, Amit Sethi, Peter H. Gann
ICIP3
2016 Structure-Preserving Color Normalization and Sparse Stain Separation for Histological Images
abstract
Staining and scanning of tissue samples for microscopic examination is fraught with undesirable color variations arising from differences in raw materials and manufacturing techniques of stain vendors, staining protocols of labs, and color responses of digital scanners. When comparing tissue samples, color normalization and stain separation of the tissue images can be helpful for both pathologists and software. Techniques that are used for natural images fail to utilize structural properties of stained tissue samples and produce undesirable color distortions. The stain concentration cannot be negative. Tissue samples are stained with only a few stains and most tissue regions are characterized by at most one effective stain. We model these physical phenomena that define the tissue structure by first decomposing images in an unsupervised manner into stain density maps that are sparse and non-negative. For a given image, we combine its stain density maps with stain color basis of a pathologist-preferred target image, thus altering only its color while preserving its structure described by the maps. Stain density correlation with ground truth and preference by pathologists were higher for images normalized using our method when compared to other alternatives. We also propose a computationally faster extension of this technique for large whole-slide images that selects an appropriate patch sample instead of using the entire image to compute the stain color basis.
Abhishek Vahadane, Tingying Peng, Amit Sethi, Shadi Albarqouni, Maximilian Baust, Katja Steiger, Anna Melissa Schlitter, Irene Esposito, Nassir Navab
IEEE Trans. Medical Imaging3
2016 Fast Learning-Based Single Image Super-Resolution
abstract
We present a learning-based single image super-resolution (SISR) method to obtain a high resolution (HR) image from a single given low resolution (LR) image. Our method gives more accurate results while also testing (runs) and training faster with a smaller number of training samples compared to other methods. We posed SISR as a problem of estimating a function to predict the pixels of an HR patch using its corresponding LR pixels and their spatial neighborhood. We studied the impact of varying the input LR and output HR patch sizes and gained the following insights: reconstruction accuracy for a given output HR patch size improves when input LR patch size is increased, but the improvement saturates after including a few extra layers of LR pixels. Moreover, HR reconstruction accuracy is the highest when the output HR patch is restricted to only that which corresponds to one LR pixel. We used zero component analysis as a pre-processing step to enhance the estimation optimization energy on perceptually salient features such as edges. We tapped into the ability of polynomial neural networks to hierarchically learn refinements of a function that maps LR to HR patches. Accurate HR reconstruction with small input and output patch sizes not only makes learning more efficient, it also indicates that SISR is a highly local problem. In contrast, a recently proposed and related technique using convolutional neural networks needs much larger training set and longer training time because of larger input-output patch sizes and a computationally expensive learning algorithm.
Neeraj Kumar 0002, Amit Sethi
IEEE Trans. Multim.2
2015 On spatial neighborhood of patch-based super resolution
abstract
We propose an intuitive formulation for single image super resolution (SISR), and an algorithm based on it that outperforms the state of the art. We model the SISR problem around an aspect of natural images that is often overlooked - sharp edges - which are important for perceptual quality, and troublesome for interpolation. We justify the use of low resolution (LR) Markovian neighborhoods to estimate high resolution (HR) pixels corresponding to only the central pixel of the LR neighborhood based on our formulation. The formulation also lends itself to learning LR to HR mapping based on their pairs as training examples. We propose a learning algorithm based on polynomial neural networks to learn this mapping. Our formulation and algorithm provide further insight into performance of various single image super resolution methods.
Neeraj Kumar 0002, Amit Sethi
ICIP2
2013 Towards generalized nuclear segmentation in histological images
abstract
Computer aided diagnosis in cancer pathology (computational pathology) using histological images of biopsies is an emerging field. Segmentation of cell nuclei can be an important step in such image processing pipelines. Although seeded watershed segmentation is a simple and computationally efficient segmentation technique, it is prone to errors like over-segmentation when applied to histological images. We report specific enhancements to this technique to improve segmentation of cell nuclei in histological images. Foreground seeds were generated by fast radial symmetry transform (FRST). Otsu thresholding was used on enhanced image to estimate tentative foreground map. Background markers were computed from the tentative foreground map. False detections in the segmented output were removed by logical AND with the tentative foreground map. Using these enhancements nuclear segmentation was significantly improved on histological images (H&E stained breast and intestinal tissue images, Feulgen stained images of prostate tissues).
Abhishek Vahadane, Amit Sethi
BIBE2
2012 Learning to predict super resolution wavelet coefficients
Neeraj Kumar 0002, Naveen Kumar Rai, Amit Sethi
ICPR3
2006 Robust Structure and Motion from Outlines of Smooth Curved Surfaces
abstract
This paper addresses the problem of estimating the motion of a camera as it observes the outline (or apparent contour) of a solid bounded by a smooth surface in successive image frames. In this context, the surface points that project onto the outline of an object depend on the viewpoint and the only true correspondences between two outlines of the same object are the projections of frontier points where the viewing rays intersect in the tangent plane of the surface. In turn, the epipolar geometry is easily estimated once these correspondences have been identified. Given the apparent contours detected in an image sequence, a robust procedure based on RANSAC and a voting strategy is proposed to simultaneously estimate the camera configurations and a consistent set of frontier point projections by enforcing the redundancy of multiview epipolar geometry. The proposed approach is, in principle, applicable to orthographic, weak-perspective, and affine projection models. Experiments with nine real image sequences are presented for the orthographic projection case, including a quantitative comparison with the ground-truth data for the six data sets for which the latter information is available. Sample visual hulls have been computed from all image sequences for qualitative evaluation.
Yasutaka Furukawa, Amit Sethi, Jean Ponce, David J. Kriegman
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Variable module graphs: a framework for inference and learning in modular vision systems
abstract
We present a novel and intuitive framework for building modular vision systems for complex tasks such as surveillance applications. Inspired by graphical models, especially factor graphs, the framework allows capturing the dependencies between different variables in form of a graph. This enforces principled coordination and exchange of information between different modules. Breaking away from the traditional probabilistic graphical models the framework allows flexibility of design in individual modules by allowing different learning and inference mechanisms to work in a common setting. It also allows easy integration of more modules into an already functional system. We demonstrate the ease of building a complex vision system within this framework by designing a fully automatic multi-target tracking system for a video surveillance scenario. Favorable results are obtained for the tracking application.
Amit Sethi, Mandar Rahurkar, Thomas S. Huang
ICIP (2)1
2004 Structure and Motion from Images of Smooth Textureless Objects
Yasutaka Furukawa, Amit Sethi, Jean Ponce, David J. Kriegman
ECCV (2)2
2004 A detection-based multiple object tracking method
abstract
In this paper we describe a method for tracking multiple objects whose number is unknown and varies during tracking. Based on preliminary results of object detection in each image which may have missing and/or false detection, the multiple object tracking method keeps a graph structure where it maintains multiple hypotheses about the number and the trajectories of the objects in the video. The image information drives the process of extending and pruning the graph, and determines the best hypothesis to explain the video. While the image-based object detection makes a local decision, the tracking process confirms and validates the detection through time, therefore, it can be regarded as temporal detection which makes a global decision across time. The multiple object tracking method gives feedbacks which are predictions of object locations to the object detection module. Therefore, the method integrates object detection and tracking tightly. The most possible hypothesis provides the multiple object tracking result. The experimental results are presented.
Amit Sethi, Yihong Gong
ICIP2
2004 Curve and Surface Duals and the Recognition of Curved 3D Objects from their Silhouettes
Amit Sethi, David Renaudie, David J. Kriegman, Jean Ponce
Int. J. Comput. Vis.1
2002 On Pencils of Tangent Planes and the Recognition of Smooth 3D Shapes from Silhouettes
Svetlana Lazebnik, Amit Sethi, Cordelia Schmid, David J. Kriegman, Jean Ponce, Martial Hebert
ECCV (3)2