EDBT 2026 Demo / reviewers in the wild / expert
Saket Anand
dblp:11/4747
· DBLP profile ↗
36ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0002-6229-3940ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 13 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Active Learning for Animal Re-Identification with Ambiguity-Aware SamplingabstractAnimal re-identification (Re-ID) has recently gained substantial attention in the AI research community due to its high impact on biodiversity monitoring and unique research challenges arising from environmental factors. The subtle distinguishing patterns like stripes or spots, handling new species and the inherent open-set nature make the problem even harder. To address these complexities, foundation models trained on labeled, large-scale and multi-species animal Re-ID datasets have recently been introduced to enable zero-shot Re-ID. However, our benchmarking reveals significant gaps in their zero-shot Re-ID performance for both known and unknown species. While this highlights the need for collecting labeled data in new domains, exhaustive annotation for Re-ID is laborious and requires domain expertise. Our analyses also show that existing unsupervised (USL) and active learning (AL) Re-ID methods underperform for animal Re-ID. To address these limitations, we introduce a novel AL Re-ID framework that leverages complementary clustering methods to uncover and target structurally ambiguous regions in the embedding space for mining pairs of samples that are both informative and broadly representative. Oracle feedback on these pairs, in the form of must-link and cannot-link constraints, facilitates a simple annotation interface, which naturally integrates with existing USL methods through our proposed constrained clustering refinement algorithm. Through extensive experiments, we demonstrate that, by utilizing only 0.033% of all possible annotations, our approach consistently outperforms existing foundational, USL and AL baselines. Specifically, we report an average improvement of 10.49%, 11.19% and 3.99% (mAP) on 13 wildlife datasets over foundational, USL and AL methods, respectively, while attaining state-of-the-art performance on each dataset. Furthermore, we also show an improvement of 11.09%, 8.2% and 2.06% (AUC ROC) for unknown individuals in an open-world setting. We also present results on 2 publicly available person Re-ID datasets, showing average gains of 7.96% and 2.86% (mAP) over existing USL and AL Re-ID methods. Depanshu Sani, Mehar Khurana, Saket Anand |
AAAI | 3 |
| 2026 | Learning to Communicate over an Unknown Shared NetworkabstractAs robots (edge-devices, agents) find uses in an increasing number of settings and edge-cloud resources become pervasive, wireless networks will often be shared by flows of data traffic that result from communication between agents and their corresponding edge-cloud nodes (cloud compute or data resource accessed by an agent). In such a setting, any agent communicating with the edge-cloud is unaware of the state of the network resource, which evolves in response to not just the agent’s own communication at any given time but also to communication by the other agents, which stays unknown to the agent. We address the challenge of an agent learning a policy that allows it to decide whether or not to communicate with its cloud node, using limited feedback it obtains from its own attempts to communicate, with the goal of optimizing its utility. The policy must generalize well to any number of other agents sharing the network and must not be trained for any particular network configuration. Our proposed policy is a deep reinforcement learning model Query Net (QNet) that we train using a proposed simulation-to-real framework. Our simulation model has just one parameter and is agnostic to specific configurations of any wireless network. It however allows training an agent’s policy over a wide range of outcomes that an agent’s communication with its edge-cloud node may face when using a shared network, by suitably randomizing the simulation parameter. We propose a learning algorithm that addresses the challenges we observe in training QNet. We validate our simulation-to-real driven approach through experiments conducted on real wireless networks including WiFi and cellular. We compare QNet with other policies to demonstrate its efficacy. Our WiFi experiments involved as few as five agents, resulting in barely any contention for the network, to as many as 50 agents, resulting in severe contention. The cellular experiments spanned a broad range of network conditions, with baseline network round-trip times ranging from a low of 0.07 s to a high of 0.83 s. Shivangi Agarwal, Adi Asija, Sanjit Krishnan Kaul, Arani Bhattacharya, Saket Anand |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2025 | Sim-to-Real Transfer for Estimation over Wireless NetworksabstractData-driven models for state estimation using measurements obtained over a wireless network are essential to cyber-physical systems. Learning a data-driven model for an estimator using real wireless network deployments is, however, impractical as it would require data that captures varied wireless network conditions and their impact on estimation. We propose and evaluate a simulation to real-world transfer of a data-driven model for state estimation. Specifically, we train an estimator model using only data generated by a low-fidelity simulation of networks. Our choice of network simulation is a first-come-first-served single server queue, which is a network model with the two parameters of arrival rate of measurement packets into the queue and packet service rates. We employ domain randomization to bridge the gap between simulation and the real world, appropriately randomizing the network model parameters during training. The efficacy of the resulting estimator model is demonstrated by testing it over two deployments of real wireless networks. In one, the estimator model estimates vehicles’ positions and speeds using data from vehicular trajectories received by it over a shared WiFi network, with up to seventy sources sending the measurements. In the other, GPS coordinates are communicated by public transit buses over city-wide cellular networks. The estimator uses the received measurements to estimate locations of the buses. Shivangi Agarwal, Adi Asija, Sanjit Krishnan Kaul, Saket Anand |
ACM Trans. Cyber Phys. Syst. | 4 |
| 2024 | BirdCollect: A Comprehensive Benchmark for Analyzing Dense Bird Flock AttributesabstractAutomatic recognition of bird behavior from long-term, un controlled outdoor imagery can contribute to conservation efforts by enabling large-scale monitoring of bird populations. Current techniques in AI-based wildlife monitoring have focused on short-term tracking and monitoring birds individually rather than in species-rich flocks. We present Bird-Collect, a comprehensive benchmark dataset for monitoring dense bird flock attributes. It includes a unique collection of more than 6,000 high-resolution images of Demoiselle Cranes (Anthropoides virgo) feeding and nesting in the vicinity of Khichan region of Rajasthan. Particularly, each image contains an average of 190 individual birds, illustrating the complex dynamics of densely populated bird flocks on a scale that has not previously been studied. In addition, a total of 433 distinct pictures captured at Keoladeo National Park, Bharatpur provide a comprehensive representation of 34 distinct bird species belonging to various taxonomic groups. These images offer details into the diversity and the behaviour of birds in vital natural ecosystem along the migratory flyways. Additionally, we provide a set of 2,500 point-annotated samples which serve as ground truth for benchmarking various computer vision tasks like crowd counting, density estimation, segmentation, and species classification. The benchmark performance for these tasks highlight the need for tailored approaches for specific wildlife applications, which include varied conditions including views, illumination, and resolutions. With around 46.2 GBs in size encompassing data collected from two distinct nesting ground sets, it is the largest birds dataset containing detailed annotations, showcasing a substantial leap in bird research possibilities. We intend to publicly release the dataset to the research community. The database is available at: https://iab-rubric.org/resources/wildlife-dataset/birdcollect Kshitiz, Sonu Sreshtha, Bikash Dutta, Muskan Dosi, Mayank Vatsa, Richa Singh 0001, Saket Anand, Sudeep Sarkar, Sevaram Mali Parihar |
AAAI | 7 |
| 2024 | Learning Geometry of Pose Image Manifolds in Latent Spaces Using Geometry-Preserving GANs
Shenyuan Liang, Benjamin Beaudett, Pavan Turaga, Saket Anand, Anuj Srivastava |
ICPR (27) | 4 |
| 2024 | Sensor-Agnostic Graph-Aware Kalman Filter for Multi-Modal Multi-Object Tracking
Depanshu Sani, Anirudh Iyer, Prakhar Rai, Saket Anand, Anuj Srivastava, Kaushik Kalyanaraman |
ICPR (16) | 4 |
| 2024 | Scalable and Sustainable Video Analytics on Edge using Sensor ClusteringabstractThe proliferation of video analytics in applications like autonomous driving, traffic surveillance, and teleoperated vehicles requires on-premise (on edge) execution of deep learning models to meet latency requirements and curb bandwidth usage by limiting frequent offloading of inference tasks. However, constrained by the compute and power availability on the edge, a cheaper model is typically deployed. These shallower models have two major associated problems: 1) using the same model for all cameras/vehicles gives inconsistent accuracy, and 2) trained models are prone to data drift. Shubham Chaudhary 0006, Arani Bhattacharya, Saket Anand, Aruna Balasubramanian |
MobiCom | 3 |
| 2024 | Army of Thieves: Enhancing Black-Box Model Extraction via Ensemble based sample selectionabstractMachine Learning (ML) models become vulnerable to Model Stealing Attacks (MSA) when they are deployed as a service. In such attacks, the deployed model is queried repeatedly to build a labelled dataset. This dataset allows the attacker to train a thief model that mimics the original model. To maximize query efficiency, the attacker has to select the most informative subset of data points from the pool of available data. Existing attack strategies utilize approaches like Active Learning and Semi-Supervised learning to minimize costs. However, in the black-box setting, these approaches may select sub-optimal samples as they train only one thief model. Depending on the thief model’s capacity and the data it was pretrained on, the model might even select noisy samples that harm the learning process. In this work, we explore the usage of an ensemble of deep learning models as our thief model. We call our attack Army of Thieves(AOT) as we train multiple models with varying complexities to leverage the crowd’s wisdom. Based on the ensemble’s collective decision, uncertain samples are selected for querying, while the most confident samples are directly included in the training data. Our approach is the first one to utilize an ensemble of thief models to perform model extraction. We outperform the base approaches of existing state-of-the-art methods by at least 3% and achieve a 21% higher adversarial sample transferability than previous work for models trained on the CIFAR-10 dataset. Code is available at: https://github.com/akshitjindal1/AOT_WACV. Akshit Jindal, Vikram Goyal, Saket Anand, Chetan Arora 0001 |
WACV | 3 |
| 2024 | SICKLE: A Multi-Sensor Satellite Imagery Dataset Annotated with Multiple Key Cropping ParametersabstractThe availability of well-curated datasets has driven the success of Machine Learning (ML) models. Despite greater access to earth observation data in agriculture, there is a scarcity of curated and labelled datasets, which limits the potential of its use in training ML models for remote sensing (RS) in agriculture. To this end, we introduce a first-of-its-kind dataset called SICKLE, which constitutes a time-series of multi-resolution imagery from 3 distinct satellites: Landsat-8, Sentinel-1 and Sentinel-2. Our dataset constitutes multi-spectral, thermal and microwave sensors during January 2018 - March 2021 period. We construct each temporal sequence by considering the cropping practices followed by farmers primarily engaged in paddy cultivation in the Cauvery Delta region of Tamil Nadu, India; and annotate the corresponding imagery with key cropping parameters at multiple resolutions (i.e. 3m, 10m and 30m). Our dataset comprises 2, 370 season-wise samples from 388 unique plots, having an average size of 0.38 acres, for classifying 21 crop types across 4 districts in the Delta, which amounts to approximately 209,000 satellite images. Out of the 2,370 samples, 351 paddy samples from 145 plots are annotated with multiple crop parameters; such as the variety of paddy, its growing season and productivity in terms of per-acre yields. Ours is also one among the first studies that consider the growing season activities pertinent to crop phenology (spans sowing, transplanting and harvesting dates) as parameters of interest. We benchmark SICKLE on three tasks: crop type, crop phenology (sowing, transplanting, harvesting), and yield prediction. Depanshu Sani, Sandeep Mahato, Sourabh Saini, Harsh Kumar Agarwal, Charu Chandra Devshali, Saket Anand, Gaurav Arora, Thiagarajan Jayaraman |
WACV | 6 |
| 2023 | Long-term Monitoring of Bird Flocks in the WildabstractMonitoring and analysis of wildlife are key to conservation planning and conflict management. The widespread use of camera traps coupled with AI-based analysis tools serves as an excellent example of successful and non-invasive use of technology for design, planning, and evaluation of conservation policies. As opposed to the typical use of camera traps that capture still images or short videos, in this project, we propose to analyze longer term videos monitoring a large flock of birds. This project, which is part of the NSF-TIH Indo-US joint R&D partnership, focuses on solving challenges associated with the analysis of long-term videos captured at feeding grounds and nesting sites, among other such locations that host large flocks of migratory birds. We foresee that the objectives of this project would lead to datasets and benchmarking tools as well as novel algorithms that would be instrumental in developing automated video analysis tools that could in turn help understand individual and social behavior of birds. The first of the key outcomes of this research will include the curation of challenging, real-world datasets for benchmarking various image and video analytics algorithms for tasks such as counting, detection, segmentation, and tracking. Our recent efforts towards this outcome is a curated dataset of 812 high-resolution, point-annotated, images (4K - 32MP) of a flock of Demoiselle cranes (Anthropoides virgo) taken from their feeding site at Khichan, Rajasthan, India. The average number of birds in each image is about 207, with a maximum count of 1500. The benchmark experiments show that state-of-the-art vision techniques struggle with tasks such as segmentation, detection, localization, and density estimation for the proposed dataset. Over the execution of this open science research, we will be scaling this dataset for segmentation and tracking in videos, as well as developing novel techniques for video analytics for wildlife monitoring. Kshitiz, Sonu Sreshtha, Ramy Mounir, Mayank Vatsa, Richa Singh 0001, Saket Anand, Sudeep Sarkar, Sevaram Mali Parihar |
IJCAI | 6 |
| 2023 | Reducing Annotation Effort by Identifying and Labeling Contextually Diverse Classes for Semantic Segmentation Under Domain ShiftabstractIn Active Domain Adaptation (ADA), one uses Active Learning (AL) to select a subset of images from the target domain, which are then annotated and used for supervised domain adaptation (DA). Given the large performance gap between supervised and unsupervised DA techniques, ADA allows for an excellent trade-off between annotation cost and performance. Prior art makes use of measures of uncertainty or disagreement of models to identify ‘regions' to be annotated by the human oracle. However, these regions frequently comprise of pixels at object boundaries which are hard and tedious to annotate. Hence, even if the fraction of image pixels annotated reduces, the overall annotation time and the resulting cost still remain high. In this work, we propose an ADA strategy, which given a frame, identifies a set of classes that are hardest for the model to predict accurately, thereby recommending semantically meaningful regions to be annotated in a selected frame. We show that these set of ‘hard' classes are context-dependent and typically vary across frames, and when annotated help the model generalize better. We propose two ADA techniques: the Anchor-based and Augmentation-based approaches to select complementary and diverse regions in the context of the current training set. Our approach achieves 66.6 mIoU on GTA5 →Cityscapes dataset with an annotation budget of 4.7% in comparison to 64.9 mIoU by MADA [22] using 5% of annotations. Our technique can also be used as a decorator for any existing frame-based AL technique, e.g., we report 1.5% performance improvement for CDAL [1] on Cityscapes using our approach. Sharat Agarwal, Saket Anand, Chetan Arora 0001 |
WACV | 2 |
| 2022 | Learning Hierarchy Aware Features for Reducing Mistake Severity
Ashima Garg, Depanshu Sani, Saket Anand |
ECCV (24) | 3 |
| 2022 | Does Data Repair Lead to Fair Models? Curating Contextually Fair Data To Reduce Model BiasabstractContextual information is a valuable cue for Deep Neural Networks (DNNs) to learn better representations and improve accuracy. However, co-occurrence bias in the training dataset may hamper a DNNmodel’s generalizabil- ity to unseen scenarios in the real world. For example, in COCO [26], many object categories have a much higher cooccurrence with men compared to women, which can bias a DNN’s prediction in favor of men. Recent works have focused on task-specific training strategies to handle bias in such scenarios, but fixing the available data is often ignored. In this paper, we propose a novel and more generic solution to address the contextual bias in the datasets by selecting a subset of the samples, which is fair in terms of the co-occurrence with various classes for a protected attribute. We introduce a data repair algorithm using the coefficient of variation( cv), which can curate fair and contextually balanced data for a protected class(es). This helps in training a fair model irrespective of the task, architecture or training methodology. Our proposed solution is simple, effective and can even be used in an active learning setting where the data labels are not present or being generated incrementally. We demonstrate the effectiveness of our algorithm for the task of object detection and multi-label image classification across different datasets. Through a series of experiments, we validate that curating contextually fair data helps make model predictions fair by balancing the true positive rate for the protected class across groups without compromising on the model’s overall performance. Code: https://github.com/sumanyumuku98/contextual-bias Sharat Agarwal, Sumanyu Muku, Saket Anand, Chetan Arora 0001 |
WACV | 3 |
| 2022 | HierMatch: Leveraging Label Hierarchies for Improving Semi-Supervised LearningabstractSemi-supervised learning approaches have emerged as an active area of research to combat the challenge of obtaining large amounts of annotated data. Towards the goal of improving the performance of semi-supervised learning methods, we propose a novel framework, HierMatch, a semi-supervised approach that leverages hierarchical information to reduce labeling costs and performs as well as a vanilla semi-supervised learning method. Hierarchical information is often available as prior knowledge in the form of coarse labels (e.g., woodpeckers) for images with fine-grained labels (e.g., downy woodpeckers or golden-fronted woodpeckers). However, the use of supervision using coarse-category labels to improve semi-supervised techniques has not been explored. In the absence of fine-grained labels, HierMatch exploits the label hierarchy and uses coarse class labels as a weak supervisory signal. Additionally, HierMatch is a generic-approach to improve any semi-supervised learning framework, we demonstrate this using our results on recent state-of-the-art techniques MixMatch and FixMatch. We evaluate the efficacy of HierMatch on two benchmark datasets, namely CIFAR-100 and NABirds. HierMatch can reduce the usage of fine-grained labels by 50% on CIFAR-100 with only a marginal drop of 0.59% in top-1 accuracy as compared to MixMatch. Ashima Garg, Shaurya Bagga, Yashvardhan Singh, Saket Anand |
WACV | 4 |
| 2022 | Intelligent Camera Selection Decisions for Target Tracking in a Camera NetworkabstractCamera Selection Decisions (CSD) are highly useful for several applications in a multi-camera network. For example, CSD benefit multi-camera target tracking by reducing the number of candidate cameras to look for the target’s next location. The correct candidate cameras, decreases the number of false Re-ID queries as well as the computation time. Also, in multi-camera trajectory forecasting (MCTF) to predict where a person will re-appear in the camera network along with the transition time. These applications require a large amount of annotated data for training. In this paper, we use state-representation learning with a reinforcement learning based policy to effectively and efficiently make camera selection decisions. We further demonstrate that by using learned state representations, as opposed to hand-crafted state variables, we are able to achieve state-of-the-art results on camera selection, while reducing the training time for the RL policy. Along with this, we use a reward function that helps to reduce the amount of supervision in training the policy in a semi-supervised way. We report our results on four datasets: NLPR_MCT, DukeMTMC, CityFlow, and WNMF dataset. We show that an RL policy reduces unnecessary Re-ID queries and therefore the false alarms, scales well to larger camera networks, and is target-agnostic. Anil Sharma, Saket Anand, Sanjit Krishnan Kaul |
WACV | 2 |
| 2022 | REGroup: Rank-aggregating Ensemble of Generative Classifiers for Robust PredictionsabstractDeep Neural Networks (DNNs) are often criticized for being susceptible to adversarial attacks. Most successful defense strategies adopt adversarial training or random input transformations that typically require retraining or finetuning the model to achieve reasonable performance. In this work, our investigations of intermediate representations of a pre-trained DNN lead to an interesting discovery pointing to intrinsic robustness to adversarial attacks. We find that we can learn a generative classifier by statistically characterizing the neural response of an intermediate layer to clean training samples. The predictions of multiple such intermediate-layer based classifiers, when aggregated, show unexpected robustness to adversarial attacks. Specifically, we devise an ensemble of these generative classifiers that rank-aggregates their predictions via a Borda count-based consensus. Our proposed approach uses a subset of the clean training data and a pre-trained model, and yet is agnostic to network architectures or the adversarial attack generation method. We show extensive experiments to establish that our defense strategy achieves state-of-the-art performance on the ImageNet validation set. Lokender Tiwari, Anish Madan, Saket Anand, Subhashis Banerjee |
WACV | 3 |
| 2022 | Modeling Functional Similarity in Source Code With Graph-Based Siamese NetworksabstractCode clones are duplicate code fragments that share (nearly) similar syntax or semantics. Code clone detection plays an important role in software maintenance, code refactoring, and reuse. A substantial amount of research has been conducted in the past to detect clones. A majority of these approaches use lexical and syntactic information to detect clones. However, only a few of them target semantic clones. Recently, motivated by the success of deep learning models in other fields, including natural language processing and computer vision, researchers have attempted to adopt deep learning techniques to detect code clones. These approaches use lexical information (tokens) and(or) syntactic structures like abstract syntax trees (ASTs) to detect code clones. However, they do not make sufficient use of the available structural and semantic information, hence limiting their capabilities. This paper addresses the problem of semantic code clone detection using program dependency graphs and geometric neural networks, leveraging the structured syntactic and semantic information. We have developed a prototype toolHolmes, based on our novel approach and empirically evaluated it on popular code clone benchmarks. Our results show thatHolmesperforms considerably better than the other state-of-the-art tool, TBCCD. We also assessedHolmeson unseen projects and performed cross dataset experiments to evaluate the generalizability ofHolmes. Our results affirm thatHolmesoutperforms TBCCD since most of the pairs thatHolmesdetected were either undetected or suboptimally reported by TBCCD. Nikita Mehrotra, Navdha Agarwal, Saket Anand, David Lo 0001, Rahul Purandare |
IEEE Trans. Software Eng. | 4 |
| 2020 | Contextual Diversity for Active Learning
Sharat Agarwal, Himanshu Arora, Saket Anand, Chetan Arora 0001 |
ECCV (16) | 3 |
| 2020 | Pseudo RGB-D for Self-improving Monocular SLAM and Depth Prediction
Lokender Tiwari, Pan Ji, Quoc-Huy Tran, Bingbing Zhuang, Saket Anand, Manmohan Krishna Chandraker |
ECCV (11) | 5 |
| 2020 | BIRDSAI: A Dataset for Detection and Tracking in Aerial Thermal Infrared VideosabstractMonitoring of protected areas to curb illegal activities like poaching and animal trafficking is a monumental task. To augment existing manual patrolling efforts, unmanned aerial surveillance using visible and thermal infrared (TIR) cameras is increasingly being adopted. Automated data acquisition has become easier with advances in unmanned aerial vehicles (UAVs) and sensors like TIR cameras, which allow surveillance at night when poaching typically occurs. However, it is still a challenge to accurately and quickly process large amounts of the resulting TIR data. In this paper, we present the first large dataset collected using a TIR camera mounted on a fixed-wing UAV in multiple African protected areas. This dataset includes TIR videos of humans and animals with several challenging scenarios like scale variations, background clutter due to thermal reflections, large camera rotations, and motion blur. Additionally, we provide another dataset with videos synthetically generated with the publicly available Microsoft AirSim simulation platform using a 3D model of an African savanna and a TIR camera model. Through our benchmarking experiments on state-of-the-art detectors, we demonstrate that leveraging the synthetic data in a domain adaptive setting can significantly improve detection performance. We also evaluate various recent approaches for single and multi-object tracking. With the increasing popularity of aerial imagery for monitoring and surveillance purposes, we anticipate this unique dataset to be used to develop and evaluate techniques for object detection, tracking, and domain adaptation for aerial, TIR videos. Elizabeth Bondi-Kelly, Raghav Jain, Palash Aggrawal, Saket Anand, Robert Hannaford, Ashish Kapoor, James Piavis, Shital Shah, Lucas Joppa, Bistra Dilkina, Milind Tambe |
WACV | 4 |
| 2020 | Intelligent querying for target tracking in camera networks using deep Q-learning with n-step bootstrapping
Anil Sharma, Saket Anand, Sanjit Krishnan Kaul |
Image Vis. Comput. | 2 |
| 2019 | PrOSe: Product of Orthogonal Spheres Parameterization for Disentangled Representation Learning
Ankita Shukla, Sarthak Bhagat, Shagun Uppal, Saket Anand, Pavan Turaga |
BMVC | 4 |
| 2019 | Primate Face Identification in the Wild
Ankita Shukla, Gullal Singh Cheema, Saket Anand, Qamar Qureshi, Yadvendradev Jhala |
PRICAI (3) | 3 |
| 2018 | Disentangling Factors of Variation with Cycle-Consistent Variational Auto-encoders
Ananya Harsh Jha, Saket Anand, Maneesh Kumar Singh 0001, V. S. R. Veeravasarapu |
ECCV (3) | 2 |
| 2018 | Adversarial Learning of Raw Speech Features for Domain Invariant Speech RecognitionabstractRecent advances in neural network based acoustic modelling have shown significant improvements in automatic speech recognition (ASR) performance. In order for acoustic models to be able to handle large acoustic variability, large amounts of labeled data is necessary, which are often expensive to obtain. This paper explores the application of adversarial training to learn features from raw speech that are invariant to acoustic variability. This acoustic variability is referred to as a domain shift in this paper. The experimental study presented in this paper leverages the architecture of Domain Adversarial Neural Networks (DANNs) [1] which uses data from two different domains. The DANN is a Y-shaped network that consists of a multi-layer CNN feature extractor module that is common to a label (senone) classifier and a so-called domain classifier. The utility of DANNs is evaluated on multiple datasets with domain shifts caused due to differences in gender and speaker accents. Promising empirical results indicate the strength of adversarial training for unsupervised domain adaptation in ASR, thereby emphasizing the ability of DANNs to learn domain invariant features from raw speech. Aditay Tripathi, Aanchan Mohan, Saket Anand, Maneesh Kumar Singh 0001 |
ICASSP | 3 |
| 2018 | DGSAC: Density Guided Sampling and ConsensusabstractIn this paper, we present an automatic multi-model fitting pipeline that can robustly fit multiple geometric models present in the corrupted and noisy data. Our approach can handle large data corruption and requires no user input, unlike most state-of-the-art approaches. The pipeline can be used as an independent block in many geometric vision applications like 3D reconstruction, motion and planar segmentation. We use residual density as the primary tool to guide hypothesis generation, estimate the fraction of inliers, and perform model selection. We show results for a diverse set of geometric models like planar homographies, fundamental matrices and vanishing points, which often arise in various computer vision applications. Despite being fully automatic, our approach achieves competitive performance compared to state-of-the-art approaches in terms of accuracy and computational time. Lokender Tiwari, Saket Anand |
WACV | 2 |
| 2017 | A Parallel Architecture for High Frame Rate Stereo using Semi-Global Matching
Akshay Jain 0003, Alexander Fell, Saket Anand |
BMVC | 3 |
| 2017 | Automatic Detection and Recognition of Individuals in Patterned Species
Gullal Singh Cheema, Saket Anand |
ECML/PKDD (3) | 2 |
| 2016 | Robust Multi-Model Fitting Using Density and Preference Analysis
Lokender Tiwari, Saket Anand, Sushil Mittal |
ACCV (4) | 2 |
| 2016 | Metric learning based automatic segmentation of patterned speciesabstractMany species in the wild exhibit a visual pattern that can be used to uniquely identify an individual. This observation has recently led to visual animal biometrics become a rapidly growing application area of computer vision. Customized software tools for animal biometrics already employ vision based techniques to recognize individuals in images taken in uncontrolled environments. However, most existing tools require the user to localize the animals for accurate identification. In this work, we propose a figure/ground segmentation method that automatically extracts out the animal in an image. Our method relies on a semi-supervised metric learning algorithm that uses a small amount of training data without compromising generalization performance. We design a simple pipeline comprising of superpixel segmentation, texture based feature extraction followed by mean shift clustering using the learned metric. We show that our approach can yield competitive results for figure/ground segmentation of patterned animals in images taken in the wild, often under extreme illumination conditions. Ankita Shukla, Saket Anand |
ICIP | 2 |
| 2016 | Fast hypothesis filtering for multi-structure geometric model fittingabstractWe propose a fast and efficient two-stage hypothesis filtering technique that can improve performance of clustering based robust multi-model fitting algorithms. Sampling based hypothesis generation is nondeterministic and permits little control over generating poor model hypotheses, often leading to a significant proportion of bad hypotheses. Our novel filtering approach leverages the asymmetry in the distributions of points around the inlier/outlier boundary via the sample skewness computed in the residual space. The output is a set of promising hypotheses which aid multi-model fitting algorithms in improving accuracy as well as running time. We validate our approach on the AdelaideRMF dataset and show favorable results along with comparisons to state-of-the-art. Lokender Tiwari, Saket Anand |
ICIP | 2 |
| 2016 | Stacked Robust Autoencoder for Classification
Janki Mehta, Kavya Gupta, Anupriya Gogna, Angshul Majumdar, Saket Anand |
ICONIP (3) | 5 |
| 2014 | Semi-Supervised Kernel Mean Shift ClusteringabstractMean shift clustering is a powerful nonparametric technique that does not require prior knowledge of the number of clusters and does not constrain the shape of the clusters. However, being completely unsupervised, its performance suffers when the original distance metric fails to capture the underlying cluster structure. Despite recent advances in semi-supervised clustering methods, there has been little effort towards incorporating supervision into mean shift. We propose a semi-supervised framework for kernel mean shift clustering (SKMS) that uses only pairwise constraints to guide the clustering procedure. The points are first mapped to a high-dimensional kernel space where the constraints are imposed by a linear transformation of the mapped points. This is achieved by modifying the initial kernel matrix by minimizing a log det divergence-based objective function. We show the advantages of SKMS by evaluating its performance on various synthetic and real datasets while comparing with state-of-the-art semi-supervised clustering algorithms. Saket Anand, Sushil Mittal, Oncel Tuzel, Peter Meer |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Generalized Projection-Based M-EstimatorabstractWe propose a novel robust estimation algorithm—the generalized projection-based M-estimator (gpbM), which does not require the user to specify any scale parameters. The algorithm is general and can handle heteroscedastic data with multiple linear constraints for single and multicarrier problems. The gpbM has three distinct stages—scale estimation, robust model estimation, and inlier/outlier dichotomy. In contrast, in its predecessor pbM, each model hypotheses was associated with a different scale estimate. For data containing multiple inlier structures with generally different noise covariances, the estimator iteratively determines one structure at a time. The model estimation can be further optimized by using Grassmann manifold theory. We present several homoscedastic and heteroscedastic synthetic and real-world computer vision problems with single and multiple carriers. Sushil Mittal, Saket Anand, Peter Meer |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Generalized projection based M-estimator: Theory and applicationsabstractWe introduce a robust estimator called generalized projection based M-estimator (gpbM) which does not require the user to specify any scale parameters. For multiple inlier structures, with different noise covariances, the estimator iteratively determines one inlier structure at a time. Unlike pbM, where the scale of the inlier noise is estimated simultaneously with the model parameters, gpbM has three distinct stages-scale estimation, robust model estimation and inlier/outlier dichotomy. We evaluate our performance on challenging synthetic data, face image clustering upto ten different faces from Yale Face Database B and multi-body projective motion segmentation problem on Hopkins155 dataset. Results of state-of-the-art methods are presented for comparison. Sushil Mittal, Saket Anand, Peter Meer |
CVPR | 2 |
| 2006 | Experimental Analysis of Sequential Decision Making Algorithms for Port of Entry Inspection Procedures
Saket Anand, David Madigan, Richard J. Mammone, Saumitr Pathak, Fred S. Roberts |
ISI | 1 |