VLDB 2026 Research / reviewers in the wild / expert
Subash Khanal
dblp:289/3186
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-7666-8603ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-EmbeddingsabstractThe choice of representation for geographic location significantly impacts the accuracy of models for a broad range of geospatial tasks, including fine-grained species classification, population density estimation, and biome classification. Recent works like SatCLIP and GeoCLIP learn such representations by contrastively aligning geolocation with co-located images. While these methods work exceptionally well, in this paper, we posit that the current training strategies fail to fully capture the important visual features. We provide an information theoretic perspective on why the resulting embeddings from these methods discard crucial visual information that is important for many downstream tasks. To solve this problem, we propose a novel retrieval-augmented strategy called RANGE. We build our method on the intuition that the visual features of a location can be estimated by combining the visual features from multiple similar-looking locations. We evaluate our method across a wide variety of tasks. Our results show that RANGE outperforms the existing state-of-the-art models with significant margins in most tasks. We show gains of up to 13.1% on classification tasks and 0.145 R2on regression tasks. All our code and models will be made available at: https://github.com/mvrl/RANGE. Aayush Dhakal, Srikumar Sastry, Subash Khanal, Eric Xing 0002, Nathan Jacobs |
CVPR | 3 |
| 2025 | Global and Local Entailment Learning for Natural World ImageryabstractLearning the hierarchical structure of data in vision-language models is a significant challenge. Previous works have attempted to address this challenge by employing entailment learning. However, these approaches fail to model the transitive nature of entailment explicitly, which establishes the relationship between order and semantics within a representation space. In this work, we introduce Radial Cross-Modal Embeddings (RCME), a framework that enables the explicit modeling of transitivity-enforced entailment. Our proposed framework optimizes for the partial order of concepts within vision-language models. By leveraging our framework, we develop a hierarchical vision-language foundation model capable of representing the hierarchy in the Tree of Life. Our experiments on hierarchical species classification and hierarchical retrieval tasks demonstrate the enhanced performance of our models compared to the existing state-of-the-art models. Our code and models are open-sourced at https://vishu26.github.io/RCME/index.html. Srikumar Sastry, Aayush Dhakal, Eric Xing 0002, Subash Khanal, Nathan Jacobs |
ICCV | 4 |
| 2025 | TaxaBind: A Unified Embedding Space for Ecological ApplicationsabstractWe present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio, and environmental features, useful for solving eco-logical problems. To learn this joint embedding space, we leverage ground-level images of species as a binding modality. We propose multimodal patching, a technique for effectively distilling the knowledge from various modalities into the binding modality. We construct two large datasets for pretraining: iSatNat with species images and satellite images, and iSoundNat with species images and audio. Additionally, we introduce TaxaBench-8k, a diverse multimodal dataset with six paired modalities for evaluating deep learning models on ecological tasks. Experiments with TaxaBind demonstrate its strong zero-shot and emer-gent capabilities on a range of tasks including species classification, cross-model retrieval, and audio classification. The datasets and models are made available at https://github.com/mvr1/TaxaBind. Srikumar Sastry, Subash Khanal, Aayush Dhakal, Nathan Jacobs |
WACV | 2 |
| 2024 | GeoBind: Binding Text, Image, and Audio through Satellite ImagesabstractIn remote sensing, we are interested in modeling various modalities for some geographic location. Several works have focused on learning the relationship between a location and type of landscape, habitability, audio, textual descriptions, etc. Recently, a common way to approach these problems is to train a deep-learning model that uses satellite images to infer some unique characteristics of the location. In this work, we present a deep-learning model, GeoBind, that can infer about multiple modalities, specifically text, image, and audio, from satellite imagery of a location. To do this, we use satellite images as the binding element and contrastively align all other modalities to the satellite image data. Our training results in a joint embedding space with multiple types of data: satellite image, ground-level image, audio, and text. Furthermore, our approach does not require a single complex dataset that contains all the modalities mentioned above. Rather it only requires multiple satellite-image paired data. While we only align three modalities in this paper, we present a general framework that can be used to create an embedding space with any number of modalities by using satellite images as the binding element. Our results show that, unlike traditional unimodal models, GeoBind is versatile and can reason about multiple modalities for a given satellite image input. Aayush Dhakal, Subash Khanal, Srikumar Sastry, Nathan Jacobs |
IGARSS | 2 |
| 2024 | PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape MappingabstractA soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial scales, we represent locations with multi-scale satellite imagery and learn a joint representation among this imagery, audio, and text. To capture the inherent uncertainty in the soundscape of a location, we design the representation space to be probabilistic. We also fuse ubiquitous metadata (including geolocation, time, and data source) to enable learning of spatially and temporally dynamic representations of soundscapes. We demonstrate the utility of our framework by creating large-scale soundscape maps integrating both audio and text with temporal control. To facilitate future research on this task, we also introduce a large-scale dataset, GeoSound, containing over 300k geotagged audio samples paired with both low- and high-resolution satellite imagery. We demonstrate that our method outperforms the existing state-of-the-art on both GeoSound and the existing SoundingEarth dataset. Our dataset and code is available at https://github.com/mvrl/PSM. Subash Khanal, Eric Xing 0002, Srikumar Sastry, Aayush Dhakal, Zhexiao Xiong, Nathan Jacobs |
ACM Multimedia | 1 |
| 2024 | BirdSAT: Cross-View Contrastive Masked Autoencoders for Bird Species Classification and MappingabstractWe propose a metadata-aware self-supervised learning (SSL) framework useful for fine-grained classification and ecological mapping of bird species around the world. Our framework unifies two SSL strategies: Contrastive Learning (CL) and Masked Image Modeling (MIM), while also enriching the embedding space with metadata available with ground-level imagery of birds. We separately train uni-modal and cross-modal ViT on a novel cross-view global bird species dataset containing ground-level imagery, metadata (location, time), and corresponding satellite imagery. We demonstrate that our models learn fine-grained and geographically conditioned features of birds, by evaluating on two downstream tasks: fine-grained visual classification (FGVC) and cross-modal retrieval. Pre-trained models learned using our framework achieve SotA performance on FGVC of iNAT-2021 birds and in transfer learning settings for CUB-200-2011 and NABirds datasets. Moreover, the impressive cross-modal retrieval performance of our model enables the creation of species distribution maps across any geographic region. The dataset and source code will be released at https://github.com/mvrl/BirdSAT. Srikumar Sastry, Subash Khanal, Aayush Dhakal, Nathan Jacobs |
WACV | 2 |
| 2023 | Learning Tri-modal Embeddings for Zero-Shot Soundscape Mapping
Subash Khanal, Srikumar Sastry, Aayush Dhakal, Nathan Jacobs |
BMVC | 1 |
| 2021 | Alzheimer's Disease Classification Using Genetic DataabstractThere has been a recent surge of interest in using genetic data to build ML-based accurate and interpretable disease classification models. In this line of research, we separately assess the potential of the peripheral blood gene expression data as well as the Single Nucleotide Polymorphism (SNP) data in building ML models for AD classification. We present a systematic approach on feature selection and ML model design using both types of genetic data provided by the Alzheimer’s Disease Neuroimaging Initiatives (ADNI). Our two-step feature selection produced a curated list of important genes. In addition to these selected genetic features, to examine the role of non-genetic covariates, we included age and number of education years (EDU) as extra features. In the Control (CN) vs. AD classification, the best performing classifier, XGBoost, trained with gene expression features only and that with extra features included had Area Under Curve (AUC) of 0.64 and 0.65 respectively. However, AUC for the same task using SNP data only and that with extra features included was 0.56 and 0.64 respectively. The just above chance results of classifier trained with SNP features and the improvement when used along with additional covariates indicate low potential of SNP data in AD classification when used alone while also indicating the importance of non-genetic factors associated with AD. Nevertheless, with well above chance performance, gene expression features show great potential especially between groups of AD progression, i.e., CN vs. AD, CN vs. EMCI, EMCI vs. AD and LMCI vs. AD. The source code and manual are available at https://github.com/mvrl/ADNI_Genetics. Subash Khanal, Jin Chen 0004, Nathan Jacobs, Ai-Ling Lin |
BIBM | 1 |
| 2021 | Hierarchical Probabilistic Embeddings for Multi-View Image ClassificationabstractWe address the task of image classification, when the available spectral bands can vary from image to image. We propose a model that learns to represent uncertainty over latent features in a way that is conditioned on the available bands. We expect that images with fewer bands will generally be more difficult to classify and hence have higher uncertainty. We compare two strategies for training such a model, one which uses explicit hierarchical constraints and one which relies on implicit constraints. We evaluate both using RGB and multispectral imagery from the EuroSat dataset and find that the hierarchical approach improves the compatibility of the resulting distributions without sacrificing accuracy. Benjamin Brodie, Subash Khanal, Muhammad Usman Rafique, Connor Greenwell, Nathan Jacobs |
IGARSS | 2 |
| 2021 | Articulatory Comparison of L1 and L2 Speech for Mispronunciation DiagnosisabstractThis paper compares the difference in articulation patterns between native (L1) and non-native (L2) Mandarin speakers of English, for the purpose of providing an understanding of mispronunciation behaviors of L2 learners. Consensus transcriptions from the Electromagnetic Articulography Mandarin Accented English (EMA-MAE) corpus are used to identify commonly occurring substitution errors for consonants and vowels. Phoneme level alignments of the utterances produced by speech recognition models are used to extract articulatory feature vectors representing correct and substituted sounds from L1 and L2 speaker groups respectively. The articulatory features that are significantly different between the two groups are identified along with the direction of error for the L2 speaker group. Experimental results provide information about which types of substitutions are most common and which specific articulators are the most significant contributors to those errors. Subash Khanal, Michael T. Johnson, Narjes Bozorg |
SLT | 1 |