VLDB 2026 Research / reviewers in the wild / expert
Aayush Dhakal
dblp:353/1867
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0003-4431-0628ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-EmbeddingsabstractThe choice of representation for geographic location significantly impacts the accuracy of models for a broad range of geospatial tasks, including fine-grained species classification, population density estimation, and biome classification. Recent works like SatCLIP and GeoCLIP learn such representations by contrastively aligning geolocation with co-located images. While these methods work exceptionally well, in this paper, we posit that the current training strategies fail to fully capture the important visual features. We provide an information theoretic perspective on why the resulting embeddings from these methods discard crucial visual information that is important for many downstream tasks. To solve this problem, we propose a novel retrieval-augmented strategy called RANGE. We build our method on the intuition that the visual features of a location can be estimated by combining the visual features from multiple similar-looking locations. We evaluate our method across a wide variety of tasks. Our results show that RANGE outperforms the existing state-of-the-art models with significant margins in most tasks. We show gains of up to 13.1% on classification tasks and 0.145 R2on regression tasks. All our code and models will be made available at: https://github.com/mvrl/RANGE. Aayush Dhakal, Srikumar Sastry, Subash Khanal, Eric Xing 0002, Nathan Jacobs |
CVPR | 1 |
| 2025 | Global and Local Entailment Learning for Natural World ImageryabstractLearning the hierarchical structure of data in vision-language models is a significant challenge. Previous works have attempted to address this challenge by employing entailment learning. However, these approaches fail to model the transitive nature of entailment explicitly, which establishes the relationship between order and semantics within a representation space. In this work, we introduce Radial Cross-Modal Embeddings (RCME), a framework that enables the explicit modeling of transitivity-enforced entailment. Our proposed framework optimizes for the partial order of concepts within vision-language models. By leveraging our framework, we develop a hierarchical vision-language foundation model capable of representing the hierarchy in the Tree of Life. Our experiments on hierarchical species classification and hierarchical retrieval tasks demonstrate the enhanced performance of our models compared to the existing state-of-the-art models. Our code and models are open-sourced at https://vishu26.github.io/RCME/index.html. Srikumar Sastry, Aayush Dhakal, Eric Xing 0002, Subash Khanal, Nathan Jacobs |
ICCV | 2 |
| 2025 | TaxaBind: A Unified Embedding Space for Ecological ApplicationsabstractWe present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio, and environmental features, useful for solving eco-logical problems. To learn this joint embedding space, we leverage ground-level images of species as a binding modality. We propose multimodal patching, a technique for effectively distilling the knowledge from various modalities into the binding modality. We construct two large datasets for pretraining: iSatNat with species images and satellite images, and iSoundNat with species images and audio. Additionally, we introduce TaxaBench-8k, a diverse multimodal dataset with six paired modalities for evaluating deep learning models on ecological tasks. Experiments with TaxaBind demonstrate its strong zero-shot and emer-gent capabilities on a range of tasks including species classification, cross-model retrieval, and audio classification. The datasets and models are made available at https://github.com/mvr1/TaxaBind. Srikumar Sastry, Subash Khanal, Aayush Dhakal, Nathan Jacobs |
WACV | 3 |
| 2024 | FroSSL: Frobenius Norm Minimization for Efficient Multiview Self-supervised Learning
Oscar Skean, Aayush Dhakal, Nathan Jacobs, Luis Gonzalo Sánchez Giraldo |
ECCV (89) | 2 |
| 2024 | GeoBind: Binding Text, Image, and Audio through Satellite ImagesabstractIn remote sensing, we are interested in modeling various modalities for some geographic location. Several works have focused on learning the relationship between a location and type of landscape, habitability, audio, textual descriptions, etc. Recently, a common way to approach these problems is to train a deep-learning model that uses satellite images to infer some unique characteristics of the location. In this work, we present a deep-learning model, GeoBind, that can infer about multiple modalities, specifically text, image, and audio, from satellite imagery of a location. To do this, we use satellite images as the binding element and contrastively align all other modalities to the satellite image data. Our training results in a joint embedding space with multiple types of data: satellite image, ground-level image, audio, and text. Furthermore, our approach does not require a single complex dataset that contains all the modalities mentioned above. Rather it only requires multiple satellite-image paired data. While we only align three modalities in this paper, we present a general framework that can be used to create an embedding space with any number of modalities by using satellite images as the binding element. Our results show that, unlike traditional unimodal models, GeoBind is versatile and can reason about multiple modalities for a given satellite image input. Aayush Dhakal, Subash Khanal, Srikumar Sastry, Nathan Jacobs |
IGARSS | 1 |
| 2024 | Aligning Geo-Tagged Clip Representations and Satellite Imagery for Few-Shot Land Use ClassificationabstractA major difference between ground-level and satellite imagery of landscapes lies in their semantic granularity: ground-level images tend to offer details on objects and human activities, while satellite images provide broader geographic context but, typically, with coarser semantics. This study aims to leverage this complementary information by integrating fine-grained insights from a ground-level view into the analysis of satellite image data. To achieve this integration, we propose to align a satellite image representation with co-located geo-tagged ground-level image CLIP representations. This method focuses on enriching satellite image visual features by leveraging the inherent visual characteristics found in ground-level images as a reference in a contrastive manner, without relying on additional textual information to guide the learning process. We evaluate the quality of the learned representations on the EuroSAT benchmark in various few-shot settings. Pallavi Jain 0004, Diego Marcos, Dino Ienco, Roberto Interdonato, Aayush Dhakal, Nathan Jacobs, Tristan Berchoux |
IGARSS | 5 |
| 2024 | PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape MappingabstractA soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial scales, we represent locations with multi-scale satellite imagery and learn a joint representation among this imagery, audio, and text. To capture the inherent uncertainty in the soundscape of a location, we design the representation space to be probabilistic. We also fuse ubiquitous metadata (including geolocation, time, and data source) to enable learning of spatially and temporally dynamic representations of soundscapes. We demonstrate the utility of our framework by creating large-scale soundscape maps integrating both audio and text with temporal control. To facilitate future research on this task, we also introduce a large-scale dataset, GeoSound, containing over 300k geotagged audio samples paired with both low- and high-resolution satellite imagery. We demonstrate that our method outperforms the existing state-of-the-art on both GeoSound and the existing SoundingEarth dataset. Our dataset and code is available at https://github.com/mvrl/PSM. Subash Khanal, Eric Xing 0002, Srikumar Sastry, Aayush Dhakal, Zhexiao Xiong, Nathan Jacobs |
ACM Multimedia | 4 |
| 2024 | BirdSAT: Cross-View Contrastive Masked Autoencoders for Bird Species Classification and MappingabstractWe propose a metadata-aware self-supervised learning (SSL) framework useful for fine-grained classification and ecological mapping of bird species around the world. Our framework unifies two SSL strategies: Contrastive Learning (CL) and Masked Image Modeling (MIM), while also enriching the embedding space with metadata available with ground-level imagery of birds. We separately train uni-modal and cross-modal ViT on a novel cross-view global bird species dataset containing ground-level imagery, metadata (location, time), and corresponding satellite imagery. We demonstrate that our models learn fine-grained and geographically conditioned features of birds, by evaluating on two downstream tasks: fine-grained visual classification (FGVC) and cross-modal retrieval. Pre-trained models learned using our framework achieve SotA performance on FGVC of iNAT-2021 birds and in transfer learning settings for CUB-200-2011 and NABirds datasets. Moreover, the impressive cross-modal retrieval performance of our model enables the creation of species distribution maps across any geographic region. The dataset and source code will be released at https://github.com/mvrl/BirdSAT. Srikumar Sastry, Subash Khanal, Aayush Dhakal, Nathan Jacobs |
WACV | 3 |
| 2023 | Learning Tri-modal Embeddings for Zero-Shot Soundscape Mapping
Subash Khanal, Srikumar Sastry, Aayush Dhakal, Nathan Jacobs |
BMVC | 3 |