Rangel Daroya

dblp:239/3946 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
4since 2021 · last 2026
0009-0007-5309-6359ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 56% Segmentation and scene understanding · 17% Learning paradigms · 13%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Environmental and earth informatics · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 57% Recommender systems · 43%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
1.012026
RiverScope: High-Resolution River Masking Dataset · AAAI 2026
Environmental and earth informatics
hydrology
1.012026
RiverScope: High-Resolution River Masking Dataset · AAAI 2026
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
WildSAT: Learning Satellite Image Representations from Wildlife Observations · ICCV 2025
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning
0.912025
WildSAT: Learning Satellite Image Representations from Wildlife Observations · ICCV 2025
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › geometric embedding
box embedding
0.812024
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships · CVPR 2024
Machine learning › Learning paradigms › multi-task learning
task relationship modeling
0.812024
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships · CVPR 2024
Machine learning › Representation and self-supervised learning › representation learning › joint representation learning › multi-task representation learning
task representation
0.812024
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships · CVPR 2024
Machine learning › Transfer learning and domain adaptation
transferability estimation
0.812024
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships · CVPR 2024
Data mining
dataset construction
0.312026
RiverScope: High-Resolution River Masking Dataset · AAAI 2026
Environmental and earth informatics
biodiversity monitoring
0.312025
WildSAT: Learning Satellite Image Representations from Wildlife Observations · ICCV 2025
Environmental and earth informatics › remote sensing
remote sensing image analysis
0.312025
WildSAT: Learning Satellite Image Representations from Wildlife Observations · ICCV 2025
Recommender systems › representation learning for recommendation
embedding-based recommendation
0.212024
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships · CVPR 2024

Methods — techniques the papers use, named apart from their topics

transformer · 3.0transfer learning · 3.0self-supervised pretraining · 3.0CNN · 3.0zero-shot retrieval · 1.7contrastive learning · 1.7task2vec · 1.5t-SNE · 1.5CLIP · 1.5
YearPublicationVenuePosition
2026 RiverScope: High-Resolution River Masking Dataset
abstract
Surface water dynamics play a critical role in Earth’s climate system, influencing ecosystems, agriculture, disaster resilience, and sustainable development. Yet monitoring rivers and surface water at fine spatial and temporal scales remains challenging---especially for narrow or sediment-rich rivers that are poorly captured by low-resolution satellite data. To address this, we introduce RiverScope, a high-resolution dataset developed through collaboration between computer science and hydrology experts. RiverScope comprises 1,145 high-resolution images (covering 2,577 square kilometers) with expert-labeled river and surface water masks, requiring over 100 hours of manual annotation. Each image is co-registered with Sentinel-2, SWOT, and the SWOT River Database (SWORD), enabling the evaluation of cost-accuracy trade-offs across sensors---a key consideration for operational water monitoring. We also establish the first global, high-resolution benchmark for river width estimation, achieving a median error of 7.2 meters---significantly outperforming existing satellite-derived methods. We extensively evaluate deep networks across multiple architectures (e.g., CNNs and transformers), pretraining strategies (e.g., supervised and self-supervised), and training datasets (e.g., ImageNet and satellite imagery). Our best-performing models combine the benefits of transfer learning with the use of all the multispectral PlanetScope channels via learned adaptors. RiverScope provides a valuable resource for fine-scale and multi-sensor hydrological modeling, supporting climate adaptation and sustainable water management.
Rangel Daroya, Taylor Rowley, Jonathan Acero Flores, Elisa Friedmann, Fiona Bennitt, Heejin An, Travis Simmons, Marissa Jean Hughes, Camryn L. Kluetmeier, Solomon Kica, J. Daniel Vélez, Sarah E. Esenther, Thomas E. Howard, Yanqi Ye, Audrey Turcotte, Colin J. Gleason, Subhransu Maji
AAAI1
2026 SuperRivolution: Fine-Scale Rivers from Coarse Temporal Satellite Imagery
abstract
Satellite missions provide valuable optical data for monitoring rivers at diverse spatial and temporal scales. However, accessibility remains a challenge: high-resolution imagery is ideal for fine-grained monitoring but is typically scarce and expensive compared to low-resolution imagery. To address this gap, we introduce SuperRivolution, a framework that improves river segmentation resolution by leveraging information from time series of low-resolution satellite images. We contribute a new benchmark dataset of 9, 810 low-resolution temporal images paired with high-resolution labels from an existing river monitoring dataset. Using this benchmark, we investigate multiple strategies for river segmentation, including ensembling single-image models, applying image super-resolution, and developing end-to-end models trained on temporal sequences. SuperRivolution significantly outperforms single-image methods and baseline temporal approaches, narrowing the gap with supervised high-resolution models. For example, the F1 score for river segmentation improves from 60.9% to 80.5%, while the state-of-the-art model operating on high-resolution images achieves 94.1%. Similar improvements are also observed in river width estimation tasks. Our results highlight the potential of publicly available low-resolution satellite archives for fine-scale river monitoring.
Rangel Daroya, Subhransu Maji
WACV1
2025 WildSAT: Learning Satellite Image Representations from Wildlife Observations
abstract
Species distributions encode valuable ecological and environmental information, yet their potential for guiding representation learning in remote sensing remains underexplored. We introduce WildSAT, which pairs satellite images with millions of geo-tagged wildlife observations readily-available on citizen science platforms. WildSAT employs a contrastive learning approach that jointly leverages satellite images, species occurrence maps, and textual habitat descriptions to train or fine-tune models. This approach significantly improves performance on diverse satellite image recognition tasks, outperforming both ImageNet-pretrained models and satellite-specific baselines. Additionally, by aligning visual and textual information, WildSAT enables zero-shot retrieval, allowing users to search geographic locations based on textual descriptions. WildSAT surpasses recent cross-modal learning methods, including approaches that align satellite images with ground imagery or wildlife photos, demonstrating the advantages of our approach. Finally, we analyze the impact of key design choices and highlight the broad applicability of WildSAT to remote sensing and biodiversity monitoring.
Rangel Daroya, Elijah Cole, Oisin Mac Aodha, Grant Van Horn, Subhransu Maji
ICCV1
2024 Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships
abstract
Modeling and visualizing relationships between tasks or datasets is an important step towards solving various meta-tasks such as dataset discovery, multi-tasking, and transfer learning. However, many relationships, such as containment and transferability, are naturally asymmetric and current approaches for representation and visualization (e.g., t-SNE [44]) do not readily support this. We propose TASK2Box, an approach to represent tasks using box embeddings-axis-aligned hyperrectangles in low dimensional spaces-that can capture asymmetric relation-ships between them through volumetric overlaps. We show that TASK2Box accurately predicts unseen hierarchical relationships between nodes in ImageNet and iNaturalist datasets, as well as transferability between tasks in the Taskonomy benchmark. We also show that box embeddings estimatedfrom task representations (e.g., CLIP [36], Task2Vec [4], or attribute based [15]) can be used to pre-dict relationships between unseen tasks more accurately than classifiers trained on the same representations, as well as handcrafted asymmetric distances (e.g., KL divergence). This suggests that low-dimensional box embeddings can effectively capture these task relationships and have the added advantage of being interpretable. We use the approach to visualize relationships among publicly available image classification datasets on popular dataset hosting platform called Hugging Face.
Rangel Daroya, Aaron Sun, Subhransu Maji
CVPR1
2018 Alphabet Sign Language Image Classification Using Deep Learning
abstract
Sign language is very important for people who have impaired hearing and speaking inabilities. In this work, we present a method to classify RGB images of static letter hand poses in Sign Language using a Convolutional Neural Netowrk (CNN) inspired by Densely Connected Convolutional Neural Networks (DenseNet). It was further implemented to classify sign languages in real time using a web camera. DenseNet has been widely used for classification tasks due to the advantages it introduces such as alleviating the vanishing gradient - a common problem encountered with deep networks. Since a deep network is proposed to be used for our sign language classification task, this characteristic is useful. Our proposed network was able to achieve an accuracy of 90.3 % which is comparable to other works including those that used depth images in addition to RGB images. Our network was also able to achieve prediction rates of 50 to 100 Hz which makes it capable of real-time prediction.
Rangel Daroya, Daryl Peralta, Prospero C. Naval Jr.
TENCON1