EDBT 2026 Demo / reviewers in the wild / expert
Nathan Jacobs
dblp:82/3140 · also Nathan B. Jacobs
· DBLP profile ↗
101ranked-venue papers
13as first author
39since 2021 · last 2026
0000-0002-4242-8967ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 9 first-author · 23 since 2021Artificial intelligence and machine learning · 47 · 8 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beta Distribution Learning for Reliable Roadway Crash Risk AssessmentabstractRoadway traffic accidents represent a global health crisis, responsible for over a million deaths annually and costing many countries up to 3% of their GDP. Traditional traffic safety studies often examine risk factors in isolation, overlooking the spatial complexity and contextual interactions inherent in the built environment. Furthermore, conventional Neural Network-based risk estimators typically generate point estimates without conveying model uncertainty, limiting their utility in critical decision-making. To address these shortcomings, we introduce a novel geospatial deep learning framework that leverages satellite imagery as a comprehensive spatial input. This approach enables the model to capture the nuanced spatial patterns and embedded environmental risk factors that contribute to fatal crash risks. Rather than producing a single deterministic output, our model estimates a full Beta probability distribution over fatal crash risk, yielding accurate and uncertainty-aware predictions--a critical feature for trustworthy AI in safety-critical applications. Our model outperforms baselines by achieving a 17-23% improvement in recall, a key metric for flagging potential dangers, while delivering superior calibration. By providing reliable and interpretable risk assessments from satellite imagery alone, our method enables safer autonomous navigation and offers a highly scalable tool for urban planners and policymakers to enhance roadway safety equitably and cost-effectively. Ahmad Elallaf, Nathan Jacobs, Xinyue Ye, Gongbo Liang |
AAAI | 2 |
| 2026 | VectorSynth: Fine-Grained Satellite Image Synthesis with Structured SemanticsabstractWe introduce VectorSynth, a diffusion-based framework for pixel-accurate satellite image synthesis conditioned on polygonal geographic annotations with semantic attributes. Unlike prior text- or layout-conditioned models, VectorSynth learns dense cross-modal correspondences that align imagery and semantic vector geometry, enabling fine-grained, spatially grounded edits. A vision language alignment module produces pixel-level embeddings from polygon semantics; these embeddings guide a conditional image generation framework to respect both spatial extents and semantic cues. VectorSynth supports interactive workflows that mix language prompts with geometry-aware conditioning, allowing rapid what-if simulations, spatial edits, and map-informed content generation. For training and evaluation, we assemble a collection of satellite scenes paired with pixel-registered polygon annotations spanning diverse urban scenes with both built and natural features. We observe strong improvements over prior methods in semantic fidelity and structural realism, and show that our trained vision language model demonstrates fine-grained spatial grounding. The code and data are available at https://github.com/mvrl/VectorSynth. Daniel Cher, Brian Wei, Srikumar Sastry, Nathan Jacobs |
WACV | 4 |
| 2026 | Towards Unconstrained Cross-View Pose EstimationabstractCross-view pose estimation entails predicting the relative 3 Degrees-of-Freedom (3DoF) pose of an image within an aerial view. Existing work focuses on imagery in controlled settings featuring highly constrained parameters. In contrast, a wide variety of camera parameterizations are seen in-the-wild across tasks where such estimation is useful. To address this gap, we propose a method capable of performing cross-view pose estimation in these less constrained scenarios with ground-view images of unknown FoV, pitch, roll, and projection type (panoramic or rectilinear). Namely, our method avoids common assumptions—such as gravity/horizon alignment needed for geometric-based projections—and purely relies on a transformer to learn the cross-view relationships in a data-driven manner, paired with prediction modules to enable continuous querying of the pose search space. Evaluations of our approach demonstrates it’s ability to perform competitively with the state-of-the-art over the VIGOR benchmark, while maintaining performance in those harder less constrained scenarios. This supports our work as the first generalized approach to this task that is capable of operating with less-constrained imagery. Alexander Wollam, Kyle Ashley, Maxim Shugaev, Oliver Arend, Ilya Semenov, Hadis Dashtestani, Sumved Ravi, Nathan Jacobs |
WACV | 8 |
| 2025 | Fields of The World: A Machine Learning Benchmark Dataset for Global Agricultural Field Boundary SegmentationabstractCrop field boundaries are foundational datasets for agricultural monitoring and assessments but are expensive to collect manually. Machine learning (ML) methods for automatically extracting field boundaries from remotely sensed images could help realize the demand for these datasets at a global scale. However, current ML methods for field instance segmentation lack sufficient geographic coverage, accuracy, and generalization capabilities. Further, research on improving ML methods is restricted by the lack of labeled datasets representing the diversity of global agricultural fields. We present Fields of The World (FTW)---a novel ML benchmark dataset for agricultural field instance segmentation spanning 24 countries on four continents (Europe, Africa, Asia, and South America). FTW is an order of magnitude larger than previous datasets with 70,462 samples, each containing instance and semantic segmentation masks paired with multi-date, multi-spectral Sentinel-2 satellite images. We provide results from baseline models for the new FTW benchmark, show that models trained on FTW have better zero-shot and fine-tuning performance in held-out countries than models that aren't pre-trained with diverse datasets, and show positive qualitative zero-shot results of FTW models in a real-world scenario -- running on Sentinel-2 scenes over Ethiopia. Hannah Kerner, Snehal Chaudhari, Aninda Ghosh, Caleb Robinson, Eddie Choi, Nathan Jacobs, Matthias Mohr, Rahul Dodhia, Juan M. Lavista Ferres, Jennifer Marcus |
AAAI | 7 |
| 2025 | Active Geospatial Search for Efficient Tenant Eviction OutreachabstractTenant evictions threaten housing stability and are a major concern for many cities. An open question concerns whether data-driven methods enhance outreach programs that target at-risk tenants to mitigate their risk of eviction. We propose a novel active geospatial search (AGS) modeling framework for this problem. AGS integrates property-level information in a search policy that identifies a sequence of rental units to canvas to both determine their eviction risk and provide support if needed. We propose a hierarchical reinforcement learning approach to learn a search policy for AGS that scales to large urban areas containing thousands of parcels, balancing exploration and exploitation and accounting for travel costs and a budget constraint. Crucially, the search policy adapts online to newly discovered information about evictions. Evaluation using eviction data for a large urban area demonstrates that the proposed framework and algorithmic approach are considerably more effective at sequentially identifying eviction cases than baseline methods. Anindya Sarkar, Alex DiChristofano, Sanmay Das, Patrick J. Fowler, Nathan Jacobs, Yevgeniy Vorobeychik |
AAAI | 5 |
| 2025 | RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-EmbeddingsabstractThe choice of representation for geographic location significantly impacts the accuracy of models for a broad range of geospatial tasks, including fine-grained species classification, population density estimation, and biome classification. Recent works like SatCLIP and GeoCLIP learn such representations by contrastively aligning geolocation with co-located images. While these methods work exceptionally well, in this paper, we posit that the current training strategies fail to fully capture the important visual features. We provide an information theoretic perspective on why the resulting embeddings from these methods discard crucial visual information that is important for many downstream tasks. To solve this problem, we propose a novel retrieval-augmented strategy called RANGE. We build our method on the intuition that the visual features of a location can be estimated by combining the visual features from multiple similar-looking locations. We evaluate our method across a wide variety of tasks. Our results show that RANGE outperforms the existing state-of-the-art models with significant margins in most tasks. We show gains of up to 13.1% on classification tasks and 0.145 R2on regression tasks. All our code and models will be made available at: https://github.com/mvrl/RANGE. Aayush Dhakal, Srikumar Sastry, Subash Khanal, Eric Xing 0002, Nathan Jacobs |
CVPR | 6 |
| 2025 | ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalabstractComposed image retrieval (CIR) is the task of retrieving a target image specified by a query image and a relative text that describes a semantic modification to the query image. Existing methods in CIR struggle to accurately represent the image and the text modification, resulting in subpar performance. To address this limitation, we introduce a CIR framework, ConText-CIR, trained with a Text Concept-Consistency loss that encourages the representations of noun phrases in the text modification to better attend to the relevant parts of the query image. To support training with this loss function, we also propose a synthetic data generation pipeline that creates training data from existing CIR datasets or unlabeled images. We show that these components together enable stronger performance on CIR tasks, setting a new state-of-the-art in composed image retrieval in both the supervised and zero-shot settings on multiple benchmark datasets, including CIRR and CIRCO. Source code, model checkpoints, and our new datasets are available at https://github.com/mvrl/ConText-CIR. Eric Xing 0002, Pranavi Kolouju, Robert Pless, Abby Stylianou, Nathan Jacobs |
CVPR | 5 |
| 2025 | Towards Open-World Generation of Stereo Images and Unsupervised Matching
Feng Qiao 0001, Zhexiao Xiong, Eric Xing 0002, Nathan Jacobs |
ICCV | 4 |
| 2025 | Global and Local Entailment Learning for Natural World ImageryabstractLearning the hierarchical structure of data in vision-language models is a significant challenge. Previous works have attempted to address this challenge by employing entailment learning. However, these approaches fail to model the transitive nature of entailment explicitly, which establishes the relationship between order and semantics within a representation space. In this work, we introduce Radial Cross-Modal Embeddings (RCME), a framework that enables the explicit modeling of transitivity-enforced entailment. Our proposed framework optimizes for the partial order of concepts within vision-language models. By leveraging our framework, we develop a hierarchical vision-language foundation model capable of representing the hierarchy in the Tree of Life. Our experiments on hierarchical species classification and hierarchical retrieval tasks demonstrate the enhanced performance of our models compared to the existing state-of-the-art models. Our code and models are open-sourced at https://vishu26.github.io/RCME/index.html. Srikumar Sastry, Aayush Dhakal, Eric Xing 0002, Subash Khanal, Nathan Jacobs |
ICCV | 5 |
| 2025 | QuARI: Query Adaptive Retrieval ImprovementabstractMassive-scale pretraining has made vision-language models increasingly popular for image-to-image and text-to-image retrieval across a broad collection of domains. However, these models do not perform well when used for challenging retrieval tasks, such as instance retrieval in very large-scale image collections. Recent work has shown that linear transformations of VLM features trained for instance retrieval can improve performance by emphasizing subspaces that relate to the domain of interest. In this paper, we explore a more extreme version of this specialization by learning to map a given query to a query-specific feature space transformation. Because this transformation is linear, it can be applied with minimal computational cost to millions of image embeddings, making it effective for large-scale retrieval or re-ranking. Results show that this method consistently outperforms state-of-the-art alternatives, including those that require many orders of magnitude more computation at query time. Eric Xing 0002, Abby Stylianou, Robert Pless, Nathan Jacobs |
NeurIPS | 4 |
| 2025 | TaxaBind: A Unified Embedding Space for Ecological ApplicationsabstractWe present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio, and environmental features, useful for solving eco-logical problems. To learn this joint embedding space, we leverage ground-level images of species as a binding modality. We propose multimodal patching, a technique for effectively distilling the knowledge from various modalities into the binding modality. We construct two large datasets for pretraining: iSatNat with species images and satellite images, and iSoundNat with species images and audio. Additionally, we introduce TaxaBench-8k, a diverse multimodal dataset with six paired modalities for evaluating deep learning models on ecological tasks. Experiments with TaxaBind demonstrate its strong zero-shot and emer-gent capabilities on a range of tasks including species classification, cross-model retrieval, and audio classification. The datasets and models are made available at https://github.com/mvr1/TaxaBind. Srikumar Sastry, Subash Khanal, Aayush Dhakal, Nathan Jacobs |
WACV | 5 |
| 2024 | FroSSL: Frobenius Norm Minimization for Efficient Multiview Self-supervised Learning
Oscar Skean, Aayush Dhakal, Nathan Jacobs, Luis Gonzalo Sánchez Giraldo |
ECCV (89) | 3 |
| 2024 | Learning Interpretable Policies in Hindsight-Observable POMDPs Through Partially Supervised Reinforcement LearningabstractDeep reinforcement learning has demonstrated remarkable achievements across diverse domains such as video games, robotic control, autonomous driving, and drug discovery. Common methodologies in partially observable domains largely lean on end-to-end learning from high-dimensional observations, such as images, without explicitly reasoning about true state. We suggest an alternative direction, introducing the Partially Supervised Reinforcement Learning (PSRL) framework. At the heart of PSRL is the fusion of both supervised and unsupervised learning. The approach leverages a state estimator to distill supervised semantic state information from high-dimensional observations which are often fully observable at training time. This yields more interpretable policies that compose state predictions with control. In parallel, it captures an unsupervised latent representation. These two—the semantic state and the latent state—are then fused and utilized as inputs to a policy network. This juxtaposition offers practitioners a flexible and dynamic spectrum: from emphasizing supervised state information to integrating richer, latent insights. Extensive experimental results indicate that by merging these dual representations, PSRL offers a balance, enhancing interpretability while preserving, and often significantly outperforming, the performance benchmarks set by traditional methods in terms of reward and convergence speed. Michael Lanier, Nathan Jacobs, Chongjie Zhang, Yevgeniy Vorobeychik |
ICMLA | 3 |
| 2024 | GeoBind: Binding Text, Image, and Audio through Satellite ImagesabstractIn remote sensing, we are interested in modeling various modalities for some geographic location. Several works have focused on learning the relationship between a location and type of landscape, habitability, audio, textual descriptions, etc. Recently, a common way to approach these problems is to train a deep-learning model that uses satellite images to infer some unique characteristics of the location. In this work, we present a deep-learning model, GeoBind, that can infer about multiple modalities, specifically text, image, and audio, from satellite imagery of a location. To do this, we use satellite images as the binding element and contrastively align all other modalities to the satellite image data. Our training results in a joint embedding space with multiple types of data: satellite image, ground-level image, audio, and text. Furthermore, our approach does not require a single complex dataset that contains all the modalities mentioned above. Rather it only requires multiple satellite-image paired data. While we only align three modalities in this paper, we present a general framework that can be used to create an embedding space with any number of modalities by using satellite images as the binding element. Our results show that, unlike traditional unimodal models, GeoBind is versatile and can reason about multiple modalities for a given satellite image input. Aayush Dhakal, Subash Khanal, Srikumar Sastry, Nathan Jacobs |
IGARSS | 5 |
| 2024 | Aligning Geo-Tagged Clip Representations and Satellite Imagery for Few-Shot Land Use ClassificationabstractA major difference between ground-level and satellite imagery of landscapes lies in their semantic granularity: ground-level images tend to offer details on objects and human activities, while satellite images provide broader geographic context but, typically, with coarser semantics. This study aims to leverage this complementary information by integrating fine-grained insights from a ground-level view into the analysis of satellite image data. To achieve this integration, we propose to align a satellite image representation with co-located geo-tagged ground-level image CLIP representations. This method focuses on enriching satellite image visual features by leveraging the inherent visual characteristics found in ground-level images as a reference in a contrastive manner, without relying on additional textual information to guide the learning process. We evaluate the quality of the learned representations on the EuroSAT benchmark in various few-shot settings. Pallavi Jain 0004, Diego Marcos, Dino Ienco, Roberto Interdonato, Aayush Dhakal, Nathan Jacobs, Tristan Berchoux |
IGARSS | 6 |
| 2024 | PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape MappingabstractA soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial scales, we represent locations with multi-scale satellite imagery and learn a joint representation among this imagery, audio, and text. To capture the inherent uncertainty in the soundscape of a location, we design the representation space to be probabilistic. We also fuse ubiquitous metadata (including geolocation, time, and data source) to enable learning of spatially and temporally dynamic representations of soundscapes. We demonstrate the utility of our framework by creating large-scale soundscape maps integrating both audio and text with temporal control. To facilitate future research on this task, we also introduce a large-scale dataset, GeoSound, containing over 300k geotagged audio samples paired with both low- and high-resolution satellite imagery. We demonstrate that our method outperforms the existing state-of-the-art on both GeoSound and the existing SoundingEarth dataset. Our dataset and code is available at https://github.com/mvrl/PSM. Subash Khanal, Eric Xing 0002, Srikumar Sastry, Aayush Dhakal, Zhexiao Xiong, Nathan Jacobs |
ACM Multimedia | 7 |
| 2024 | GOMAA-Geo: GOal Modality Agnostic Active Geo-localizationabstractWe consider the task of active geo-localization (AGL) in which an agent uses a sequence of visual cues observed during aerial navigation to find a target specified through multiple possible modalities. This could emulate a UAV involved in a search-and-rescue operation navigating through an area, observing a stream of aerial images as it goes. The AGL task is associated with two important challenges. Firstly, an agent must deal with a goal specification in one of multiple modalities (e.g., through a natural language description) while the search cues are provided in other modalities (aerial imagery). The second challenge is limited localization time (e.g., limited battery life, urgency) so that the goal must be localized as efficiently as possible, i.e. the agent must effectively leverage its sequentially observed aerial views when searching for the goal. To address these challenges, we propose GOMAA-Geo -- a goal modality agnostic active geo-localization agent -- for zero-shot generalization between different goal modalities. Our approach combines cross-modality contrastive learning to align representations across modalities with supervised foundation model pretraining and reinforcement learning to obtain highly effective navigation and localization policies. Through extensive evaluations, we show that GOMAA-Geo outperforms alternative learnable approaches and that it generalizes across datasets -- e.g., to disaster-hit areas without seeing a single disaster scenario during training -- and goal modalities -- e.g., to ground-level imagery or textual descriptions, despite only being trained with goals specified as aerial views. Our code is available at: https://github.com/mvrl/GOMAA-Geo. Anindya Sarkar, Srikumar Sastry, Aleksis Pirinen, Chongjie Zhang, Nathan Jacobs, Yevgeniy Vorobeychik |
NeurIPS | 5 |
| 2024 | WATCH: Wide-Area Terrestrial Change HypercubeabstractMonitoring Earth activity using data collected from multiple satellite imaging platforms in a unified way is a significant challenge, especially with large variability in image resolution, spectral bands, and revisit rates. Further, the availability of sensor data varies across time as new platforms are launched. In this work, we introduce an adaptable framework and network architecture capable of predicting on subsets of the available platforms, bands, or temporal ranges it was trained on. Our system, called WATCH, is highly general and can be applied to a variety of geospatial tasks. In this work, we analyze the performance of WATCH using the recent IARPA SMART public dataset and metrics. We focus primarily on the problem of broad area search for heavy construction sites. Experiments validate the robustness of WATCH during inference to limited sensor availability, as well the the ability to alter inference-time spatial or temporal sampling. WATCH is open source and available for use on this or other remote sensing problems. Code and model weights are available at: https://gitlab.kitware.com/computer-vision/geowatch Connor Greenwell, Jon Crall, Matthew Purri, Kristin J. Dana, Nathan Jacobs, Armin Hadzic, Scott Workman, Matthew J. Leotta |
WACV | 5 |
| 2024 | A Visual Active Search Framework for Geospatial ExplorationabstractMany problems can be viewed as forms of geospatial search aided by aerial imagery, with examples ranging from detecting poaching activity to human trafficking. We model this class of problems in a visual active search (VAS) framework, which has three key inputs: (1) an image of the entire search area, which is subdivided into regions, (2) a local search function, which determines whether a previously unseen object class is present in a given region, and (3) a fixed search budget, which limits the number of times the local search function can be evaluated. The goal is to maximize the number of objects found within the search budget. We propose a reinforcement learning approach for VAS that learns a meta-search policy from a collection of fully annotated search tasks. This meta-search policy is then used to dynamically search for a novel target-object class, leveraging the outcome of any previous queries to determine where to query next. Through extensive experiments on several large-scale satellite imagery datasets, we show that the proposed approach significantly outperforms several strong baselines. We also propose novel domain adaptation techniques that improve the policy at decision time when there is a significant domain gap with the training data. Code is publicly available at this link. Anindya Sarkar, Michael Lanier, Scott Alfeld, Jiarui Feng, Roman Garnett, Nathan Jacobs, Yevgeniy Vorobeychik |
WACV | 6 |
| 2024 | BirdSAT: Cross-View Contrastive Masked Autoencoders for Bird Species Classification and MappingabstractWe propose a metadata-aware self-supervised learning (SSL) framework useful for fine-grained classification and ecological mapping of bird species around the world. Our framework unifies two SSL strategies: Contrastive Learning (CL) and Masked Image Modeling (MIM), while also enriching the embedding space with metadata available with ground-level imagery of birds. We separately train uni-modal and cross-modal ViT on a novel cross-view global bird species dataset containing ground-level imagery, metadata (location, time), and corresponding satellite imagery. We demonstrate that our models learn fine-grained and geographically conditioned features of birds, by evaluating on two downstream tasks: fine-grained visual classification (FGVC) and cross-modal retrieval. Pre-trained models learned using our framework achieve SotA performance on FGVC of iNAT-2021 birds and in transfer learning settings for CUB-200-2011 and NABirds datasets. Moreover, the impressive cross-modal retrieval performance of our model enables the creation of species distribution maps across any geographic region. The dataset and source code will be released at https://github.com/mvrl/BirdSAT. Srikumar Sastry, Subash Khanal, Aayush Dhakal, Nathan Jacobs |
WACV | 5 |
| 2024 | ArcGeo: Localizing Limited Field-of-View Images using Cross-view MatchingabstractCross-view matching techniques for image geo-localization attempt to match features in ground-level query images against a collection of satellite images to determine their positions of origin. We present ArcGeo, a novel cross-view image matching approach which introduces a batch-all angular margin loss and several train-time strategies including large-scale pretraining and FoV-based data augmentation. This allows our model to perform well even in challenging cases with limited field-of-view (FoV). Further, we evaluate multiple model architectures, data augmentation approaches and optimization strategies to train a deep cross-view matching network, specifically optimized for limited FoV cases. In low FoV experiments (FoV = 90°) our method improves top-1 image recall rate on the CVUSA dataset from 30.12% to 43.08%. We also demonstrate improved performance over the state-of-the-art techniques for panoramic cross-view retrieval, improving top-1 recall from 95.43% to 96.06% on the CVUSA dataset and from 64.52% to 79.88% on the CVACT test dataset. Lastly, we evaluate the role of large-scale pretraining for improved robustness. With appropriate pretraining on external data, our model improves top-1 recall dramatically to 66.83% for the FoV = 90° test case on CVUSA, an increase of over twice what is reported by existing approaches. Maxim Shugaev, Ilya Semenov, Kyle Ashley, Michael Klaczynski, Naresh P. Cuntoor, Mun Wai Lee, Nathan Jacobs |
WACV | 7 |
| 2023 | Learning Tri-modal Embeddings for Zero-Shot Soundscape Mapping
Subash Khanal, Srikumar Sastry, Aayush Dhakal, Nathan Jacobs |
BMVC | 4 |
| 2023 | StereoFlowGAN: Co-training for Stereo and Flow with Unsupervised Domain Adaptation
Zhexiao Xiong, Feng Qiao 0001, Yu Zhang 0094, Nathan Jacobs |
BMVC | 4 |
| 2023 | Fine-Grained Property Value Assessment Using Probabilistic DisaggregationabstractThe monetary value of a given piece of real estate, a parcel, is often readily available from a geographic information system. However, for many applications, such as insurance and urban planning, it is useful to have estimates of property value at much higher spatial resolutions. We propose a method to estimate the distribution over property value at the pixel level from remote sensing imagery. We evaluate on a real-world dataset of a major urban area. Our results show that the proposed approaches are capable of generating fine-level estimates of property values, significantly improving upon a diverse collection of baseline approaches. Cohen Archbold, Benjamin Brodie, Aram Ansary Ogholbake, Nathan Jacobs |
IGARSS | 4 |
| 2023 | Task Agnostic Cost Prediction Module for Semantic Labeling in Active LearningabstractWe consider the problem of cost effective active learning for semantic segmentation, which aims at reducing the efforts of semantically annotating images. Current studies have ignored the inclusion of cost of labeling into their active learning frameworks. To this end, we first present a novel cost prediction module based on what we call the M-Net. M-Net combines the power of unsupervised W-Net and supervised U-Net to compute a refined segmentation map. The refined segmentation map is used to estimate the cost of annotations. The cost of annotation is estimated by the number of clicks required to annotate an image. To solve this task, we make use of the harris corner detector algorithm to estimate the location of the clicks required to annotate an image. Finally, we employ a multi armed bandit setting to minimize the cost of annotations while maximizing the performance of the semantic segmentation task. The M-Net outperforms fully supervised U-Net with +4.37 Acc and +3.75 mIoU. The proposed active learning framework also outperforms the existing baselines to prove the relevance of the approach in the current paradigm. Srikumar Sastry, Nathan Jacobs, Mariana Belgiu, Raian Vargas Maretto |
IGARSS | 2 |
| 2023 | CrossAdapt: Cross-Scene Adaptation for Multi-Domain Depth EstimationabstractWe address the task of monocular depth estimation in the multi-domain setting. Given a large dataset (source) with ground-truth depth maps, and a set of unlabeled datasets (targets), our goal is to create a model that works well on unlabeled target datasets across different scenes. This is a challenging problem when there is a significant domain shift, often resulting in poor performance on the target datasets. We propose to address this task with a unified approach that includes adversarial knowledge distillation and uncertainty-guided self-supervised reconstruction. We provide both quantitative and qualitative evaluations on four datasets: KITTI, Virtual KITTI, UAVid China, and UAVid Germany. These datasets contain widely varying viewpoints, including ground-level and overhead perspectives, which is more challenging than is typically considered in prior work on domain adaptation for single-image depth. Our approach significantly improves upon conventional domain adaptation baselines and does not require additional memory as the number of target sets increases. Yu Zhang 0094, Muhammad Usman Rafique, Gordon A. Christie, Nathan Jacobs |
IGARSS | 4 |
| 2023 | Crossseg: Cross-Scene Few-Shot Aerial Segmentation Using Probabilistic PrototypesabstractIn this work, we propose a novel framework called CrossSeg that addresses the task of few-shot semantic segmentation for different aerial imagery. Conventional semantic segmentation approaches struggle to generalize well to unseen object categories, making them a significant limitation for modern intelligent systems, especially those deployed in realistic real-time settings, such as unmanned aerial vehicles (UAVs). CrossSeg overcomes this limitation and generalizes well in a cross-scene setting with only a few labeled samples. Unlike traditional methods that use a set of fixed prototypes for each class, CrossSeg utilizes high-quality probabilistic prototypes that can not only represent different semantic classes but also handle significant variations in different scenes. Experiments show that our approach significantly improves upon conventional few-shot segmentation baselines and does not require extensive tuning. Yu Zhang 0094, Muhammad Usman Rafique, Nathan Jacobs |
IGARSS | 3 |
| 2023 | A Partially-Supervised Reinforcement Learning Framework for Visual Active SearchabstractVisual active search (VAS) has been proposed as a modeling framework in which visual cues are used to guide exploration, with the goal of identifying regions of interest in a large geospatial area. Its potential applications include identifying hot spots of rare wildlife poaching activity, search-and-rescue scenarios, identifying illegal trafficking of weapons, drugs, or people, and many others. State of the art approaches to VAS include applications of deep reinforcement learning (DRL), which yield end-to-end search policies, and traditional active search, which combines predictions with custom algorithmic approaches. While the DRL framework has been shown to greatly outperform traditional active search in such domains, its end-to-end nature does not make full use of supervised information attained either during training, or during actual search, a significant limitation if search tasks differ significantly from those in the training distribution. We propose an approach that combines the strength of both DRL and conventional active search approaches by decomposing the search policy into a prediction module, which produces a geospatial distribution of regions of interest based on task embedding and search history, and a search module, which takes the predictions and search history as input and outputs the search distribution. In addition, we develop a novel meta-learning approach for jointly learning the resulting combined policy that can make effective use of supervised information obtained both at training and decision time. Our extensive experiments demonstrate that the proposed representation and meta-learning frameworks significantly outperform state of the art in visual active search on several problem domains. Anindya Sarkar, Nathan Jacobs, Yevgeniy Vorobeychik |
NeurIPS | 2 |
| 2022 | AssocFormer: Association Transformer for Multi-label Classification
Xin Xing 0002, Yu Zhang 0094, Ai-Ling Lin, Nathan Jacobs |
BMVC | 5 |
| 2022 | Revisiting Near/Remote Sensing with Geospatial AttentionabstractThis work addresses the task of overhead image segmentation when auxiliary ground-level images are available. Recent work has shown that performing joint inference over these two modalities, often called near/remote sensing, can yield significant accuracy improvements. Extending this line of work, we introduce the concept of geospatial attention, a geometry-aware attention mechanism that explicitly considers the geospatial relationship between the pixels in a ground-level image and a geographic location. We propose an approach for computing geospatial attention that incorporates geometric features and the appearance of the overhead and ground-level imagery. We introduce a novel architecture for near/remote sensing that is based on geospatial attention and demonstrate its use for five segmentation tasks. The results demonstrate that our method significantly outperforms the previous state-of-the-art methods. Scott Workman, Muhammad Usman Rafique, Hunter Blanton, Nathan Jacobs |
CVPR | 4 |
| 2022 | Neural Network Decision-Making Criteria Consistency Analysis via Inputs SensitivityabstractNeural networks (NNs) have demonstrated exciting results on various tasks within the last decade. For example, the performance on image classification tasks has been improved dramatically. However, the performance evaluations are often based on a black-box performance, such as accuracy, while insightful analysis of the black-box, such as the prediction formation mechanism, is often missing. Empirically, a NN usually produces a stable overall performance on the same task across multiple training trials when treating it as a black-box. However, when unveiling the black-box, the performance is usually volatile. The decision-making criteria learned by the training trials are often significantly different, which is problematic in many ways. We believe achieving consistent criteria between different training trials is equally important to achieving high performance, if not more. This work, firstly, evaluates the decision-making criteria of NNs via inputs sensitivity using feature-attribution explanation methods in combination with computational analysis and clustering analysis. Through intensive experimentation, we find that decision-making criteria are easily distinguishable between training trials of the same architecture and task, suggesting the criteria learned between training trials are significantly inconsistent. To mitigate this inconsistency, we propose three general training schemes. Our demonstration result shows that the proposed methods effectively reduce the inconsistency of the decision-making criteria learned by different training trials while maintaining the overall performance. Eric Xing 0002, Liangliang Liu 0001, Xin Xing 0002, Yunni Qu, Nathan Jacobs, Gongbo Liang |
ICPR | 5 |
| 2022 | A Structure-Aware Method for Direct Pose EstimationabstractEstimating camera pose from a single image is a fundamental problem in computer vision. Existing methods for solving this task fall into two distinct categories, which we refer to as direct and indirect. Direct methods, such as PoseNet, regress pose from the image as a fixed function, for example using a feed-forward convolutional network. Such methods are desirable because they are deterministic and run in constant time. Indirect methods for pose regression are often non-deterministic, with various external dependencies such as image retrieval and hypothesis sampling. We propose a direct method that takes inspiration from structure-based approaches to incorporate explicit 3D constraints into the network. Our approach maintains the desirable qualities of other direct methods while achieving much lower error in general. Code is available at https://github.com/mvrl/structure-aware-pose-estimation. Hunter Blanton, Scott Workman, Nathan Jacobs |
WACV | 3 |
| 2022 | Content-Aware Detection of Temporal Metadata ManipulationabstractMost pictures shared online are accompanied by temporal metadata (i.e., the day and time they were taken), which makes it possible to associate an image content with real-world events. Maliciously manipulating this metadata can convey a distorted version of reality. In this work, we present the emerging problem of detecting timestamp manipulation. We propose an end-to-end approach to verify whether the purported time of capture of an outdoor image is consistent with its content and geographic location. We consider manipulations done in the hour and/or month of capture of a photograph. The central idea is the use of supervised consistency verification, in which we predict the probability that the image content, capture time, and geographical location are consistent. We also include a pair of auxiliary tasks, which can be used to explain the network decision. Our approach improves upon previous work on a large benchmark dataset, increasing the classification accuracy from 59.0% to 81.1%. We perform an ablation study that highlights the importance of various components of the method, showing what types of tampering are detectable using our approach. Finally, we demonstrate how the proposed method can be employed to estimate a possible time-of-capture in scenarios in which the timestamp is missing from the metadata. Rafael Padilha, Tawfiq Salem, Scott Workman, Fernanda A. Andaló, Anderson Rocha 0001, Nathan Jacobs |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2022 | Contrastive Cross-Modal Pre-Training: A General Strategy for Small Sample Medical ImagingabstractA key challenge in training neural networks for a given medical imaging task is the difficulty of obtaining a sufficient number of manually labeled examples. In contrast, textual imaging reports are often readily available in medical records and contain rich but unstructured interpretations written by experts as part of standard clinical practice. We propose using these textual reports as a form of weak supervision to improve the image interpretation performance of a neural network without requiring additional manually labeled examples. We use an image-text matching task to train a feature extractor and then fine-tune it in a transfer learning setting for a supervised task using a small labeled dataset. The end result is a neural network that automatically interprets imagery without requiring textual reports during inference. We evaluate our method on three classification tasks and find consistent performance improvements, reducing the need for labeled data by 67%-98%. Gongbo Liang, Connor Greenwell, Yu Zhang 0094, Xin Xing 0002, Ramakanth Kavuluru, Nathan Jacobs |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Alzheimer's Disease Classification Using Genetic DataabstractThere has been a recent surge of interest in using genetic data to build ML-based accurate and interpretable disease classification models. In this line of research, we separately assess the potential of the peripheral blood gene expression data as well as the Single Nucleotide Polymorphism (SNP) data in building ML models for AD classification. We present a systematic approach on feature selection and ML model design using both types of genetic data provided by the Alzheimer’s Disease Neuroimaging Initiatives (ADNI). Our two-step feature selection produced a curated list of important genes. In addition to these selected genetic features, to examine the role of non-genetic covariates, we included age and number of education years (EDU) as extra features. In the Control (CN) vs. AD classification, the best performing classifier, XGBoost, trained with gene expression features only and that with extra features included had Area Under Curve (AUC) of 0.64 and 0.65 respectively. However, AUC for the same task using SNP data only and that with extra features included was 0.56 and 0.64 respectively. The just above chance results of classifier trained with SNP features and the improvement when used along with additional covariates indicate low potential of SNP data in AD classification when used alone while also indicating the importance of non-genetic factors associated with AD. Nevertheless, with well above chance performance, gene expression features show great potential especially between groups of AD progression, i.e., CN vs. AD, CN vs. EMCI, EMCI vs. AD and LMCI vs. AD. The source code and manual are available at https://github.com/mvrl/ADNI_Genetics. Subash Khanal, Jin Chen 0004, Nathan Jacobs, Ai-Ling Lin |
BIBM | 3 |
| 2021 | Dynamic Feature Alignment for Semi-supervised Domain Adaptation
Yu Zhang 0094, Gongbo Liang, Nathan Jacobs |
BMVC | 3 |
| 2021 | Hierarchical Probabilistic Embeddings for Multi-View Image ClassificationabstractWe address the task of image classification, when the available spectral bands can vary from image to image. We propose a model that learns to represent uncertainty over latent features in a way that is conditioned on the available bands. We expect that images with fewer bands will generally be more difficult to classify and hence have higher uncertainty. We compare two strategies for training such a model, one which uses explicit hierarchical constraints and one which relies on implicit constraints. We evaluate both using RGB and multispectral imagery from the EuroSat dataset and find that the hierarchical approach improves the compatibility of the resulting distributions without sacrificing accuracy. Benjamin Brodie, Subash Khanal, Muhammad Usman Rafique, Connor Greenwell, Nathan Jacobs |
IGARSS | 5 |
| 2021 | Intensity Harmonization for Airborne LiDARabstractConstructing a point cloud for a large geographic region, such as a state or country, can require multiple years of effort. Often several vendors will be used to acquire LiDAR data, and a single region may be captured by multiple LiDAR scans. A key challenge is maintaining consistency between these scans, which includes point density, number of returns, and intensity. Intensity in particular can be very different between scans, even in areas that are overlapping. Harmonizing the intensity between scans to remove these discrepancies is expensive and time consuming. In this paper, we propose a novel method for point cloud harmonization based on deep neural networks. We evaluate our method quantitatively and qualitatively using a high quality real world LiDAR dataset. We compare our method to several baselines, including standard interpolation methods as well as histogram matching. We show that our method performs as well as the best baseline in areas with similar intensity distributions, and outperforms all baselines in areas with different intensity distributions. Source code is available at https://github.com/mvrl/lidar-harmonization. Nathan Jacobs |
IGARSS | 2 |
| 2021 | Spatio-Temporal Deep Learning Approach to Map Deforestation in Amazon RainforestabstractWe address the task of mapping deforested areas in the Brazilian Amazon. Accurate maps are an important tool for informing effective deforestation containment policies. The main existing approaches to this task are largely manual, requiring significant effort by trained experts. To reduce this effort, we propose a fully automatic approach based on spatio-temporal deep convolutional neural networks. We introduce several domain-specific components, including approaches for: image preprocessing; handling image noise, such as clouds and shadow; and constructing the training data set. We show that our preprocessing protocol reduces the impact of noise in the training data set. Furthermore, we propose two spatio-temporal variations of the U-Net architecture, which make it possible to incorporate both spatial and temporal contexts. Using a large, real-world data set, we show that our method outperforms a traditional U-Net architecture, thus achieving approximately 95% accuracy. Raian Vargas Maretto, Leila M. G. Fonseca, Nathan Jacobs, Thales Sehn Körting, Hugo N. Bendini, Leandro Parente |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Improved Trainable Calibration Method for Neural Networks
Gongbo Liang, Yu Zhang 0094, Nathan Jacobs |
BMVC | 4 |
| 2020 | Generative Appearance Flow: A Hybrid Approach for Outdoor View Synthesis
Muhammad Usman Rafique, Hunter Blanton, Noah Snavely, Nathan Jacobs |
BMVC | 4 |
| 2020 | Learning a Dynamic Map of Visual AppearanceabstractThe appearance of the world varies dramatically not only from place to place but also from hour to hour and month to month. Every day billions of images capture this complex relationship, many of which are associated with precise time and location metadata. We propose to use these images to construct a global-scale, dynamic map of visual appearance attributes. Such a map enables fine-grained understanding of the expected appearance at any geographic location and time. Our approach integrates dense overhead imagery with location and time metadata into a general framework capable of mapping a wide variety of visual attributes. A key feature of our approach is that it requires no manual data annotation. We demonstrate how this approach can support various applications, including image-driven mapping, image geolocalization, and metadata verification. Tawfiq Salem, Scott Workman, Nathan Jacobs |
CVPR | 3 |
| 2020 | Dynamic Traffic Modeling From Overhead ImageryabstractOur goal is to use overhead imagery to understand patterns in traffic flow, for instance answering questions such as how fast could you traverse Times Square at 3am on a Sunday. A traditional approach for solving this problem would be to model the speed of each road segment as a function of time. However, this strategy is limited in that a significant amount of data must first be collected before a model can be used and it fails to generalize to new areas. Instead, we propose an automatic approach for generating dynamic maps of traffic speeds using convolutional neural networks. Our method operates on overhead imagery, is conditioned on location and time, and outputs a local motion model that captures likely directions of travel and corresponding travel speeds. To train our model, we take advantage of historical traffic data collected from New York City. Experimental results demonstrate that our method can be applied to generate accurate city-scale traffic models. Scott Workman, Nathan Jacobs |
CVPR | 2 |
| 2020 | Multi-Branch Attention Networks for Classifying Galaxy ClustersabstractThis paper addresses the task of classifying galaxy clusters, which are the largest known objects in the Universe. Galaxy clusters can be categorized as cool-core (CC), weak-cool-core (WCC), and non-cool-core (NCC), depending on their central cooling times. Traditional classification approaches used in astrophysics are inaccurate and rely on measuring surface brightness concentrations or central gas densities. In this work, we propose a multi-branch attention network that uses spatial attention to classify a given cluster. To evaluate our network, we use a database of simulated X-ray emissivity images, which contains 954 projections of 318 clusters. Experimental results show that our network outperforms several strong baseline methods and achieves a macro-averaged F1 score of 0.83. We highlight the value of our proposed spatial attention module through an ablation study. Yu Zhang 0094, Gongbo Liang, Yuanyuan Su, Nathan Jacobs |
ICPR | 4 |
| 2020 | Surface Modeling for Airborne LidarabstractRepeat-visit airborne lidar is a powerful tool for change detection in urban and rural environments. In this work, we present a learning-based approach that addresses one of the key challenges in comparing point cloud scans of the same region: handling geometric differences caused by varying sensor position. Our approach is to perform shape modeling through ray casting with a point cloud neural network. Recent work on learning-based shape modeling has been based on the assumption that an explicit surface representation is available, which is not the case for airborne lidar datasets. Our key insight is that by using a ray casting approach we can perform shape modeling directly with lidar measurements. We evaluate our method both quantitatively and qualitatively on learned surface accuracy and show that our method correctly predicts surface intersection even in sparse regions of the input cloud. Hunter Blanton, Sean Grate, Nathan Jacobs |
IGARSS | 3 |
| 2020 | Estimating Displaced Populations from OverheadabstractWe introduce a deep learning approach to perform fine-grained population estimation for displacement camps using high-resolution overhead imagery. We train and evaluate our approach on drone imagery cross-referenced with population data for refugee camps in Cox's Bazar, Bangladesh in 2018 and 2019. Our proposed approach achieves 7.41% mean absolute percent error on sequestered camp imagery. We believe our experiments with real-world displacement camp data constitute an important step towards the development of tools that enable the humanitarian community to effectively and rapidly respond to the global displacement crisis. Armin Hadzic, Gordon A. Christie, Jeffrey Freeman, Amber Dismer, Stevan Bullard, Ashley Greiner, Nathan Jacobs, Ryan Mukherjee |
IGARSS | 7 |
| 2020 | Single Image Cloud Detection via Multi-Image FusionabstractArtifacts in imagery captured by remote sensing, such as clouds, snow, and shadows, present challenges for various tasks, including semantic segmentation and object detection. A primary challenge in developing algorithms for identifying such artifacts is the cost of collecting annotated training data. In this work, we explore how recent advances in multi-image fusion can be leveraged to bootstrap single image cloud detection. We demonstrate that a network optimized to estimate image quality also implicitly learns to detect clouds. To support the training and evaluation of our approach, we collect a large dataset of Sentinel-2 images along with a per-pixel semantic labelling for land cover. Through various experiments, we demonstrate that our method reduces the need for annotated training data and improves cloud detection performance. Scott Workman, Muhammad Usman Rafique, Hunter Blanton, Connor Greenwell, Nathan Jacobs |
IGARSS | 5 |
| 2019 | Joint 2D-3D Breast Cancer ClassificationabstractBreast cancer is the malignant tumor that causes the highest number of cancer deaths in females. Digital mammograms (DM or 2D mammogram) and digital breast tomosynthesis (DBT or 3D mammogram) are the two types of mammography imagery that are used in clinical practice for breast cancer detection and diagnosis. Radiologists usually read both imaging modalities in combination; however, existing computer-aided diagnosis tools are designed using only one imaging modality. Inspired by clinical practice, we propose an innovative convolutional neural network (CNN) architecture for breast cancer classification, which uses both 2D and 3D mammograms, simultaneously. Our experiment shows that the proposed method significantly improves the performance of breast cancer classification. By assembling three CNN classifiers, the proposed model achieves 0.97 AUC, which is 34.72% higher than the methods using only one imaging modality. Gongbo Liang, Yu Zhang 0094, Xin Xing 0002, Hunter Blanton, Tawfiq Salem, Nathan Jacobs |
BIBM | 7 |
| 2019 | 2D Convolutional Neural Networks for 3D Digital Breast Tomosynthesis ClassificationabstractAutomated methods for breast cancer detection have focused on 2D mammography and have largely ignored 3D digital breast tomosynthesis (DBT), which is frequently used in clinical practice. The two key challenges in developing automated methods for DBT classification are handling the variable number of slices and retaining slice-to-slice changes. We propose a novel deep 2D convolutional neural network (CNN) architecture for DBT classification that simultaneously overcomes both challenges. Our approach operates on the full volume, regardless of the number of slices, and allows the use of pre-trained 2D CNNs for feature extraction, which is important given the limited amount of annotated training data. In an extensive evaluation on a real-world clinical dataset, our approach achieves 0.854 auROC, which is 28.80% higher than approaches based on 3D CNNs. We also find that these improvements are stable across a range of model configurations. Yu Zhang 0094, Hunter Blanton, Gongbo Liang, Xin Xing 0002, Nathan Jacobs |
BIBM | 6 |
| 2019 | Defense-PointNet: Protecting PointNet Against Adversarial AttacksabstractDespite remarkable performance across a broad range of tasks, neural networks have been shown to be vulnerable to adversarial attacks. Many works focus on adversarial attacks and defenses on 2D images, but few focus on 3D point clouds. In this paper, our goal is to enhance the adversarial robustness of PointNet, which is one of the most widely used models for 3D point clouds. We apply the fast gradient sign attack method (FGSM) on 3D point clouds and find that FGSM can be used to generate not only adversarial images but also adversarial point clouds. To minimize the vulnerability of PointNet to adversarial attacks, we propose Defense-PointNet. We compare our model with two baseline approaches and show that Defense-PointNet significantly improves the robustness of the network against adversarial samples. Yu Zhang 0094, Gongbo Liang, Tawfiq Salem, Nathan Jacobs |
IEEE BigData | 4 |
| 2019 | Weakly Supervised Building Segmentation from Aerial ImagesabstractWe propose a novel framework for weakly supervised semantic segmentation from aerial images. Instead of requiring labels for every pixel, our method only requires a bounding box for each building and leverages domain information to translate these into pixel-level predictions. We convert the bounding boxes into probabilistic masks, each represented using a bivariate Gaussian distribution. We propose a loss function that encompasses our domain knowledge that the bounding box is an upper bound for the object it contains. Combining these two elements significantly improves over many baseline methods. We show extensive results on a recent, large-scale dataset prepared by the United Nations Global Pulse and compare with several baselines. Muhammad Usman Rafique, Nathan Jacobs |
IGARSS | 2 |
| 2019 | Learning to Map Nearly AnythingabstractLooking at the world from above, it is possible to estimate many properties of a given location, including the type of land cover and the expected land use. Historically, such tasks have relied on relatively coarse-grained categories due to the difficulty of obtaining fine-grained annotations. In this work, we propose an easily extensible approach that makes it possible to estimate fine-grained properties from overhead imagery. In particular, we propose a cross-modal distillation strategy to learn to predict the distribution of fine-grained properties from overhead imagery, without requiring any manual annotation of overhead imagery. We show that our learned models can be used directly for applications in mapping and image localization. Tawfiq Salem, Connor Greenwell, Hunter Blanton, Nathan Jacobs |
IGARSS | 4 |
| 2019 | Remote Estimation of Free-Flow SpeedsabstractWe propose an automated method to estimate a road segment's free-flow speed from overhead imagery and road meta-data. The free-flow speed of a road segment is the average observed vehicle speed in ideal conditions, without congestion or adverse weather. Standard practice for estimating free-flow speeds depends on several road attributes, including grade, curve, and width of the right of way. Unfortunately, many of these fine-grained labels are not always readily available and are costly to manually annotate. To compensate, our model uses a small, easy to obtain subset of road features along with aerial imagery to directly estimate free-flow speed with a deep convolutional neural network (CNN). We evaluate our approach on a large dataset, and demonstrate that using imagery alone performs nearly as well as the road features and that the combination of imagery with road features leads to the highest accuracy. Weilian Song, Tawfiq Salem, Hunter Blanton, Nathan Jacobs |
IGARSS | 4 |
| 2019 | A Generative Model of Worldwide Facial AppearanceabstractHuman appearance depends on many proximate factors, including age, gender, ethnicity, and personal style choices. In this work, we model the relationship between human appearance and geographic location, which can impact these factors in complex ways. We propose GPS2Face, a dual-component generative network architecture that enables flexible facial generation with fine-grained control of latent factors. We use facial landmarks as a guide to synthesize likely faces for locations around in the world. We train our model on a large-scale dataset of geotagged faces and evaluate our proposed model, both qualitatively and quantitatively, against previous work. Zachary Bessinger, Nathan Jacobs |
WACV | 2 |
| 2019 | Motion and appearance based background subtraction for freely moving cameras
Hasan Sajid, Sen-Ching S. Cheung, Nathan Jacobs |
Signal Process. Image Commun. | 3 |
| 2018 | Automatic Hand Skeletal Shape Estimation from Radiographs
Radu Paul Mihail, Nathan Jacobs |
BIBM | 2 |
| 2018 | Learning Geo-Temporal Image Features
Menghua Zhai, Tawfiq Salem, Connor Greenwell, Scott Workman, Robert Pless, Nathan Jacobs |
BMVC | 6 |
| 2018 | Learning to Look around Objects for Top-View Representations of Outdoor Scenes
Samuel Schulter, Menghua Zhai, Nathan Jacobs, Manmohan Krishna Chandraker |
ECCV (15) | 3 |
| 2018 | A weakly supervised approach for estimating spatial density functions from high-resolution satellite imageryabstractWe propose a neural network component, the regional aggregation layer, that makes it possible to train a pixel-level density estimator using only coarse-grained density aggregates, which reflect the number of objects in an image region. Our approach is simple to use and does not require domain-specific assumptions about the nature of the density function. We evaluate our approach on several synthetic datasets. In addition, we use this approach to learn to estimate high-resolution population and housing density from satellite imagery. In all cases, we find that our approach results in better density estimates than a commonly used baseline. We also show how our housing density estimator can be used to classify buildings as residential or non-residential. Nathan Jacobs, Adam Kraft, Muhammad Usman Rafique, Ranti Dev Sharma |
SIGSPATIAL/GIS | 1 |
| 2018 | What Goes Where: Predicting Object Distributions from AboveabstractIn this work, we propose a cross-view learning approach, in which images captured from a ground-level view are used as weakly supervised annotations for interpreting overhead imagery. The outcome is a convolutional neural network for overhead imagery that is capable of predicting the type and count of objects that are likely to be seen from a ground-level perspective. We demonstrate our approach on a large dataset of geotagged ground-level and overhead imagery and find that our network captures semantically meaningful features, despite being trained without manual annotations. Connor Greenwell, Scott Workman, Nathan Jacobs |
IGARSS | 3 |
| 2018 | A Multimodal Approach to Mapping SoundscapesabstractWe explore the problem of mapping soundscapes, that is, predicting the types of sounds that are likely to be heard at a given geographic location. Using a novel dataset, which includes geo-tagged audio and overhead imagery, we develop an approach for constructing an aural atlas, which captures the geospatial distribution of soundscapes. We build on previous work relating sound to ground-level imagery but incorporate overhead imagery to overcome the limitations of sparsely distributed geo-tagged audio. In the end, all that we require to construct an aural atlas is overhead imagery of the region of interest. We show examples of aural atlases at multiple spatial scales, from block-level to country. Tawfiq Salem, Menghua Zhai, Scott Workman, Nathan Jacobs |
IGARSS | 4 |
| 2018 | FARSA: Fully Automated Roadway Safety AssessmentabstractThis paper addresses the task of road safety assessment. An emerging approach for conducting such assessments in the United States is through the US Road Assessment Program (usRAP), which rates roads from highest risk (1 star) to lowest (5 stars). Obtaining these ratings requires manual, fine-grained labeling of roadway features in streetlevel panoramas, a slow and costly process. We propose to automate this process using a deep convolutional neural network that directly estimates the star rating from a street-level panorama, requiring milliseconds per image at test time. Our network also estimates many other roadlevel attributes, including curvature, roadside hazards, and the type of median. To support this, we incorporate taskspecific attention layers so the network can focus on the panorama regions that are most useful for a particular task. We evaluated our approach on a large dataset of real-world images from two US states. We found that incorporating additional tasks, and using a semi-supervised training approach, significantly reduced overfitting problems, allowed us to optimize more layers of the network, and resulted in higher accuracy. Weilian Song, Scott Workman, Armin Hadzic, Eric Green, Reginald R. Souleyrette, Nathan Jacobs |
WACV | 8 |
| 2017 | Whole mammogram image classification with convolutional neural networksabstractDue to the high variability in tumor morphology and the low signal-to-noise ratio inherent to mammography, manual classification of mammogram yields a significant number of patients being called back, and subsequent large number of biopsies performed to reduce the risk of missing cancer. The convolutional neural network (CNN) is a popular deep-learning construct used in image classification. This technique has achieved significant advancements in large-set image-classification challenges in recent years. In this study, we had obtained over 3000 high-quality original mammograms with approval from an institutional review board at the University of Kentucky. Different classifiers based on CNNs were built, and each classifier was evaluated based on its performance relative to truth values generated by histology results from biopsy and two-year negative mammogram follow-up confirmed by expert radiologists. Our results showed that CNN model we had built and optimized via data augmentation and transfer learning have a great potential for automatic breast cancer detection using mammograms. Yi Zhang 0078, Erik Y. Han, Nathan Jacobs, Qiong Han, Jinze Liu |
BIBM | 4 |
| 2017 | Predicting Ground-Level Scene Layout from Aerial Imagery
Menghua Zhai, Zachary Bessinger, Scott Workman, Nathan Jacobs |
CVPR | 4 |
| 2017 | Revisiting IM2GPS in the Deep Learning EraabstractImage geolocalization, inferring the geographic location of an image, is a challenging computer vision problem with many potential applications. The recent state-of-the-art approach to this problem is a deep image classification approach in which the world is spatially divided into cells and a deep network is trained to predict the correct cell for a given image. We propose to combine this approach with the original Im2GPS approach in which a query image is matched against a database of geotagged images and the location is inferred from the retrieved set. We estimate the geographic location of a query image by applying kernel density estimation to the locations of its nearest neighbors in the reference database. Interestingly, we find that the best features for our retrieval task are derived from networks trained with classification loss even though we do not use a classification approach at test time. Training with classification loss outperforms several deep feature learning methods (e.g. Siamese networks with contrastive of triplet loss) more typical for retrieval applications. Our simple approach achieves state-of-the-art geolocalization accuracy while also requiring significantly less training data. Nam N. Vo, Nathan Jacobs, James Hays |
ICCV | 2 |
| 2017 | Understanding and Mapping Natural BeautyabstractWhile natural beauty is often considered a subjective property of images, in this paper, we take an objective approach and provide methods for quantifying and predicting the scenicness of an image. Using a dataset containing hundreds of thousands of outdoor images captured throughout Great Britain with crowdsourced ratings of natural beauty, we propose an approach to predict scenicness which explicitly accounts for the variance of human ratings. We demonstrate that quantitative measures of scenicness can benefit semantic image understanding, content-aware image processing, and a novel application of cross-view mapping, where the sparsity of ground-level images can be addressed by incorporating unlabeled overhead images in the training and prediction steps. For each application, our methods for scenicness prediction result in quantitative and qualitative improvements over baseline approaches. Scott Workman, Richard Souvenir, Nathan Jacobs |
ICCV | 3 |
| 2017 | A Unified Model for Near and Remote SensingabstractWe propose a novel convolutional neural network architecture for estimating geospatial functions such as population density, land cover, or land use. In our approach, we combine overhead and ground-level images in an end-toend trainable neural network, which uses kernel regression and density estimation to convert features extracted from the ground-level images into a dense feature map. The output of this network is a dense estimate of the geospatial function in the form of a pixel-level labeling of the overhead image. To evaluate our approach, we created a large dataset of overhead and ground-level images from a major urban area with three sets of labels: land use, building function, and building age. We find that our approach is more accurate for all tasks, in some cases dramatically so. Scott Workman, Menghua Zhai, David Crandall, Nathan Jacobs |
ICCV | 4 |
| 2016 | Horizon Lines in the Wild
Scott Workman, Menghua Zhai, Nathan Jacobs |
BMVC | 3 |
| 2016 | Detecting Vanishing Points Using Global Image Context in a Non-ManhattanWorldabstractWe propose a novel method for detecting horizontal vanishing points and the zenith vanishing point in man-made environments. The dominant trend in existing methods is to first find candidate vanishing points, then remove outliers by enforcing mutual orthogonality. Our method reverses this process: we propose a set of horizon line candidates and score each based on the vanishing points it contains. A key element of our approach is the use of global image context, extracted with a deep convolutional network, to constrain the set of candidates under consideration. Our method does not make a Manhattan-world assumption and can operate effectively on scenes with only a single horizontal vanishing point. We evaluate our approach on three benchmark datasets and achieve state-of the-art performance on each. In addition, our approach is significantly faster than the previous best method. Menghua Zhai, Scott Workman, Nathan Jacobs |
CVPR | 3 |
| 2016 | Who goes there?: approaches to mapping facial appearance diversityabstractGeotagged imagery, from satellite, aerial, and ground-level cameras, provides a rich record of how the appearance of scenes and objects differ across the globe. Modern web- based mapping software makes it easy to see how different places around the world look, both from satellite and ground-level views. Unfortunately, interfaces for exploring how the appearance of objects depend on geographic location are quite limited. In this work, we focus on a particularly common object, the human face, and propose learning generative models that relate facial appearance and geographic location. We train these models using a novel dataset of geotagged face imagery we constructed for this task. We present qualitative and quantitative results that demonstrate that these models capture meaningful trends in appearance. We also describe a framework for constructing a web-based visualization that captures the geospatial distribution of human facial appearance. Zachary Bessinger, Chris Stauffer, Nathan Jacobs |
SIGSPATIAL/GIS | 3 |
| 2016 | Quantifying curb appealabstractThe curb appeal of a home, which refers to how attractive it is when viewed from the street, is an important decisionmaking factor for many home buyers. Existing models for automatically estimating the price of a home ignore this factor, instead focusing exclusively on objective attributes, such as number of bedrooms, the square footage, and the age. We propose to use street-level imagery of a home, in addition to the objective attributes, to estimate the price of the home, thereby quantifying curb appeal. Our method uses deep convolutional neural networks to extract informative image features. We introduce a large dataset to support an extensive evaluation of several approaches. We find that using images and objective attributes together results in more accurate home price estimates than using either in isolation. We also find that representations learned for scene classification tasks are more discriminative for home price estimation than those learned for other tasks. Zachary Bessinger, Nathan Jacobs |
ICIP | 2 |
| 2016 | Camera geo-calibration using an MCMC approachabstractWe address the problem of single-image geo-calibration, in which an estimate of the geographic location, viewing direction and field of view is sought for the camera that captured an image. The dominant approach to this problem is to match features of the query image, using color and texture, against a reference database of nearby ground imagery. However, this fails when such imagery is not available. We propose to overcome this limitation by matching against a geographic database that contains the locations of known objects, such as houses, roads and bodies of water. Since we are unable to find one-to-one correspondences between image locations and objects in our database, we model the problem probabilistically based on the geometric configuration of multiple such weak correspondences. We propose a Markov Chain Monte Carlo (MCMC) sampling approach to approximate the underlying probability distribution over the full geo-calibration of the camera. Menghua Zhai, Scott Workman, Nathan Jacobs |
ICIP | 3 |
| 2016 | A fast method for estimating transient scene attributesabstractWe propose the use of deep convolutional neural networks to estimate the transient attributes of a scene from a single image. Transient scene attributes describe both the objective conditions, such as the weather, time of day, and the season, and subjective properties of a scene, such as whether or not the scene seems busy. Recently, convolutional neural networks have been used to achieve state-of-the-art results for many vision problems, from object detection to scene classification, but have not previously been used for estimating transient attributes. We compare several methods for adapting an existing network architecture and present state-of-the-art results on two benchmark datasets. Our method is more accurate and significantly faster than previous methods, enabling real-world applications. Ryan Baltenberger, Menghua Zhai, Connor Greenwell, Scott Workman, Nathan Jacobs |
WACV | 5 |
| 2016 | Sky segmentation in the wild: An empirical studyabstractAutomatically determining which pixels in an image view the sky, the problem of sky segmentation, is a critical preprocessing step for a wide variety of outdoor image interpretation problems, including horizon estimation, robot navigation and image geolocalization. Many methods for this problem have been proposed with recent work achieving significant improvements on benchmark datasets. However, such datasets are often constructed to contain images captured in favorable conditions and, therefore, do not reflect the broad range of conditions with which a real-world vision system must cope. This paper presents the results of a large-scale empirical evaluation of the performance of three state-of-the-art approaches on a new dataset, which consists of roughly 100k images captured "in the wild". The results show that the performance of these methods can be dramatically degraded by the local lighting and weather conditions. We propose a deep learning based variant of an ensemble solution that outperforms the methods we tested, in some cases achieving above 50% relative reduction in misclassified pixels. While our results show there is room for improvement, our hope is that this dataset will encourage others to improve the real-world performance of their algorithms. Radu Paul Mihail, Scott Workman, Zachary Bessinger, Nathan Jacobs |
WACV | 4 |
| 2016 | Analyzing human appearance as a cue for dating imagesabstractGiven an image, we propose to use the appearance of people in the scene to estimate when the picture was taken. There are a wide variety of cues that can be used to address this problem. Most previous work has focused on low-level image features, such as color and vignetting. Recent work on image dating has used more semantic cues, such as the appearance of automobiles and buildings. We extend this line of research by focusing on human appearance. Our approach, based on a deep convolutional neural network, allows us to more deeply explore the relationship between human appearance and time. We find that clothing, hair styles, and glasses can all be informative features. To support our analysis, we have collected a new dataset containing images of people from many high school yearbooks, covering the years 1912-2014. While not a complete solution to the problem of image dating, our results show that human appearance is strongly related to time and that semantic information can be a useful cue. Tawfiq Salem, Scott Workman, Menghua Zhai, Nathan Jacobs |
WACV | 4 |
| 2016 | Cloudmaps from static ground-view video
Nathan Jacobs, Scott Workman, Richard Souvenir |
Image Vis. Comput. | 1 |
| 2016 | Appearance based background subtraction for PTZ cameras
Hasan Sajid, Sen-Ching S. Cheung, Nathan Jacobs |
Signal Process. Image Commun. | 3 |
| 2015 | Building Dynamic Cloud Maps from the Ground UpabstractSatellite imagery of cloud cover is extremely important for understanding and predicting weather. We demonstrate how this imagery can be constructed "from the ground up" without requiring expensive geo-stationary satellites. This is accomplished through a novel approach to approximate continental-scale cloud maps using only ground-level imagery from publicly-available webcams. We collected a year's worth of satellite data and simultaneously-captured, geo-located outdoor webcam images from 4388 sparsely distributed cameras across the continental USA. The satellite data is used to train a dynamic model of cloud motion alongside 4388 regression models (one for each camera) to relate ground-level webcam data to the satellite data at the camera's location. This novel application of large-scale computer vision to meteorology and remote sensing is enabled by a smoothed, hierarchically-regularized dynamic texture model whose system dynamics are driven to remain consistent with measurements from the geo-located webcams. We show that our hierarchical model is better able to incorporate sparse webcam measurements resulting in more accurate cloud maps in comparison to a standard dynamic textures implementation. Finally, we demonstrate that our model can be successfully applied to other natural image sequences from the DynTex database, suggesting a broader applicability of our method. Calvin Murdock, Nathan Jacobs, Robert Pless |
ICCV | 2 |
| 2015 | Wide-Area Image Geolocalization with Aerial Reference ImageryabstractWe propose to use deep convolutional neural networks to address the problem of cross-view image geolocalization, in which the geolocation of a ground-level query image is estimated by matching to georeferenced aerial images. We use state-of-the-art feature representations for ground-level images and introduce a cross-view training approach for learning a joint semantic feature representation for aerial images. We also propose a network architecture that fuses features extracted from aerial images at multiple spatial scales. To support training these networks, we introduce a massive database that contains pairs of aerial and ground-level images from across the United States. Our methods significantly out-perform the state of the art on two benchmark datasets. We also show, qualitatively, that the proposed feature representations are discriminative at both local and continental spatial scales. Scott Workman, Richard Souvenir, Nathan Jacobs |
ICCV | 3 |
| 2015 | FACE2GPS: Estimating geographic location from facial featuresabstractThe facial appearance of a person is a product of many factors, including their gender, age, and ethnicity. Methods for estimating these latent factors directly from an image of a face have been extensively studied for decades. We extend this line of work to include estimating the location where the image was taken. We propose a deep network architecture for making such predictions and demonstrate its superiority to other approaches in an extensive set of quantitative experiments on the GeoFaces dataset. Our experiments show that in 26% of the cases the ground truth location is the topmost prediction, and if we allow ourselves to consider the top five predictions, the accuracy increases to 47%. In both cases, the deep learning based approach significantly outperforms random chance as well as another baseline method. Mohammad T. Islam 0001, Scott Workman, Nathan Jacobs |
ICIP | 3 |
| 2015 | DEEPFOCAL: A method for direct focal length estimationabstractEstimating the focal length of an image is an important preprocessing step for many applications. Despite this, existing methods for single-view focal length estimation are limited in that they require particular geometric calibration objects, such as orthogonal vanishing points, co-planar circles, or a calibration grid, to occur in the field of view. In this work, we explore the application of a deep convolutional neural network, trained on natural images obtained from Internet photo collections, to directly estimate the focal length using only raw pixel intensities as input features. We present quantitative results that demonstrate the ability of our technique to estimate the focal length with comparisons against several baseline methods, including an automatic method which uses orthogonal vanishing points. Scott Workman, Connor Greenwell, Menghua Zhai, Ryan Baltenberger, Nathan Jacobs |
ICIP | 5 |
| 2015 | Scene shape estimation from multiple partly cloudy days
Scott Workman, Richard Souvenir, Nathan Jacobs |
Comput. Vis. Image Underst. | 3 |
| 2014 | A Pot of Gold: Rainbows as a Calibration Cue
Scott Workman, Radu Paul Mihail, Nathan Jacobs |
ECCV (5) | 3 |
| 2014 | MPCA: EM-based PCA for mixed-size image datasetsabstractPrincipal component analysis (PCA) is a widely used technique for dimensionality reduction which assumes that the input data can be represented as a collection of fixed-length vectors. Many real-world datasets, such as those constructed from Internet photo collections, do not satisfy this assumption. A natural approach to addressing this problem is to first coerce all input data to a fixed size, and then use standard PCA techniques. This approach is problematic because it either introduces artifacts when we must upsample an image, or loses information when we must downsample an image. We propose MPCA, an approach for estimating the PCA decomposition from multi-sized input data which avoids this initial resizing step. We demonstrate the effectiveness of this approach on simulated and real-world datasets. Feiyu Shi, Menghua Zhai, Drew Duncan, Nathan Jacobs |
ICIP | 4 |
| 2014 | Covariance-Based PCA for Multi-size DataabstractPrincipal component analysis (PCA) is used in diverse settings for dimensionality reduction. If data elements are all the same size, there are many approaches to estimating the PCA decomposition of the dataset. However, many datasets contain elements of different sizes that must be coerced into a fixed size before analysis. Such approaches introduce errors into the resulting PCA decomposition. We introduce CO-MPCA, a nonlinear method of directly estimating the PCA decomposition from datasets with elements of different sizes. We compare our method with two baseline approaches on three datasets: a synthetic vector dataset, a synthetic image dataset, and a real dataset of color histograms extracted from surveillance video. We provide quantitative and qualitative evidence that using CO-MPCA gives a more accurate estimate of the PCA basis. Menghua Zhai, Feiyu Shi, Drew Duncan, Nathan Jacobs |
ICPR | 4 |
| 2014 | Exploring the geo-dependence of human face appearanceabstractThe expected appearance of a human face depends strongly on age, ethnicity and gender. While these relationships are well-studied, our work explores the little-studied dependence of facial appearance on geographic location. To support this effort, we constructed GeoFaces, a large dataset of geotagged face images. We examine the geo-dependence of Eigenfaces and use two supervised methods for extracting geo-informative features. The first, canonical correlation analysis, is used to find location-dependent component images as well as the spatial direction of most significant face appearance change. The second, linear discriminant analysis, is used to find countries with relatively homogeneous, yet distinctive, facial appearance. Mohammad T. Islam 0001, Scott Workman, Hui Wu 0006, Nathan Jacobs, Richard Souvenir |
WACV | 4 |
| 2014 | Estimating cloudmaps from outdoor image sequencesabstractCloud shadows dramatically affect the appearance of outdoor scenes. We describe two approaches that use video of cloud shadows to estimate a cloudmap, a spatio-temporal function that represents the clouds passing over the scene. Our first method makes strong assumptions about the camera geometry and estimates the cloud motion direction. Our second method uses techniques from manifold learning and does not require known geometry. Neither method requires directly viewing the clouds, but instead uses the pattern of intensity changes caused by the cloud shadows. We show renderings of cloudmaps extracted using both methods from videos of real outdoor scenes as well as quantitative results on synthetic datasets. An accurate estimate of the cloudmap has potential applications in surveillance and graphics, as well as scientific studies that depend on solar radiation. Nathan Jacobs, Joshua King, Daniel Bowers, Richard Souvenir |
WACV | 1 |
| 2014 | A CRF approach to fitting a generalized hand skeleton modelabstractWe present a new point distribution model capable of modeling joint subluxation (shifting) in rheumatoid arthritis (RA) patients and an approach to fitting this model to posteroanterior view hand radiographs. We formulate this shape fitting problem as inference in a conditional random field. This model combines potential functions that focus on specific anatomical structures and a learned shape prior. We evaluate our approach on two datasets: one containing relatively healthy hands and one containing hands of rheumatoid arthritis patients. We provide an empirical analysis of the relative value of different potential functions. We also show how to use the fitted hand skeleton to initialize a process for automatically estimating bone contours, which is a challenging, but important, problem in RA disease progression assessment. Radu Paul Mihail, Gustav Blomquist, Nathan Jacobs |
WACV | 3 |
| 2013 | Cloud Motion as a Calibration CueabstractWe propose cloud motion as a natural scene cue that enables geometric calibration of static outdoor cameras. This work introduces several new methods that use observations of an outdoor scene over days and weeks to estimate radial distortion, focal length and geo-orientation. Cloud-based cues provide strong constraints and are an important alternative to methods that require specific forms of static scene geometry or clear sky conditions. Our method makes simple assumptions about cloud motion and builds upon previous work on motion-based and line-based calibration. We show results on real scenes that highlight the effectiveness of our proposed methods. Nathan Jacobs, Mohammad T. Islam 0001, Scott Workman |
CVPR | 1 |
| 2013 | Webcam2Satellite: Estimating cloud maps from webcam imageryabstractWe consider the problem of estimating the current satellite cloud map from a collection of broadly distributed, ground-based webcams. The approach uses historical, geo-referenced satellite imagery to learn a mapping between the satellite image and the ground imagery. We explore representational choices for inferring the cloud status based on the ground-level imagery and consider several alternatives for spatially interpolating these sparse measurements to give a complete map. Proof of concept results show that this gives plausible estimates of satellite imagery. Calvin Murdock, Nathan Jacobs, Robert Pless |
WACV | 2 |
| 2013 | Two Cloud-Based Cues for Estimating Scene Structure and Camera CalibrationabstractWe describe algorithms that use cloud shadows as a form of stochastically structured light to support 3D scene geometry estimation. Taking video captured from a static outdoor camera as input, we use the relationship of the time series of intensity values between pairs of pixels as the primary input to our algorithms. We describe two cues that relate the 3D distance between a pair of points to the pair of intensity time series. The first cue results from the fact that two pixels that are nearby in the world are more likely to be under a cloud at the same time than two distant points. We describe methods for using this cue to estimate focal length and scene structure. The second cue is based on the motion of cloud shadows across the scene; this cue results in a set of linear constraints on scene structure. These constraints have an inherent ambiguity, which we show how to overcome by combining the cloud motion cue with the spatial cue. We evaluate our method on several time lapses of real outdoor scenes. Nathan Jacobs, Austin Abrams, Robert Pless |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | LOST: Longterm Observation of Scenes (with Tracks)abstractWe introduce the Longterm Observation of Scenes (with Tracks) dataset. This dataset comprises videos taken from streaming outdoor webcams, capturing the same half hour, each day, for over a year. LOST contains rich metadata, including geolocation, day-by-day weather annotation, object detections, and tracking results. We believe that sharing this dataset opens opportunities for computer vision research involving very long-term outdoor surveillance, robust anomaly detection, and scene analysis methods based on trajectories. Efficient analysis of changes in behavior in a scene at very long time scale requires features that summarize large amounts of trajectory data in an economical way. We describe a trajectory clustering algorithm and aggregate statistics about these exemplars through time and show that these statistics exhibit strong correlations with external meta-data, such as weather signals and day of the week. Austin Abrams, Jim Tucek, Joshua Little, Nathan Jacobs, Robert Pless |
WACV | 4 |
| 2011 | On analyzing video with very small motionsabstractWe characterize a class of videos consisting of very small but potentially complicated motions. We find that in these scenes, linear appearance variations have a direct relationship to scene motions. We show how to interpret appearance variations captured through a PCA decomposition of the image set as a scene-specific non-parametric motion basis. We propose fast, robust tools for dense flow estimates that are effective in scenes with small motions and potentially large image noise. We show example results in a variety of applications, including motion segmentation and long-term point tracking. Michael Dixon, Austin Abrams, Nathan Jacobs, Robert Pless |
CVPR | 3 |
| 2011 | Webcam geo-localization using aggregate light levelsabstractWe consider the problem of geo-locating static cameras from long-term time-lapse imagery. This problem has received significant attention recently, with most methods making strong assumptions on the geometric structure of the scene. We explore a simple, robust cue that relates overall image intensity to the zenith angle of the sun (which need not be visible). We characterize the accuracy of geolocation based on this cue as a function of different models of the zenith-intensity relationship and the amount of imagery available. We evaluate our algorithm on a dataset of more than 60 million images captured from outdoor webcams located around the globe. We find that using our algorithm with images sampled every 30 minutes, yields localization errors of less than 100 km for the majority of cameras. Nathan Jacobs, Kylia Miskell, Robert Pless |
WACV | 1 |
| 2010 | Using cloud shadows to infer scene structure and camera calibrationabstractWe explore the use of clouds as a form of structured lighting to capture the 3D structure of outdoor scenes observed over time from a static camera. We derive two cues that relate 3D distances to changes in pixel intensity due to clouds shadows. The first cue is primarily spatial, works with low frame-rate time lapses, and supports estimating focal length and scene structure, up to a scale ambiguity. The second cue depends on cloud motion and has a more complex, but still linear, ambiguity. We describe a method that uses the spatial cue to estimate a depth map and a method that combines both cues. Results on time lapses of several outdoor scenes show that these cues enable estimating scene geometry and camera focal length. Nathan Jacobs, Brian Bies, Robert Pless |
CVPR | 1 |
| 2010 | Compressive sensing and differential image-motion estimationabstractCompressive-sensing cameras are an important new class of sensors that have different design constraints than standard cameras. Surprisingly, little work has explored the relationship between compressive-sensing measurements and differential image motion. We show that, given modest constraints on the measurements and image motions, we can omit the computationally expensive compressive-sensing reconstruction step and obtain more accurate motion estimates with significantly less computation time. We also formulate a compressive-sensing reconstruction problem that incorporates known image motion and show that this method outperforms the state-of-the-art in compressive-sensing video reconstruction. Nathan Jacobs, S. Schuh, Robert Pless |
ICASSP | 1 |
| 2009 | The global network of outdoor webcams: properties and applicationsabstractThere are thousands of outdoor webcams which offer live images freely over the Internet. We report on methods for discovering and organizing this already existing and massively distributed global sensor, and argue that it provides an interesting alternative to satellite imagery for global-scale remote sensing applications. In particular, we characterize the live imaging capabilities that are freely available as of the summer of 2009 in terms of the spatial distribution of the cameras, their update rate, and characteristics of the scene in view. We offer algorithms that exploit the fact that webcams are typically static to simplify the tasks of inferring relevant environmental and weather variables directly from image data. Finally, we show that organizing and exploiting the large, ad-hoc, set of cameras attached to the web can dramatically increase the data available for studying particular problems in phenology. Nathan Jacobs, Walker Burgin, Nick Fridrich, Austin Abrams, Kylia Miskell, Bobby H. Braswell, Andrew D. Richardson, Robert Pless |
GIS | 1 |
| 2008 | Toward Fully Automatic Geo-Location and Geo-Orientation of Static Outdoor CamerasabstractAutomating tools for geo-locating and geo-orienting static cameras is a key step in creating a useful global imaging network from cameras attached to the Internet. We present algorithms for partial camera calibration that rely on access to accurately time-stamped images captured over time from cameras that do not move. To support these algorithms we also offer a method of camera viewpoint change detection, or "tamper detection", which determines if a camera has moved in the challenging case when images are only captured every half hour. These algorithms are tested on a subset of the AMOS (Archive of Many Outdoor Scenes) database, and we present preliminary results that highlight the promise of these approaches. Nathan Jacobs, Nathaniel Roman, Robert Pless |
WACV | 1 |
| 2008 | Time Scales in Video SurveillanceabstractEvents in surveillance video occur over many time scales, but common approaches to background subtraction and video representation are implicitly based on a single temporal scale. In this work, we derive a set of causal filters which define a temporal scale-space representation for the activity at each pixel. This scale-space can be maintained and continuously updated in real time and, for static cameras viewing dynamic scenes, has several interesting properties. In particular, it directly characterizes interesting temporal features and supports approximate reconstruction of the video history under challenging noise conditions. The temporal scale-space grounds novel approaches to several applications, including a natural visualization tool to summarize recent video behavior in a single image, and a tool to directly report how long the object has been present in a scene without reexamining any video data. Nathan Jacobs, Robert Pless |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Consistent Temporal Variations in Many Outdoor ScenesabstractThis paper details an empirical study of large image sets taken by static cameras. These images have consistent correlations over the entire image and over time scales of days to months. Simple second-order statistics of such image sets show vastly more structure than exists in generic natural images or video from moving cameras. Using a slight variant to PCA, we can decompose all cameras into comparable components and annotate images with respect to surface orientation, weather, and seasonal change. Experiments are based on a data set from 538 cameras across the United States which have collected more than 17 million images over the the last 6 months. Nathan Jacobs, Nathaniel Roman, Robert Pless |
CVPR | 1 |
| 2007 | Geolocating Static CamerasabstractA key problem in widely distributed camera networks is locating the cameras. This paper considers three scenarios for camera localization: localizing a camera in an unknown environment, adding a new camera in a region with many other cameras, and localizing a camera by finding correlations with satellite imagery. We find that simple summary statistics (the time course of principal component coefficients) are sufficient to geolocate cameras without determining correspondences between cameras or explicitly reasoning about weather in the scene. We present results from a database of images from 538 cameras collected over the course of a year. We find that for cameras that remain stationary and for which we have accurate image times- tamps, we can localize most cameras to within 50 miles of the known location. In addition, we demonstrate the use of a distributed camera network in the construction a map of weather conditions. Nathan Jacobs, Scott Satkin, Nathaniel Roman, Robert Speyer, Robert Pless |
ICCV | 1 |