VLDB 2026 Research / reviewers in the wild / expert
Scott Workman
dblp:135/4920
· DBLP profile ↗
31ranked-venue papers
13as first author
7since 2021 · last 2024
0000-0002-7145-7484ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 11 first-author · 6 since 2021Artificial intelligence and machine learning · 16 · 10 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Probabilistic Image-Driven Traffic Modeling via Remote Sensing
Scott Workman, Armin Hadzic |
ECCV (60) | 1 |
| 2024 | WATCH: Wide-Area Terrestrial Change HypercubeabstractMonitoring Earth activity using data collected from multiple satellite imaging platforms in a unified way is a significant challenge, especially with large variability in image resolution, spectral bands, and revisit rates. Further, the availability of sensor data varies across time as new platforms are launched. In this work, we introduce an adaptable framework and network architecture capable of predicting on subsets of the available platforms, bands, or temporal ranges it was trained on. Our system, called WATCH, is highly general and can be applied to a variety of geospatial tasks. In this work, we analyze the performance of WATCH using the recent IARPA SMART public dataset and metrics. We focus primarily on the problem of broad area search for heavy construction sites. Experiments validate the robustness of WATCH during inference to limited sensor availability, as well the the ability to alter inference-time spatial or temporal sampling. WATCH is open source and available for use on this or other remote sensing problems. Code and model weights are available at: https://gitlab.kitware.com/computer-vision/geowatch Connor Greenwell, Jon Crall, Matthew Purri, Kristin J. Dana, Nathan Jacobs, Armin Hadzic, Scott Workman, Matthew J. Leotta |
WACV | 7 |
| 2023 | Handling Image and Label Resolution Mismatch in Remote SensingabstractThough semantic segmentation has been heavily explored in vision literature, unique challenges remain in the remote sensing domain. One such challenge is how to handle resolution mismatch between overhead imagery and ground-truth label sources, due to differences in ground sample distance. To illustrate this problem, we introduce a new dataset and use it to showcase weaknesses inherent in existing strategies that naively upsample the target label to match the image resolution. Instead, we present a method that is supervised using low-resolution labels (without upsampling), but takes advantage of an exemplar set of highresolution labels to guide the learning process. Our method incorporates region aggregation, adversarial learning, and self-supervised pretraining to generate fine-grained predictions, without requiring high-resolution annotations. Extensive experiments demonstrate the real-world applicability of our approach. Scott Workman, Armin Hadzic, Muhammad Usman Rafique |
WACV | 1 |
| 2022 | Revisiting Near/Remote Sensing with Geospatial AttentionabstractThis work addresses the task of overhead image segmentation when auxiliary ground-level images are available. Recent work has shown that performing joint inference over these two modalities, often called near/remote sensing, can yield significant accuracy improvements. Extending this line of work, we introduce the concept of geospatial attention, a geometry-aware attention mechanism that explicitly considers the geospatial relationship between the pixels in a ground-level image and a geographic location. We propose an approach for computing geospatial attention that incorporates geometric features and the appearance of the overhead and ground-level imagery. We introduce a novel architecture for near/remote sensing that is based on geospatial attention and demonstrate its use for five segmentation tasks. The results demonstrate that our method significantly outperforms the previous state-of-the-art methods. Scott Workman, Muhammad Usman Rafique, Hunter Blanton, Nathan Jacobs |
CVPR | 1 |
| 2022 | A Structure-Aware Method for Direct Pose EstimationabstractEstimating camera pose from a single image is a fundamental problem in computer vision. Existing methods for solving this task fall into two distinct categories, which we refer to as direct and indirect. Direct methods, such as PoseNet, regress pose from the image as a fixed function, for example using a feed-forward convolutional network. Such methods are desirable because they are deterministic and run in constant time. Indirect methods for pose regression are often non-deterministic, with various external dependencies such as image retrieval and hypothesis sampling. We propose a direct method that takes inspiration from structure-based approaches to incorporate explicit 3D constraints into the network. Our approach maintains the desirable qualities of other direct methods while achieving much lower error in general. Code is available at https://github.com/mvrl/structure-aware-pose-estimation. Hunter Blanton, Scott Workman, Nathan Jacobs |
WACV | 2 |
| 2022 | Content-Aware Detection of Temporal Metadata ManipulationabstractMost pictures shared online are accompanied by temporal metadata (i.e., the day and time they were taken), which makes it possible to associate an image content with real-world events. Maliciously manipulating this metadata can convey a distorted version of reality. In this work, we present the emerging problem of detecting timestamp manipulation. We propose an end-to-end approach to verify whether the purported time of capture of an outdoor image is consistent with its content and geographic location. We consider manipulations done in the hour and/or month of capture of a photograph. The central idea is the use of supervised consistency verification, in which we predict the probability that the image content, capture time, and geographical location are consistent. We also include a pair of auxiliary tasks, which can be used to explain the network decision. Our approach improves upon previous work on a large benchmark dataset, increasing the classification accuracy from 59.0% to 81.1%. We perform an ablation study that highlights the importance of various components of the method, showing what types of tampering are detectable using our approach. Finally, we demonstrate how the proposed method can be employed to estimate a possible time-of-capture in scenarios in which the timestamp is missing from the metadata. Rafael Padilha, Tawfiq Salem, Scott Workman, Fernanda A. Andaló, Anderson Rocha 0001, Nathan Jacobs |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Augmenting Depth Estimation with Geospatial ContextabstractModern cameras are equipped with a wide array of sensors that enable recording the geospatial context of an image. Taking advantage of this, we explore depth estimation under the assumption that the camera is geocalibrated, a problem we refer to as geo-enabled depth estimation. Our key insight is that if capture location is known, the corresponding overhead viewpoint offers a valuable resource for understanding the scale of the scene. We propose an end-to-end architecture for depth estimation that uses geospatial context to infer a synthetic ground-level depth map from a co-located overhead image, then fuses it inside of an encoder/decoder style segmentation network. To support evaluation of our methods, we extend a recently released dataset with overhead imagery and corresponding height maps. Results demonstrate that integrating geospatial context significantly reduces error compared to baselines, both at close ranges and when evaluating at much larger distances than existing benchmarks consider. Scott Workman, Hunter Blanton |
ICCV | 1 |
| 2020 | Learning a Dynamic Map of Visual AppearanceabstractThe appearance of the world varies dramatically not only from place to place but also from hour to hour and month to month. Every day billions of images capture this complex relationship, many of which are associated with precise time and location metadata. We propose to use these images to construct a global-scale, dynamic map of visual appearance attributes. Such a map enables fine-grained understanding of the expected appearance at any geographic location and time. Our approach integrates dense overhead imagery with location and time metadata into a general framework capable of mapping a wide variety of visual attributes. A key feature of our approach is that it requires no manual data annotation. We demonstrate how this approach can support various applications, including image-driven mapping, image geolocalization, and metadata verification. Tawfiq Salem, Scott Workman, Nathan Jacobs |
CVPR | 2 |
| 2020 | Dynamic Traffic Modeling From Overhead ImageryabstractOur goal is to use overhead imagery to understand patterns in traffic flow, for instance answering questions such as how fast could you traverse Times Square at 3am on a Sunday. A traditional approach for solving this problem would be to model the speed of each road segment as a function of time. However, this strategy is limited in that a significant amount of data must first be collected before a model can be used and it fails to generalize to new areas. Instead, we propose an automatic approach for generating dynamic maps of traffic speeds using convolutional neural networks. Our method operates on overhead imagery, is conditioned on location and time, and outputs a local motion model that captures likely directions of travel and corresponding travel speeds. To train our model, we take advantage of historical traffic data collected from New York City. Experimental results demonstrate that our method can be applied to generate accurate city-scale traffic models. Scott Workman, Nathan Jacobs |
CVPR | 1 |
| 2020 | Single Image Cloud Detection via Multi-Image FusionabstractArtifacts in imagery captured by remote sensing, such as clouds, snow, and shadows, present challenges for various tasks, including semantic segmentation and object detection. A primary challenge in developing algorithms for identifying such artifacts is the cost of collecting annotated training data. In this work, we explore how recent advances in multi-image fusion can be leveraged to bootstrap single image cloud detection. We demonstrate that a network optimized to estimate image quality also implicitly learns to detect clouds. To support the training and evaluation of our approach, we collect a large dataset of Sentinel-2 images along with a per-pixel semantic labelling for land cover. Through various experiments, we demonstrate that our method reduces the need for annotated training data and improves cloud detection performance. Scott Workman, Muhammad Usman Rafique, Hunter Blanton, Connor Greenwell, Nathan Jacobs |
IGARSS | 1 |
| 2018 | Learning Geo-Temporal Image Features
Menghua Zhai, Tawfiq Salem, Connor Greenwell, Scott Workman, Robert Pless, Nathan Jacobs |
BMVC | 4 |
| 2018 | What Goes Where: Predicting Object Distributions from AboveabstractIn this work, we propose a cross-view learning approach, in which images captured from a ground-level view are used as weakly supervised annotations for interpreting overhead imagery. The outcome is a convolutional neural network for overhead imagery that is capable of predicting the type and count of objects that are likely to be seen from a ground-level perspective. We demonstrate our approach on a large dataset of geotagged ground-level and overhead imagery and find that our network captures semantically meaningful features, despite being trained without manual annotations. Connor Greenwell, Scott Workman, Nathan Jacobs |
IGARSS | 2 |
| 2018 | A Multimodal Approach to Mapping SoundscapesabstractWe explore the problem of mapping soundscapes, that is, predicting the types of sounds that are likely to be heard at a given geographic location. Using a novel dataset, which includes geo-tagged audio and overhead imagery, we develop an approach for constructing an aural atlas, which captures the geospatial distribution of soundscapes. We build on previous work relating sound to ground-level imagery but incorporate overhead imagery to overcome the limitations of sparsely distributed geo-tagged audio. In the end, all that we require to construct an aural atlas is overhead imagery of the region of interest. We show examples of aural atlases at multiple spatial scales, from block-level to country. Tawfiq Salem, Menghua Zhai, Scott Workman, Nathan Jacobs |
IGARSS | 3 |
| 2018 | FARSA: Fully Automated Roadway Safety AssessmentabstractThis paper addresses the task of road safety assessment. An emerging approach for conducting such assessments in the United States is through the US Road Assessment Program (usRAP), which rates roads from highest risk (1 star) to lowest (5 stars). Obtaining these ratings requires manual, fine-grained labeling of roadway features in streetlevel panoramas, a slow and costly process. We propose to automate this process using a deep convolutional neural network that directly estimates the star rating from a street-level panorama, requiring milliseconds per image at test time. Our network also estimates many other roadlevel attributes, including curvature, roadside hazards, and the type of median. To support this, we incorporate taskspecific attention layers so the network can focus on the panorama regions that are most useful for a particular task. We evaluated our approach on a large dataset of real-world images from two US states. We found that incorporating additional tasks, and using a semi-supervised training approach, significantly reduced overfitting problems, allowed us to optimize more layers of the network, and resulted in higher accuracy. Weilian Song, Scott Workman, Armin Hadzic, Eric Green, Reginald R. Souleyrette, Nathan Jacobs |
WACV | 2 |
| 2017 | Predicting Ground-Level Scene Layout from Aerial Imagery
Menghua Zhai, Zachary Bessinger, Scott Workman, Nathan Jacobs |
CVPR | 3 |
| 2017 | Understanding and Mapping Natural BeautyabstractWhile natural beauty is often considered a subjective property of images, in this paper, we take an objective approach and provide methods for quantifying and predicting the scenicness of an image. Using a dataset containing hundreds of thousands of outdoor images captured throughout Great Britain with crowdsourced ratings of natural beauty, we propose an approach to predict scenicness which explicitly accounts for the variance of human ratings. We demonstrate that quantitative measures of scenicness can benefit semantic image understanding, content-aware image processing, and a novel application of cross-view mapping, where the sparsity of ground-level images can be addressed by incorporating unlabeled overhead images in the training and prediction steps. For each application, our methods for scenicness prediction result in quantitative and qualitative improvements over baseline approaches. Scott Workman, Richard Souvenir, Nathan Jacobs |
ICCV | 1 |
| 2017 | A Unified Model for Near and Remote SensingabstractWe propose a novel convolutional neural network architecture for estimating geospatial functions such as population density, land cover, or land use. In our approach, we combine overhead and ground-level images in an end-toend trainable neural network, which uses kernel regression and density estimation to convert features extracted from the ground-level images into a dense feature map. The output of this network is a dense estimate of the geospatial function in the form of a pixel-level labeling of the overhead image. To evaluate our approach, we created a large dataset of overhead and ground-level images from a major urban area with three sets of labels: land use, building function, and building age. We find that our approach is more accurate for all tasks, in some cases dramatically so. Scott Workman, Menghua Zhai, David Crandall, Nathan Jacobs |
ICCV | 1 |
| 2016 | Horizon Lines in the Wild
Scott Workman, Menghua Zhai, Nathan Jacobs |
BMVC | 1 |
| 2016 | Detecting Vanishing Points Using Global Image Context in a Non-ManhattanWorldabstractWe propose a novel method for detecting horizontal vanishing points and the zenith vanishing point in man-made environments. The dominant trend in existing methods is to first find candidate vanishing points, then remove outliers by enforcing mutual orthogonality. Our method reverses this process: we propose a set of horizon line candidates and score each based on the vanishing points it contains. A key element of our approach is the use of global image context, extracted with a deep convolutional network, to constrain the set of candidates under consideration. Our method does not make a Manhattan-world assumption and can operate effectively on scenes with only a single horizontal vanishing point. We evaluate our approach on three benchmark datasets and achieve state-of the-art performance on each. In addition, our approach is significantly faster than the previous best method. Menghua Zhai, Scott Workman, Nathan Jacobs |
CVPR | 2 |
| 2016 | Camera geo-calibration using an MCMC approachabstractWe address the problem of single-image geo-calibration, in which an estimate of the geographic location, viewing direction and field of view is sought for the camera that captured an image. The dominant approach to this problem is to match features of the query image, using color and texture, against a reference database of nearby ground imagery. However, this fails when such imagery is not available. We propose to overcome this limitation by matching against a geographic database that contains the locations of known objects, such as houses, roads and bodies of water. Since we are unable to find one-to-one correspondences between image locations and objects in our database, we model the problem probabilistically based on the geometric configuration of multiple such weak correspondences. We propose a Markov Chain Monte Carlo (MCMC) sampling approach to approximate the underlying probability distribution over the full geo-calibration of the camera. Menghua Zhai, Scott Workman, Nathan Jacobs |
ICIP | 2 |
| 2016 | A fast method for estimating transient scene attributesabstractWe propose the use of deep convolutional neural networks to estimate the transient attributes of a scene from a single image. Transient scene attributes describe both the objective conditions, such as the weather, time of day, and the season, and subjective properties of a scene, such as whether or not the scene seems busy. Recently, convolutional neural networks have been used to achieve state-of-the-art results for many vision problems, from object detection to scene classification, but have not previously been used for estimating transient attributes. We compare several methods for adapting an existing network architecture and present state-of-the-art results on two benchmark datasets. Our method is more accurate and significantly faster than previous methods, enabling real-world applications. Ryan Baltenberger, Menghua Zhai, Connor Greenwell, Scott Workman, Nathan Jacobs |
WACV | 4 |
| 2016 | Sky segmentation in the wild: An empirical studyabstractAutomatically determining which pixels in an image view the sky, the problem of sky segmentation, is a critical preprocessing step for a wide variety of outdoor image interpretation problems, including horizon estimation, robot navigation and image geolocalization. Many methods for this problem have been proposed with recent work achieving significant improvements on benchmark datasets. However, such datasets are often constructed to contain images captured in favorable conditions and, therefore, do not reflect the broad range of conditions with which a real-world vision system must cope. This paper presents the results of a large-scale empirical evaluation of the performance of three state-of-the-art approaches on a new dataset, which consists of roughly 100k images captured "in the wild". The results show that the performance of these methods can be dramatically degraded by the local lighting and weather conditions. We propose a deep learning based variant of an ensemble solution that outperforms the methods we tested, in some cases achieving above 50% relative reduction in misclassified pixels. While our results show there is room for improvement, our hope is that this dataset will encourage others to improve the real-world performance of their algorithms. Radu Paul Mihail, Scott Workman, Zachary Bessinger, Nathan Jacobs |
WACV | 2 |
| 2016 | Analyzing human appearance as a cue for dating imagesabstractGiven an image, we propose to use the appearance of people in the scene to estimate when the picture was taken. There are a wide variety of cues that can be used to address this problem. Most previous work has focused on low-level image features, such as color and vignetting. Recent work on image dating has used more semantic cues, such as the appearance of automobiles and buildings. We extend this line of research by focusing on human appearance. Our approach, based on a deep convolutional neural network, allows us to more deeply explore the relationship between human appearance and time. We find that clothing, hair styles, and glasses can all be informative features. To support our analysis, we have collected a new dataset containing images of people from many high school yearbooks, covering the years 1912-2014. While not a complete solution to the problem of image dating, our results show that human appearance is strongly related to time and that semantic information can be a useful cue. Tawfiq Salem, Scott Workman, Menghua Zhai, Nathan Jacobs |
WACV | 2 |
| 2016 | Cloudmaps from static ground-view video
Nathan Jacobs, Scott Workman, Richard Souvenir |
Image Vis. Comput. | 2 |
| 2015 | Wide-Area Image Geolocalization with Aerial Reference ImageryabstractWe propose to use deep convolutional neural networks to address the problem of cross-view image geolocalization, in which the geolocation of a ground-level query image is estimated by matching to georeferenced aerial images. We use state-of-the-art feature representations for ground-level images and introduce a cross-view training approach for learning a joint semantic feature representation for aerial images. We also propose a network architecture that fuses features extracted from aerial images at multiple spatial scales. To support training these networks, we introduce a massive database that contains pairs of aerial and ground-level images from across the United States. Our methods significantly out-perform the state of the art on two benchmark datasets. We also show, qualitatively, that the proposed feature representations are discriminative at both local and continental spatial scales. Scott Workman, Richard Souvenir, Nathan Jacobs |
ICCV | 1 |
| 2015 | FACE2GPS: Estimating geographic location from facial featuresabstractThe facial appearance of a person is a product of many factors, including their gender, age, and ethnicity. Methods for estimating these latent factors directly from an image of a face have been extensively studied for decades. We extend this line of work to include estimating the location where the image was taken. We propose a deep network architecture for making such predictions and demonstrate its superiority to other approaches in an extensive set of quantitative experiments on the GeoFaces dataset. Our experiments show that in 26% of the cases the ground truth location is the topmost prediction, and if we allow ourselves to consider the top five predictions, the accuracy increases to 47%. In both cases, the deep learning based approach significantly outperforms random chance as well as another baseline method. Mohammad T. Islam 0001, Scott Workman, Nathan Jacobs |
ICIP | 2 |
| 2015 | DEEPFOCAL: A method for direct focal length estimationabstractEstimating the focal length of an image is an important preprocessing step for many applications. Despite this, existing methods for single-view focal length estimation are limited in that they require particular geometric calibration objects, such as orthogonal vanishing points, co-planar circles, or a calibration grid, to occur in the field of view. In this work, we explore the application of a deep convolutional neural network, trained on natural images obtained from Internet photo collections, to directly estimate the focal length using only raw pixel intensities as input features. We present quantitative results that demonstrate the ability of our technique to estimate the focal length with comparisons against several baseline methods, including an automatic method which uses orthogonal vanishing points. Scott Workman, Connor Greenwell, Menghua Zhai, Ryan Baltenberger, Nathan Jacobs |
ICIP | 1 |
| 2015 | Scene shape estimation from multiple partly cloudy days
Scott Workman, Richard Souvenir, Nathan Jacobs |
Comput. Vis. Image Underst. | 1 |
| 2014 | A Pot of Gold: Rainbows as a Calibration Cue
Scott Workman, Radu Paul Mihail, Nathan Jacobs |
ECCV (5) | 1 |
| 2014 | Exploring the geo-dependence of human face appearanceabstractThe expected appearance of a human face depends strongly on age, ethnicity and gender. While these relationships are well-studied, our work explores the little-studied dependence of facial appearance on geographic location. To support this effort, we constructed GeoFaces, a large dataset of geotagged face images. We examine the geo-dependence of Eigenfaces and use two supervised methods for extracting geo-informative features. The first, canonical correlation analysis, is used to find location-dependent component images as well as the spatial direction of most significant face appearance change. The second, linear discriminant analysis, is used to find countries with relatively homogeneous, yet distinctive, facial appearance. Mohammad T. Islam 0001, Scott Workman, Hui Wu 0006, Nathan Jacobs, Richard Souvenir |
WACV | 2 |
| 2013 | Cloud Motion as a Calibration CueabstractWe propose cloud motion as a natural scene cue that enables geometric calibration of static outdoor cameras. This work introduces several new methods that use observations of an outdoor scene over days and weeks to estimate radial distortion, focal length and geo-orientation. Cloud-based cues provide strong constraints and are an important alternative to methods that require specific forms of static scene geometry or clear sky conditions. Our method makes simple assumptions about cloud motion and builds upon previous work on motion-based and line-based calibration. We show results on real scenes that highlight the effectiveness of our proposed methods. Nathan Jacobs, Mohammad T. Islam 0001, Scott Workman |
CVPR | 3 |