Gordon A. Christie

dblp:215/3335 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
1since 2021 · last 2023
0000-0001-6537-8061ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 32% Vision and language · 18% Image recognition and object detection · 12%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Environmental and earth informatics · 100%
Human-computer interaction and pervasive computing
1 paper
Usability and user experience research · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Environmental and earth informatics
earth observation
0.522020
Functional Map of the World · CVPR 2018
Learning Geocentric Object Pose in Oblique Monocular Images · CVPR 2020
Computer vision › 3D vision
depth estimation
0.412020
Learning Geocentric Object Pose in Oblique Monocular Images · CVPR 2020
Computer vision › Image recognition and object detection › image classification › remote sensing image classification
land use classification
0.312018
Functional Map of the World · CVPR 2018
Environmental and earth informatics › remote sensing
satellite imagery analysis
0.312018
Functional Map of the World · CVPR 2018
Computer vision › Vision and language
grounded language understanding
0.212016
Resolving Language and Vision Ambiguities Together: Joint Segmentation & Prepositional Attachment Resolution in Captioned Scenes · EMNLP 2016
Natural language and speech › Information extraction and text analysis › syntactic parsing › syntactic disambiguation
prepositional phrase attachment
0.212016
Resolving Language and Vision Ambiguities Together: Joint Segmentation & Prepositional Attachment Resolution in Captioned Scenes · EMNLP 2016
Natural language and speech › Question answering and dialogue systems
question understanding
0.212016
Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions · EMNLP 2016
Computer vision › Vision and language
visual question answering
0.212016
Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions · EMNLP 2016
Machine learning › Trustworthy machine learning
interpretability
0.212014
Predicting User Annoyance Using Visual Attributes · CVPR 2014
Computer vision › Video understanding and tracking › spatio-temporal modeling
image sequence analysis
0.112018
Functional Map of the World · CVPR 2018
Information retrieval
image retrieval
0.112014
Predicting User Annoyance Using Visual Attributes · CVPR 2014

Methods — techniques the papers use, named apart from their topics

lidar supervision · 0.9image rectification · 0.9deep network · 0.9temporal view fusion · 0.7metadata reasoning · 0.7attribute-based representation · 0.6transfer learning · 0.4model uncertainty · 0.2joint inference · 0.2caption-question similarity · 0.2LSTM-RNN · 0.2
YearPublicationVenuePosition
2023 CrossAdapt: Cross-Scene Adaptation for Multi-Domain Depth Estimation
abstract
We address the task of monocular depth estimation in the multi-domain setting. Given a large dataset (source) with ground-truth depth maps, and a set of unlabeled datasets (targets), our goal is to create a model that works well on unlabeled target datasets across different scenes. This is a challenging problem when there is a significant domain shift, often resulting in poor performance on the target datasets. We propose to address this task with a unified approach that includes adversarial knowledge distillation and uncertainty-guided self-supervised reconstruction. We provide both quantitative and qualitative evaluations on four datasets: KITTI, Virtual KITTI, UAVid China, and UAVid Germany. These datasets contain widely varying viewpoints, including ground-level and overhead perspectives, which is more challenging than is typically considered in prior work on domain adaptation for single-image depth. Our approach significantly improves upon conventional domain adaptation baselines and does not require additional memory as the number of target sets increases.
Yu Zhang 0094, Muhammad Usman Rafique, Gordon A. Christie, Nathan Jacobs
IGARSS3
2020 Learning Geocentric Object Pose in Oblique Monocular Images
abstract
An object's geocentric pose, defined as the height above ground and orientation with respect to gravity, is a powerful representation of real-world structure for object detection, segmentation, and localization tasks using RGBD images. For close-range vision tasks, height and orientation have been derived directly from stereo-computed depth and more recently from monocular depth predicted by deep networks. For long-range vision tasks such as Earth observation, depth cannot be reliably estimated with monocular images. Inspired by recent work in monocular height above ground prediction and optical flow prediction from static images, we develop an encoding of geocentric pose to address this challenge and train a deep network to compute the representation densely, supervised by publicly available airborne lidar. We exploit these attributes to rectify oblique images and remove observed object parallax to dramatically improve the accuracy of localization and to enable accurate alignment of multiple images taken from very different oblique viewpoints. We demonstrate the value of our approach by extending two large-scale public datasets for semantic segmentation in oblique satellite images. All of our data and code are publicly available.
Gordon A. Christie, Rodrigo Rene Rai Munoz Abujder, Kevin Foster, Shea Hagstrom, Gregory D. Hager, Myron Z. Brown
CVPR1
2020 Estimating Displaced Populations from Overhead
abstract
We introduce a deep learning approach to perform fine-grained population estimation for displacement camps using high-resolution overhead imagery. We train and evaluate our approach on drone imagery cross-referenced with population data for refugee camps in Cox's Bazar, Bangladesh in 2018 and 2019. Our proposed approach achieves 7.41% mean absolute percent error on sequestered camp imagery. We believe our experiments with real-world displacement camp data constitute an important step towards the development of tools that enable the humanitarian community to effectively and rapidly respond to the global displacement crisis.
Armin Hadzic, Gordon A. Christie, Jeffrey Freeman, Amber Dismer, Stevan Bullard, Ashley Greiner, Nathan Jacobs, Ryan Mukherjee
IGARSS2
2019 Sensor Adaptation for Improved Semantic Segmentation of Overhead Imagery
abstract
Semantic segmentation is a powerful method to facilitate visual scene understanding. Each pixel is assigned a label according to a pre-defined list of object classes and semantic entities. This becomes very useful as a means to summarize large scale overhead imagery. In this paper we present our work on semantic segmentation with applications to overhead imagery. We propose an algorithm that builds and extends upon the DeepLab framework to be able to refine and resolve small objects (relative to the image size) such as vehicles. We have also investigated sensor adaptation as a means to augment available training data to be able to reduce some of the shortcomings of neural networks when deployed in new environments and to new sensors. We report results on several datasets and compare performance with other state-of-the-art architectures.
Marc Bosch, Gordon A. Christie, Christopher M. Gifford
WACV2
2019 Semantic Stereo for Incidental Satellite Images
abstract
The increasingly common use of incidental satellite images for stereo reconstruction versus rigidly tasked binocular or trinocular coincident collection is helping to enable timely global-scale 3D mapping; however, reliable stereo correspondence from multi-date image pairs remains very challenging due to seasonal appearance differences and scene change. Promising recent work suggests that semantic scene segmentation can provide a robust regularizing prior for resolving ambiguities in stereo correspondence and reconstruction problems. To enable research for pairwise semantic stereo and multi-view semantic 3D reconstruction with incidental satellite images, we have established a large-scale public dataset including multi-view, multi-band satellite images and ground truth geometric and semantic labels for two large cities. To demonstrate the complementary nature of the stereo and segmentation tasks, we present lightweight public baselines adapted from recent state of the art convolutional neural network models and assess their performance.
Marc Bosch, Kevin Foster, Gordon A. Christie, Sean Wang 0001, Gregory D. Hager, Myron Z. Brown
WACV3
2018 Functional Map of the World
abstract
We present a new dataset, Functional Map of the World (fMoW), which aims to inspire the development of machine learning models capable of predicting the functional purpose of buildings and land use from temporal sequences of satellite images and a rich set of metadata features. The metadata provided with each image enables reasoning about location, time, sun angles, physical sizes, and other features when making predictions about objects in the image. Our dataset consists of over 1 million images from over 200 countries. For each image, we provide at least one bounding box annotation containing one of 63 categories, including a "false detection" category. We present an analysis of the dataset along with baseline approaches that reason about metadata and temporal views. Our data, code, and pretrained models have been made publicly available.
Gordon A. Christie, Neil Fendley, Ryan Mukherjee
CVPR1
2017 Resolving vision and language ambiguities together: Joint segmentation & prepositional attachment resolution in captioned scenes
Gordon A. Christie, Ankit Laddha, Aishwarya Agrawal, Stanislaw Antol, Yash Goyal, Kevin Kochersberger, Dhruv Batra
Comput. Vis. Image Underst.1
2016 Resolving Language and Vision Ambiguities Together: Joint Segmentation & Prepositional Attachment Resolution in Captioned Scenes
abstract
Gordon Christie, Ankit Laddha, Aishwarya Agrawal, Stanislaw Antol, Yash Goyal, Kevin Kochersberger, Dhruv Batra. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016.
Gordon A. Christie, Ankit Laddha, Aishwarya Agrawal, Stanislaw Antol, Yash Goyal, Kevin Kochersberger, Dhruv Batra
EMNLP1
2016 Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions
abstract
Visual Question Answering (VQA) is the task of answering natural-language questions about images.We introduce the novel problem of determining the relevance of questions to images in VQA.Current VQA models do not reason about whether a question is even related to the given image (e.g., What is the capital of Argentina?) or if it requires information from external resources to answer correctly.This can break the continuity of a dialogue in human-machine interaction.Our approaches for determining relevance are composed of two stages.Given an image and a question, (1) we first determine whether the question is visual or not, (2) if visual, we determine whether the question is relevant to the given image or not.Our approaches, based on LSTM-RNNs, VQA model uncertainty, and caption-question similarity, are able to outperform strong baselines on both relevance tasks.We also present human studies showing that VQA models augmented with such question relevance reasoning are perceived as more intelligent, reasonable, and human-like.
Arijit Ray, Gordon A. Christie, Mohit Bansal, Dhruv Batra, Devi Parikh
EMNLP2
2015 Fast inspection for size-based analysis in aggregate processing
Gordon A. Christie, Kevin Kochersberger, A. Lynn Abbott
Mach. Vis. Appl.1
2014 Predicting User Annoyance Using Visual Attributes
abstract
Computer Vision algorithms make mistakes. In human-centric applications, some mistakes are more annoying to users than others. In order to design algorithms that minimize the annoyance to users, we need access to an annoyance or cost matrix that holds the annoyance of each type of mistake. Such matrices are not readily available, especially for a wide gamut of human-centric applications where annoyance is tied closely to human perception. To avoid having to conduct extensive user studies to gather the annoyance matrix for all possible mistakes, we propose predicting the annoyance of previously unseen mistakes by learning from example mistakes and their corresponding annoyance. We promote the use of attribute-based representations to transfer this knowledge of annoyance. Our experimental results with faces and scenes demonstrate that our approach can predict annoyance more accurately than baselines. We show that as a result, our approach makes less annoying mistakes in a real-world image retrieval application.
Gordon A. Christie, Amar Parkash, Ujwal Krothapalli, Devi Parikh
CVPR1