VLDB 2026 Research / reviewers in the wild / expert
Jeffrey Byrne
dblp:25/6285
· DBLP profile ↗
14ranked-venue papers
6as first author
2since 2021 · last 2025
0000-0001-8973-0322ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
3D vision · 66% Trustworthy machine learning · 9% Face, body and person analysis · 9% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 19 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features · CVPR 2025 |
Computer vision › 3D vision
structure from motion |
0.9 | 1 | 2025 | Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features · CVPR 2025 |
Computer vision › 3D vision › structure from motion
visual disambiguation |
0.9 | 1 | 2025 | Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features · CVPR 2025 |
Computer vision › Face, body and person analysis
face recognition |
0.4 | 1 | 2020 | Explainable Face Recognition · ECCV (11) 2020 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2020 | Explainable Face Recognition · ECCV (11) 2020 |
Computer vision › Video understanding and tracking
activity recognition |
0.2 | 1 | 2015 | Nested motion descriptors · CVPR 2015 |
Computer vision › Video understanding and tracking › motion representation
motion descriptor |
0.2 | 1 | 2015 | Nested motion descriptors · CVPR 2015 |
Image and video processing › image representation
spatiotemporal representation |
0.2 | 1 | 2015 | Nested motion descriptors · CVPR 2015 |
Computer vision › 3D vision
feature matching |
0.2 | 2 | 2013 | Nested Shape Descriptors · ICCV 2013 Inertial aided SIFT for time to collision estimation · ICRA 2009 |
Computer vision › 3D vision › motion estimation
time-to-collision estimation |
0.2 | 2 | 2009 | Inertial aided SIFT for time to collision estimation · ICRA 2009 Expansion segmentation for visual collision detection and estimation · ICRA 2009 |
Computer vision › 3D vision
local feature descriptor |
0.2 | 1 | 2013 | Nested Shape Descriptors · ICCV 2013 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.2 | 2 | 2009 | Expansion segmentation for visual collision detection and estimation · ICRA 2009 Stereo based Obstacle Detection for an Unmanned Air Vehicle · ICRA 2006 |
Image and video processing › image representation
image descriptor |
0.1 | 1 | 2015 | Nested motion descriptors · CVPR 2015 |
Computer vision › Segmentation and scene understanding › image segmentation
graph-based segmentation |
0.1 | 1 | 2006 | Stereo based Obstacle Detection for an Unmanned Air Vehicle · ICRA 2006 |
Robotics › Robot navigation and mapping
obstacle detection |
0.1 | 1 | 2006 | Stereo based Obstacle Detection for an Unmanned Air Vehicle · ICRA 2006 |
Computer vision › 3D vision
stereo vision |
0.1 | 1 | 2006 | Stereo based Obstacle Detection for an Unmanned Air Vehicle · ICRA 2006 |
Robotics › Legged, aerial and field robots › aerial robots
unmanned aerial vehicle |
0.0 | 2 | 2009 | Expansion segmentation for visual collision detection and estimation · ICRA 2009 Stereo based Obstacle Detection for an Unmanned Air Vehicle · ICRA 2006 |
Robotics › Robot navigation and mapping › mobile robot navigation
safe navigation |
0.0 | 1 | 2009 | Expansion segmentation for visual collision detection and estimation · ICRA 2009 |
Robotics › Motion planning and robot control
collision avoidance |
0.0 | 1 | 2006 | Stereo based Obstacle Detection for an Unmanned Air Vehicle · ICRA 2006 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9MASt3R · 0.93d-aware features · 0.9quadrature steerable pyramid · 0.4log-spiral normalization · 0.4oriented gradient pooling · 0.2nesting distance · 0.2uncertainty modeling · 0.1inertial-aided estimation · 0.1inertial-aided correspondence · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Doppelgangers++: Improved Visual Disambiguation with Geometric 3D FeaturesabstractAccurate 3D reconstruction is frequently hindered by visual aliasing, where visually similar but distinct surfaces (aka, doppelgangers), are incorrectly matched. These spurious matches distort the structure-from-motion (SfM) process, leading to misplaced model elements and reduced accuracy. Prior efforts addressed this with CNN classifiers trained on curated datasets, but these approaches struggle to generalize across diverse real-world scenes and can require extensive parameter tuning. In this work, we present Doppelgangers++, a method to enhance doppelganger detection and improve 3D reconstruction accuracy. Our contributions include a diversified training dataset that incorporates geo-tagged images from everyday scenes to expand robustness beyond landmark-based datasets. We further propose a Transformer-based classifier that leverages 3D-aware features from the MASt3R model, achieving superior precision and recall across both in-domain and out-of-domain tests. Doppelgangers++ integrates seamlessly into standard SfM and MASt3R-SfM pipelines, offering efficiency and adaptability across varied scenes. To evaluate SfM accuracy, we introduce an automated, geotag-based method for validating reconstructed models, eliminating the need for manual inspection. Through extensive experiments, we demonstrate that Doppelgangers++ significantly enhances pairwise vi sual disambiguation and improves 3D reconstruction quality in complex and diverse scenarios. Yuanbo Xiangli, Ruojin Cai, Hanyu Chen 0002, Jeffrey Byrne, Noah Snavely |
CVPR | 4 |
| 2023 | Fine-grained Activities of People WorldwideabstractEvery day, humans perform many closely related activities that involve subtle discriminative motions, such as putting on a shirt vs. putting on a jacket, or shaking hands vs. giving a high five. Activity recognition by ethical visual AI could provide insights into our patterns of daily life, however existing activity recognition datasets do not capture the massive diversity of these human activities around the world. To address this limitation, we introduce Collector, a free mobile app to record video while simultaneously annotating objects and activities of consented subjects. This new data collection platform was used to curate the Consented Activities of People (CAP) dataset, the first large-scale, fine-grained activity dataset of people worldwide. The CAP dataset contains 1.45M video clips of 512 fine grained activity labels of daily life, collected by 780 subjects in 33 countries. We provide activity classification and activity detection benchmarks for this dataset, and analyze baseline results to gain insight into how people around with world perform common activities. The dataset, benchmarks, evaluation tools, public leaderboards and mobile apps are available for use at https://visym.github.io/cap. Jeffrey Byrne, Greg Castañón, Zhongheng Li, Gil J. Ettinger |
WACV | 1 |
| 2020 | Key-Nets: Optical Transformation Convolutional Networks for Privacy Preserving Vision Sensors
Jeffrey Byrne, Brian DeCann, Scott Bloom |
BMVC | 1 |
| 2020 | Inducing Predictive Uncertainty Estimation for Face Verification
Weidi Xie, Jeffrey Byrne, Andrew Zisserman |
BMVC | 2 |
| 2020 | Explainable Face Recognition
Jonathan R. Williford, Brandon B. May, Jeffrey Byrne |
ECCV (11) | 3 |
| 2020 | CANOPIC: Pre-Digital Privacy-Enhancing Encodings for Computer VisionabstractThe standard pipeline for many vision tasks uses a conventional camera to capture an image that is then passed to a digital processor for information extraction. In some deployments, such as private locations, the captured digital imagery contains sensitive information exposed to digital vulnerabilities such as spyware, Trojans, etc. However, in many applications, the full imagery is unnecessary for the vision task at hand. In this paper we propose an optical and analog system that preprocesses the light from the scene before it reaches the digital imager to destroy sensitive information. We explore analog and optical encodings consisting of easily implementable operations such as convolution, pooling, and quantization. We perform a case study to evaluate how such encodings can destroy face identity information while preserving enough information for face detection. The encoding parameters are learned via an alternating optimization scheme based on adversarial learning with deep neural networks. We name our system CAnOPIC (Camera with Analog and Optical Privacy-Integrating Computations) and show that it has better performance in terms of both privacy and utility than conventional optical privacy-enhancing methods such as blurring and pixelation. Jasper Tan, Salman Siddique Khan, Vivek Boominathan, Jeffrey Byrne, Richard G. Baraniuk, Kaushik Mitra, Ashok Veeraraghavan |
ICME | 4 |
| 2018 | Visualizing and Quantifying Discriminative Features for Face RecognitionabstractDeep convolutional networks have generated significant performance improvements in the domain of face recognition. However, these improvements do not provide insight into which facial features lead to classification decisions. In this paper, we explore the problem of visualizing discriminative information in faces, to show which properties of images and subjects influence classification. We compare six different techniques for computing a network saliency map, which identifies influential local features in an image, using a metric called the "hiding game" to directly evaluate these techniques on classification performance. Results show that contrastive excitation backprop (cEBP) [26] best localizes features that lead to face identification. However, these maps are nearly identical across subjects, which can result in an unstable network saliency map. We introduce a robust improvement called truncated cEBP and demonstrate the capability to predict the performance of a given map. Our evaluation provides the first application of network saliency to face recognition, and we provide a robust new tool for face recognition analysts to explore which facial regions lead to changes in match scores. Greg Castañón, Jeffrey Byrne |
FG | 2 |
| 2018 | Template adaptation for face verification and identification
Nate Crosswhite, Jeffrey Byrne, Chris Stauffer, Omkar M. Parkhi, Qiong Cao, Andrew Zisserman |
Image Vis. Comput. | 2 |
| 2017 | Template Adaptation for Face Verification and Identification
Nate Crosswhite, Jeffrey Byrne, Chris Stauffer, Omkar M. Parkhi, Qiong Cao, Andrew Zisserman |
FG | 2 |
| 2015 | Nested motion descriptorsabstractA nested motion descriptor is a spatiotemporal representation of motion that is invariant to global camera translation, without requiring an explicit estimate of optical flow or camera stabilization. This descriptor is a natural spatiotemporal extension of the nested shape descriptor [2] to the representation of motion. We demonstrate that the quadrature steerable pyramid can be used to pool phase, and that pooling phase rather than magnitude provides an estimate of camera motion. This motion can be removed using the log-spiral normalization as introduced in the nested shape descriptor. Furthermore, this structure enables an elegant visualization of salient motion using the reconstruction properties of the steerable pyramid. We compare our descriptor to local motion descriptors, HOG-3D and HOG-HOF, and show improvements on three activity recognition datasets. Jeffrey Byrne |
CVPR | 1 |
| 2013 | Nested Shape DescriptorsabstractIn this paper, we propose a new family of binary local feature descriptors called nested shape descriptors. These descriptors are constructed by pooling oriented gradients over a large geometric structure called the Hawaiian earring, which is constructed with a nested correlation structure that enables a new robust local distance function called the nesting distance. This distance function is unique to the nested descriptor and provides robustness to outliers from order statistics. In this paper, we define the nested shape descriptor family and introduce a specific member called the seed-of-life descriptor. We perform a trade study to determine optimal descriptor parameters for the task of image matching. Finally, we evaluate performance compared to state-of-the-art local feature descriptors on the VGG-Affine image matching benchmark, showing significant performance gains. Our descriptor is the first binary descriptor to outperform SIFT on this benchmark. Jeffrey Byrne, Jianbo Shi |
ICCV | 1 |
| 2009 | Expansion segmentation for visual collision detection and estimationabstractCollision detection and estimation from a monocular visual sensor is an important enabling technology for safe navigation of small or micro air vehicles in near earth flight. In this paper, we introduce a new approach called expansion segmentation, which simultaneously detects “collision danger regions” of significant positive divergence in inertial aided video, and estimates maximum likelihood time to collision (TTC) in a correspondenceless framework within the danger regions. This approach was motivated from a literature review which showed that existing approaches make strong assumptions about scene structure or camera motion, or pose collision detection without determining obstacle boundaries, both of which limit the operational envelope of a deployable system. Expansion segmentation is based on a new formulation of 6-DOF inertial aided TTC estimation, and a new derivation of a first order TTC uncertainty model due to subpixel quantization error and epipolar geometry uncertainty. Proof of concept results are shown in a custom designed urban flight simulator and on operational flight data from a small air vehicle. Jeffrey Byrne, Camillo J. Taylor |
ICRA | 1 |
| 2009 | Inertial aided SIFT for time to collision estimationabstractVisual time to collision estimation for small or micro air vehicles is challenging due to aggressive 6-DOF motion, real time performance requirements and significant size, weight and power constraints of the platform. Recent work in collision detection using insect inspired optical flow based methods have been demonstrated in low power hardware implementations [1][2][3][4], but have not achieved the obstacle detection and false alarm rate performance necessary for practical deployment. This performance is sensitive to correspondence errors in the optical flow field, so one approach to improving performance is to use a richer feature set for correspondence, along with calibrated inertial information from the platform to aid correspondence. In this video, we show proof of concept results for such an approach. Estimation results are noisy, but encouraging, and given that SIFT feature correspondence has been demonstrated in real time on low power GPUs, it has the potential for future small UAV integration. Benjamin J. Cohen, Jeffrey Byrne |
ICRA | 2 |
| 2006 | Stereo based Obstacle Detection for an Unmanned Air VehicleabstractThis paper presents the visual threat awareness (VISTA) system for real time collision obstacle detection for an unmanned air vehicle (UAV). Computational stereo performance has progressed such that several commercial or open source implementations are available which operate at frame rate, but suffer from well known correspondence errors. We show that introducing a global segmentation step after commodity stereo can increase robustness and leverage existing stereo software. The global segmentation step is based on a graph structure appropriate for collision detection, human vision inspired foveation, perceptual organization and graph partitioning using the minimum s-t graph cut. This system has been prototyped using the Sarnoff Acadia I vision processor to enable processing of 640 times 480 resolution imagery at 5-10 Hz operation on embedded avionics. We describe system theory, demonstrate segmentation results on scenes of increasing complexity, and show flight experiment results on Georgia Tech's GT-Max autonomous helicopter against real collision obstacles Jeffrey Byrne, Martin Cosgrove, Raman K. Mehra |
ICRA | 1 |