Garrick Brazil

dblp:202/2306 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 55% Image recognition and object detection · 28% Segmentation and scene understanding · 7%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
monocular 3d object detection
2.142023
Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild · CVPR 2023
DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection · ECCV (9) 2022
GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection · CVPR 2021
Computer vision › 3D vision
3d object detection
2.042023
Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild · CVPR 2023
GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection · CVPR 2021
Kinematic 3D Object Detection in Monocular Video · ECCV (23) 2020
Computer vision › Image recognition and object detection
object detection
1.332023
Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild · CVPR 2023
M3D-RPN: Monocular 3D Region Proposal Network for Object Detection · ICCV 2019
Illuminating Pedestrians via Simultaneous Detection and Segmentation · ICCV 2017
Computer vision › 3D vision
depth estimation
1.022022
DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection · ECCV (9) 2022
The Edge of Depth: Explicit Constraints Between Segmentation and Depth · CVPR 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
0.722020
The Edge of Depth: Explicit Constraints Between Segmentation and Depth · CVPR 2020
Illuminating Pedestrians via Simultaneous Detection and Segmentation · ICCV 2017
Computer vision › Image recognition and object detection
pedestrian detection
0.722019
Pedestrian Detection With Autoregressive Network Phases · CVPR 2019
Illuminating Pedestrians via Simultaneous Detection and Segmentation · ICCV 2017
Machine learning › Transfer learning and domain adaptation
domain generalization
0.712023
Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild · CVPR 2023
Computer vision › Image recognition and object detection › object detection › object detection post-processing
non-maximum suppression
0.512021
GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection · CVPR 2021
Computer vision › 3D vision › depth estimation › self-supervised depth estimation
self-supervised monocular depth estimation
0.412020
The Edge of Depth: Explicit Constraints Between Segmentation and Depth · CVPR 2020
Computer vision › Image recognition and object detection › object detection
cascaded detection
0.412019
Pedestrian Detection With Autoregressive Network Phases · CVPR 2019
Machine learning › Deep learning architectures and training
convolutional neural network
0.412019
Pedestrian Detection With Autoregressive Network Phases · CVPR 2019
Computer vision › 3D vision
monocular video
0.112020
Kinematic 3D Object Detection in Monocular Video · ECCV (23) 2020

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 0.7Cube R-CNN · 0.7geometric equivariance · 0.6depth equivariant network · 0.6end-to-end training · 0.5differentiable NMS · 0.5kinematic modeling · 0.4greedy iterative supervision · 0.4border consistency constraint · 0.4cascaded detection · 0.4
YearPublicationVenuePosition
2023 Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
abstract
Recognizing scenes and objects in 3D from a single image is a longstanding goal of computer vision with applications in robotics and AR/VR. For 2D recognition, large datasets and scalable solutions have led to unprecedented advances. In 3D, existing benchmarks are small in size and approaches specialize in few object categories and specific domains, e.g. urban driving scenes. Motivated by the success of 2D recognition, we revisit the task of 3D object detection by introducing a large benchmark, called Omni3D. Omni3DRE-purposes and combines existing datasets resulting in 234k images annotated with more than 3 million instances and 98 categories. 3D detection at such scale is challenging due to variations in camera intrinsics and the rich diversity of scene and object types. We propose a model, called Cube R-CNN, designed to generalize across camera and scene types with a unified approach. We show that Cube R-CNN outperforms prior works on the larger Omni3D and existing benchmarks. Finally, we prove that Omni3D is a powerful dataset for 3D object recognition and show that it improves single-dataset performance and can accelerate learning on new smaller datasets via pre-training.11We release the Omni3D benchmark and Cube R-CNN models at https://github.com/facebookresearch/omni3d.
Garrick Brazil, Abhinav Kumar 0004, Julian Straub, Nikhila Ravi, Justin Johnson 0001, Georgia Gkioxari
CVPR1
2023 Camera Self-Calibration Using Human Faces
abstract
Despite recent advancements in depth estimation and face alignment, it remains difficult to predict the distance to a human face in arbitrary videos due to the lack of camera calibration. A typical pipeline is to perform calibration with a checkerboard before the video capture, but this is inconvenient to users or impossible for unknown cameras. This work proposes to use the human face as the calibration object to estimate metric depth information and camera intrinsics. Our novel approach alternates between optimizing the 3D face and the camera intrinsics parameterized by a neural network. Compared to prior work, our method performs camera calibration on a larger variety of videos captured by unknown cameras. Further, due to the face prior, our method is more robust to noise in 2D observations compared to previous self-calibration methods. We show that our method improves calibration and depth prediction accuracy over prior works on both synthetic and real data. Code will be available at https://github.com/yhu9/FaceCalibration.
Masa Hu, Garrick Brazil, Nanxiang Li, Liu Ren 0001, Xiaoming Liu 0002
FG2
2022 DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection
Abhinav Kumar 0004, Garrick Brazil, Enrique Corona, Armin Parchami, Xiaoming Liu 0002
ECCV (9)2
2021 GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection
abstract
Modern 3D object detectors have immensely benefited from the end-to-end learning idea. However, most of them use a post-processing algorithm called Non-Maximal Suppression (NMS) only during inference. While there were attempts to include NMS in the training pipeline for tasks such as 2D object detection, they have been less widely adopted due to a non-mathematical expression of the NMS. In this paper, we present and integrate GrooMeD-NMS – a novel Grouped Mathematically Differentiable NMS for monocular 3D object detection, such that the network is trained end-to-end with a loss on the boxes after NMS. We first formulate NMS as a matrix operation and then group and mask the boxes in an unsupervised manner to obtain a simple closed-form expression of the NMS. GrooMeD-NMS addresses the mismatch between training and inference pipelines and, therefore, forces the network to select the best 3D box in a differentiable manner. As a result, GrooMeD-NMS achieves state-of-the-art monocular 3D object detection results on the KITTI benchmark dataset performing comparably to monocular video-based methods.
Abhinav Kumar 0004, Garrick Brazil, Xiaoming Liu 0002
CVPR2
2020 The Edge of Depth: Explicit Constraints Between Segmentation and Depth
abstract
In this work we study the mutual benefits of two common computer vision tasks, self-supervised depth estimation and semantic segmentation from images. For example, to help unsupervised monocular depth estimation, constraint from semantic segmentation has been explored implicitly such as sharing and transforming features. In contrast, we propose to explicitly measure the border consistency between segmentation and depth and minimize it in a greedy manner by iteratively supervising the network towards a locally optimal solution. Partially this is motivated by our observation that semantic segmentation even trained with limited ground truth (200 images of KITTI) can offer more accurate border than that of any (monocular or stereo) image-based depth estimation. Through extensive experiments, our proposed approach advance the state of the art on unsupervised monocular depth estimation in the KITTI benchmark.
Garrick Brazil, Xiaoming Liu 0002
CVPR2
2020 Kinematic 3D Object Detection in Monocular Video
Garrick Brazil, Gerard Pons-Moll, Xiaoming Liu 0002, Bernt Schiele
ECCV (23)1
2019 Pedestrian Detection With Autoregressive Network Phases
abstract
We present an autoregressive pedestrian detection framework with cascaded phases designed to progressively improve precision. The proposed framework utilizes a novel lightweight stackable decoder-encoder module which uses convolutional re-sampling layers to improve features while maintaining efficient memory and runtime cost. Unlike previous cascaded detection systems, our proposed framework is designed within a region proposal network and thus retains greater context of nearby detections compared to independently processed RoI systems. We explicitly encourage increasing levels of precision by assigning strict labeling policies to each consecutive phase such that early phases develop features primarily focused on achieving high recall and later on accurate precision. In consequence, the final feature maps form more peaky radial gradients emulating from the centroids of unique pedestrians. Using our proposed autoregressive framework leads to new state-of-the-art performance on the reasonable and occlusion settings of the Caltech pedestrian dataset, and achieves competitive state-of-the-art performance on the KITTI dataset.
Garrick Brazil, Xiaoming Liu 0002
CVPR1
2019 M3D-RPN: Monocular 3D Region Proposal Network for Object Detection
abstract
Understanding the world in 3D is a critical component of urban autonomous driving. Generally, the combination of expensive LiDAR sensors and stereo RGB imaging has been paramount for successful 3D object detection algorithms, whereas monocular image-only methods experience drastically reduced performance. We propose to reduce the gap by reformulating the monocular 3D detection problem as a standalone 3D region proposal network. We leverage the geometric relationship of 2D and 3D perspectives, allowing 3D boxes to utilize well-known and powerful convolutional features generated in the image-space. To help address the strenuous 3D parameter estimations, we further design depth-aware convolutional layers which enable location specific feature development and in consequence improved 3D scene understanding. Compared to prior work in monocular 3D detection, our method consists of only the proposed 3D region proposal network rather than relying on external networks, data, or multiple stages. M3D-RPN is able to significantly improve the performance of both monocular 3D Object Detection and Bird's Eye View tasks within the KITTI urban autonomous driving dataset, while efficiently using a shared multi-class model.
Garrick Brazil, Xiaoming Liu 0002
ICCV1
2019 Recurrent Flow-Guided Semantic Forecasting
abstract
Understanding the world around us and making decisions about the future is a critical component to human intelligence. As autonomous systems continue to develop, their ability to reason about the future will be the key to their success. Semantic anticipation is a relatively under-explored area for which autonomous vehicles could take advantage of (e.g., forecasting pedestrian trajectories). Motivated by the need for real-time prediction in autonomous systems, we propose to decompose the challenging semantic forecasting task into two subtasks: current frame segmentation and future optical flow prediction. Through this decomposition, we built an efficient, effective, low overhead model with three main components: flow prediction network, feature-flow aggregation LSTM, and end-to-end learnable warp layer. Our proposed method achieves state-of-the-art accuracy on short-term and moving objects semantic forecasting while simultaneously reducing model parameters by up to 95% and increasing efficiency by greater than 40x.
Adam M. Terwilliger, Garrick Brazil, Xiaoming Liu 0002
WACV2
2017 Illuminating Pedestrians via Simultaneous Detection and Segmentation
abstract
Pedestrian detection is a critical problem in computer vision with significant impact on safety in urban autonomous driving. In this work, we explore how semantic segmentation can be used to boost pedestrian detection accuracy while having little to no impact on network efficiency. We propose a segmentation infusion network to enable joint supervision on semantic segmentation and pedestrian detection. When placed properly, the additional supervision helps guide features in shared layers to become more sophisticated and helpful for the downstream pedestrian detector. Using this approach, we find weakly annotated boxes to be sufficient for considerable performance gains. We provide an in-depth analysis to demonstrate how shared layers are shaped by the segmentation supervision. In doing so, we show that the resulting feature maps become more semantically meaningful and robust to shape and occlusion. Overall, our simultaneous detection and segmentation framework achieves a considerable gain over the state-of-the-art on the Caltech pedestrian dataset, competitive performance on KITTI, and executes 2 × faster than competitive methods.
Garrick Brazil, Xi Yin 0001, Xiaoming Liu 0002
ICCV1