EDBT 2026 Demo / reviewers in the wild / expert
Garrick Brazil
dblp:202/2306
· DBLP profile ↗
10ranked-venue papers
5as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
3D vision · 55% Image recognition and object detection · 28% Segmentation and scene understanding · 7% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
monocular 3d object detection |
2.1 | 4 | 2023 | Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild · CVPR 2023 DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection · ECCV (9) 2022 GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection · CVPR 2021 |
Computer vision › 3D vision
3d object detection |
2.0 | 4 | 2023 | Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild · CVPR 2023 GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection · CVPR 2021 Kinematic 3D Object Detection in Monocular Video · ECCV (23) 2020 |
Computer vision › Image recognition and object detection
object detection |
1.3 | 3 | 2023 | Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild · CVPR 2023 M3D-RPN: Monocular 3D Region Proposal Network for Object Detection · ICCV 2019 Illuminating Pedestrians via Simultaneous Detection and Segmentation · ICCV 2017 |
Computer vision › 3D vision
depth estimation |
1.0 | 2 | 2022 | DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection · ECCV (9) 2022 The Edge of Depth: Explicit Constraints Between Segmentation and Depth · CVPR 2020 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.7 | 2 | 2020 | The Edge of Depth: Explicit Constraints Between Segmentation and Depth · CVPR 2020 Illuminating Pedestrians via Simultaneous Detection and Segmentation · ICCV 2017 |
Computer vision › Image recognition and object detection
pedestrian detection |
0.7 | 2 | 2019 | Pedestrian Detection With Autoregressive Network Phases · CVPR 2019 Illuminating Pedestrians via Simultaneous Detection and Segmentation · ICCV 2017 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.7 | 1 | 2023 | Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild · CVPR 2023 |
Computer vision › Image recognition and object detection › object detection › object detection post-processing
non-maximum suppression |
0.5 | 1 | 2021 | GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection · CVPR 2021 |
Computer vision › 3D vision › depth estimation › self-supervised depth estimation
self-supervised monocular depth estimation |
0.4 | 1 | 2020 | The Edge of Depth: Explicit Constraints Between Segmentation and Depth · CVPR 2020 |
Computer vision › Image recognition and object detection › object detection
cascaded detection |
0.4 | 1 | 2019 | Pedestrian Detection With Autoregressive Network Phases · CVPR 2019 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.4 | 1 | 2019 | Pedestrian Detection With Autoregressive Network Phases · CVPR 2019 |
Computer vision › 3D vision
monocular video |
0.1 | 1 | 2020 | Kinematic 3D Object Detection in Monocular Video · ECCV (23) 2020 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 0.7Cube R-CNN · 0.7geometric equivariance · 0.6depth equivariant network · 0.6end-to-end training · 0.5differentiable NMS · 0.5kinematic modeling · 0.4greedy iterative supervision · 0.4border consistency constraint · 0.4cascaded detection · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Omni3D: A Large Benchmark and Model for 3D Object Detection in the WildabstractRecognizing scenes and objects in 3D from a single image is a longstanding goal of computer vision with applications in robotics and AR/VR. For 2D recognition, large datasets and scalable solutions have led to unprecedented advances. In 3D, existing benchmarks are small in size and approaches specialize in few object categories and specific domains, e.g. urban driving scenes. Motivated by the success of 2D recognition, we revisit the task of 3D object detection by introducing a large benchmark, called Omni3D. Omni3DRE-purposes and combines existing datasets resulting in 234k images annotated with more than 3 million instances and 98 categories. 3D detection at such scale is challenging due to variations in camera intrinsics and the rich diversity of scene and object types. We propose a model, called Cube R-CNN, designed to generalize across camera and scene types with a unified approach. We show that Cube R-CNN outperforms prior works on the larger Omni3D and existing benchmarks. Finally, we prove that Omni3D is a powerful dataset for 3D object recognition and show that it improves single-dataset performance and can accelerate learning on new smaller datasets via pre-training.11We release the Omni3D benchmark and Cube R-CNN models at https://github.com/facebookresearch/omni3d. Garrick Brazil, Abhinav Kumar 0004, Julian Straub, Nikhila Ravi, Justin Johnson 0001, Georgia Gkioxari |
CVPR | 1 |
| 2023 | Camera Self-Calibration Using Human FacesabstractDespite recent advancements in depth estimation and face alignment, it remains difficult to predict the distance to a human face in arbitrary videos due to the lack of camera calibration. A typical pipeline is to perform calibration with a checkerboard before the video capture, but this is inconvenient to users or impossible for unknown cameras. This work proposes to use the human face as the calibration object to estimate metric depth information and camera intrinsics. Our novel approach alternates between optimizing the 3D face and the camera intrinsics parameterized by a neural network. Compared to prior work, our method performs camera calibration on a larger variety of videos captured by unknown cameras. Further, due to the face prior, our method is more robust to noise in 2D observations compared to previous self-calibration methods. We show that our method improves calibration and depth prediction accuracy over prior works on both synthetic and real data. Code will be available at https://github.com/yhu9/FaceCalibration. Masa Hu, Garrick Brazil, Nanxiang Li, Liu Ren 0001, Xiaoming Liu 0002 |
FG | 2 |
| 2022 | DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection
Abhinav Kumar 0004, Garrick Brazil, Enrique Corona, Armin Parchami, Xiaoming Liu 0002 |
ECCV (9) | 2 |
| 2021 | GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object DetectionabstractModern 3D object detectors have immensely benefited from the end-to-end learning idea. However, most of them use a post-processing algorithm called Non-Maximal Suppression (NMS) only during inference. While there were attempts to include NMS in the training pipeline for tasks such as 2D object detection, they have been less widely adopted due to a non-mathematical expression of the NMS. In this paper, we present and integrate GrooMeD-NMS – a novel Grouped Mathematically Differentiable NMS for monocular 3D object detection, such that the network is trained end-to-end with a loss on the boxes after NMS. We first formulate NMS as a matrix operation and then group and mask the boxes in an unsupervised manner to obtain a simple closed-form expression of the NMS. GrooMeD-NMS addresses the mismatch between training and inference pipelines and, therefore, forces the network to select the best 3D box in a differentiable manner. As a result, GrooMeD-NMS achieves state-of-the-art monocular 3D object detection results on the KITTI benchmark dataset performing comparably to monocular video-based methods. Abhinav Kumar 0004, Garrick Brazil, Xiaoming Liu 0002 |
CVPR | 2 |
| 2020 | The Edge of Depth: Explicit Constraints Between Segmentation and DepthabstractIn this work we study the mutual benefits of two common computer vision tasks, self-supervised depth estimation and semantic segmentation from images. For example, to help unsupervised monocular depth estimation, constraint from semantic segmentation has been explored implicitly such as sharing and transforming features. In contrast, we propose to explicitly measure the border consistency between segmentation and depth and minimize it in a greedy manner by iteratively supervising the network towards a locally optimal solution. Partially this is motivated by our observation that semantic segmentation even trained with limited ground truth (200 images of KITTI) can offer more accurate border than that of any (monocular or stereo) image-based depth estimation. Through extensive experiments, our proposed approach advance the state of the art on unsupervised monocular depth estimation in the KITTI benchmark. Garrick Brazil, Xiaoming Liu 0002 |
CVPR | 2 |
| 2020 | Kinematic 3D Object Detection in Monocular Video
Garrick Brazil, Gerard Pons-Moll, Xiaoming Liu 0002, Bernt Schiele |
ECCV (23) | 1 |
| 2019 | Pedestrian Detection With Autoregressive Network PhasesabstractWe present an autoregressive pedestrian detection framework with cascaded phases designed to progressively improve precision. The proposed framework utilizes a novel lightweight stackable decoder-encoder module which uses convolutional re-sampling layers to improve features while maintaining efficient memory and runtime cost. Unlike previous cascaded detection systems, our proposed framework is designed within a region proposal network and thus retains greater context of nearby detections compared to independently processed RoI systems. We explicitly encourage increasing levels of precision by assigning strict labeling policies to each consecutive phase such that early phases develop features primarily focused on achieving high recall and later on accurate precision. In consequence, the final feature maps form more peaky radial gradients emulating from the centroids of unique pedestrians. Using our proposed autoregressive framework leads to new state-of-the-art performance on the reasonable and occlusion settings of the Caltech pedestrian dataset, and achieves competitive state-of-the-art performance on the KITTI dataset. Garrick Brazil, Xiaoming Liu 0002 |
CVPR | 1 |
| 2019 | M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionabstractUnderstanding the world in 3D is a critical component of urban autonomous driving. Generally, the combination of expensive LiDAR sensors and stereo RGB imaging has been paramount for successful 3D object detection algorithms, whereas monocular image-only methods experience drastically reduced performance. We propose to reduce the gap by reformulating the monocular 3D detection problem as a standalone 3D region proposal network. We leverage the geometric relationship of 2D and 3D perspectives, allowing 3D boxes to utilize well-known and powerful convolutional features generated in the image-space. To help address the strenuous 3D parameter estimations, we further design depth-aware convolutional layers which enable location specific feature development and in consequence improved 3D scene understanding. Compared to prior work in monocular 3D detection, our method consists of only the proposed 3D region proposal network rather than relying on external networks, data, or multiple stages. M3D-RPN is able to significantly improve the performance of both monocular 3D Object Detection and Bird's Eye View tasks within the KITTI urban autonomous driving dataset, while efficiently using a shared multi-class model. Garrick Brazil, Xiaoming Liu 0002 |
ICCV | 1 |
| 2019 | Recurrent Flow-Guided Semantic ForecastingabstractUnderstanding the world around us and making decisions about the future is a critical component to human intelligence. As autonomous systems continue to develop, their ability to reason about the future will be the key to their success. Semantic anticipation is a relatively under-explored area for which autonomous vehicles could take advantage of (e.g., forecasting pedestrian trajectories). Motivated by the need for real-time prediction in autonomous systems, we propose to decompose the challenging semantic forecasting task into two subtasks: current frame segmentation and future optical flow prediction. Through this decomposition, we built an efficient, effective, low overhead model with three main components: flow prediction network, feature-flow aggregation LSTM, and end-to-end learnable warp layer. Our proposed method achieves state-of-the-art accuracy on short-term and moving objects semantic forecasting while simultaneously reducing model parameters by up to 95% and increasing efficiency by greater than 40x. Adam M. Terwilliger, Garrick Brazil, Xiaoming Liu 0002 |
WACV | 2 |
| 2017 | Illuminating Pedestrians via Simultaneous Detection and SegmentationabstractPedestrian detection is a critical problem in computer vision with significant impact on safety in urban autonomous driving. In this work, we explore how semantic segmentation can be used to boost pedestrian detection accuracy while having little to no impact on network efficiency. We propose a segmentation infusion network to enable joint supervision on semantic segmentation and pedestrian detection. When placed properly, the additional supervision helps guide features in shared layers to become more sophisticated and helpful for the downstream pedestrian detector. Using this approach, we find weakly annotated boxes to be sufficient for considerable performance gains. We provide an in-depth analysis to demonstrate how shared layers are shaped by the segmentation supervision. In doing so, we show that the resulting feature maps become more semantically meaningful and robust to shape and occlusion. Overall, our simultaneous detection and segmentation framework achieves a considerable gain over the state-of-the-art on the Caltech pedestrian dataset, competitive performance on KITTI, and executes 2 × faster than competitive methods. Garrick Brazil, Xi Yin 0001, Xiaoming Liu 0002 |
ICCV | 1 |