Rodrigo Benenson

dblp:89/2011 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-authorSystems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
22 papers
Image recognition and object detection · 29% Segmentation and scene understanding · 28% Face, body and person analysis · 19%
Network and information security
2 papers
Privacy and data protection · 87% Cryptographic protocols and secure computation · 13%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
pedestrian detection
1.882018
Towards Reaching Human Performance in Pedestrian Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018
CityPersons: A Diverse Dataset for Pedestrian Detection · CVPR 2017
How Far are We from Solving Pedestrian Detection? · CVPR 2016
Computer vision › Image recognition and object detection
object detection
1.062018
Towards Reaching Human Performance in Pedestrian Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Learning Non-maximum Suppression · CVPR 2017
Seeking the Strongest Rigid Detector · CVPR 2013
Computer vision › Segmentation and scene understanding
instance segmentation
0.932019
Large-Scale Interactive Object Segmentation With Human Annotators · CVPR 2019
Simple Does It: Weakly Supervised Instance and Semantic Segmentation · CVPR 2017
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Computer vision › Face, body and person analysis
person identification
0.932020
Person Recognition in Personal Photo Collections · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Faceless Person Recognition: Privacy Implications in Social Media · ECCV (3) 2016
Person Recognition in Personal Photo Collections · ICCV 2015
Computer vision › Segmentation and scene understanding
semantic segmentation
0.832017
Exploiting Saliency for Object Segmentation from Image Level Labels · CVPR 2017
Simple Does It: Weakly Supervised Instance and Semantic Segmentation · CVPR 2017
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Computer vision › Video understanding and tracking
video object segmentation
0.722019
Lucid Data Dreaming for Video Object Segmentation · Int. J. Comput. Vis. 2019
Learning Video Object Segmentation from Static Images · CVPR 2017
Computer vision › Face, body and person analysis › person identification
person recognition in photo collections
0.722020
Person Recognition in Personal Photo Collections · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Person Recognition in Personal Photo Collections · ICCV 2015
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation
0.622017
Exploiting Saliency for Object Segmentation from Image Level Labels · CVPR 2017
Simple Does It: Weakly Supervised Instance and Semantic Segmentation · CVPR 2017
Computer vision › Face, body and person analysis
face recognition
0.412020
Person Recognition in Personal Photo Collections · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Computer vision › Face, body and person analysis
person re-identification
0.412020
Person Recognition in Personal Photo Collections · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Computer vision › 3D vision
stereo vision
0.422016
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Pedestrian detection at 100 frames per second · CVPR 2012
Computer vision › Segmentation and scene understanding
interactive segmentation
0.412019
Large-Scale Interactive Object Segmentation With Human Annotators · CVPR 2019
Machine learning › Deep learning architectures and training › data-centric deep learning
training data generation
0.412019
Lucid Data Dreaming for Video Object Segmentation · Int. J. Comput. Vis. 2019
Human-AI interaction › human-in-the-loop
human-in-the-loop annotation
0.412019
Large-Scale Interactive Object Segmentation With Human Annotators · CVPR 2019
Natural language and speech › Information extraction and text analysis › dataset construction
benchmark dataset construction
0.312017
CityPersons: A Diverse Dataset for Pedestrian Detection · CVPR 2017
Computer vision › Image recognition and object detection › object detection › deep learning object detection
CNN-based detection
0.312017
CityPersons: A Diverse Dataset for Pedestrian Detection · CVPR 2017
Computer vision › Segmentation and scene understanding › semantic segmentation › weakly supervised semantic segmentation
image-level label segmentation
0.312017
Exploiting Saliency for Object Segmentation from Image Level Labels · CVPR 2017
Computer vision › Image recognition and object detection › object detection › object detection post-processing
non-maximum suppression
0.312017
Learning Non-maximum Suppression · CVPR 2017
Computer vision › Segmentation and scene understanding
object segmentation
0.312017
Exploiting Saliency for Object Segmentation from Image Level Labels · CVPR 2017
Computer vision › Video understanding and tracking
object tracking
0.312017
Learning Video Object Segmentation from Static Images · CVPR 2017
Computer vision › Segmentation and scene understanding › instance segmentation
weakly supervised instance segmentation
0.312017
Simple Does It: Weakly Supervised Instance and Semantic Segmentation · CVPR 2017
Computer vision › Segmentation and scene understanding › boundary detection
object boundary detection
0.212016
Weakly Supervised Object Boundaries · CVPR 2016
Computer vision › Segmentation and scene understanding › dense prediction
pixel labeling
0.212016
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Computer vision › 3D vision
stereo video dataset
0.212016
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Robotics › Autonomous driving
urban scene understanding
0.212016
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Machine learning › Deep learning architectures and training
convolutional neural network
0.212015
Taking a deeper look at pedestrians · CVPR 2015
Computer vision › Face, body and person analysis
face detection
0.212014
Face Detection without Bells and Whistles · ECCV (4) 2014
Machine learning › Learning paradigms › supervised learning
classifier training
0.212013
Handling Occlusions with Franken-Classifiers · ICCV 2013
Computer vision › Image recognition and object detection › object detection › part-based object detection
deformable part model
0.212013
Seeking the Strongest Rigid Detector · CVPR 2013
Computer vision › Image recognition and object detection › pedestrian detection
occluded pedestrian detection
0.212013
Handling Occlusions with Franken-Classifiers · ICCV 2013

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 1.0deep interactive segmentation · 0.8human baseline · 0.6optical flow · 0.4convolutional neural network features · 0.4DeepID2+ · 0.4in-domain training data synthesis · 0.4data dreaming · 0.4convolutional network · 0.4annotation sanitization · 0.3social media analysis · 0.2short-range radio · 0.2secure multiparty computation · 0.2person recognition · 0.2
YearPublicationVenuePosition
2020 Person Recognition in Personal Photo Collections
abstract
People nowadays share large parts of their personal lives through social media. Being able to automatically recognise people in personal photos may greatly enhance user convenience by easing photo album organisation. For human identification task, however, traditional focus of computer vision has been face recognition and pedestrian re-identification. Person recognition in social media photos sets new challenges for computer vision, including non-cooperative subjects (e.g. backward viewpoints, unusual poses) and great changes in appearance. To tackle this problem, we build a simple person recognition framework that leverages convnet features from multiple image regions (head, body, etc.). We propose new recognition scenarios that focus on the time and appearance gap between training and testing samples. We present an in-depth analysis of the importance of different features according to time and viewpoint generalisability. In the process, we verify that our simple approach achieves the state of the art result on the PIPA [1] benchmark, arguably the largest social media based benchmark for person recognition to date with diverse poses, viewpoints, social groups, and events. Compared the conference version of the paper [2], this paper additionally presents (1) analysis of a face recogniser (DeepID2+ [3]), (2) new method naeil2 that combines the conference version method naeil and DeepID2+ to achieve state of the art results even compared to post-conference works, (3) discussion of related work since the conference version, (4) additional analysis including the head viewpoint-wise breakdown of performance, and (5) results on the open-world setup.
Seong Joon Oh, Rodrigo Benenson, Mario Fritz, Bernt Schiele
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Large-Scale Interactive Object Segmentation With Human Annotators
abstract
Manually annotating object segmentation masks is very time consuming. Interactive object segmentation methods offer a more efficient alternative where a human annotator and a machine segmentation model collaborate. In this paper we make several contributions to interactive segmentation: (1) we systematically explore in simulation the design space of deep interactive segmentation models and report new insights and caveats; (2) we execute a large-scale annotation campaign with real human annotators, producing masks for 2.5M instances on the OpenImages dataset. We released this data publicly, forming the largest existing dataset for instance segmentation. Moreover, by re-annotating part of the COCO dataset, we show that we can produce instance masks 3x faster than traditional polygon drawing tools while also providing better quality. (3) We present a technique for automatically estimating the quality of the produced masks which exploits indirect signals from the annotation process.
Rodrigo Benenson, Stefan Popov, Vittorio Ferrari
CVPR1
2019 Lucid Data Dreaming for Video Object Segmentation
abstract
Convolutional networks reach top quality in pixel-level video object segmentation but require a large amount of training data (1k–100k) to deliver such results. We propose a new training strategy which achieves state-of-the-art results across three evaluation datasets while using \(20\,\times \) – \(1000\,\times \) less annotated data than competing methods. Our approach is suitable for both single and multiple object segmentation. Instead of using large training sets hoping to generalize across domains, we generate in-domain training data using the provided annotation on the first frame of each video to synthesize—“lucid dream” (in a lucid dream the sleeper is aware that he or she is dreaming and is sometimes able to control the course of the dream)—plausible future video frames. In-domain per-video training data allows us to train high quality appearance- and motion-based models, as well as tune the post-processing stage. This approach allows to reach competitive results even when training from only a single annotated frame, without ImageNet pre-training. Our results indicate that using a larger training set is not automatically better, and that for the video object segmentation task a smaller training set that is closer to the target domain is more effective. This changes the mindset regarding how many training samples and general “objectness” knowledge are required for the video object segmentation task.
Anna Khoreva, Rodrigo Benenson, Eddy Ilg, Thomas Brox, Bernt Schiele
Int. J. Comput. Vis.2
2018 Towards Reaching Human Performance in Pedestrian Detection
abstract
Encouraged by the recent progress in pedestrian detection, we investigate the gap between current state-of-the-art methods and the "perfect single frame detector". We enable our analysis by creating a human baseline for pedestrian detection (over the Caltech pedestrian dataset). After manually clustering the frequent errors of a top detector, we characterise both localisation and background-versus-foreground errors. To address localisation errors we study the impact of training annotation noise on the detector performance, and show that we can improve results even with a small portion of sanitised training data. To address background/foreground discrimination, we study convnets for pedestrian detection, and discuss which factors affect their performance. Other than our in-depth analysis, we report top performance on the Caltech pedestrian dataset, and provide a new sanitised set of training and test annotations.
Shanshan Zhang 0001, Rodrigo Benenson, Mohamed Omran, Jan Hosang, Bernt Schiele
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Simple Does It: Weakly Supervised Instance and Semantic Segmentation
abstract
Semantic labelling and instance segmentation are two tasks that require particularly costly annotations. Starting from weak supervision in the form of bounding box detection annotations, we propose a new approach that does not require modification of the segmentation training procedure. We show that when carefully designing the input labels from given bounding boxes, even a single round of training is enough to improve over previously reported weakly supervised results. Overall, our weak supervision approach reaches ~95% of the quality of the fully supervised model, both for semantic labelling and instance segmentation.
Anna Khoreva, Rodrigo Benenson, Jan Hosang, Matthias Hein 0001, Bernt Schiele
CVPR2
2017 Exploiting Saliency for Object Segmentation from Image Level Labels
abstract
There have been remarkable improvements in the semantic labelling task in the recent years. However, the state of the art methods rely on large-scale pixel-level annotations. This paper studies the problem of training a pixel-wise semantic labeller network from image-level annotations of the present object classes. Recently, it has been shown that high quality seeds indicating discriminative object regions can be obtained from image-level labels. Without additional information, obtaining the full extent of the object is an inherently ill-posed problem due to co-occurrences. We propose using a saliency model as additional information and hereby exploit prior knowledge on the object extent and image statistics. We show how to combine both information sources in order to recover 80% of the fully supervised performance - which is the new state of the art in weakly supervised training for pixel-wise semantic labelling.
Seong Joon Oh, Rodrigo Benenson, Anna Khoreva, Zeynep Akata, Mario Fritz, Bernt Schiele
CVPR2
2017 Learning Video Object Segmentation from Static Images
abstract
Inspired by recent advances of deep learning in instance segmentation and object tracking, we introduce the concept of convnet-based guidance applied to video object segmentation. Our model proceeds on a per-frame basis, guided by the output of the previous frame towards the object of interest in the next frame. We demonstrate that highly accurate object segmentation in videos can be enabled by using a convolutional neural network (convnet) trained with static images only. The key component of our approach is a combination of offline and online learning strategies, where the former produces a refined mask from the previous frame estimate and the latter allows to capture the appearance of the specific object instance. Our method can handle different types of input annotations such as bounding boxes and segments while leveraging an arbitrary amount of annotated frames. Therefore our system is suitable for diverse applications with different requirements in terms of accuracy and efficiency. In our extensive evaluation, we obtain competitive results on three different datasets, independently from the type of input annotation.
Federico Perazzi, Anna Khoreva, Rodrigo Benenson, Bernt Schiele, Alexander Sorkine-Hornung
CVPR3
2017 CityPersons: A Diverse Dataset for Pedestrian Detection
abstract
Convnets have enabled significant progress in pedestrian detection recently, but there are still open questions regarding suitable architectures and training data. We revisit CNN design and point out key adaptations, enabling plain FasterRCNN to obtain state-of-the-art results on the Caltech dataset. To achieve further improvement from more and better data, we introduce CityPersons, a new set of person annotations on top of the Cityscapes dataset. The diversity of CityPersons allows us for the first time to train one single CNN model that generalizes well over multiple benchmarks. Moreover, with additional training with CityPersons, we obtain top results using FasterRCNN on Caltech, improving especially for more difficult cases (heavy occlusion and small scale) and providing higher localization quality.
Shanshan Zhang 0001, Rodrigo Benenson, Bernt Schiele
CVPR2
2017 Learning Non-maximum Suppression
abstract
Object detectors have hugely profited from moving towards an end-to-end learning paradigm: proposals, features, and the classifier becoming one neural network improved results two-fold on general object detection. One indispensable component is non-maximum suppression (NMS), a post-processing algorithm responsible for merging all detections that belong to the same object. The de facto standard NMS algorithm is still fully hand-crafted, suspiciously simple, and - being based on greedy clustering with a fixed distance threshold - forces a trade-off between recall and precision. We propose a new network architecture designed to perform NMS, using only boxes and their score. We report experiments for person detection on PETS and for general object categories on the COCO dataset. Our approach shows promise providing improved localization and occlusion handling.
Jan Hosang, Rodrigo Benenson, Bernt Schiele
CVPR2
2016 The Cityscapes Dataset for Semantic Urban Scene Understanding
abstract
Visual understanding of complex urban street scenes is an enabling factor for a wide range of applications. Object detection has benefited enormously from large-scale datasets, especially in the context of deep learning. For semantic urban scene understanding, however, no current dataset adequately captures the complexity of real-world urban scenes. To address this, we introduce Cityscapes, a benchmark suite and large-scale dataset to train and test approaches for pixel-level and instance-level semantic labeling. Cityscapes is comprised of a large, diverse set of stereo video sequences recorded in streets from 50 different cities. 5000 of these images have high quality pixel-level annotations, 20 000 additional images have coarse annotations to enable methods that leverage large volumes of weakly-labeled data. Crucially, our effort exceeds previous attempts in terms of dataset size, annotation richness, scene variability, and complexity. Our accompanying empirical study provides an in-depth analysis of the dataset characteristics, as well as a performance evaluation of several state-of-the-art approaches based on our benchmark.
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth 0001, Bernt Schiele
CVPR6
2016 Weakly Supervised Object Boundaries
abstract
State-of-the-art learning based boundary detection methods require extensive training data. Since labelling object boundaries is one of the most expensive types of annotations, there is a need to relax the requirement to carefully annotate images to make both the training more affordable and to extend the amount of training data. In this paper we propose a technique to generate weakly supervised annotations and show that bounding box annotations alone suffice to reach high-quality object boundaries without using any object-specific boundary annotations. With the proposed weak supervision techniques we achieve the top performance on the object boundary detection task, outperforming by a large margin the current fully supervised state-of-theart methods.
Anna Khoreva, Rodrigo Benenson, Mohamed Omran, Matthias Hein 0001, Bernt Schiele
CVPR2
2016 How Far are We from Solving Pedestrian Detection?
abstract
Encouraged by the recent progress in pedestrian detection, we investigate the gap between current state-of-the-art methods and the "perfect single frame detector". We enable our analysis by creating a human baseline for pedestrian detection (over the Caltech dataset), and by manually clustering the recurrent errors of a top detector. Our results characterise both localisation and background-versusforeground errors. To address localisation errors we study the impact of training annotation noise on the detector performance, and show that we can improve even with a small portion of sanitised training data. To address background/foreground discrimination, we study convnets for pedestrian detection, and discuss which factors affect their performance. Other than our in-depth analysis, we report top performance on the Caltech dataset, and provide a new sanitised set of training and test annotations.
Shanshan Zhang 0001, Rodrigo Benenson, Mohamed Omran, Jan Hosang, Bernt Schiele
CVPR2
2016 Faceless Person Recognition: Privacy Implications in Social Media
Seong Joon Oh, Rodrigo Benenson, Mario Fritz, Bernt Schiele
ECCV (3)2
2016 I-Pic: A Platform for Privacy-Compliant Image Capture
abstract
The ubiquity of portable mobile devices equipped with built-in cameras have led to a transformation in how and when digital images are captured, shared, and archived. Photographs and videos from social gatherings, public events, and even crime scenes are commonplace online. While the spontaneity afforded by these devices have led to new personal and creative outlets, privacy concerns of bystanders (and indeed, in some cases, unwilling subjects) have remained largely unaddressed. We present I-Pic, a trusted software platform that integrates digital capture with user-defined privacy. In I-Pic, users choose alevel of privacy (e.g., image capture allowed or not) based upon social context (e.g., out in public vs. with friends vs. at workplace). Privacy choices of nearby users are advertised via short-range radio, and I-Pic-compliant capture platforms generate edited media to conform to privacy choices of image subjects. I-Pic uses secure multiparty computation to ensure that users' visual features and privacy choices are not revealed publicly, regardless of whether they are the subjects of an image capture. Just as importantly, I-Pic preserves the ease-of-use and spontaneous nature of capture and sharing between trusted users. Our evaluation of I-Pic shows that a practical, energy-efficient system that conforms to the privacy choices of many users within a scene can be built and deployed using current hardware.
Paarijaat Aditya, Rijurekha Sen, Peter Druschel, Seong Joon Oh, Rodrigo Benenson, Mario Fritz, Bernt Schiele, Bobby Bhattacharjee, Tong Tong Wu
MobiSys5
2016 What Makes for Effective Detection Proposals?
abstract
Current top performing object detectors employ detection proposals to guide the search for objects, thereby avoiding exhaustive sliding window search across images. Despite the popularity and widespread use of detection proposals, it is unclear which trade-offs are made when using them during object detection. We provide an in-depth analysis of twelve proposal methods along with four baselines regarding proposal repeatability, ground truth annotation recall on PASCAL, ImageNet, and MS COCO, and their impact on DPM, R-CNN, and Fast R-CNN detection performance. Our analysis shows that for object detection improving proposal localisation accuracy is as important as improving recall. We introduce a novel metric, the average recall (AR), which rewards both high recall and good localisation and correlates surprisingly well with detection performance. Our findings show common strengths and weaknesses of existing methods, and provide insights and metrics for selecting and tuning proposal methods.
Jan Hosang, Rodrigo Benenson, Piotr Dollár, Bernt Schiele
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Taking a deeper look at pedestrians
abstract
In this paper we study the use of convolutional neural networks (convnets) for the task of pedestrian detection. Despite their recent diverse successes, convnets historically underperform compared to other pedestrian detectors. We deliberately omit explicitly modelling the problem into the network (e.g. parts or occlusion modelling) and show that we can reach competitive performance without bells and whistles. In a wide range of experiments we analyse small and big convnets, their architectural choices, parameters, and the influence of different training data, including pretraining on surrogate tasks. We present the best convnet detectors on the Caltech and KITTI dataset. On Caltech our convnets reach top performance both for the Caltech1x and Caltech10x training setup. Using additional data at training time our strongest convnet model is competitive even to detectors that use additional data (optical flow) at test time.
Jan Hosang, Mohamed Omran, Rodrigo Benenson, Bernt Schiele
CVPR3
2015 Filtered channel features for pedestrian detection
abstract
This paper starts from the observation that multiple top performing pedestrian detectors can be modelled by using an intermediate layer filtering low-level features in combination with a boosted decision forest. Based on this observation we propose a unifying framework and experimentally explore different filter families. We report extensive results enabling a systematic analysis. Using filtered channel features we obtain top performance on the challenging Caltech and KITTI datasets, while using only HOG+LUV as low-level features. When adding optical flow features we further improve detection quality and report the best known results on the Caltech dataset, reaching 93% recall at 1 FPPI.
Shanshan Zhang 0001, Rodrigo Benenson, Bernt Schiele
CVPR2
2015 Person Recognition in Personal Photo Collections
abstract
Recognising persons in everyday photos presents major challenges (occluded faces, different clothing, locations, etc.) for machine vision. We propose a convnet based person recognition system on which we provide an in-depth analysis of informativeness of different body cues, impact of training data, and the common failure modes of the system. In addition, we discuss the limitations of existing benchmarks and propose more challenging ones. Our method is simple and is built on open source and open data, yet it improves the state of the art results on a large dataset of social media photos (PIPA).
Seong Joon Oh, Rodrigo Benenson, Mario Fritz, Bernt Schiele
ICCV2
2015 Detecting Surgical Tools by Modelling Local Appearance and Global Shape
abstract
Detecting tools in surgical videos is an important ingredient for context-aware computer-assisted surgical systems. To this end, we present a new surgical tool detection dataset and a method for joint tool detection and pose estimation in 2d images. Our two-stage pipeline is data-driven and relaxes strong assumptions made by previous works regarding the geometry, number, and position of tools in the image. The first stage classifies each pixel based on local appearance only, while the second stage evaluates a tool-specific shape template to enforce global shape. Both local appearance and global shape are learned from training data. Our method is validated on a new surgical tool dataset of 2 476 images from neurosurgical microscopes, which is made freely available. It improves over existing datasets in size, diversity and detail of annotation. We show that our method significantly improves over competitive baselines from the computer vision field. We achieve 15% detection miss-rate at 10(-1) false positives per image (for the suction tube) over our surgical tool dataset. Results indicate that performing semantic labelling as an intermediate task is key for high quality detection.
David Bouget, Rodrigo Benenson, Mohamed Omran, Laurent Riffaud, Bernt Schiele, Pierre Jannin
IEEE Trans. Medical Imaging2
2014 How good are detection proposals, really?
Jan Hosang, Rodrigo Benenson, Bernt Schiele
BMVC2
2014 Face Detection without Bells and Whistles
Markus Mathias, Rodrigo Benenson, Marco Pedersoli, Luc Van Gool
ECCV (4)2
2013 Seeking the Strongest Rigid Detector
abstract
The current state of the art solutions for object detection describe each class by a set of models trained on discovered sub-classes (so called "components"), with each model itself composed of collections of interrelated parts (deformable models). These detectors build upon the now classic Histogram of Oriented Gradients+linear SVM combo. In this paper we revisit some of the core assumptions in HOG+SVM and show that by properly designing the feature pooling, feature selection, preprocessing, and training methods, it is possible to reach top quality, at least for pedestrian detections, using a single rigid component. Abstract We provide experiments for a large design space, that give insights into the design of classifiers, as well as relevant information for practitioners. Our best detector is fully feed-forward, has a single unified architecture, uses only histograms of oriented gradients and colour information in monocular static images, and improves over 23 other methods on the INRIA, ETH and Caltech-USA datasets, reducing the average miss-rate over HOG+SVM by more than 30%.
Rodrigo Benenson, Markus Mathias, Tinne Tuytelaars, Luc Van Gool
CVPR1
2013 Handling Occlusions with Franken-Classifiers
abstract
Detecting partially occluded pedestrians is challenging. A common practice to maximize detection quality is to train a set of occlusion-specific classifiers, each for a certain amount and type of occlusion. Since training classifiers is expensive, only a handful are typically trained. We show that by using many occlusion-specific classifiers, we outperform previous approaches on three pedestrian datasets, INRIA, ETH, and Caltech USA. We present a new approach to train such classifiers. By reusing computations among different training stages, 16 occlusion-specific classifiers can be trained at only one tenth the cost of one full training. We show that also test time cost grows sub-linearly.
Markus Mathias, Rodrigo Benenson, Radu Timofte, Luc Van Gool
ICCV2
2013 Traffic sign recognition - How far are we from the solution?
abstract
Traffic sign recognition has been a recurring application domain for visual objects detection. The public datasets have only recently reached large enough size and variety to enable proper empirical studies. We revisit the topic by showing how modern methods perform on two large detection and classification datasets (thousand of images, tens of categories) captured in Belgium and Germany. We show that, without any application specific modification, existing methods for pedestrian detection, and for digit and face classification; can reach performances in the range of 95% ~ 99% of the perfect solution. We show detailed experiments and discuss the trade-off of different options. Our top performing methods use modern variants of HOG features for detection, and sparse representations for classification.
Markus Mathias, Radu Timofte, Rodrigo Benenson, Luc Van Gool
IJCNN3
2012 Pedestrian detection at 100 frames per second
abstract
We present a new pedestrian detector that improves both in speed and quality over state-of-the-art. By efficiently handling different scales and transferring computation from test time to training time, detection speed is improved. When processing monocular images, our system provides high quality detections at 50 fps. We also propose a new method for exploiting geometric context extracted from stereo images. On a single CPU+GPU desktop machine, we reach 135 fps, when processing street scenes, from rectified input to detections output.
Rodrigo Benenson, Markus Mathias, Radu Timofte, Luc Van Gool
CVPR1
2012 Stixels Motion Estimation without Optical Flow Computation
Bertan Günyel, Rodrigo Benenson, Radu Timofte, Luc Van Gool
ECCV (6)2
2008 Achievable safety of driverless ground vehicles
abstract
Safety is an important issue of driverless car. Yet, most current approaches fail to ensure safety even in a fully informed situation. In this paper we discuss how the safety criteria apply when the robot uses its on board sensors to evolve in an environment populated with static and moving obstacles. The sensors can only provide a partial and uncertain knowledge of the surroundings. We show that the usual safety notion does not apply for this relevant case and discuss which safety guarantees can be given and how to achieve them.
Rodrigo Benenson, Thierry Fraichard, Michel Parent
ICARCV1
2006 Integrating Perception and Planning for Autonomous Navigation of Urban Vehicles
abstract
The paper addresses the problem of autonomous navigation of a car-like robot evolving in an urban environment. Such an environment exhibits an heterogeneous geometry and is cluttered with moving obstacles. Furthermore, in this context, motion safety is a critical issue. The proposed approach to the problem lies in the coupling of two crucial robotic capabilities, namely perception and planning. The main contributions of this work are the development and integration of these modules into one single application, considering explicitly the constraints related to the environment and the system
Rodrigo Benenson, Stéphane Petti, Thierry Fraichard, Michel Parent
IROS1