EDBT 2026 Demo / reviewers in the wild / expert
José M. Álvarez 0004
dblp:59/6703 · also José Manuel Álvarez 0001
· DBLP profile ↗
105ranked-venue papers
17as first author
58since 2021 · last 2026
0000-0002-7535-6322ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 85 · 8 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 56 · 8 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DriveSuprim: Towards Precise Trajectory Selection for End-to-End PlanningabstractAutonomous vehicles must navigate safely in complex driving environments. Imitating a single expert trajectory, as in regression-based approaches, usually does not explicitly assess the safety of the predicted trajectory. Selection-based methods address this by generating and scoring multiple trajectory candidates and predicting the safety score for each. However, they face optimization challenges in precisely selecting the best option from thousands of candidates and distinguishing subtle but safety-critical differences, especially in rare and challenging scenarios. We propose DriveSuprim to overcome these challenges and advance the selection-based paradigm through a coarse-to-fine paradigm for progressive candidate filtering, a rotation-based augmentation method to improve robustness in out-of-distribution scenarios, and a self-distillation framework to stabilize training. DriveSuprim achieves state-of-the-art performance, reaching 93.5% PDMS in NAVSIM v1 and 87.1% EPDMS in NAVSIM v2 without extra data, with 83.02 Driving Score and 60.00 Success Rate on Bench2Drive, demonstrating superior planning capabilities in various driving scenarios. Wenhao Yao, Zhenxin Li, Shiyi Lan, Xinglong Sun, José M. Álvarez 0004, Zuxuan Wu |
AAAI | 6 |
| 2026 | GHOST: Getting to the Bottom of Hallucinations with A Multi-round Consistency Benchmark
Vibashan VS, Nadine Chang, Jenny Schmalfuss, Vishal M. Patel, Zhiding Yu, José M. Álvarez 0004 |
WACV | 6 |
| 2025 | Joint Optimization of Neural Radiance Fields and Continuous Camera Motion from a Monocular VideoabstractNeural Radiance Fields (NeRF) has demonstrated its superior capability to represent 3D geometry but require accurately precomputed camera poses during training. To mitigate this requirement, existing methods jointly optimize camera poses and NeRF often relying on good pose initialisation or depth priors. However, these approaches struggle in challenging scenarios, such as large rotations, as they map each camera to a world co-ordinate system. We propose a novel method that eliminates prior dependencies by modeling continuous camera motions as time-dependent angular velocity and velocity. Relative motions between cameras are learned first via velocity integration, while camera poses can be obtained by aggregating such relative motions up to a world coordinate system defined at a single time step within the video. Specifically, accurate continuous camera movements are learned through a time-dependent NeRF, which captures local scene geometry and motion by training from neighboring frames for each time step. The learned motions enable fine-tuning the NeRF to represent the full scene geometry. Experiments on Co3D and Scannet show our approach achieves superior camera pose and depth estimation and comparable novel-view synthesis performance compared to state-of-the-art methods. Our code is available at https: //github.com/HoangChuongNguyen/cope-nerf. Hoang Chuong Nguyen, Wei Mao 0001, José M. Álvarez 0004, Miaomiao Liu 0001 |
CVPR | 3 |
| 2025 | PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language ModelsabstractVision language models (VLMs) respond to user-crafted text prompts and visual inputs, and are applied to numerous real-world problems. VLMs integrate visual modalities with large language models (LLMs), which are well known to be prompt-sensitive. Hence, it is crucial to determine whether VLMs inherit this instability to varying prompts. We therefore investigate which prompt variations VLMs are most sensitive to and which VLMs are most agnostic to prompt variations. To this end, we introduce PARC (Prompt Analysis via Reliability and Calibration), a VLM prompt sensitivity analysis framework built on three pillars: (1) plausible prompt variations in both the language and vision domain, (2) a novel model reliability score with built-in guarantees, and (3) a calibration step that enables dataset-and prompt-spanning prompt variation analysis. Regarding prompt variations, PARC’s evaluation shows that VLMs mirror LLM language prompt sensitivity in the vision domain, and most destructive variations change the expected answer. Regarding models, outstandingly robust VLMs among 22 evaluated models come from the InternVL2 family. We further find indications that prompt sensitivity is linked to training data. https://github.com/NVlabs/PARC Jenny Schmalfuss, Nadine Chang, Vibashan VS, Maying Shen, Andrés Bruhn, José M. Álvarez 0004 |
CVPR | 6 |
| 2025 | MDP: Multidimensional Vision Model Pruning with Latency ConstraintabstractCurrent structural pruning methods face two significant limitations: (i) they often limit pruning to finer-grained levels like channels, making aggressive parameter reduction challenging, and (ii) they focus heavily on parameter and FLOP reduction, with existing latency-aware methods frequently relying on simplistic, suboptimal linear models that fail to generalize well to transformers, where multiple interacting dimensions impact latency. In this paper, we address both limitations by introducing Multi-Dimensional Pruning(MDP), a novel paradigm that jointly optimizes across a variety of pruning granularities—including channels, query/key, heads, embeddings, and blocks. MDP employs an advanced latency modeling technique to accurately capture latency variations across all prunable dimensions, achieving an optimal balance between latency and accuracy. By reformulating pruning as a Mixed-Integer Nonlinear Program (MINLP), MDP efficiently identifies the optimal pruned structure across all prunable dimensions while respecting latency constraints. This versatile framework supports both CNNs and transformers. Extensive experiments demonstrate that MDP significantly outperforms previous methods, especially at high pruning ratios. On ImageNet, MDP achieves a 28% speed increase with a +1.4 Top-1 accuracy improvement over prior work like HALP for ResNet50 pruning. Against the latest transformer pruning method, Isomorphic, MDP delivers an additional 37% acceleration with a +0.7 Top-1 accuracy improvement. Xinglong Sun, Barath Lakshmanan, Maying Shen, Shiyi Lan, Jingde Chen, José M. Álvarez 0004 |
CVPR | 6 |
| 2025 | OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual ReasoningabstractThe advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for real-world applications. To address this challenge, we propose OmniDrive, a holistic vision-language dataset that aligns agent models with 3D driving tasks through counter-factual reasoning. This approach enhances decision-making by evaluating potential scenarios and their outcomes, similar to human drivers considering alternative actions. Our counterfactual-based synthetic data annotation process generates large-scale, high-quality datasets, providing denser supervision signals that bridge planning trajectories and language-based reasoning. Futher, we explore two advanced OmniDrive-Agent frameworks, namely Omni-L and Omni-Q, to assess the importance of vision-language alignment versus 3D perception, revealing critical insights into designing effective LLM-agents. Significant improvements on the DriveLM Q&A benchmark and nuScenes open-loop planning demonstrate the effectiveness of our dataset and methods. Zhiding Yu, Xiaohui Jiang, Shiyi Lan, Nadine Chang, Jan Kautz, José M. Álvarez 0004 |
CVPR | 9 |
| 2025 | Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
Zhenxin Li, Shiyi Lan, Zhiding Yu, Zuxuan Wu, José M. Álvarez 0004 |
ICCV | 6 |
| 2025 | Enhancing Autonomous Driving Safety with Collision Scenario IntegrationabstractAutonomous vehicle safety is crucial for the successful deployment of self-driving cars. However, most existing planning methods rely heavily on imitation learning, which limits their ability to leverage collision data effectively. Moreover, collecting collision or near-collision data is inherently challenging, as it involves risks and raises ethical and practical concerns. In this paper, we propose SafeFusion, a training framework to learn from collision data. Instead of over-relying on imitation learning, SafeFusion integrates safety-oriented metrics during training to enable collision avoidance learning. In addition, to address the scarcity of collision data, we propose CollisionGen, a scalable data generation pipeline to generate diverse, high-quality scenarios using natural language prompts, generative models, and rule-based filtering. Experimental results show that our approach improves planning performance in collision-prone scenarios by 56% over previous state-of-the-art planners while maintaining effectiveness in regular driving situations. Our work provides a scalable and effective solution for advancing the safety of autonomous driving systems. Shiyi Lan, Xinglong Sun, Nadine Chang, Zhenxin Li, Zhiding Yu, José M. Álvarez 0004 |
IROS | 7 |
| 2025 | SSE: Multimodal Semantic Data Selection and Enrichment for Industrial-scale Data Assimilation
Maying Shen, Nadine Chang, Sifei Liu, José M. Álvarez 0004 |
KDD (1) | 4 |
| 2025 | IBGS: Image-Based Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has recently emerged as a fast, high-quality method for novel view synthesis (NVS). However, its use of low-degree spherical harmonics limits its ability to capture spatially varying color and view-dependent effects such as specular highlights. Existing works augment Gaussians with either a global texture map, which struggles with complex scenes, or per-Gaussian texture maps, which introduces high storage overhead. We propose Image-Based Gaussian Splatting, an efficient alternative that leverages high-resolution source images for fine details and view-specific color modeling. Specifically, we model each pixel color as a combination of a base color from standard 3DGS rendering and a learned residual inferred from neighboring training images. This promotes accurate surface alignment and enables rendering images of high-frequency details and accurate view-dependent effects. Experiments on standard NVS benchmarks show that our method significantly outperforms prior Gaussian Splatting approaches in rendering quality, without increasing the storage footprint. Our project page is available at https://hoangchuongnguyen.github.io/ibgs. Hoang Chuong Nguyen, Wei Mao 0001, José M. Álvarez 0004, Miaomiao Liu 0001 |
NeurIPS | 3 |
| 2025 | Advancing Weight and Channel Sparsification with Enhanced SaliencyabstractPruning aims to accelerate and compress models by removing redundant parameters, identified by specifically designed importance scores which are usually imperfect. This removal is irreversible, often leading to subpar performance in pruned models. Dynamic sparse training, while attempting to adjust sparse structures during training for continual reassessment and refinement, has several limitations including criterion inconsistency between pruning and growth, unsuitability for structured sparsity, and short-sighted growth strategies. Our paper introduces an efficient, innovative paradigm to enhance a given importance criterion for either unstructured or structured sparsity. Our method separates the model into an active structure for exploitation and an exploration space for potential updates. During exploitation, we optimize the active structure, whereas in exploration, we reevaluate and reintegrate parameters from the exploration space through a pruning and growing step consistently guided by the same given importance criterion. To prepare for exploration, we briefly “reactivate” all parameters in the exploration space and train them for a few iterations while keeping the active part frozen, offering a preview of the potential performance gains from reintegrating these parameters. We show on various datasets and configurations that existing importance criterion even simple as magnitude can be enhanced with ours to achieve state-of-the-art performance and training cost reductions. Notably, on ImageNet with ResNet50, ours achieves an$+1.3$increase in Top-1 accuracy over prior art at 90% ERK [49] sparsity. Compared with the SOTA latency pruning method HALP [58], we reduced its training cost by over 70% while attaining a faster and more accurate pruned model. Xinglong Sun, Maying Shen, Hongxu Yin, Pavlo Molchanov 0001, José M. Álvarez 0004 |
WACV | 6 |
| 2025 | Optimizing Data Collection for Machine LearningabstractModern deep learning systems require huge data sets to achieve impressive performance, but there is little guidance on how much or what kind of data to collect. Over-collecting data incurs unnecessary present costs, while under-collecting may incur future costs and delay workflows. We propose a new paradigm to model the data collection workflow as a formal optimal data collection problem that allows designers to specify performance targets, collection costs, a time horizon, and penalties for failing to meet the targets. This formulation generalizes to tasks with multiple data sources, such as labeled and unlabeled data used in semi-supervised learning, and can be easily modified to customized analyses such as how to introduce data from new classes to an existing model. To solve our problem, we develop Learn-Optimize-Collect (LOC), which minimizes expected future collection costs. Finally, we numerically compare our framework to the conventional baseline of estimating data requirements by extrapolating from neural scaling laws. We significantly reduce the risks of failing to meet desired performance targets on several classification, segmentation, and detection tasks, while maintaining low total collection costs. Rafid Mahmood, James Lucas, José M. Álvarez 0004, Sanja Fidler, Marc T. Law |
J. Mach. Learn. Res. | 3 |
| 2024 | BEVNeXt: Reviving Dense BEV Frameworks for 3D Object DetectionabstractRecently, the rise of query-based Transformer decoders is reshaping camera-based 3D object detection. These query-based decoders are surpassing the traditional dense BEV (Bird's Eye View)-based methods. However, we argue that dense BEV frameworks remain important due to their out-standing abilities in depth estimation and object localization, depicting 3D scenes accurately and comprehensively. This paper aims to address the drawbacks of the existing dense BEV-based 3D object detectors by introducing our proposed enhanced components, including a CRF-modulated depth estimation module enforcing object-level consistencies, a long-term temporal aggregation module with extended receptive fields, and a two-stage object decoder combining perspective techniques with CRF-modulated depth embedding. These enhancements lead to a “modernized” dense BEV framework dubbed BEVNeXt. On the nuScenes benchmark, BEVNeXt outperforms both BEV-based and query-based frameworks under various settings, achieving a state-of-the-art result of 64.2 NDS on the nuScenes test set. Zhenxin Li, Shiyi Lan, José M. Álvarez 0004, Zuxuan Wu |
CVPR | 3 |
| 2024 | Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?abstractEnd-to-end autonomous driving recently emerged as a promising research direction to target autonomy from a full-stack perspective. Along this line, many of the latest works follow an open-loop evaluation setting on nuScenes to study the planning behavior. In this paper, we delve deeper into the problem by conducting thorough analyses and demystifying more devils in the details. We initially observed that the nuScenes dataset, characterized by relatively simple driving scenarios, leads to an under-utilization of perception information in end-to-end models incorporating ego status, such as the ego vehicle's velocity. These models tend to rely predominantly on the ego vehicle's status for future path planning. Beyond the limitations of the dataset, we also note that current metrics do not comprehensively assess the planning quality, leading to potentially biased conclusions drawn from existing benchmarks. To address this issue, we introduce a new metric to evaluate whether the predicted trajectories adhere to the road. We further propose a simple baseline able to achieve competitive results without relying on perception annotations. Given the current limitations on the benchmark and metrics, we suggest the community reassess relevant prevailing research and be cautious about whether the continued pursuit of state-of-the-art would yield convincing and universal conclusions. Code and models are available at https://github.com/NVlabs/BEV-Planner. Zhiding Yu, Shiyi Lan, Jiahan Li, Jan Kautz, Tong Lu 0002, José M. Álvarez 0004 |
CVPR | 7 |
| 2024 | Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth EstimationabstractThis paper focuses on self-supervised monocular depth estimation in dynamic scenes trained on monocular videos. Existing methods jointly estimate pixel-wise depth and motion, relying mainly on an image reconstruction loss. Dynamic regions11Dynamic regions indicate regions covered by moving objects. remain a critical challenge for these methods due to the inherent ambiguity in depth and motion estimation, resulting in inaccurate depth estimation. This paper proposes a self-supervised training framework exploiting pseudo depth labels for dynamic regions from training data. The key contribution of our framework is to decouple depth estimation for static and dynamic regions of images in the training data. We start with an unsupervised depth estimation approach, which provides reliable depth estimates for static regions and motion cues for dynamic regions and allows us to extract moving object information at the instance level. In the next stage, we use an object network to estimate the depth of those moving objects assuming rigid motions. Then, we propose a new scale alignment module to address the scale ambiguity between estimated depths for static and dynamic regions. We can then use the depth labels generated to train an end-to-end depth estimation network and improve its performance. Extensive experiments on the Cityscapes and KITTI datasets show that our self-training strategy consistently outperforms existing self-/unsupervised depth estimation methods. Our code is available at https://github.com/HoangChuongNguyen/mono-consistent-depth.git Hoang Chuong Nguyen, Tianyu Wang 0035, José M. Álvarez 0004, Miaomiao Liu 0001 |
CVPR | 3 |
| 2024 | Improving Distant 3D Object Detection Using 2D Box SupervisionabstractImproving the detection of distant 3d objects is an impor-tant yet challenging task. For camera-based 3D perception, the annotation of 3d bounding relies heavily on LiDAR for accurate depth information. As such, the distance of anno-tation is often limited due to the sparsity of LiDAR points on distant objects, which hampers the capability of existing de-tectors for long-range scenarios. We address this challenge by considering only 2D box supervision for distant objects since they are easy to annotate. We propose LR3D, a frame-work that learns to recover the missing depth of distant ob-jects. LR3D adopts an implicit projection head to learn the generation of mapping between 2D boxes and depth using the 3D supervision on close objects. This mapping allows the depth estimation of distant objects conditioned on their 2D boxes, making long-range 3D detection with 2D super-vision feasible. Experiments show that without distant 3D annotations, LR3D allows camera-based methods to detect distant objects (over 200m) with comparable accuracy to full 3D supervision. Our framework is general, and could widely benefit 3D detection methods to a large extent. Zetong Yang, Zhiding Yu, Christopher B. Choy, Renhao Wang, Anima Anandkumar, José M. Álvarez 0004 |
CVPR | 6 |
| 2024 | SegIC: Unleashing the Emergent Correspondence for In-Context Segmentation
Lingchen Meng, Shiyi Lan, Hengduo Li, José M. Álvarez 0004, Zuxuan Wu, Yu-Gang Jiang 0001 |
ECCV (38) | 4 |
| 2024 | Adaptive Sharpness-Aware Pruning for Robust Sparse NetworksabstractRobustness and compactness are two essential attributes of deep learning models that are deployed in the real world.
The goals of robustness and compactness may seem to be at odds, since robustness requires generalization across domains, while the process of compression exploits specificity in one domain.
We introduce \textit{Adaptive Sharpness-Aware Pruning (AdaSAP)}, which unifies these goals through the lens of network sharpness.
The AdaSAP method produces sparse networks that are robust to input variations which are \textit{unseen at training time}.
We achieve this by strategically incorporating weight perturbations in order to optimize the loss landscape. This allows the model to be both primed for pruning and regularized for improved robustness.
AdaSAP improves the robust accuracy of pruned models on image classification by up to +6\% on ImageNet C and +4\% on ImageNet V2, and on object detection by +4\% on a corrupted Pascal VOC dataset, over a wide range of compression ratios, pruning criteria, and network architectures, outperforming recent pruning art by large margins. Anna Bair, Hongxu Yin, Maying Shen, Pavlo Molchanov 0001, José M. Álvarez 0004 |
ICLR | 5 |
| 2024 | FasterViT: Fast Vision Transformers with Hierarchical AttentionabstractWe design a new family of hybrid CNN-ViT neural networks, named FasterViT, with a focus on high image throughput for computer vision (CV) applications. FasterViT combines the benefits of fast local representation learning in CNNs and global modeling properties in ViT. Our newly introduced Hierarchical Attention (HAT) approach decomposes global self-attention with quadratic complexity into a multi-level attention with reduced computational costs. We benefit from efficient window-based self-attention. Each window has access to dedicated carrier tokens that participate in local and global representation learning. At a high level, global self-attentions enable the efficient cross-window communication at lower costs. FasterViT achieves a SOTA Pareto-front in terms of accuracy and image throughput. We have extensively validated its effectiveness on various CV tasks including classification, object detection and segmentation. We also show that HAT can be used as a plug-and-play module for existing networks and enhance them. We further demonstrate significantly faster and more accurate performance than competitive counterparts for images with high resolution. Code is available at https://github.com/NVlabs/FasterViT. Ali Hatamizadeh, Greg Heinrich, Hongxu Yin, Andrew Tao, José M. Álvarez 0004, Jan Kautz, Pavlo Molchanov 0001 |
ICLR | 5 |
| 2024 | SF3D: SlowFast Temporal 3D Object DetectionabstractLeveraging inputs over multiple consecutive frames has been shown to benefit 3D object detection. However, existing approaches often demonstrate unsatisfactory scaling with increasing temporal histories. In this work, we propose SF3D, a late fusion module which addresses this issue by better modeling temporal relationships via a two-stream factorization. Concretely, SF3D operates on an input sequence of consecutive bird’s-eye view (BEV) features, which is partitioned into "short-term" and "long-term" frames. A more heavily parameterized short-term branch using adapters and deformable attention aggregates features closer to the current timestep. In parallel, a long-term branch composed of efficiently implemented global convolution layers aggregates a larger window of temporally distant historical features. This two-stream paradigm allows SF3D to effectively consume near-term information, while scaling to efficiently leverage longer historical windows. We show that SF3D works with arbitrary upstream BEV encoders and downstream detectors, achieving improvements over recent state-of-the-art on the Waymo Open and nuScenes benchmarks. Renhao Wang, Zhiding Yu, Shiyi Lan, Enze Xie, Anima Anandkumar, José M. Álvarez 0004 |
IV | 7 |
| 2024 | Memorize What Matters: Emergent Scene Decomposition from MultitraverseabstractHumans naturally retain memories of permanent elements, while ephemeral moments often slip through the cracks of memory. This selective retention is crucial for robotic perception, localization, and mapping. To endow robots with this capability, we introduce 3D Gaussian Mapping (3DGM), a self-supervised, camera-only offline mapping framework grounded in 3D Gaussian Splatting. 3DGM converts multitraverse RGB videos from the same region into a Gaussian-based environmental map while concurrently performing 2D ephemeral object segmentation. Our key observation is that the environment remains consistent across traversals, while objects frequently change. This allows us to exploit self-supervision from repeated traversals to achieve environment-object decomposition. More specifically, 3DGM formulates multitraverse environmental mapping as a robust 3D representation learning problem, treating pixels of the environment and objects as inliers and outliers, respectively. Using robust feature distillation, feature residual mining, and robust optimization, 3DGM simultaneously performs 2D segmentation and 3D mapping without human intervention. We build the Mapverse benchmark, sourced from the Ithaca365 and nuPlan datasets, to evaluate our method in unsupervised 2D segmentation, 3D reconstruction, and neural rendering. Extensive results verify the effectiveness and potential of our method for self-driving and robotics. Yiming Li 0003, Zehong Wang, Yue Wang 0041, Zhiding Yu, Zan Gojcic, Marco Pavone 0001, Chen Feng 0002, José M. Álvarez 0004 |
NeurIPS | 8 |
| 2023 | Knowledge Distillation for 6D Pose Estimation by Aligning Distributions of Local PredictionsabstractKnowledge distillation facilitates the training of a compact student network by using a deep teacher one. While this has achieved great success in many tasks, it remains completely unstudied for image-based 6D object pose estimation. In this work, we introduce the first knowledge distillation method driven by the 6D pose estimation task. To this end, we observe that most modern 6D pose estimation frameworks output local predictions, such as sparse 2D keypoints or dense representations, and that the compact student network typically struggles to predict such local quantities precisely. Therefore, instead of imposing prediction-to-prediction supervision from the teacher to the student, we propose to distill the teacher's distribution of local predictions into the student network, facilitating its training. Our experiments on several benchmarks show that our distillation method yields state-of-the-art results with different compact student models and for both keypoint-based and dense prediction-based architectures. Shuxuan Guo, Yinlin Hu, José M. Álvarez 0004, Mathieu Salzmann |
CVPR | 3 |
| 2023 | Vision Transformers are Good Mask Auto-LabelersabstractWe propose Mask Auto-Labeler (MAL), a high-quality Transformer-based mask auto-labeling framework for instance segmentation using only box annotations. MAL takes box-cropped images as inputs and conditionally generates their mask pseudo-labels. We show that Vision Transformers are good mask auto-labelers. Our method significantly reduces the gap between auto-labeling and human annotation regarding mask quality. Instance segmentation models trained using the MAL-generated masks can nearly match the performance of their fully-supervised counterparts, retaining up to 97.4% performance of fully supervised models. The best model achieves 44.1% mAP on COCO instance segmentation (test-dev 2017), outperforming state-of-the-art box-supervised methods by significant margins. Qualitative results indicate that masks produced by MAL are, in some cases, even better than human annotations. Shiyi Lan, Xitong Yang, Zhiding Yu, Zuxuan Wu, José M. Álvarez 0004, Anima Anandkumar |
CVPR | 5 |
| 2023 | VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionabstractHumans can easily imagine the complete 3D geometry of occluded objects and scenes. This appealing ability is vital for recognition and understanding. To enable such capability in AI systems, we propose VoxFormer, a Transformer-based semantic scene completion framework that can output complete 3D volumetric semantics from only 2D images. Our framework adopts a two-stage design where we start from a sparse set of visible and occupied voxel queries from depth estimation, followed by a densification stage that generates dense 3D voxels from the sparse ones. A key idea of this design is that the visual features on 2D images correspond only to the visible scene structures rather than the occluded or empty spaces. Therefore, starting with the fea-turization and prediction of the visible structures is more reliable. Once we obtain the set of sparse queries, we apply a masked autoencoder design to propagate the information to all the voxels by self-attention. Experiments on SemanticKITTI show that VoxFormer outperforms the state of the art with a relative improvement of 20.0% in geometry and 18.1% in semantics and reduces GPU memory during training to less than 16GB. Our code is available on https://github.com/NV1abs/VoxFormer. Yiming Li 0003, Zhiding Yu, Christopher B. Choy, Chaowei Xiao, José M. Álvarez 0004, Sanja Fidler, Chen Feng 0002, Anima Anandkumar |
CVPR | 5 |
| 2023 | FocalFormer3D : Focusing on Hard Instance for 3D Object DetectionabstractFalse negatives (FN) in 3D object detection, e.g., missing predictions of pedestrians, vehicles, or other obstacles, can lead to potentially dangerous situations in autonomous driving. While being fatal, this issue is understudied in many current 3D detection methods. In this work, we propose Hard Instance Probing (HIP), a general pipeline that identifies FN in a multi-stage manner and guides the models to focus on excavating difficult instances. For 3D object detection, we instantiate this method as FocalFormer3D, a simple yet effective detector that excels at excavating difficult objects and improving prediction recall. FocalFormer3D features a multi-stage query generation to discover hard objects and a box-level transformer decoder to efficiently distinguish objects from massive object candidates. Experimental results on the nuScenes and Waymo datasets validate the superior performance of FocalFormer3D. The advantage leads to strong performance on both detection and tracking, in both LiDAR and multi-modal settings. Notably, FocalFormer3D achieves a 70.5 mAP and 73.9 NDS on nuScenes detection benchmark, while the nuScenes tracking benchmark shows 72.1 AMOTA, both ranking 1st place on the nuScenes LiDAR leaderboard. Our code is available at https://github.com/NVlabs/FocalFormer3D. Zhiding Yu, Yukang Chen, Shiyi Lan, Anima Anandkumar, Jiaya Jia, José M. Álvarez 0004 |
ICCV | 7 |
| 2023 | Towards Viewpoint Robustness in Bird's Eye View SegmentationabstractAutonomous vehicles (AV) require that neural networks used for perception be robust to different viewpoints if they are to be deployed across many types of vehicles without the repeated cost of data collection and labeling for each. AV companies typically focus on collecting data from diverse scenarios and locations, but not camera rig configurations, due to cost. As a result, only a small number of rig variations exist across most fleets. In this paper, we study how AV perception models are affected by changes in camera viewpoint and propose a way to scale them across vehicle types without repeated data collection and labeling. Using bird’s eye view (BEV) segmentation as a motivating task, we find through extensive experiments that existing perception models are surprisingly sensitive to changes in camera viewpoint. When trained with data from one camera rig, small changes to pitch, yaw, depth, or height of the camera at inference time lead to large drops in performance. We introduce a technique for novel view synthesis and use it to transform collected data to the viewpoint of target rigs, allowing us to train BEV segmentation models for diverse target rigs without any additional data collection or labeling cost. To analyze the impact of viewpoint changes, we leverage synthetic data to mitigate other gaps (content, ISP, etc). Our approach is then trained on real data and evaluated on synthetic data, enabling evaluation on diverse target rigs. We release all data for use in future work. Our method is able to recover an average of 14.7% of the IoU that is otherwise lost when deploying to new rigs. Tzofi Klinghoffer, Jonah Philion, Wenzheng Chen, Or Litany, Zan Gojcic, Jungseock Joo, Ramesh Raskar, Sanja Fidler, José M. Álvarez 0004 |
ICCV | 9 |
| 2023 | FB-BEV: BEV Representation from Forward-Backward View TransformationsabstractView Transformation Module (VTM), where transformations happen between multi-view image features and Bird-Eye-View (BEV) representation, is a crucial step in camera-based BEV perception systems. Currently, the two most prominent VTM paradigms are forward projection and backward projection. Forward projection, represented by Lift-Splat-Shoot, leads to sparsely projected BEV features without post-processing. Backward projection, with BEV-Former being an example, tends to generate false-positive BEV features from incorrect projections due to the lack of utilization on depth. To address the above limitations, we propose a novel forward-backward view transformation module. Our approach compensates for the deficiencies in both existing methods, allowing them to enhance each other to obtain higher quality BEV representations mutually. We instantiate the proposed module with FB-BEV, which achieves a new state-of-the-art result of 62.4% NDS on the nuScenes test set. Code and models are available at https://github.com/NVlabs/FB-BEV. Zhiding Yu, Wenhai Wang, Anima Anandkumar, Tong Lu 0002, José M. Álvarez 0004 |
ICCV | 6 |
| 2023 | Parametric Depth Based Feature Representation Learning for Object Detection and Segmentation in Bird's-Eye ViewabstractRecent vision-only perception models for autonomous driving achieved promising results by encoding multi-view image features into Bird’s-Eye-View (BEV) space. A critical step and the main bottleneck of these methods is transforming image features into the BEV coordinate frame. This paper focuses on leveraging geometry information, such as depth, to model such feature transformation. Existing works rely on non-parametric depth distribution modeling leading to significant memory consumption, or ignore the geometry information to address this problem. In contrast, we propose to use parametric depth distribution modeling for feature transformation. We first lift the 2D image features to the 3D space defined for the ego vehicle via a predicted parametric depth distribution for each pixel in each view. Then, we aggregate the 3D feature volume based on the 3D space occupancy derived from depth to the BEV frame. Finally, we use the transformed features for downstream tasks such as object detection and semantic segmentation. Existing semantic segmentation methods do also suffer from an hallucination problem as they do not take visibility information into account. This hallucination can be particularly problematic for subsequent modules such as control and planning. To mitigate the issue, our method provides depth uncertainty and reliable visibility-aware estimations. We further leverage our parametric depth modeling to present a novel visibility-aware evaluation metric that, when taken into account, can mitigate the hallucination problem. Ex tensive experiments on object detection and semantic segmentation on the nuScenes datasets demonstrate that our method outperforms existing methods on both tasks. Enze Xie, Miaomiao Liu 0001, José M. Álvarez 0004 |
ICCV | 4 |
| 2023 | Fully Attentional Networks with Self-emerging Token LabelingabstractRecent studies indicate that Vision Transformers (ViTs) are robust against out-of-distribution scenarios. In particular, the Fully Attentional Network (FAN) - a family of ViT backbones, has achieved state-of-the-art robustness. In this paper, we revisit the FAN models and improve their pretraining with a self-emerging token labeling (STL) framework. Our method contains a two-stage training framework. Specifically, we first train a FAN token labeler (FAN-TL) to generate semantically meaningful patch token labels, followed by a FAN student model training stage that uses both the token labels and the original class label. With the proposed STL framework, our best model based on FANL-Hybrid (77.3M parameters) achieves 84.8% Top-1 accuracy and 42.1% mCE on ImageNet-1K and ImageNetC, and sets a new state-of-the-art for ImageNet-A (46.1%) and ImageNet-R (56.6%) without using extra data, outperforming the original FAN counterpart by significant margins. The proposed framework also demonstrates significantly enhanced performance on downstream tasks such as semantic segmentation, with up to 1.7% improvement in robustness over the counterpart model. Bingyin Zhao, Zhiding Yu, Shiyi Lan, Yutao Cheng, Anima Anandkumar, Yingjie Lao, José M. Álvarez 0004 |
ICCV | 7 |
| 2023 | Augmenting Legacy Networks for Flexible InferenceabstractOn intelligent vehicles, Deep Neural Networks (DNNs) may run on devices whose computational load varies over time. Within the context of variable network architectures, that can be used to constrain the inference latency for real-time deployment with varying system resources, we introduce LeAF (Legacy Augmentation for Flexible inference), a novel paradigm to augment a pre-trained DNN with trainable, shallow execution paths that can run in place of the legacy ones. While preserving the legacy DNN weights, LeAF allows changing the DNN architecture with minimal overhead to effectively adapt to different system performance targets. LeAF-ResNet-50 has less than 14% storage overhead over the legacy DNN; its accuracy varies from the legacy 76.1% to 70.15% (up to 5% better than Slimmable [1] with a latency that is 37% better than OFA [2] on an A100 GPU with batch size 256). Our analysis shows the importance of considering not only the target device, but also the batch size and the temporal dynamic of the DNN configuration to optimize the performances of variable architecture DNNs, LeAF in particular. Jason Clemons, Iuri Frosio, Maying Shen, José M. Álvarez 0004, Stephen W. Keckler |
IV | 4 |
| 2023 | Hardware-Aware Latency Pruning for Real-Time 3D Object Detectionabstract3D Object detection is a fundamental task in vision-based autonomous driving. Deep learning perception models achieve an outstanding performance at the expense of continuously increasing resource needs and, as such, increasing training costs. As inference time is still a priority, developers usually adopt a training pipeline where they first start using a compact architecture that yields a good trade-off between accuracy and latency. This architecture is usually found either by searching manually or by using neural architecture search approaches. Then, train the model and use light optimization techniques such as quantization to boost the model’s performance. In contrast, in this paper, we advocate for starting on a much larger model and then applying aggressive optimization to adapt the model to the resource-constraints. Our results on large-scale settings for 3D object detection demonstrate the benefits of initially focusing on maximizing the model’s accuracy and then achieving the latency requirements using network pruning. Maying Shen, Joshua Chen, Justin Hsu, Xinglong Sun, Oliver Knieps, Carmen Maxim, José M. Álvarez 0004 |
IV | 8 |
| 2022 | Privacy Vulnerability of Split Computing to Data-Free Model Inversion Attacks
Xin Dong 0009, Hongxu Yin, José M. Álvarez 0004, Jan Kautz, Pavlo Molchanov 0001, H. T. Kung 0001 |
BMVC | 3 |
| 2022 | Not All Labels Are Equal: Rationalizing The Labeling Costs for Training Object DetectionabstractDeep neural networks have reached high accuracy on object detection but their success hinges on large amounts of labeled data. To reduce the labels dependency, various active learning strategies have been proposed, based on the confidence of the detector. However, these methods are biased towards high-performing classes and lead to acquired datasets that are not good representatives of the testing set data. In this work, we propose a unified frame-work for active learning, that considers both the uncertainty and the robustness of the detector, ensuring that the network performs well in all classes. Furthermore, our method leverages auto-labeling to suppress a potential distribution drift while boosting the performance of the model. Experiments on PASCAL VOC07+12 and MS-COCO show that our method consistently outperforms a wide range of active learning methods, yielding up to a 7.7% improvement in mAP, or up to 82% reduction in labeling cost. Code is available at https://github.com/NVlabs/AL-SSL. Ismail Elezi, Zhiding Yu, Anima Anandkumar, Laura Leal-Taixé, José M. Álvarez 0004 |
CVPR | 5 |
| 2022 | Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with TransformersabstractPanoptic segmentation involves a combination of joint semantic segmentation and instance segmentation, where image contents are divided into two types: things and stuff. We present Panoptic SegFormer, a general framework for panoptic segmentation with transformers. It contains three innovative components: an efficient deeply-supervised mask decoder, a query decoupling strategy, and an improved postprocessing method. We also use Deformable DETR to efficiently process multiscale features, which is a fast and efficient version of DETR. Specifically, we supervise the attention modules in the mask decoder in a layer-wise manner. This deep supervision strategy lets the attention modules quickly focus on meaningful semantic regions. It improves performance and reduces the number of required training epochs by half compared to Deformable DETR. Our query decoupling strategy decouples the responsibilities of the query set and avoids mutual interference between things and stuff. In addition, our post-processing strategy improves performance without additional costs by jointly considering classification and segmentation qualities to resolve conflicting mask overlaps. Our approach increases the accuracy 6.2% PQ over the baseline DETR model. Panoptic SegFormer achieves state-of-the-art results on COCO testdev with 56.2% PQ. It also shows stronger zero-shot robustness over existing methods. Wenhai Wang, Enze Xie, Zhiding Yu, Anima Anandkumar, José M. Álvarez 0004, Ping Luo 0002, Tong Lu 0002 |
CVPR | 6 |
| 2022 | How Much More Data Do I Need? Estimating Requirements for Downstream TasksabstractGiven a small training data set and a learning algorithm, how much more data is necessary to reach a target validation or test performance? This question is of critical importance in applications such as autonomous driving or medical imaging where collecting data is expensive and time-consuming. Overestimating or underestimating data requirements incurs substantial costs that could be avoided with an adequate budget. Prior work on neural scaling laws suggest that the power-law function can fit the validation performance curve and extrapolate it to larger data set sizes. We find that this does not immediately translate to the more difficult downstream task of estimating the required data set size to meet a target performance. In this work, we consider a broad class of computer vision tasks and systematically investigate a family of functions that generalize the power-law function to allow for better estimation of data requirements. Finally, we show that incorporating a tuned correction factor and collecting over multiple rounds significantly improves the performance of the data estimators. Using our guidelines, practitioners can accurately estimate data requirements of machine learning systems to gain savings in both development time and data acquisition costs. Rafid Mahmood, James Lucas, David Acuna, Daiqing Li, Jonah Philion, José M. Álvarez 0004, Zhiding Yu, Sanja Fidler, Marc T. Law |
CVPR | 6 |
| 2022 | When to Prune? A Policy towards Early Structural PruningabstractPruning enables appealing reductions in network memory footprint and time complexity. Conventional post-training pruning techniques lean towards efficient inference while overlooking the heavy computation for training. Recent exploration of pre-training pruning at initialization hints on training cost reduction via pruning, but suffers noticeable performance degradation. We attempt to combine the benefits of both directions and propose a policy that prunes as early as possible during training without hurting performance. Instead of pruning at initialization, our method exploits initial dense training for few epochs to quickly guide the architecture, while constantly evaluating dominant sub-networks via neuron importance ranking. This unveils dominant sub-networks whose structures turn stable, allowing conventional pruning to be pushed earlier into the training. To do this early, we further introduce an Early Pruning Indicator (EPI) that relies on sub-network architectural similarity and quickly triggers pruning when the sub-network's architecture stabilizes. Through extensive experiments on ImageNet, we show that EPI empowers a quick tracking of early training epochs suitable for pruning, offering same efficacy as an otherwise “oracle” grid-search that scans through epochs and requires orders of magnitude more compute. Our method yields 1.4% top-l accuracy boost over state-of-the-art pruning counterparts, cuts down training cost on GPU by 2.4x, hence offers a new efficiency-accuracy boundary for network pruning during training. Maying Shen, Pavlo Molchanov 0001, Hongxu Yin, José M. Álvarez 0004 |
CVPR | 4 |
| 2022 | FreeSOLO: Learning to Segment Objects without AnnotationsabstractInstance segmentation is a fundamental vision task that aims to recognize and segment each object in an image. However, it requires costly annotations such as bounding boxes and segmentation masks for learning. In this work, we propose a fully unsupervised learning method that learns class-agnostic instance segmentation without any annotations. We present FreeSOLO, a self-supervised instance segmentation framework built on top of the simple instance segmentation method SOLO. Our method also presents a novel localization-aware pre-training framework, where objects can be discovered from complicated scenes in an unsupervised manner. FreeSOLO achieves 9.8%$AP_{50}$on the challenging COCO dataset, which even outperforms several segmentation proposal methods that use manual annotations. For the first time, we demonstrate unsupervised class-agnostic instance segmen-tation successfully. FreeSOLO's box localization significantly outperforms state-of-the-art unsupervised object de-tection/discovery methods, with about 100% relative improvements in COCO AP. FreeSOLO further demonstrates superiority as a strong pre-training method, outperforming state-of-the-art self-supervised pre-training methods by$+9.8\%$AP when fine-tuning instance segmentation with only 5% COCO masks. Code is available at: github.com/NVlabs/FreeSOLO Zhiding Yu, Shalini De Mello, Jan Kautz, Anima Anandkumar, Chunhua Shen, José M. Álvarez 0004 |
CVPR | 7 |
| 2022 | Non-parametric Depth Distribution Modelling based Depth Inference for Multi-view StereoabstractRecent cost volume pyramid based deep neural networks have unlocked the potential of efficiently leveraging high-resolution images for depth inference from multi-view stereo. In general, those approaches assume that the depth of each pixel follows a unimodal distribution. Boundary pixels usually follow a multi-modal distribution as they represent different depths; Therefore, the assumption results in an erroneous depth prediction at the coarser level of the cost volume pyramid and can not be corrected in the refinement levels leading to wrong depth predictions. In contrast, we propose constructing the cost volume by non-parametric depth distribution modeling to handle pixels with unimodal and multi-modal distributions. Our approach outputs multiple depth hypotheses at the coarser level to avoid errors in the early stage. As we perform local search around these multiple hypotheses in subsequent levels, our approach does not maintain the rigid depth spatial ordering and, therefore, we introduce a sparse cost aggregation network to derive information within each volume. We evaluate our approach extensively on two benchmark datasets: DTU and Tanks & Temples. Our experimental results show that our model outperforms existing methods by a large margin and achieves superior performance on boundary regions. Code is available at https://github.com/NVlabs/NP-CVP-MVSNet José M. Álvarez 0004, Miaomiao Liu 0001 |
CVPR | 2 |
| 2022 | A-ViT: Adaptive Tokens for Efficient Vision TransformerabstractWe introduce A - ViT, a method that adaptively adjusts the inference cost of vision transformer (ViT) for images of different complexity. A - ViT achieves this by automatically reducing the number of tokens in vision transformers that are processed in the network as inference proceeds. We refor-mulate Adaptive Computation Time (ACT [17]) for this task, extending halting to discard redundant spatial tokens. The appealing architectural properties of vision transformers enables our adaptive token reduction mechanism to speed up inference without modifying the network architecture or inference hardware. We demonstrate that A - ViT requires no extra parameters or sub-network for halting, as we base the learning of adaptive halting on the original network parameters. We further introduce distributional prior regularization that stabilizes training compared to prior ACT approaches. On the image classification task (ImageNet1K), we show that our proposed A - ViT yields high efficacy in filtering informative spatial features and cutting down on the overall compute. The proposed method improves the throughput of DeiT-Tiny by 62% and DeiT-Small by 38% with only 0.3% accuracy drop, outperforming prior art by a large margin. Hongxu Yin, Arash Vahdat, José M. Álvarez 0004, Arun Mallya, Jan Kautz, Pavlo Molchanov 0001 |
CVPR | 3 |
| 2022 | Soft Masking for Cost-Constrained Channel Pruning
Ryan Humble, Maying Shen, Jorge Albericio Latorre, Eric Darve, José M. Álvarez 0004 |
ECCV (11) | 5 |
| 2022 | Understanding The Robustness in Vision TransformersabstractRecent studies show that Vision Transformers (ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of an explanatory framework towards a more systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of self-attention in visual grouping which indicate that self-attention could promote improved mid-level representation and robustness. We thus propose a family of fully attentional networks (FANs) that incorporate self-attention in both token mixing and channel processing. We validate the design comprehensively on various hierarchical backbones. Our model with a DeiT architecture achieves a state-of-the-art 47.6% mCE on ImageNet-C with 29M parameters. We also demonstrate significantly improved robustness in two downstream tasks: semantic segmentation and object detection Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Anima Anandkumar, Jiashi Feng, José M. Álvarez 0004 |
ICML | 7 |
| 2022 | Object-Level Targeted Selection via Deep Template MatchingabstractRetrieving images with objects that are semantically similar to objects of interest (OOI) in a query image has many practical use cases. A few examples include fixing failures like false negatives/positives of a learned model or mitigating class imbalance in a dataset. The targeted selection task requires finding the relevant data from a large-scale pool of unlabeled data. Manual mining at this scale is infeasible. Further, the OOI are often small and occupy less than 1% of image area, are occluded, and co-exist with many semantically different objects in cluttered scenes. Existing semantic image retrieval methods often focus on mining for larger sized geographical landmarks, and/or require extra labeled data, such as images/image-pairs with similar objects, for mining images with generic objects. We propose a fast and robust template matching algorithm in the DNN feature space, that retrieves semantically similar images at the object-level from a large unlabeled pool of data. We project the region(s) around the OOI in the query image to the DNN feature space for use as the template. This enables our method to focus on the semantics of the OOI without requiring extra labeled data. In the context of autonomous driving, we evaluate our system for targeted selection by using failure cases of object detectors as OOI. We demonstrate its efficacy on a large unlabeled dataset with 2.2M images and show high recall in mining for images with small-sized OOI. We compare our method against a well-known semantic image retrieval method, which also does not require extra labeled data. Lastly, we show that our method is flexible and retrieves images with one or more semantically different co-occurring OOI seamlessly. Suraj Kothawade, Donna Roy, Michele Fenzi, Elmar Haussmann, José M. Álvarez 0004, Christoph Angerer |
IV | 5 |
| 2022 | Optimizing Data Collection for Machine LearningabstractModern deep learning systems require huge data sets to achieve impressive performance, but there is little guidance on how much or what kind of data to collect. Over-collecting data incurs unnecessary present costs, while under-collecting may incur future costs and delay workflows. We propose a new paradigm for modeling the data collection workflow as a formal optimal data collection problem that allows designers to specify performance targets, collection costs, a time horizon, and penalties for failing to meet the targets. Additionally, this formulation generalizes to tasks requiring multiple data sources, such as labeled and unlabeled data used in semi-supervised learning. To solve our problem, we develop Learn-Optimize-Collect (LOC), which minimizes expected future collection costs. Finally, we numerically compare our framework to the conventional baseline of estimating data requirements by extrapolating from neural scaling laws. We significantly reduce the risks of failing to meet desired performance targets on several classification, segmentation, and detection tasks, while maintaining low total collection costs. Rafid Mahmood, James Lucas, José M. Álvarez 0004, Sanja Fidler, Marc T. Law |
NeurIPS | 3 |
| 2022 | Structural Pruning via Latency-Saliency KnapsackabstractStructural pruning can simplify network architecture and improve inference speed. We propose Hardware-Aware Latency Pruning (HALP) that formulates structural pruning as a global resource allocation optimization problem, aiming at maximizing the accuracy while constraining latency under a predefined budget on targeting device. For filter importance ranking, HALP leverages latency lookup table to track latency reduction potential and global saliency score to gauge accuracy drop. Both metrics can be evaluated very efficiently during pruning, allowing us to reformulate global structural pruning under a reward maximization problem given target constraint. This makes the problem solvable via our augmented knapsack solver, enabling HALP to surpass prior work in pruning efficacy and accuracy-efficiency trade-off. We examine HALP on both classification and detection tasks, over varying networks, on ImageNet and VOC datasets, on different platforms. In particular, for ResNet-50/-101 pruning on ImageNet, HALP improves network throughput by $1.60\times$/$1.90\times$ with $+0.3\%$/$-0.2\%$ top-1 accuracy changes, respectively. For SSD pruning on VOC, HALP improves throughput by $1.94\times$ with only a $0.56$ mAP drop. HALP consistently outperforms prior art, sometimes by large margins. Project page at \url{https://halp-neurips.github.io/}. Maying Shen, Hongxu Yin, Pavlo Molchanov 0001, Jianna Liu, José M. Álvarez 0004 |
NeurIPS | 6 |
| 2022 | Cost Volume Pyramid Based Depth Inference for Multi-View StereoabstractWe propose a cost volume-based neural network for depth inference from multi-view images. We demonstrate that building a cost volume pyramid in a coarse-to-fine manner instead of constructing a cost volume at a fixed resolution leads to a compact, lightweight network and allows us inferring high resolution depth maps to achieve better reconstruction results. To this end, we first build a cost volume based on uniform sampling of fronto-parallel planes across the entire depth range at the coarsest resolution of an image. Then, given current depth estimate, we construct new cost volumes iteratively to perform depth map refinement. We show that working on cost volume pyramid can lead to a more compact, yet efficient network structure compared with existing works. We further show that the (residual) depth sampling can be fully determined by analytical geometric derivation, which serves as a principle for building compact cost volume pyramid. To demonstrate the effectiveness of our proposed framework, we extend our cost volume pyramid structure to handle the unsupervised depth inference scenario. Experimental results on benchmark datasets show that our model can perform 6x faster with similar performance as state-of-the-art methods for supervised scenario and demonstrates superior performance on unsupervised scenario. Code is available at https://github.com/JiayuYANG/CVP-MVSNet. Wei Mao 0001, José M. Álvarez 0004, Miaomiao Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Training Data Subset Search With Ensemble Active LearningabstractDeep Neural Networks (DNNs) often rely on vast datasets for training. Given the large size of such datasets, it is conceivable that they contain specific samples that either do not contribute or negatively impact the DNN’s optimization. Modifying the training distribution to exclude such samples could provide an effective solution to improve performance and reduce training time. This paper proposes to scale up ensemble Active Learning (AL) methods to perform acquisition at a large scale (10k to 500k samples at a time). We do this with ensembles of hundreds of models, obtained at a minimal computational cost by reusing intermediate training checkpoints. This allows us to automatically and efficiently perform a training data subset search for large labeled datasets. We observe that our approach obtains favorable subsets of training data, which can be used to train more accurate DNNs than training with the entire dataset. We perform an extensive experimental study of this phenomenon on three image classification benchmarks (CIFAR-10, CIFAR-100, and ImageNet), as well as an internal object detection benchmark for prototyping perception models for autonomous driving. Unlike existing studies, our experiments on object detection are at the scale required for production-ready autonomous driving systems. We provide insights on the impact of different initialization schemes, acquisition functions, and ensemble configurations at this scale. Our results provide strong empirical evidence that optimizing the training data distribution can significantly benefit large-scale vision tasks. Kashyap Chitta, José M. Álvarez 0004, Elmar Haussmann, Clément Farabet |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Capitalizing on RGB-FIR Hybrid Imaging for Road DetectionabstractTraditionally, road detection approaches mostly capitalize on RGB images, 3D LiDAR point cloud or their fusion. However, RGB camera is sensitive to light conditions, while LiDAR point cloud is sparse compared with dense image pixels. In this work, a new hybrid image dataset is provided for the task of road detection based on cameras. In this dataset, the hybrid images are acquired by an optically aligned hybrid imaging device, consisting of a far-infrared (FIR) imager and an RGB camera to output pixel-wise registration of thermal and RGB frames. Then we investigate on three methods based on fully convolutional neural network (F-CNN) to demonstrate the advantages by fusing RGB-FIR images in road detection. First, a middle-fusion based model is built, where the output feature maps of encoder branches from RGB and FIR images are directly concatenated into a single-fusion branch as the decoder. Next, the originally discarded layers after fusion operation for both RGB and FIR branches are recovered as the mimic branches to imitate the distributions of the fusion outputs, which constitutes an extended cross model (ECM). Moreover, the outputs of mimic branches at different scales are also used to imitate the corresponding outputs in the fusion branch, called a hierarchical cross model (HCM). The experimental results demonstrate the effectiveness and efficiency of our fusion strategies. Yigong Zhang, Jin Xie 0001, José M. Álvarez 0004, Cheng-Zhong Xu 0001, Jian Yang 0003, Hui Kong 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Optimal Quantization Using Scaled CodebookabstractWe study the problem of quantizing N sorted, scalar datapoints with a fixed codebook containing K entries that are allowed to be rescaled. The problem is defined as finding the optimal scaling factor α and the datapoint assignments into the α-scaled codebook to minimize the squared error between original and quantized points. Previously, the globally optimal algorithms for this problem were derived only for certain codebooks (binary and ternary) or under the assumption of certain distributions (Gaussian, Laplacian). By studying the properties of the optimal quantizer, we derive an $\mathcal{O}\left( {NK\log K} \right)$ algorithm that is guaranteed to find the optimal quantization parameters for any fixed codebook regardless of data distribution. We apply our algorithm to synthetic and real-world neural network quantization problems and demonstrate the effectiveness of our approach. Yerlan Idelbayev, Pavlo Molchanov 0001, Maying Shen, Hongxu Yin, Miguel Á. Carreira-Perpiñán, José M. Álvarez 0004 |
CVPR | 6 |
| 2021 | Self-Supervised Learning of Depth Inference for Multi-View StereoabstractRecent supervised multi-view depth estimation networks have achieved promising results. Similar to all supervised approaches, these networks require ground-truth data during training. However, collecting a large amount of multi-view depth data is very challenging. Here, we propose a self-supervised learning framework for multi-view stereo that exploit pseudo labels from the input data. We start by learning to estimate depth maps as initial pseudo labels under an unsupervised learning framework relying on image reconstruction loss as supervision. We then refine the initial pseudo labels using a carefully designed pipeline leveraging depth information inferred from a higher resolution image and neighboring views. We use these high-quality pseudo labels as the supervision signal to train the network and improve, iteratively, its performance by self-training. Extensive experiments on the DTU dataset show that our proposed self-supervised learning framework outperforms existing unsupervised multi-view stereo networks by a large margin and performs on par compared to the supervised counterpart. Code is available at https://github.com/JiayuYANG/Self-supervised-CVP-MVSNet. José M. Álvarez 0004, Miaomiao Liu 0001 |
CVPR | 2 |
| 2021 | See Through Gradients: Image Batch Recovery via GradInversionabstractTraining deep neural networks requires gradient estimation from data batches to update parameters. Gradients per parameter are averaged over a set of data and this has been presumed to be safe for privacy-preserving training in joint, collaborative, and federated learning applications. Prior work only showed the possibility of recovering input data given gradients under very restrictive conditions – a single input point, or a network with no non-linearities, or a small 32 × 32 px input batch. Therefore, averaging gradients over larger batches was thought to be safe. In this work, we introduce GradInversion, using which input images from a larger batch (8 – 48 images) can also be recovered for large networks such as ResNets (50 layers), on complex datasets such as ImageNet (1000 classes, 224 × 224 px). We formulate an optimization task that converts random noise into natural images, matching gradients while regularizing image fidelity. We also propose an algorithm for target class label recovery given gradients. We further propose a group consistency regularization framework, where multiple agents starting from different random seeds work together to find an enhanced reconstruction of the original data batch. We show that gradients encode a surprisingly large amount of information, such that all the individual images can be recovered with high fidelity via GradInversion, even for complex datasets, deep networks, and large batch sizes. Hongxu Yin, Arun Mallya, Arash Vahdat, José M. Álvarez 0004, Jan Kautz, Pavlo Molchanov 0001 |
CVPR | 4 |
| 2021 | Active Learning for Deep Object Detection via Probabilistic ModelingabstractActive learning aims to reduce labeling costs by selecting only the most informative samples on a dataset. Few existing works have addressed active learning for object detection. Most of these methods are based on multiple models or are straightforward extensions of classification methods, hence estimate an image’s informativeness using only the classification head. In this paper, we propose a novel deep active learning approach for object detection. Our approach relies on mixture density networks that estimate a probabilistic distribution for each localization and classification head’s output. We explicitly estimate the aleatoric and epistemic uncertainty in a single forward pass of a single model. Our method uses a scoring function that aggregates these two types of uncertainties for both heads to obtain every image’s informativeness score. We demonstrate the efficacy of our approach in PASCAL VOC and MS-COCO datasets. Our approach outperforms single-model based methods and performs on par with multi-model based methods at a fraction of the computing cost. Code is available at https://github.com/NVlabs/AL-MDN. Jiwoong Choi, Ismail Elezi, Clément Farabet, José M. Álvarez 0004 |
ICCV | 5 |
| 2021 | Contrastive Syn-to-Real Generalization
Wuyang Chen 0001, Zhiding Yu, Shalini De Mello, Sifei Liu, José M. Álvarez 0004, Zhangyang Wang, Anima Anandkumar |
ICLR | 5 |
| 2021 | Personalized Federated Learning with First Order Model Optimization
Karan Sapra, Sanja Fidler, Serena Yeung-Levy, José M. Álvarez 0004 |
ICLR | 5 |
| 2021 | Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed DetectionabstractTraining on datasets with long-tailed distributions has been challenging for major recognition tasks such as classification and detection. To deal with this challenge, image resampling is typically introduced as a simple but effective approach. However, we observe that long-tailed detection differs from classification since multiple classes may be present in one image. As a result, image resampling alone is not enough to yield a sufficiently balanced distribution at the object-level. We address object-level resampling by introducing an object-centric sampling strategy based on a dynamic, episodic memory bank. Our proposed strategy has two benefits: 1) convenient object-level resampling without significant extra computation, and 2) implicit feature-level augmentation from model updates. We show that image-level and object-level resamplings are both important, and thus unify them with a joint resampling strategy. Our method achieves state-of-the-art performance on the rare categories of LVIS, with 1.89% and 3.13% relative improvements over Forest R-CNN on detection and instance segmentation. Nadine Chang, Zhiding Yu, Yu-Xiong Wang, Anima Anandkumar, Sanja Fidler, José M. Álvarez 0004 |
ICML | 6 |
| 2021 | Boosting Supervised Learning Performance with Co-trainingabstractDeep learning perception models require a massive amount of labeled training data to achieve good performance. While unlabeled data is easy to acquire, the cost of labeling is prohibitive and could create a tremendous burden on companies or individuals. Recently, self-supervision has emerged as an alternative to leveraging unlabeled data. In this paper, we propose a new light-weight self-supervised learning framework that could boost supervised learning performance with minimum additional computation cost. Here, we introduce a simple and flexible multi-task co-training framework that integrates a self-supervised task into any supervised task. Our approach exploits pretext tasks to incur minimum compute and parameter overheads and minimal disruption to existing training pipelines. We demonstrate the effectiveness of our framework by using two self-supervised tasks, object detection and panoptic segmentation, on different perception models. Our results show that both self-supervised tasks can improve the accuracy of the supervised task and, at the same time, demonstrates strong domain adaption capability when used with additional unlabeled data. Xinnan Du, José M. Álvarez 0004 |
IV | 3 |
| 2021 | Distilling Image Classifiers in Object DetectorsabstractKnowledge distillation constitutes a simple yet effective way to improve the performance of a compact student network by exploiting the knowledge of a more powerful teacher. Nevertheless, the knowledge distillation literature remains limited to the scenario where the student and the teacher tackle the same task. Here, we investigate the problem of transferring knowledge not only across architectures but also across tasks. To this end, we study the case of object detection and, instead of following the standard detector-to-detector distillation approach, introduce a classifier-to-detector knowledge transfer framework. In particular, we propose strategies to exploit the classification teacher to improve both the detector's recognition accuracy and localization performance. Our experiments on several detectors with different backbones demonstrate the effectiveness of our approach, allowing us to outperform the state-of-the-art detector-to-detector distillation methods. Shuxuan Guo, José M. Álvarez 0004, Mathieu Salzmann |
NeurIPS | 2 |
| 2021 | SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersabstractWe present SegFormer, a simple, efficient yet powerful semantic segmentation framework which unifies Transformers with lightweight multilayer perceptron (MLP) decoders. SegFormer has two appealing features: 1) SegFormer comprises a novel hierarchically structured Transformer encoder which outputs multiscale features. It does not need positional encoding, thereby avoiding the interpolation of positional codes which leads to decreased performance when the testing resolution differs from training. 2) SegFormer avoids complex decoders. The proposed MLP decoder aggregates information from different layers, and thus combining both local attention and global attention to render powerful representations. We show that this simple and lightweight design is the key to efficient segmentation on Transformers. We scale our approach up to obtain a series of models from SegFormer-B0 to Segformer-B5, which reaches much better performance and efficiency than previous counterparts.For example, SegFormer-B4 achieves 50.3% mIoU on ADE20K with 64M parameters, being 5x smaller and 2.2% better than the previous best method. Our best model, SegFormer-B5, achieves 84.0% mIoU on Cityscapes validation set and shows excellent zero-shot robustness on Cityscapes-C. Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, José M. Álvarez 0004, Ping Luo 0002 |
NeurIPS | 5 |
| 2021 | Data-free Knowledge Distillation for Object DetectionabstractWe present DeepInversion for Object Detection (DIODE) to enable data-free knowledge distillation for neural networks trained on the object detection task. From a data-free perspective, DIODE synthesizes images given only an off-the-shelf pre-trained detection network and without any prior domain knowledge, generator network, or pre-computed activations. DIODE relies on two key components-first, an extensive set of differentiable augmentations to improve image fidelity and distillation effectiveness. Second, a novel automated bounding box and category sampling scheme for image synthesis enabling generating a large number of images with a diverse set of spatial and category objects. The resulting images enable data-free knowledge distillation from a teacher to a student detector, initialized from scratch. In an extensive set of experiments, we demonstrate that DIODE's ability to match the original training distribution consistently enables more effective knowledge distillation than out-of-distribution proxy datasets, which unavoidably occur in a data-free setup given the absence of the original domain knowledge. Akshay Chawla, Hongxu Yin, Pavlo Molchanov 0001, José M. Álvarez 0004 |
WACV | 4 |
| 2020 | Cost Volume Pyramid Based Depth Inference for Multi-View StereoabstractWe propose a cost volume-based neural network for depth inference from multi-view images. We demonstrate that building a cost volume pyramid in a coarse-to-fine manner instead of constructing a cost volume at a fixed resolution leads to a compact, lightweight network and allows us inferring high resolution depth maps to achieve better reconstruction results. To this end, we first build a cost volume based on uniform sampling of fronto-parallel planes across the entire depth range at the coarsest resolution of an image. Then, given current depth estimate, we construct new cost volumes iteratively on the pixelwise depth residual to perform depth map refinement. While sharing similar insight with Point-MVSNet as predicting and refining depth iteratively, we show that working on cost volume pyramid can lead to a more compact, yet efficient network structure compared with the Point-MVSNet on 3D points. We further provide detailed analyses of the relation between (residual) depth sampling and image resolution, which serves as a principle for building compact cost volume pyramid. Experimental results on benchmark datasets show that our model can perform 6× faster and has similar performance as state-of-the-art methods. Code is available at https://github.com/JiayuYANG/CVP-MVSNet. Wei Mao 0001, José M. Álvarez 0004, Miaomiao Liu 0001 |
CVPR | 3 |
| 2020 | Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversionabstractWe introduce DeepInversion, a new method for synthesizing images from the image distribution used to train a deep neural network. We ``invert'' a trained network (teacher) to synthesize class-conditional input images starting from random noise, without using any additional information about the training dataset. Keeping the teacher fixed, our method optimizes the input while regularizing the distribution of intermediate feature maps using information stored in the batch normalization layers of the teacher. Further, we improve the diversity of synthesized images using Adaptive DeepInversion, which maximizes the Jensen-Shannon divergence between the teacher and student network logits. The resulting synthesized images from networks trained on the CIFAR-10 and ImageNet datasets demonstrate high fidelity and degree of realism, and help enable a new breed of data-free applications - ones that do not require any real images or labeled data. We demonstrate the applicability of our proposed method to three tasks of immense practical importance - (i) data-free network pruning, (ii) data-free knowledge transfer, and (iii) data-free continual learning. Hongxu Yin, Pavlo Molchanov 0001, José M. Álvarez 0004, Zhizhong Li 0001, Arun Mallya, Derek Hoiem, Niraj K. Jha, Jan Kautz |
CVPR | 3 |
| 2020 | Scalable Active Learning for Object DetectionabstractDeep Neural Networks trained in a fully supervised fashion are the dominant technology in perception-based autonomous driving systems. While collecting large amounts of unlabeled data is already a major undertaking, only a subset of it can be labeled by humans due to the effort needed for high-quality annotation. Therefore, finding the right data to label has become a key challenge. Active learning is a powerful technique to improve data efficiency for supervised learning methods, as it aims at selecting the smallest possible training set to reach a required performance. We have built a scalable production system for active learning in the domain of autonomous driving. In this paper, we describe the resulting high-level design, sketch some of the challenges and their solutions, present our current results at scale, and briefly describe the open problems and future directions. Elmar Haussmann, Michele Fenzi, Kashyap Chitta, Jan Ivanecky, Hanson Xu, Donna Roy, Akshita Mittel, Nicolas Koumchatzky, Clément Farabet, José M. Álvarez 0004 |
IV | 10 |
| 2020 | ExpandNets: Linear Over-parameterization to Train Compact Convolutional NetworksabstractWe introduce an approach to training a given compact network. To this end, we leverage over-parameterization, which typically improves both neural network optimization and generalization. Specifically, we propose to expand each linear layer of the compact network into multiple consecutive linear layers, without adding any nonlinearity. As such, the resulting expanded network, or ExpandNet, can be contracted back to the compact one algebraically at inference. In particular, we introduce two convolutional expansion strategies and demonstrate their benefits on several tasks, including image classification, object detection, and semantic segmentation. As evidenced by our experiments, our approach outperforms both training the compact network from scratch and performing knowledge distillation from a teacher. Furthermore, our linear over-parameterization empirically reduces gradient confusion during training and improves the network generalization. Shuxuan Guo, José M. Álvarez 0004, Mathieu Salzmann |
NeurIPS | 2 |
| 2020 | Quadtree Generating Networks: Efficient Hierarchical Scene Parsing with Sparse ConvolutionsabstractSemantic segmentation with Convolutional Neural Networks is a memory-intensive task due to the high spatial resolution of feature maps and output predictions. In this paper, we present Quadtree Generating Networks (QGNs), a novel approach able to drastically reduce the memory footprint of modern semantic segmentation networks. The key idea is to use quadtrees to represent the predictions and target segmentation masks instead of dense pixel grids. Our quadtree representation enables hierarchical processing of an input image, with the most computationally demanding layers only being used at regions in the image containing boundaries between classes. In addition, given a trained model, our representation enables flexible inference schemes to trade-off accuracy and computational cost, allowing the network to adapt in constrained situations such as embedded devices. We demonstrate the benefits of our approach on the Cityscapes, SUN-RGBD and ADE20k datasets. On Cityscapes, we obtain an relative 3% mIoU improvement compared to a dilated network with similar memory consumption; and only receive a 3% relative mIoU drop compared to a large dilated network, while reducing memory consumption by over 4×. Our code is available at https://github.com/kashyap7x/QGN. Kashyap Chitta, José M. Álvarez 0004, Martial Hebert |
WACV | 2 |
| 2020 | Context Based Emotion Recognition Using EMOTIC DatasetabstractIn our everyday lives and social interactions we often try to perceive the emotional states of people. There has been a lot of research in providing machines with a similar capacity of recognizing emotions. From a computer vision perspective, most of the previous efforts have been focusing in analyzing the facial expressions and, in some cases, also the body pose. Some of these methods work remarkably well in specific settings. However, their performance is limited in natural, unconstrained environments. Psychological studies show that the scene context, in addition to facial expression and body pose, provides important information to our perception of people's emotions. However, the processing of the context for automatic emotion recognition has not been explored in depth, partly due to the lack of proper data. In this paper we present EMOTIC, a dataset of images of people in a diverse set of natural situations, annotated with their apparent emotion. The EMOTIC dataset combines two different types of emotion representation: (1) a set of 26 discrete categories, and (2) the continuous dimensions Valence, Arousal, and Dominance. We also present a detailed statistical and algorithmic analysis of the dataset along with annotators' agreement analysis. Using the EMOTIC dataset we train different CNN models for emotion recognition, combining the information of the bounding box containing the person with the contextual information extracted from the scene. Our results show how scene context provides important information to automatically recognize emotional states and motivate further research in this direction. Ronak Kosti, José M. Álvarez 0004, Adrià Recasens, Àgata Lapedriza |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Two-View Fusion based Convolutional Neural Network for Urban Road DetectionabstractIn this paper, we propose a two-view fusion based convolutional neural network to estimate road areas in urban environments with LiDAR point clouds as input only. The proposed network takes two transformed LiDAR data representations, the LiDAR imageries and the camera-perspective maps, as inputs. It outputs pixel-wise road detection results in both the LiDAR's imagery view and the camera's perspective view simultaneously, in an end-to-end manner. To make better use of the data associations between two representations, we construct a novel mapping layer to transform features from the LiDAR's imagery view to the camera's perspective view in order to strengthen the road detection performance in the camera's perspective view. Experiments on the KITTI-Road dataset show that the proposed network can achieve the state-of-the-art performance among all LiDAR-only methods in real time. Shuo Gu, Yigong Zhang, Jian Yang 0003, José M. Álvarez 0004, Hui Kong 0001 |
IROS | 4 |
| 2019 | Bridging the Day and Night Domain Gap for Semantic SegmentationabstractPerception in autonomous vehicles has progressed exponentially in the last years thanks to the advances of vision-based methods such as Convolutional Neural Networks (CNNs). Current deep networks are both efficient and reliable, at least in standard conditions, standing as a suitable solution for the perception tasks of autonomous vehicles. However, there is a large accuracy downgrade when these methods are taken to adverse conditions such as nighttime. In this paper, we study methods to alleviate this accuracy gap by using recent techniques such as Generative Adversarial Networks (GANs). We explore diverse options such as enlarging the dataset to cover these domains in unsupervised training or adapting the images on-the-fly during inference to a comfortable domain such as sunny daylight in a pre-processing step. The results show some interesting insights and demonstrate that both proposed approaches considerably reduce the domain gap, allowing IV perception systems to work reliably also at night. Eduardo Romera, Luis Miguel Bergasa, Kailun Yang 0001, José M. Álvarez 0004, Rafael Barea |
IV | 4 |
| 2018 | Effective Use of Synthetic Data for Urban Scene Semantic Segmentation
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004 |
ECCV (2) | 5 |
| 2018 | Train Here, Deploy There: Robust Segmentation in Unseen DomainsabstractSemantic Segmentation methods play a key role in today’s Autonomous Driving research, since they provide a global understanding of the traffic scene for upper-level tasks like navigation. However, main research efforts are being put on enlarging deep architectures to achieve marginal accuracy boosts in existing datasets, forgetting that these algorithms must be deployed in a real vehicle with images that were not seen during training. On the other hand, achieving robustness in any domain is not an easy task, since deep networks are prone to overfitting even with thousands of training images. In this paper, we study in a systematic way what is the gap between the concepts of “accuracy” and “robustness”. A comprehensive set of experiments demonstrates the relevance of using data augmentation to yield models that can produce robust semantic segmentation outputs in any domain. Our results suggest that the existing domain gap can be significantly reduced when appropriate augmentation techniques regarding geometry (position and shape) and texture (color and illumination) are applied. In addition, the proposed training process results in better calibrated models, which is of special relevance to assess the robustness of current systems. Eduardo Romera, Luis Miguel Bergasa, José M. Álvarez 0004, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 3 |
| 2018 | Fusion of LiDAR and Camera by Scanning in LiDAR Imagery and Image-Guided Diffusion for Urban Road DetectionabstractThis paper proposes a new method for road detection based on a 3D LiDAR and a camera. First, the original LiDAR point cloud is re-organized in an ordered way to generate a LiDAR imagery. Then the flat region is extracted from the LiDAR imagery as the candidate road region. Next, a strategy of row- and column- scanning is given in the LiDAR imagery to detect a finer road region from the candidate region. To fuse the point cloud with image information, we transform the point cloud that corresponds to the above detected road region to the image space according to the calibration parameters between the LiDAR and camera. Then we give two image-guided diffusion schemes to conduct image segmentation of road area, respectively. Our experiments demonstrate that this training free approach detects the road region fast, accurately and robustly, and compares favorably with the state-of-the-art on the KITTI benchmark. Yigong Zhang, Shuo Gu, Jian Yang 0003, José M. Álvarez 0004, Hui Kong 0001 |
Intelligent Vehicles Symposium | 4 |
| 2018 | Frame selection for OCR from video stream of book flipping
Dibyayan Chakraborty, Partha Pratim Roy 0001, Rajkumar Saini, José M. Álvarez 0004, Umapada Pal 0001 |
Multim. Tools Appl. | 4 |
| 2018 | Expression-Invariant Age Estimation Using Structured LearningabstractIn this paper, we investigate and exploit the influence of facial expressions on automatic age estimation. Different from existing approaches, our method jointly learns the age and expression by introducing a new graphical model with a latent layer between the age/expression labels and the features. This layer aims to learn the relationship between the age and expression and captures the face changes which induce the aging and expression appearance, and thus obtaining expression-invariant age estimation. Conducted on three age-expression datasets (FACES , Lifespan and NEMO ), our experiments illustrate the improvement in performance when the age is jointly learnt with expression in comparison to expression-independent age estimation. The age estimation error is reduced by 14.43, 37.75 and 9.30 percent for the FACES, Lifespan and NEMO datasets respectively. The results obtained by our graphical model, without prior-knowledge of the expressions of the tested faces, are better than the best reported ones for all datasets. The flexibility of the proposed model to include more cues is explored by incorporating gender together with age and expression. The results show performance improvements for all cues. Zhongyu Lou, Fares Alnajar, José M. Álvarez 0004, Ninghang Hu, Theo Gevers |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Incorporating Network Built-in Priors in Weakly-Supervised Semantic SegmentationabstractPixel-level annotations are expensive and time consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recently, CNN-based methods have proposed to fine-tune pre-trained networks using image tags. Without additional information, this leads to poor localization accuracy. This problem, however, was alleviated by making use of objectness priors to generate foreground/background masks. Unfortunately these priors either require pixel-level annotations/bounding boxes, or still yield inaccurate object boundaries. Here, we propose a novel method to extract accurate masks from networks pre-trained for the task of object recognition, thus forgoing external objectness modules. We first show how foreground/background masks can be obtained from the activations of higher-level convolutional layers of a network. We then show how to obtain multi-class masks by the fusion of foreground/background ones with information extracted from a weakly-supervised localization network. Our experiments evidence that exploiting these masks in conjunction with a weakly-supervised training loss yields state-of-the-art tag-based weakly-supervised semantic segmentation results. Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004, Stephen Gould |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | ERFNet: Efficient Residual Factorized ConvNet for Real-Time Semantic SegmentationabstractSemantic segmentation is a challenging task that addresses most of the perception needs of intelligent vehicles (IVs) in an unified way. Deep neural networks excel at this task, as they can be trained end-to-end to accurately classify multiple object categories in an image at pixel level. However, a good tradeoff between high quality and computational resources is yet not present in the state-of-the-art semantic segmentation approaches, limiting their application in real vehicles. In this paper, we propose a deep architecture that is able to run in real time while providing accurate semantic segmentation. The core of our architecture is a novel layer that uses residual connections and factorized convolutions in order to remain efficient while retaining remarkable accuracy. Our approach is able to run at over 83 FPS in a single Titan X, and 7 FPS in a Jetson TX1 (embedded device). A comprehensive set of experiments on the publicly available Cityscapes data set demonstrates that our system achieves an accuracy that is similar to the state of the art, while being orders of magnitude faster to compute than other architectures that achieve top precision. The resulting tradeoff makes our model an ideal approach for scene understanding in IV applications. The code is publicly available at: https://github.com/Eromera/erfnet. Eduardo Romera, José M. Álvarez 0004, Luis Miguel Bergasa, Roberto Arroyo |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Guest Editorial Introduction to the Special Issue on Robust and Efficient Vision Techniques for Intelligent VehiclesabstractIn recent years, intelligent vehicles have been a hot topic for both research and industry communities. Since the whole system is a comprehensive integration of many advanced techniques, their respective development and improvement become fundamentally important. Qi Wang 0009, Luis Miguel Bergasa, José M. Álvarez 0004 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Emotion Recognition in ContextabstractUnderstanding what a person is experiencing from her frame of reference is essential in our everyday life. For this reason, one can think that machines with this type of ability would interact better with people. However, there are no current systems capable of understanding in detail peoples emotional states. Previous research on computer vision to recognize emotions has mainly focused on analyzing the facial expression, usually classifying it into the 6 basic emotions [11]. However, the context plays an important role in emotion perception, and when the context is incorporated, we can infer more emotional states. In this paper we present the Emotions in Context Database (EMCO), a dataset of images containing people in context in non-controlled environments. In these images, people are annotated with 26 emotional categories and also with the continuous dimensions valence, arousal, and dominance [21]. With the EMCO dataset, we trained a Convolutional Neural Network model that jointly analyzes the person and the whole scene to recognize rich information about emotional states. With this, we show the importance of considering the context for recognizing peoples emotions in images, and provide a benchmark in the task of emotion recognition in visual context. Ronak Kosti, José M. Álvarez 0004, Adrià Recasens, Àgata Lapedriza |
CVPR | 2 |
| 2017 | Domain-Adaptive Deep Network Compression
Marc Masana, Joost van de Weijer 0001, Luis Herranz, Andrew D. Bagdanov, José M. Álvarez 0004 |
ICCV | 5 |
| 2017 | Bringing Background into the Foreground: Making All Classes Equal in Weakly-Supervised Video Semantic SegmentationabstractPixel-level annotations are expensive and time-consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recent years have seen great progress in weakly-supervised semantic segmentation, whether from a single image or from videos. However, most existing methods are designed to handle a single background class. In practical applications, such as autonomous navigation, it is often crucial to reason about multiple background classes. In this paper, we introduce an approach to doing so by making use of classifier heatmaps. We then develop a two-stream deep architecture that jointly leverages appearance and motion, and design a loss based on our heatmaps to train it. Our experiments demonstrate the benefits of our classifier heatmaps and of our two-stream architecture on challenging urban scene datasets and on the YouTube-Objects benchmark, where we obtain state-of-the-art results. Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004 |
ICCV | 5 |
| 2017 | Efficient ConvNet for real-time semantic segmentationabstractSemantic segmentation is a task that covers most of the perception needs of intelligent vehicles in an unified way. ConvNets excel at this task, as they can be trained end-to-end to accurately classify multiple object categories in an image at the pixel level. However, current approaches normally involve complex architectures that are expensive in terms of computational resources and are not feasible for ITS applications. In this paper, we propose a deep architecture that is able to run in real-time while providing accurate semantic segmentation. The core of our ConvNet is a novel layer that uses residual connections and factorized convolutions in order to remain highly efficient while still retaining remarkable performance. Our network is able to run at 83 FPS in a single Titan X, and at more than 7 FPS in a Jetson TX1 (embedded GPU). A comprehensive set of experiments demonstrates that our system, trained from scratch on the challenging Cityscapes dataset, achieves a classification performance that is among the state of the art, while being orders of magnitude faster to compute than other architectures that achieve top precision. This makes our model an ideal approach for scene understanding in intelligent vehicles applications. Eduardo Romera, José M. Álvarez 0004, Luis Miguel Bergasa, Roberto Arroyo |
Intelligent Vehicles Symposium | 2 |
| 2017 | Compression-aware Training of Deep NetworksabstractIn recent years, great progress has been made in a variety of application domains thanks to the development of increasingly deeper neural networks. Unfortunately, the huge number of units of these networks makes them expensive both computationally and memory-wise. To overcome this, exploiting the fact that deep networks are over-parametrized, several compression strategies have been proposed. These methods, however, typically start from a network that has been trained in a standard manner, without considering such a future compression. In this paper, we propose to explicitly account for compression in the training process. To this end, we introduce a regularizer that encourages the parameter matrix of each layer to have low rank during training. We show that accounting for compression during training allows us to learn much more compact, yet at least as effective, models than state-of-the-art compression techniques. José M. Álvarez 0004, Mathieu Salzmann |
NIPS | 1 |
| 2016 | Learning Image Matching by Simply Watching Video
Gucan Long, Laurent Kneip, José M. Álvarez 0004, Hongdong Li |
ECCV (6) | 3 |
| 2016 | Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, Stephen Gould, José M. Álvarez 0004 |
ECCV (8) | 6 |
| 2016 | Latent structural SVM with marginal probabilities for weakly labeled structured learningabstractIn the last years, the increasing availability of annotated data has facilitated the great success of supervised learning in real-world applications such as semantic labeling. However, the vast majority of data is nowadays unlabeled or partially annotated. In this paper, we develop an Expected Marginal Latent Structural SVM (EM-LSSVM) framework for performing structured learning in the presence of weakly (partially) annotated data by incorporating the uncertainty of the unobserved data as marginals. Experimental results on semantic labeling show the potential of the proposed method. In particular, we learn the parameters of a CRF where large amounts of noisy and unobserved data are available. Comparison against state of the art demonstrates the applicability of our algorithm to practical applications. Shahin Namin, José M. Álvarez 0004, Laurent Kneip, Lars Petersson |
ICIP | 2 |
| 2016 | 2D-3D semantic segmentation using cardinality as higher-order lossabstractMulti-modal scene analysis is a growing field of importance as additional sensors, such as 3D LIDAR, is becoming a common complement to image capturing systems. However, while additional sensory data potentially can make the analysis more accurate, it also comes with a host of associated issues. For example, inconsistencies in the data between sensors resulting from, e.g., misalignment, moving objects, or parallax effects, can severely affect the performance. Additionally, real-world scenes tend to have an inherent imbalance in the number of items of each class which typically suppresses the performance of infrequent classes. In this paper, we address those two issues specifically by a) using a cardinality loss function designed to target inconsistencies at training time, and b) devising an average per class loss function addressing the imbalance issue. Shahin Namin, José M. Álvarez 0004, Lars Petersson |
ICPR | 2 |
| 2016 | Learning the Number of Neurons in Deep NetworksabstractNowadays, the number of layers and of neurons in each layer of a deep network are typically set manually. While very deep and wide networks have proven effective in general, they come at a high memory and computation cost, thus making them impractical for constrained platforms. These networks, however, are known to have many redundant parameters, and could thus, in principle, be replaced by more compact architectures. In this paper, we introduce an approach to automatically determining the number of neurons in each layer of a deep network during learning. To this end, we propose to make use of a group sparsity regularizer on the parameters of the network, where each group is defined to act on a single neuron. Starting from an overcomplete network, we show that our approach can reduce the number of parameters by up to 80\% while retaining or even improving the network accuracy. José M. Álvarez 0004, Mathieu Salzmann |
NIPS | 1 |
| 2016 | Efficient transductive semantic segmentationabstractSemantically describing the contents of images is one of the classical problems of computer vision. With huge numbers of images being made available daily, there is increasing interest in methods for semantic pixel labelling that exploit large image sets. Graph transduction provides a framework for the flexible inclusion of labeled data that can be exploited in the classification of unlabeled samples without requiring a trained classifier. Unfortunately, current approaches lack the scalability to tackle the joint segmentation of large image sets. Here we introduce an efficient flexible graph transduction approach to semantic segmentation that allows simple and efficient leveraging of large image sets without requiring separate computation of unary potentials, or a trained classifier. We demonstrate that this technique can handle far larger graphs than previous methods, and that results continue to improve as more labeled images are made available. Furthermore, we show that the method is able to benefit from dense or sparse unary labels when they are available. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
WACV | 1 |
| 2016 | Semantic labeling for prosthetic vision
Lachlan Horne, José M. Álvarez 0004, Chris McCarthy, Mathieu Salzmann, Nick Barnes |
Comput. Vis. Image Underst. | 2 |
| 2016 | Exploiting Large Image Sets for Road Scene ParsingabstractThere is an increasing interest in exploiting multiple images for scene understanding, with great progress in areas such as cosegmentation and video segmentation. Jointly analyzing the images in a large set offers the opportunity to exploit a greater source of information than when considering a single image on its own. However, this also yields challenges since, to effectively exploit all the available information, the resulting methods need to consider not just local connections, but efficiently analyze similarity between all pairs of pixels within and across all the images. In this paper, we propose to model an image set as a fully connected pairwise Conditional Random Field (CRF) defined over the image pixels, or superpixels, with Gaussian edge potentials. We show that this lets us co-label the images of a large set efficiently, thus yielding increased accuracy at no additional computational cost compared to sequential labeling of the images. Furthermore, we extend our framework to incorporate temporal dependence, thus effectively encompassing video segmentation as a special case of our approach, as well as to modeling label dependence over larger image regions. Our experimental evaluation demonstrates that our framework lets us handle over 10 000 images in a matter of seconds. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | Efficient scene parsing by sampling unary potentials in a fully-connected CRFabstractEfficient, fully-connected CRF inference enables fast semantic labelling of images. However, this requires high-quality unary potentials to be computed, which is currently time-consuming. While some recent work attempts to address this issue by only computing a subset of unary potentials, a need remains for a simple, fast way to decide which unary potentials should be computed, without sacrificing accuracy. In particular, for embedded applications, a method which avoids time or memory-intensive operations is desired. In this paper, we introduce an approach to selecting good locations to compute unary potentials. We implement an efficient morphological approach to select a small proportion of pixel locations where unary potentials will be calculated. The speed of our labelling method allows us to directly search a large parameter space to optimize our method for a given task. We show that our method can achieve comparable accuracy to what can be achieved when all unary potentials are calculated, with significant time saving. Furthermore, we show that it is possible to tune our method to yield improved accuracy for certain classes of interest. We demonstrate this over multiple datasets representing challenging applications for our approach. Lachlan Horne, José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
Intelligent Vehicles Symposium | 2 |
| 2015 | Unsupervised image transformation for outdoor semantic labellingabstractSemantic labelling of urban images is a crucial component towards autonomous driving. The accuracy of current methods is highly dependent on the training set being used and drops drastically when the distribution in the test image does not match the expected distribution of the training set. This situation will inevitably occur, as for instance, when the illumination changes from daytime to dusk. To address this problem we propose a fast unsupervised image transformation approach following a global color transfer strategy. Our proposal generalizes classical one-to-one color transfer schemes to the more suitable one-to-many scheme. In addition, our approach can naturally deal with the temporal consistency of video streams to perform a coherent transformation. We demonstrate the benefits of our proposal in two publicly available datasets using different state-of-the-art semantic labelling frameworks. Germán Ros 0001, José M. Álvarez 0004 |
Intelligent Vehicles Symposium | 2 |
| 2014 | Expression-Invariant Age Estimation
Fares Alnajar, Zhongyu Lou, José M. Álvarez 0004, Theo Gevers |
BMVC | 3 |
| 2014 | Fast road detection and tracking in aerial videosabstractWe propose a fast approach for detecting and tracking a specific road in aerial videos. It combines adaptive Gaussian Mixture Models (GMMs) to describe road colour distributions, and homography based tracking to track road geometries, where an efficient technique is developed to estimate homography transformations between two frames. Experiments are conducted on videos captured by our unmanned aerial vehicles. All the results demonstrate the effectiveness of our proposed method. We test 1755 frames from 5 videos. Our approach can achieve 0.032 seconds per frame and 2.64% segmentation error for images with 908 × 513 resolutions, on average. Hailing Zhou, Hui Kong 0001, José M. Álvarez 0004, Douglas C. Creighton, Saeid Nahavandi |
Intelligent Vehicles Symposium | 3 |
| 2014 | Large-scale semantic co-labeling of image setsabstractAs evidenced by video segmentation and cosegmentation approaches, exploiting multiple images is key to the success of visual scene understanding. With the availability of increasingly large sets of images, there is a clear need for methods that can efficiently analyze the similarities and structure across huge numbers of image pixels. Furthermore, to make effective use of this data, these similarities should not just be considered locally between neighboring pixels, but between all pairs of pixels across all images. In this paper, we tackle this challenging scenario by introducing a semantic co-labeling approach that performs efficient inference in a fully-connected CRF defined over the pixels, or superpixels, of an image set. Our experimental evaluation demonstrates that our approach yields improved accuracy while coming at no additional computation cost compared to performing segmentation sequentially on individual images. Furthermore, our formulation lets us perform inference over ten thousand images in a matter of seconds. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
WACV | 1 |
| 2014 | Data-driven road detectionabstractIn this paper, we tackle the problem of road detection from RGB images. In particular, we follow a data-driven approach to segmenting the road pixels in an image. To this end, we introduce two road detection methods: A top-down approach that builds an image-level road prior based on the traffic pattern observed in an input image, and a bottom-up technique that estimates the probability that an image superpixel belongs to the road surface in a nonparametric manner. Both our algorithms work on the principle of label transfer in the sense that the road prior is directly constructed from the ground-truth segmentations of training images. Our experimental evaluation on four different datasets shows that this approach outperforms existing top-down and bottom-up techniques, and is key to the robustness of road detection algorithms to the dataset bias. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
WACV | 1 |
| 2014 | Combining Priors, Appearance, and Context for Road DetectionabstractDetecting the free road surface ahead of a moving vehicle is an important research topic in different areas of computer vision, such as autonomous driving or car collision warning. Current vision-based road detection methods are usually based solely on low-level features. Furthermore, they generally assume structured roads, road homogeneity, and uniform lighting conditions, constraining their applicability in real-world scenarios. In this paper, road priors and contextual information are introduced for road detection. First, we propose an algorithm to estimate road priors online using geographical information, providing relevant initial information about the road location. Then, contextual cues, including horizon lines, vanishing points, lane markings, 3-D scene layout, and road geometry, are used in addition to low-level cues derived from the appearance of roads. Finally, a generative model is used to combine these cues and priors, leading to a road detection method that is, to a large degree, robust to varying imaging conditions, road types, and scenarios. José M. Álvarez 0004, Antonio M. López 0001, Theo Gevers, Felipe Lumbreras |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2013 | Learning appearance models for road detectionabstractWe introduce an approach to image-based road detection that exploits the availability of unannotated training images to learn an appearance model. Our approach allows us to remove the standard assumption that the lower part of the input image belongs to the road surface, which does not always hold and often yields strongly biased appearance models. Instead, we exploit this assumption in the training images, which yields a much more general appearance model. We then use the learned model to classify the pixels of an input image as road or background without requiring any assumptions about this image. Our experimental evaluation shows the benefits of our approach over existing methods in challenging real-world driving scenarios. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
Intelligent Vehicles Symposium | 1 |
| 2013 | Road Geometry Classification by Adaptive Shape ModelsabstractVision-based road detection is important for different applications in transportation, such as autonomous driving, vehicle collision warning, and pedestrian crossing detection. Common approaches to road detection are based on low-level road appearance (e.g., color or texture) and neglect of the scene geometry and context. Hence, using only low-level features makes these algorithms highly depend on structured roads, road homogeneity, and lighting conditions. Therefore, the aim of this paper is to classify road geometries for road detection through the analysis of scene composition and temporal coherence. Road geometry classification is proposed by building corresponding models from training images containing prototypical road geometries. We propose adaptive shape models where spatial pyramids are steered by the inherent spatial structure of road images. To reduce the influence of lighting variations, invariant features are used. Large-scale experiments show that the proposed road geometry classifier yields a high recognition rate of 73.57% ± 13.1, clearly outperforming other state-of-the-art methods. Including road shape information improves road detection results over existing appearance-based methods. Finally, it is shown that invariant features and temporal information provide robustness against disturbing imaging conditions. José M. Álvarez 0004, Theo Gevers, Ferran Diego, Antonio M. López 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2012 | Road Scene Segmentation from a Single Image
José M. Álvarez 0004, Theo Gevers, Yann LeCun, Antonio M. López 0001 |
ECCV (7) | 1 |
| 2011 | Road Detection Based on Illuminant InvarianceabstractBy using an onboard camera, it is possible to detect the free road surface ahead of the ego-vehicle. Road detection is of high relevance for autonomous driving, road departure warning, and supporting driver-assistance systems such as vehicle and pedestrian detection. The key for vision-based road detection is the ability to classify image pixels as belonging or not to the road surface. Identifying road pixels is a major challenge due to the intraclass variability caused by lighting conditions. A particularly difficult scenario appears when the road surface has both shadowed and nonshadowed areas. Accordingly, we propose a novel approach to vision-based road detection that is robust to shadows. The novelty of our approach relies on using a shadow-invariant feature space combined with a model-based classifier. The model is built online to improve the adaptability of the algorithm to the current lighting and the presence of other vehicles in the scene. The proposed algorithm works in still images and does not depend on either road shape or temporal restrictions. Quantitative and qualitative experiments on real-world road sequences with heavy traffic and shadows show that the method is robust to shadows and lighting variations. Moreover, the proposed method provides the highest performance when compared with hue-saturation-intensity (HSI)-based algorithms. José M. Álvarez 0004, Antonio M. López 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2011 | A New Framework for Stereo Sensor Pose Through Road Segmentation and RegistrationabstractThis paper proposes a new framework for real-time estimation of the onboard stereo head's position and orientation relative to the road surface, which is required for any advanced driver-assistance application. This framework can be used with all road types: highways, urban, etc. Unlike existing works that rely on feature extraction in either the image domain or 3-D space, we propose a framework that directly estimates the unknown parameters from the stream of stereo pairs' brightness. The proposed approach consists of two stages that are invoked for every stereo frame. The first stage segments the road region in one monocular view. The second stage estimates the camera pose using a featureless registration between the segmented monocular road region and the other view in the stereo pair. This paper has two main contributions. The first contribution combines a road segmentation algorithm with a registration technique to estimate the online stereo camera pose. The second contribution solves the registration using a featureless method, which is carried out using two different optimization techniques: 1) the differential evolution algorithm and 2) the Levenberg-Marquardt (LM) algorithm. We provide experiments and evaluations of performance. The results presented show the validity of our proposed framework. Fadi Dornaika, José M. Álvarez 0004, Angel Domingo Sappa, Antonio M. López 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2010 | 3D Scene priors for road detectionabstractVision-based road detection is important in different areas of computer vision such as autonomous driving, car collision warning and pedestrian crossing detection. However, current vision-based road detection methods are usually based on low-level features and they assume structured roads, road homogeneity, and uniform lighting conditions. Therefore, in this paper, contextual 3D information is used in addition to low-level cues. Low-level photometric invariant cues are derived from the appearance of roads. Contextual cues used include horizon lines, vanishing points, 3D scene layout and 3D road stages. Moreover, temporal road cues are included. All these cues are sensitive to different imaging conditions and hence are considered as weak cues. Therefore, they are combined to improve the overall performance of the algorithm. To this end, the low-level, contextual and temporal cues are combined in a Bayesian framework to classify road sequences. Large scale experiments on road sequences show that the road detection method is robust to varying imaging conditions, road types, and scenarios (tunnels, urban and highway). Further, using the combined cues outperforms all other individual cues. Finally, the proposed method provides highest road detection accuracy when compared to state-of-the-art methods. José M. Álvarez 0004, Theo Gevers, Antonio M. López 0001 |
CVPR | 1 |
| 2010 | Geographic information for vision-based road detectionabstractRoad detection is a vital task for the development of autonomous vehicles. The knowledge of the free road surface ahead of the target vehicle can be used for autonomous driving, road departure warning, as well as to support advanced driver assistance systems like vehicle or pedestrian detection. Using vision to detect the road has several advantages in front of other sensors: richness of features, easy integration, low cost or low power consumption. Common vision-based road detection approaches use low-level features (such as color or texture) as visual cues to group pixels exhibiting similar properties. However, it is difficult to foresee a perfect clustering algorithm since roads are in outdoor scenarios being imaged from a mobile platform. In this paper, we propose a novel high-level approach to vision-based road detection based on geographical information. The key idea of the algorithm is exploiting geographical information to provide a rough detection of the road. Then, this segmentation is refined at low-level using color information to provide the final result. The results presented show the validity of our approach. José M. Álvarez 0004, Felipe Lumbreras, Theo Gevers, Antonio M. López 0001 |
Intelligent Vehicles Symposium | 1 |
| 2010 | Learning Photometric Invariance for Object Detection
José M. Álvarez 0004, Theo Gevers, Antonio M. López 0001 |
Int. J. Comput. Vis. | 1 |
| 2009 | Learning photometric invariance from diversified color model ensemblesabstractColor is a powerful visual cue for many computer vision applications such as image segmentation and object recognition. However, most of the existing color models depend on the imaging conditions affecting negatively the performance of the task at hand. Often, a reflection model (e.g., Lambertian or dichromatic reflectance) is used to derive color invariant models. However, those reflection models might be too restricted to model real-world scenes in which different reflectance mechanisms may hold simultaneously. Therefore, in this paper, we aim to derive color invariance by learning from color models to obtain diversified color invariant ensembles. First, a photometrical orthogonal and non-redundant color model set is taken on input composed of both color variants and invariants. Then, the proposed method combines and weights these color models to arrive at a diversified color ensemble yielding a proper balance between invariance (repeatability) and discriminative power (distinctiveness). To achieve this, the fusion method uses a multi-view approach to minimize the estimation error. In this way, the method is robust to data uncertainty and produces properly diversified color invariant ensembles. Experiments are conducted on three different image datasets to validate the method. From the theoretical and experimental results, it is concluded that the method is robust against severe variations in imaging conditions. The method is not restricted to a certain reflection model or parameter tuning. Further, the method outperforms state-of- the-art detection techniques in the field of object, skin and road recognition. José M. Álvarez 0004, Theo Gevers, Antonio M. López 0001 |
CVPR | 1 |
| 2009 | Automatic ground-truthing using video registration for on-board detection algorithmsabstractGround-truth data is essential for the objective evaluation of object detection methods in computer vision. Many works claim their method is robust but they support it with experiments which are not quantitatively assessed with regard some ground-truth. This is one of the main obstacles to properly evaluate and compare such methods. One of the main reasons is that creating an extensive and representative ground-truth is very time consuming, specially in the case of video sequences, where thousands of frames have to be labelled. Could such a ground-truth be generated, at least in part, automatically? Though it may seem a contradictory question, we show that this is possible for the case of video sequences recorded from a moving camera. The key idea is transferring existing frame segmentations from a reference sequence into another video sequence recorded at a different time on the same track, possibly under a different ambient lighting. We have carried out experiments on several video sequence pairs and quantitatively assessed the precision of the transformed ground-truth, which prove that our approach is not only feasible but also quite accurate. José M. Álvarez 0004, Ferran Diego, Antonio M. López 0001, Joan Serrat 0002, Daniel Ponsa |
ICIP | 1 |
| 2009 | Vision-based road detection using road modelsabstractVision-based road detection is very challenging since the road is in an outdoor scenario imaged from a mobile platform. In this paper, a new top-down road detection algorithm is proposed. The method is based on scene (road) classification which provides the probability that an image contains certain type of road geometry (straight, left/right curve, etc.). During the training of the classifier a road probability map is also learned for each road geometry. Then, the proper pixel-based method is selected and fused to provide an improved road detection approach. From experiments it is concluded that the proposed method outperforms state-of-the-art algorithms in a frame by frame context. José M. Álvarez 0004, Theo Gevers, Antonio M. López 0001 |
ICIP | 1 |