EDBT 2026 Demo / reviewers in the wild / expert
Alexandre Boulch
dblp:47/9368
· DBLP profile ↗
46ranked-venue papers
13as first author
23since 2021 · last 2026
0000-0002-4196-9665ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 11 first-author · 18 since 2021Artificial intelligence and machine learning · 22 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is Clustering Enough for LiDAR Instance Segmentation? A State-of-the-Art Training-Free Baseline
Corentin Sautier, Gilles Puy, Alexandre Boulch, Renaud Marlet, Vincent Lepetit |
3DV | 3 |
| 2025 | UNIT: Unsupervised Online Instance Segmentation Through TimeabstractOnline object segmentation and tracking in Lidar point clouds enables autonomous agents to understand their surroundings and make safe decisions. Unfortunately, manual annotations for these tasks are prohibitively costly. We tackle this problem with the task of class-agnostic unsupervised online instance segmentation and tracking. To that end, we leverage an instance segmentation backbone and propose a new training recipe that enables the online tracking of objects. Our network is trained on pseudo-labels, eliminating the need for manual annotations. We conduct an evaluation using metrics adapted for temporal instance segmentation. Computing these metrics requires temporally-consistent instance labels. When unavailable, we construct these labels using the available 3D bounding boxes and semantic labels in the dataset. We compare our method against strong baselines and demonstrate its superiority across two different outdoor Lidar datasets. Project page: csautier.github.io/unit Corentin Sautier, Gilles Puy, Alexandre Boulch, Renaud Marlet, Vincent Lepetit |
3DV | 3 |
| 2025 | GaussRender: Learning 3D Occupancy with Gaussian RenderingabstractUnderstanding the 3D geometry and semantics of driving scenes is critical for safe autonomous driving. Recent advances in 3D occupancy prediction have improved scene representation but often suffer from spatial inconsistencies, leading to floating artifacts and poor surface localization. Existing voxel-wise losses (e.g., cross-entropy) fail to enforce geometric coherence. In this paper, we propose GaussRender, a module that improves 3D occupancy learning by enforcing projective consistency. Our key idea is to project both predicted and ground-truth 3D occupancy into 2D camera views, where we apply supervision. Our method penalizes 3D configurations that produce inconsistent 2D projections, thereby enforcing a more coherent 3D structure. To achieve this efficiently, we leverage differentiable rendering with Gaussian splatting. GaussRender seamlessly integrates with existing architectures while maintaining efficiency and requiring no inference-time modifications. Extensive evaluations on multiple benchmarks (SurroundOcc-nuScenes, Occ3D-nuScenes, SSCBench-KITTI360) demonstrate that GaussRender significantly improves geometric fidelity across various 3D occupancy models (TPVFormer, SurroundOcc, Symphonies), achieving state-of-the-art results, particularly on surface-sensitive metrics. The code is open-sourced at https://github.com/valeoai/GaussRender. Loïck Chambon, Eloi Zablocki, Alexandre Boulch, Mickaël Chen, Matthieu Cord |
ICCV | 3 |
| 2025 | LiDPM: Rethinking Point Diffusion for Lidar Scene CompletionabstractTraining diffusion models that work directly on lidar points at the scale of outdoor scenes is challenging due to the difficulty of generating fine-grained details from white noise over a broad field of view. The latest works addressing scene completion with diffusion models tackle this problem by reformulating the original DDPM as a local diffusion process. It contrasts with the common practice of operating at the level of objects, where vanilla DDPMs are currently used. In this work, we close the gap between these two lines of work. We identify approximations in the local diffusion formulation, show that they are not required to operate at the scene level, and that a vanilla DDPM with a well-chosen starting point is enough for completion. Finally, we demonstrate that our method, LiDPM, leads to better results in scene completion on SemanticKITTI. The project page is https://astra-vision.github.io/LiDPM. Tetiana Martyniuk, Gilles Puy, Alexandre Boulch, Renaud Marlet, Raoul de Charette |
IV | 3 |
| 2024 | SALUDA: Surface-based Automotive Lidar Unsupervised Domain AdaptationabstractLearning models on one labeled dataset that generalize well on another domain is a difficult task, as several shifts might happen between the data domains. This is notably the case for lidar data, for which models can exhibit large performance discrepancies due for instance to different lidar patterns or changes in acquisition conditions. This paper addresses the corresponding Unsupervised Domain Adaptation (UDA) task for semantic segmentation. To mitigate this problem, we introduce an unsupervised auxiliary task of learning an implicit underlying surface representation simultaneously on source and target data. As both domains share the same latent representation, the model is forced to accommodate discrepancies between the two sources of data. This novel strategy differs from classical minimization of statistical divergences or lidar-specific domain adaptation techniques. Our experiments demonstrate that our method achieves a better performance than the current state of the art, both in real-to-real and synthetic-to-real scenarios.The project repository: github.com/valeoai/SALUDA Björn Michele, Alexandre Boulch, Gilles Puy, Renaud Marlet, Nicolas Courty |
3DV | 2 |
| 2024 | BEVContrast: Self-Supervision in BEV Space for Automotive Lidar Point CloudsabstractWe present a surprisingly simple and efficient method for self-supervision of 3D backbone on automotive Lidar point clouds. We design a contrastive loss between features of Lidar scans captured in the same scene. Several such approaches have been proposed in the literature from PointConstrast [40], which uses a contrast at the level of points, to the state-of-the-art TARL [30], which uses a contrast at the level of segments, roughly corresponding to objects. While the former enjoys a great simplicity of implementation, it is surpassed by the latter, which however requires a costly pre-processing. In BEVContrast, we define our contrast at the level of 2D cells in the Bird’s Eye View plane. Resulting cell-level representations offer a good trade-off between the point-level representations exploited in PointContrast and segment-level representations exploited in TARL: we retain the simplicity of PointContrast (cell representations are cheap to compute) while surpassing the performance of TARL in downstream semantic segmentation. The code is available at github.com/valeoai/BEVContrast Corentin Sautier, Gilles Puy, Alexandre Boulch, Renaud Marlet, Vincent Lepetit |
3DV | 3 |
| 2024 | Three Pillars Improving Vision Foundation Model Distillation for LidarabstractSelf-supervised image backbones can be used to address complex 2D tasks (e.g., semantic segmentation, object discovery) very efficiently and with little or no downstream supervision. Ideally, 3D backbones for lidar should be able to inherit these properties after distillation of these powerful 2D features. The most recent methods for image-to-lidar distillation on autonomous driving data show promising results, obtained thanks to distillation methods that keep improving. Yet, we still notice a large performance gap when measuring by linear probing the quality of distilled vs fully supervised features. In this work, instead of focusing only on the distillation method, we study the effect of three pillars for distillation: the 3D backbone, the pretrained 2D backbone, and the pretraining 2D+3D dataset. In particular, thanks to our scalable distillation method named ScaLR, we show that scaling the 2D and 3D backbones and pretraining on diverse datasets leads to a substantial improvement of the feature quality. This allows us to significantly reduce the gap between the quality of distilled and fully-supervised 3D features, and to improve the robustness of the pretrained backbones to domain gaps and perturbations. The code is available at https://github.com/valeoai/ScaLR. Gilles Puy, Spyros Gidaris, Alexandre Boulch, Oriane Siméoni, Corentin Sautier, Patrick Pérez, Andrei Bursuc, Renaud Marlet |
CVPR | 3 |
| 2024 | Train Till You Drop: Towards Stable and Robust Source-Free Unsupervised 3D Domain Adaptation
Björn Michele, Alexandre Boulch, Gilles Puy, Renaud Marlet, Nicolas Courty |
ECCV (20) | 2 |
| 2023 | RangeViT: Towards Vision Transformers for 3D Semantic Segmentation in Autonomous DrivingabstractCasting semantic segmentation of outdoor LiDAR point clouds as a 2D problem, e.g., via range projection, is an effective and popular approach. These projection-based methods usually benefit from fast computations and, when combined with techniques which use other point cloud representations, achieve state-of-the-art results. Today, projection-based methods leverage 2D CNNs but recent advances in computer vision show that vision transformers (ViTs) have achieved state-of-the-art results in many image- based benchmarks. In this work, we question if projection- based methods for 3D semantic segmentation can benefit from these latest improvements on ViTs. We answer positively but only after combining them with three key ingredients: (a) ViTs are notoriously hard to train and require a lot of training data to learn powerful representations. By preserving the same backbone architecture as for RGB images, we can exploit the knowledge from long training on large image collections that are much cheaper to acquire and annotate than point clouds. We reach our best results with pre-trained ViTs on large image datasets. (b) We compensate ViTs' lack of inductive bias by substituting a tailored convolutional stem for the classical linear embedding layer. (c) We refine pixel-wise predictions with a convolutional decoder and a skip connection from the convolutional stem to combine low-level but fine-grained features of the the convolutional stem with the high-level but coarse predictions of the ViT encoder. With these ingredients, we show that our method, called RangeViT, outperforms existing projection-based methods on nuScenes and SemanticKITTI. The code is available at https://github.com/valeoai/rangevit. Angelika Ando, Spyros Gidaris, Andrei Bursuc, Gilles Puy, Alexandre Boulch, Renaud Marlet |
CVPR | 5 |
| 2023 | ALSO: Automotive Lidar Self-Supervision by Occupancy EstimationabstractWe propose a new self-supervised method for pre-training the backbone of deep perception models operating on point clouds. The core idea is to train the model on a pretext task which is the reconstruction of the surface on which the 3D points are sampled, and to use the underlying latent vectors as input to the perception head. The intuition is that if the network is able to reconstruct the scene surface, given only sparse input points, then it probably also captures some fragments of semantic information, that can be used to boost an actual perception task. This principle has a very simple formulation, which makes it both easy to implement and widely applicable to a large range of 3D sensors and deep networks performing semantic segmentation or object detection. In fact, it supports a single-stream pipeline, as opposed to most contrastive learning approaches, allowing training on limited resources. We conducted extensive experiments on various autonomous driving datasets, involving very different kinds of lidars, for both semantic segmentation and object detection. The results show the effectiveness of our method to learn useful representations without any annotation, compared to existing approaches. The code is available at github.com/valeoai/ALSO Alexandre Boulch, Corentin Sautier, Björn Michele, Gilles Puy, Renaud Marlet |
CVPR | 1 |
| 2023 | Using a Waffle Iron for Automotive Point Cloud Semantic SegmentationabstractSemantic segmentation of point clouds in autonomous driving datasets requires techniques that can process large numbers of points efficiently. Sparse 3D convolutions have become the de-facto tools to construct deep neural networks for this task: they exploit point cloud sparsity to reduce the memory and computational loads and are at the core of today's best methods. In this paper, we propose an alternative method that reaches the level of state-of-the-art methods without requiring sparse convolutions. We actually show that such level of performance is achievable by relying on tools a priori unfit for large scale and high-performing 3D perception. In particular, we propose a novel 3D backbone, WaffleIron, made almost exclusively of MLPs and dense 2D convolutions and present how to train it to reach high performance on SemanticKITTI and nuScenes. We believe that WaffleIron is a compelling alternative to backbones using sparse 3D convolutions, especially in frameworks and on hardware where those convolutions are not readily available. The code is available at https://github.com/valeoai/WaffleIron. Gilles Puy, Alexandre Boulch, Renaud Marlet |
ICCV | 2 |
| 2023 | Weakly supervised change detection using guided anisotropic diffusion
Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch, Yann Gousseau |
Mach. Learn. | 3 |
| 2022 | POCO: Point Convolution for Surface ReconstructionabstractImplicit neural networks have been successfully used for surface reconstruction from point clouds. However, many of them face scalability issues as they encode the isosurface function of a whole object or scene into a single latent vector. To overcome this limitation, a few approaches infer latent vectors on a coarse regular 3D grid or on 3D patches, and interpolate them to answer occupancy queries. In doing so, they lose the direct connection with the input points sampled on the surface of objects, and they attach information uniformly in space rather than where it matters the most, i.e., near the surface. Besides, relying on fixed patch sizes may require discretization tuning. To address these issues, we propose to use point cloud convolutions and compute latent vectors at each input point. We then perform a learning-based interpolation on nearest neighbors using inferred weights. Experiments on both object and scene datasets show that our approach significantly outperforms other methods on most classical metrics, producing finer details and better reconstructing thinner volumes. The code is available at https://github.com/valeoai/POCO. Alexandre Boulch, Renaud Marlet |
CVPR | 1 |
| 2022 | Image-to-Lidar Self-Supervised Distillation for Autonomous Driving DataabstractSegmenting or detecting objects in sparse Lidar point clouds are two important tasks in autonomous driving to allow a vehicle to act safely in its 3D environment. The best performing methods in 3D semantic segmentation or object detection rely on a large amount of annotated data. Yet annotating 3D Lidar data for these tasks is tedious and costly. In this context, we propose a self-supervised pretraining method for 3D perception models that is tailored to autonomous driving data. Specifically, we leverage the availability of synchronized and calibrated image and Lidar sensors in autonomous driving setups for distilling self-supervised pre-trained image representations into 3D models. Hence, our method does not require any point cloud nor image annotations. The keyingredient of our method is the use of superpixels which are used to pool 3D point features and 2D pixel features in visually similar regions. We then train a 3D network on the self-supervised task of matching these pooled point features with the corresponding pooled image pixel features. The advantages of contrasting regions obtained by superpixels are that: (1) grouping together pixels and points of visually coherent regions leads to a more meaningful contrastive task that produces features well adapted to 3D semantic segmentation and 3D object detection; (2) all the different regions have the same weight in the contrastive loss regardless of the number of 3D points sampled in these regions; (3) it mitigates the noise produced by incorrect matching of points and pixels due to occlusions between the different sensors. Extensive experiments on autonomous driving datasets demonstrate the ability of our image-to-Lidar distillation strategy to produce 3D representations that transfer well on semantic segmentation and object detection tasks. Corentin Sautier, Gilles Puy, Spyros Gidaris, Alexandre Boulch, Andrei Bursuc, Renaud Marlet |
CVPR | 4 |
| 2022 | VASAD: a Volume and Semantic dataset for Building Reconstruction from Point Cloudsabstract3D scene reconstruction has important applications to help to produce digital twins of existing buildings. While the community has mostly focused on surface reconstruction or semantic segmentation as separate problems, the joint reconstruction of both volumes and semantics has little been discussed, mostly due to the lack of large scale volume datasets with semantic annotations. In this work, we introduce a new dataset called VASAD for Volume And Semantic Architectural Dataset. It is composed of 6 building models, with full volume description and semantic labels. It approximately represents 62,000 m2of building floors, making it large enough for the development and evaluation of learning-based approaches. We propose several methods to jointly reconstruct both geometry and semantics and evaluate on the test set of the dataset. We show that the proposed dataset is challenging enough to stimulate research. The dataset is available at https://github.com/palanglois/vasad. Pierre-Alain Langlois, Yang Xiao 0009, Alexandre Boulch, Renaud Marlet |
ICPR | 3 |
| 2022 | Deep Surface Reconstruction from Point Clouds with Visibility InformationabstractMost current neural networks for reconstructing surfaces from point clouds ignore sensor poses and only operate on point locations. Sensor visibility, however, holds meaningful information regarding space occupancy and surface orientation. In this paper, we present two simple ways to augment point clouds with visibility information, so it can directly be leveraged by surface reconstruction networks with minimal adaptation. Our proposed modifications consistently improve the accuracy of generated surfaces as well as the generalization capability of the networks to unseen domains. Our code, data and pretrained models can be found online: https://github.com/raphaelsulzer/dsrv-data. Raphael Sulzer, Loïc Landrieu, Alexandre Boulch, Renaud Marlet, Bruno Vallet |
ICPR | 3 |
| 2022 | Semi-supervised semantic segmentation in Earth Observation: the MiniFrance suite, dataset analysis and multi-task network study
Javiera Castillo-Navarro, Bertrand Le Saux, Alexandre Boulch, Nicolas Audebert, Sébastien Lefèvre |
Mach. Learn. | 3 |
| 2022 | Energy-Based Models in Earth Observation: From Generation to Semisupervised LearningabstractDeep learning, together with the availability of large amounts of data, has transformed the way we process Earth observation (EO) tasks, such as land cover mapping or image registration. Yet, today, new models are needed to push further the revolution and enable new possibilities. This work focuses on a recent framework for generative modeling and explores its applicability to the EO images. The framework learns an energy-based model (EBM) to estimate the underlying joint distribution of the data and the categories, obtaining a neural network that is able to classify and synthesize images. On these two tasks, we show that EBMs reach comparable or better performances than convolutional networks on various public EO datasets and that they are naturally adapted to semisupervised settings, with very few labeled data. Moreover, models of this kind allow us to address high-potential applications, such as out-of-distribution analysis and land cover mapping with confidence estimation. Javiera Castillo-Navarro, Bertrand Le Saux, Alexandre Boulch, Sébastien Lefèvre |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | NeeDrop: Self-supervised Shape Representation from Sparse Point Clouds using Needle DroppingabstractThere has been recently a growing interest for implicit shape representations. Contrary to explicit representations, they have no resolution limitations and they easily deal with a wide variety of surface topologies. To learn these implicit representations, current approaches rely on a certain level of shape supervision (e.g., inside/outside information or distance-to-shape knowledge), or at least require a dense point cloud (to approximate well enough the distance-to-shape). In contrast, we introduce NeeDrop, an self-supervised method for learning shape representations from possibly extremely sparse point clouds. Like in Buffon’s needle problem, we “drop” (sample) needles on the point cloud and consider that, statistically, close to the surface, the needle end points lie on opposite sides of the surface. No shape knowledge is required and the point cloud can be highly sparse, e.g., as lidar point clouds acquired by vehicles. Previous self-supervised shape representation approaches fail to produce good-quality results on this kind of data. We obtain quantitative results on par with existing supervised approaches on shape reconstruction datasets and show promising qualitative results on hard autonomous driving datasets such as KITTI. Alexandre Boulch, Pierre-Alain Langlois, Gilles Puy, Renaud Marlet |
3DV | 1 |
| 2021 | Generative Zero-Shot Learning for Semantic Segmentation of 3D Point CloudsabstractWhile there has been a number of studies on Zero-Shot Learning (ZSL) for 2D images, its application to 3D data is still recent and scarce, with just a few methods limited to classification. We present the first generative approach for both ZSL and Generalized ZSL (GZSL) on 3D data, that can handle both classification and, for the first time, semantic segmentation. We show that it reaches or outperforms the state of the art on ModelNet40 classification for both inductive ZSL and inductive GZSL. For semantic segmentation, we created three benchmarks for evaluating this new ZSL task, using S3DIS, ScanNet and SemanticKITTI. Our experiments show that our method outperforms strong baselines, which we additionally propose for this task. Björn Michele, Alexandre Boulch, Gilles Puy, Maxime Bucher, Renaud Marlet |
3DV | 2 |
| 2021 | PCAM: Product of Cross-Attention Matrices for Rigid Registration of Point CloudsabstractRigid registration of point clouds with partial overlaps is a longstanding problem usually solved in two steps: (a) finding correspondences between the point clouds; (b) filtering these correspondences to keep only the most reliable ones to estimate the transformation. Recently, several deep nets have been proposed to solve these steps jointly. We built upon these works and propose PCAM: a neural network whose key element is a pointwise product of crossattention matrices that permits to mix both low-level geometric and high-level contextual information to find point correspondences. These cross-attention matrices also permits the exchange of context information between the point clouds, at each layer, allowing the network construct better matching features within the overlapping regions. The experiments show that PCAM achieves state-of-the-art results among methods which, like us, solve steps (a) and (b) jointly via deepnets. Anh-Quan Cao, Gilles Puy, Alexandre Boulch, Renaud Marlet |
ICCV | 3 |
| 2021 | 3D Reconstruction By Parameterized Surface MappingabstractWe introduce an approach for computing a 3D mesh from one or more views of an object by establishing dense correspondences between pixels in the views and 3D locations on a learnable parameterized surface. We propose a multi-view shape encoder that can be jointly trained with the AtlasNet surface parameterization. The shape is further refined using a novel geometric cycle-consistency loss between the learnable parameterized surface and input views. We demonstrate the efficacy of our approach on the ShapeNet-COCO dataset. Pierre-Alain Langlois, Matthew Fisher, Oliver Wang, Vladimir G. Kim, Alexandre Boulch, Renaud Marlet, Bryan C. Russell |
ICIP | 5 |
| 2021 | Classification and Generation of Earth Observation Images Using a Joint Energy-Based ModelabstractDeep learning has changed unbelievably the processing of Earth Observation tasks such as land cover mapping or image registration. Yet, today new models are needed to push further the revolution and enable new possibilities. We propose a new framework for generative modelling of Earth Observation images. It learns an energy-based model to estimate the underlying distribution of the data while jointly training a deep neural network for classification. On the varied image types of the EuroSAT benchmark, we show this model obtains classification results on par with state-of-the-art and moreover allows us to tackle a wide range of high-potential applications: image synthesis, out-of-distribution testing for domain adaptation, and image completion or denoising. Javiera Castillo-Navarro, Bertrand Le Saux, Alexandre Boulch, Sébastien Lefèvre |
IGARSS | 3 |
| 2020 | FKAConv: Feature-Kernel Alignment for Point Cloud Convolution
Alexandre Boulch, Gilles Puy, Renaud Marlet |
ACCV (1) | 1 |
| 2020 | FLOT: Scene Flow on Point Clouds Guided by Optimal Transport
Gilles Puy, Alexandre Boulch, Renaud Marlet |
ECCV (28) | 2 |
| 2020 | STaRFlow: A SpatioTemporal Recurrent Cell for Lightweight Multi-Frame Optical Flow EstimationabstractWe present a new lightweight CNN-based algorithm for multi-frame optical flow estimation. Our solution introduces a double recurrence over spatial scale and time through repeated use of a generic “STaR” (SpatioTemporal Recurrent) cell. It includes (i) a temporal recurrence based on conveying learned features rather than optical flow estimates; (ii) an occlusion detection process which is coupled with optical flow estimation and therefore uses a very limited number of extra parameters. The resulting STaRFlow algorithm gives state-of-the-art performances on MPI Sintel and Kitti2015 and involves significantly less parameters than all other methods with comparable results. Pierre Godet, Alexandre Boulch, Aurélien Plyer, Guy Le Besnerais |
ICPR | 2 |
| 2020 | ConvPoint: Continuous convolutions for point cloud processing
Alexandre Boulch |
Comput. Graph. | 1 |
| 2019 | Fast Stereo Disparity Maps Refinement By Fusion of Data-Based And Model-Based EstimationsabstractThe estimation of disparity maps from stereo pairs has many applications in robotics and autonomous driving. Stereo matching has first been solved using model-based approaches, with real-time considerations for some, but today's most recent works rely on deep convolutional neural networks and mainly focus on accuracy at the expense of computing time. In this paper, we present a new method for disparity maps estimation getting the best of both worlds: the accuracy of data-based methods and the speed of fast model-based ones. The proposed approach fuses prior disparity maps to estimate a refined version. The core of this fusion pipeline is a convolutional neural network that leverages dilated convolutions for fast context aggregation without spatial resolution loss. The resulting architecture is both very effective for the task of refining and fusing prior disparity maps and very light, allowing our fusion pipeline to produce disparity maps at rates up to 125 Hz. We obtain state-of-the-art results in terms of speed and accuracy on the KITTI benchmarks. Code and pre-trained models are available on our github: https://github.com/ferreram/FD-Fusion. Maxime Ferrera, Alexandre Boulch, Julien Moras |
3DV | 2 |
| 2019 | Surface Reconstruction from 3D Line SegmentsabstractIn man-made environments such as indoor scenes, when point-based 3D reconstruction fails due to the lack of texture, lines can still be detected and used to support surfaces. We present a novel method for watertight piecewise-planar surface reconstruction from 3D line segments with visibility information. First, planes are extracted by a novel RANSAC approach for line segments that allows multiple shape support. Then, each 3D cell of a plane arrangement is labeled full or empty based on line attachment to planes, visibility and regularization. Experiments show the robustness to sparse input data, noise and outliers. Pierre-Alain Langlois, Alexandre Boulch, Renaud Marlet |
3DV | 2 |
| 2019 | Learning to Understand Earth Observation Images with Weak and Unreliable Ground TruthabstractIn this paper we discuss the issues of using inexact and inaccurate ground truth in the context of supervised learning. To leverage large amounts of Earth observation data for training algorithms, one often has to use ground truth which was not been carefully assessed. We address both the problems of training and evaluation. We first propose a weakly supervised approach for training change classifiers which is able to detect pixel-level changes in aerial images. We then propose a data poisoning approach to get a reliable estimate of the accuracy that can be expected from a classifier, even when the only ground-truth available does not match the reality. Both are assessed on practical land use and land cover applications. Rodrigo Caye Daudt, Adrien Chan-Hon-Tong, Bertrand Le Saux, Alexandre Boulch |
IGARSS | 4 |
| 2019 | Distance transform regression for spatially-aware deep semantic segmentation
Nicolas Audebert, Alexandre Boulch, Bertrand Le Saux, Sébastien Lefèvre |
Comput. Vis. Image Underst. | 2 |
| 2019 | Multitask learning for large-scale semantic change detection
Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch, Yann Gousseau |
Comput. Vis. Image Underst. | 3 |
| 2018 | Fully Convolutional Siamese Networks for Change DetectionabstractThis paper presents three fully convolutional neural network architectures which perform change detection using a pair of coregistered images. Most notably, we propose two Siamese extensions of fully convolutional networks which use heuristics about the current problem to achieve the best results in our tests on two open change detection datasets, using both RGB and multispectral images. We show that our system is able to learn from scratch using annotated change detection images. Our architectures achieve better performance than previously proposed methods, while being at least 500 times faster than related systems. This work is a step towards efficient processing of data from large scale Earth observation systems such as Copernicus or Landsat. Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch |
ICIP | 3 |
| 2018 | Learning Speckle Suppression in Sar Images Without Ground Truth: Application to Sentinel-1 Time-SeriesabstractThis paper proposes a method of denoising SAR images, using a deep learning method, which takes advantage of the abundance of data to learn on large stacks of images of the same scene. The approach is based on the use of convolutional networks, used as auto-encoders. Learning is led on a large pile of images acquired on the same area, and assumes that the images of this stack differ only by the speckle noise. Several pairs of images are chosen randomly in the stack, and the network tries to predict the slave image from the master image. In this prediction, the network can not predict the noise because of its random nature. Also the application of this network to a new image fulfills the speckle filtering function. Results are given on Sentinel 1 images. They show that this approach is qualitatively competitive with literature. Alexandre Boulch, Pauline Trouvé-Peloux, Elise Colin, Fabrice Janez, Bertrand Le Saux |
IGARSS | 1 |
| 2018 | Urban Change Detection for Multispectral Earth Observation Using Convolutional Neural NetworksabstractThe Copernicus Sentinel-2 program now provides multispectral images at a global scale with a high revisit rate. In this paper we explore the usage of convolutional neural networks for urban change detection using such multispectral images. We first present the new change detection dataset that was used for training the proposed networks, which will be openly available to serve as a benchmark. The Onera Satellite Change Detection (OSCD) dataset is composed of pairs of multispectral aerial images, and the changes were manually annotated at pixel level. We then propose two architectures to detect changes, Siamese and Early Fusion, and compare the impact of using different numbers of spectral channels as inputs. These architectures are trained from scratch using the provided dataset. Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch, Yann Gousseau |
IGARSS | 3 |
| 2018 | Large-Scale Semantic Classification: Outcome of the First Year of Inria Aerial Image Labeling BenchmarkabstractOver the recent years, there has been an increasing interest in large-scale classification of remote sensing images. In this context, the Inria Aerial Image Labeling Benchmark has been released online in December 2016. In this paper, we discuss the outcomes of the first year of the benchmark contest, which consisted in dense labeling of aerial images into building / not building classes, covering areas of five cities not present in the training set. We present four methods with the highest numerical accuracies, all four being convolutional neural network approaches. It is remarkable that three of these methods use the U-net architecture, which has thus proven to become a new standard in image dense labeling. Bohao Huang, Kangkang Lu 0001, Nicolas Audebert, Andrew Khalel, Yuliya Tarabalka, Jordan M. Malof, Alexandre Boulch, Bertrand Le Saux, Leslie M. Collins, Kyle Bradbury, Sébastien Lefèvre, Motaz El-Saban |
IGARSS | 7 |
| 2018 | Railway Detection: From Filtering to Segmentation NetworksabstractThis paper deals with classification of remote sensing data to extract objects for industrial mapping. While land-cover or urban mapping have been extensively studied, industrial cartography remains a field yet to explore, in spite of tremendous needs. We present and compare here four approaches for railway detection in very high resolution images. They use various kind of filtering approaches, including the trained filters of fully convolutional networks. Moreover, they benefit from different a-priori and post-processing techniques to make them more robust. We evaluate all approaches on a challenging dataset captured on an operating station site with complex objects. Bertrand Le Saux, Anne Beaupère, Alexandre Boulch, Jérémie Brossard, Antoine Manier, Guilhem Villemin |
IGARSS | 3 |
| 2018 | SnapNet: 3D point cloud semantic labeling with 2D deep segmentation networks
Alexandre Boulch, Joris Guerry, Bertrand Le Saux, Nicolas Audebert |
Comput. Graph. | 1 |
| 2018 | Reducing parameter number in residual networks by sharing weights
Alexandre Boulch |
Pattern Recognit. Lett. | 1 |
| 2017 | Deep Sequence-to-Sequence Neural Networks for Ionospheric Activity Map Prediction
Noëlie Cherrier, Thibaut Castaings, Alexandre Boulch |
ICONIP (5) | 3 |
| 2016 | Deep Learning for Robust Normal Estimation in Unstructured Point CloudsabstractAbstract Normal estimation in point clouds is a crucial first step for numerous algorithms, from surface reconstruction and scene understanding to rendering. A recurrent issue when estimating normals is to make appropriate decisions close to sharp features, not to smooth edges, or when the sampling density is not uniform, to prevent bias. Rather than resorting to manually‐designed geometric priors, we propose to learn how to make these decisions, using ground‐truth data made from synthetic scenes. For this, we project a discretized Hough space representing normal directions onto a structure amenable to deep learning. The resulting normal estimation method outperforms most of the time the state of the art regarding robustness to outliers, to noise and to point density variation, in the presence of sharp edges, while remaining fast, scaling up to millions of points. Alexandre Boulch, Renaud Marlet |
Comput. Graph. Forum | 1 |
| 2015 | Benchmarking classification of earth-observation data: From learning explicit features to convolutional networksabstractIn this paper, we address the task of semantic labeling of multisource earth-observation (EO) data. Precisely, we benchmark several concurrent methods of the last 15 years, from expert classifiers, spectral support-vector classification and high-level features to deep neural networks. We establish that (1) combining multisensor features is essential for retrieving some specific classes, (2) in the image domain, deep convolutional networks obtain significantly better overall performances and (3) transfer of learning from large generic-purpose image sets is highly effective to build EO data classifiers. Adrien Lagrange, Bertrand Le Saux, Anne Beaupère, Alexandre Boulch, Adrien Chan-Hon-Tong, Stéphane Herbin, Hicham Randrianarivo, Marin Ferecatu |
IGARSS | 4 |
| 2014 | Statistical Criteria for Shape Fusion and SelectionabstractSurface reconstruction from point clouds often relies on a primitive extraction step, that may be followed by a merging step because of a possible over-segmentation. We present two statistical criteria to decide whether or not two surfaces are to be considered as the same, and thus can be merged. They are based on the statistical tests of Kolmogorov-Smirnov and Mann-Whitney for comparing distributions. Moreover, computation time can be significantly cut down using a reduced sampling based on the Dvoretzky-Keifer-Wolfowitz inequality. The strength of our approach is that it relies in practice on a single intuitive parameter (homogeneous to a distance) and that it can be applied to any shape, including meshes, not just geometric primitives. It also enables the comparison of shapes of different kinds, providing a way to choose between different shape candidates. We show several applications of our method, experimenting geometric primitive (plane and cylinder) detection, selection and fusion, both on precise laser scans and noisy photogrammetric 3D data. Alexandre Boulch, Renaud Marlet |
ICPR | 1 |
| 2014 | Piecewise-Planar 3D Reconstruction with Edge and Corner RegularizationabstractAbstract This paper presents a method for the 3D reconstruction of a piecewise‐planar surface from range images, typically laser scans with millions of points. The reconstructed surface is a watertight polygonal mesh that conforms to observations at a given scale in the visible planar parts of the scene, and that is plausible in hidden parts. We formulate surface reconstruction as a discrete optimization problem based on detected and hypothesized planes. One of our major contributions, besides a treatment of data anisotropy and novel surface hypotheses, is a regularization of the reconstructed surface w.r.t. the length of edges and the number of corners. Compared to classical area‐based regularization, it better captures surface complexity and is therefore better suited for man‐made environments, such as buildings. To handle the underlying higher‐order potentials, that are problematic for MRF optimizers, we formulate minimization as a sparse mixed‐integer linear programming problem and obtain an approximate solution using a simple relaxation. Experiments show that it is fast and reaches near‐optimal solutions. Alexandre Boulch, Martin de La Gorce, Renaud Marlet |
Comput. Graph. Forum | 1 |
| 2013 | Semantizing Complex 3D Scenes using Constrained Attribute GrammarsabstractAbstract We propose a new approach to automatically semantize complex objects in a 3D scene. For this, we define an expressive formalism combining the power of both attribute grammars and constraint. It offers a practical conceptual interface, which is crucial to write large maintainable specifications. As recursion is inadequate to express large collections of items, we introduce maximal operators, that are essential to reduce the parsing search space. Given a grammar in this formalism and a 3D scene, we show how to automatically compute a shared parse forest of all interpretations — in practice, only a few, thanks to relevant constraints. We evaluate this technique for building model semantization using CAD model examples as well as photogrammetric and simulated LiDAR data. Alexandre Boulch, S. Houllier, Renaud Marlet, Olivier Tournaire |
Comput. Graph. Forum | 1 |
| 2012 | Fast and Robust Normal Estimation for Point Clouds with Sharp FeaturesabstractAbstract This paper presents a new method for estimating normals on unorganized point clouds that preserves sharp features. It is based on a robust version of the Randomized Hough Transform (RHT). We consider the filled Hough transform accumulator as an image of the discrete probability distribution of possible normals. The normals we estimate corresponds to the maximum of this distribution. We use a fixed‐size accumulator for speed, statistical exploration bounds for robustness, and randomized accumulators to prevent discretization effects. We also propose various sampling strategies to deal with anisotropy, as produced by laser scans due to differences of incidence. Our experiments show that our approach offers an ideal compromise between precision, speed, and robustness: it is at least as precise and noise‐resistant as state‐of‐the‐art methods that preserve sharp features, while being almost an order of magnitude faster. Besides, it can handle anisotropy with minor speed and precision losses. Alexandre Boulch, Renaud Marlet |
Comput. Graph. Forum | 1 |