VLDB 2026 Research / reviewers in the wild / expert
Mickaël Chen
dblp:190/7274
· DBLP profile ↗
19ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0002-3960-2389ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GaussRender: Learning 3D Occupancy with Gaussian RenderingabstractUnderstanding the 3D geometry and semantics of driving scenes is critical for safe autonomous driving. Recent advances in 3D occupancy prediction have improved scene representation but often suffer from spatial inconsistencies, leading to floating artifacts and poor surface localization. Existing voxel-wise losses (e.g., cross-entropy) fail to enforce geometric coherence. In this paper, we propose GaussRender, a module that improves 3D occupancy learning by enforcing projective consistency. Our key idea is to project both predicted and ground-truth 3D occupancy into 2D camera views, where we apply supervision. Our method penalizes 3D configurations that produce inconsistent 2D projections, thereby enforcing a more coherent 3D structure. To achieve this efficiently, we leverage differentiable rendering with Gaussian splatting. GaussRender seamlessly integrates with existing architectures while maintaining efficiency and requiring no inference-time modifications. Extensive evaluations on multiple benchmarks (SurroundOcc-nuScenes, Occ3D-nuScenes, SSCBench-KITTI360) demonstrate that GaussRender significantly improves geometric fidelity across various 3D occupancy models (TPVFormer, SurroundOcc, Symphonies), achieving state-of-the-art results, particularly on surface-sensitive metrics. The code is open-sourced at https://github.com/valeoai/GaussRender. Loïck Chambon, Eloi Zablocki, Alexandre Boulch, Mickaël Chen, Matthieu Cord |
ICCV | 4 |
| 2025 | Halton Scheduler for Masked Generative Image TransformerabstractMasked Generative Image Transformers (MaskGIT) have emerged as a scalable
and efficient image generation framework, able to deliver high-quality visuals with
low inference costs. However, MaskGIT’s token unmasking scheduler, an essential
component of the framework, has not received the attention it deserves. We analyze
the sampling objective in MaskGIT, based on the mutual information between
tokens, and elucidate its shortcomings. We then propose a new sampling strategy
based on our Halton scheduler instead of the original Confidence scheduler. More
precisely, our method selects the token’s position according to a quasi-random,
low-discrepancy Halton sequence. Intuitively, that method spreads the tokens
spatially, progressively covering the image uniformly at each step. Our analysis
shows that it allows reducing non-recoverable sampling errors, leading to simpler
hyper-parameters tuning and better quality images. Our scheduler does not require
retraining or noise injection and may serve as a simple drop-in replacement for
the original sampling strategy. Evaluation of both class-to-image synthesis on
ImageNet and text-to-image generation on the COCO dataset demonstrates that the
Halton scheduler outperforms the Confidence scheduler quantitatively by reducing
the FID and qualitatively by generating more diverse and more detailed images.
Our code is at https://github.com/valeoai/Halton-MaskGIT. Victor Besnier, Mickaël Chen, David Hurych, Eduardo Valle, Matthieu Cord |
ICLR | 2 |
| 2025 | Annealed Winner-Takes-All for Motion ForecastingabstractIn autonomous driving, motion prediction aims at forecasting the future trajectories of nearby agents, helping the ego vehicle to anticipate behaviors and drive safely. A key challenge is generating a diverse set of future predictions, commonly addressed using data-driven models with Multiple Choice Learning (MCL) architectures and Winner-Takes-All (WTA) training objectives. However, these methods face initialization sensitivity and training instabilities. Additionally, to compensate for limited performance, some approaches rely on training with a large set of hypotheses, requiring a post-selection step during inference to significantly reduce the number of predictions. To tackle these issues, we take inspiration from annealed MCL, a recently introduced technique that improves the convergence properties of MCL methods through an annealed Winner-Takes-All loss (aWTA). In this paper, we demonstrate how the aWTA loss can be integrated with state-of-the-art motion forecasting models to enhance their performance using only a minimal set of hypotheses, eliminating the need for the cumbersome post-selection step. Our approach can be easily incorporated into any trajectory prediction model normally trained using WTA and yields significant improvements. To facilitate the application of our approach to future motion forecasting models, the code is made publicly available: https://github.com/valeoai/MF_aWTA. Victor Letzelter, Mickaël Chen, Eloi Zablocki, Matthieu Cord |
ICRA | 3 |
| 2025 | Manipulating Trajectory Prediction Models With BackdoorsabstractAutonomous vehicles depend on accurate trajectory prediction to navigate safely in complex traffic. Yet current models are vulnerable to stealthy backdoor attacks: an adversary embeds subtle, physically plausible triggers during training that remain latent until activated. To address this risk, we introduce a structured framework categorizing four trigger types—spatial, kinetic (braking), coordinated, and composite—and demonstrate on two benchmarks (nuScenes and Argoverse 2) and two state-of-the-art architectures (Autobot and Wayformer) that poisoning as little as 5% of training samples can reliably hijack future predictions. We further propose a real-time defense leveraging social attention: by encoding agent histories, computing cross-attention to the target vehicle, and filtering out agents with anomalously high weights, our method neutralizes backdoor triggers without degrading clean-data accuracy. Comprehensive experiments show our defense reduces attack success rates across diverse urban scenarios—intersections, roundabouts, multi-lane roads—highlighting both the severity of backdoor threats and a promising pathway to secure trajectory predictors in autonomous driving systems. Kaouther Messaoud, Kathrin Grosse, Mickaël Chen, Matthieu Cord, Patrick Pérez, Alexandre Alahi |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | PointBeV: A Sparse Approach to BeV PredictionsabstractBird's-eye View (BeV) representations have emerged as the de-facto shared space in driving applications, offering a unified space for sensor data fusion and supporting various downstream tasks. However, conventional models use grids with fixed resolution and range and face computational inefficiencies due to the uniform allocation of resources across all cells. To address this, we propose Point-BeV, a novel sparse BeV segmentation model operating on sparse BeV cells instead of dense grids. This approach offers precise control over memory usage, enabling the use of long temporal contexts and accommodating memory-constrained platforms. PointBeV employs an efficient two-pass strategy for training, enabling focused computation on regions of interest. At inference time, it can be used with various memory/performance trade-offs and flexibly adjusts to new specific use cases. PointBeV achieves state-of-the-art results on the nuScenes dataset for vehicle, pedes-trian, and lane segmentation, showcasing superior performance in static and temporal settings despite being trained solely with sparse signals. We release our code with two new efficient modules used in the architecture: Sparse Feature Pulling, designed for the effective extraction of features from images to BeV, and Submanifold Attention, which en-ables efficient temporal modeling. The code is available at https://github.com/valeoai/PointBeV. Loïck Chambon, Eloi Zablocki, Mickaël Chen, Florent Bartoccioni, Patrick Pérez, Matthieu Cord |
CVPR | 3 |
| 2024 | Reliability in Semantic Segmentation: Can We Use Synthetic Data?
Thibaut Loiseau, Mickaël Chen, Patrick Pérez, Matthieu Cord |
ECCV (23) | 3 |
| 2024 | Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?abstractMotion forecasting is crucial in enabling autonomous vehicles to anticipate the future trajectories of surrounding agents. To do so, it requires solving mapping, detection, tracking, and then forecasting problems, in a multi-step pipeline. In this complex system, advances in conventional forecasting methods have been made using curated data, i.e., with the assumption of perfect maps, detection, and tracking. This paradigm, however, ignores any errors from upstream modules. Meanwhile, an emerging end-to-end paradigm, that tightly integrates the perception and forecasting architectures into joint training, promises to solve this issue. However, the evaluation protocols between the two methods were so far incompatible and their comparison was not possible. In fact, conventional forecasting methods are usually not trained nor tested in real-world pipelines (e.g., with upstream detection, tracking, and mapping modules). In this work, we aim to bring forecasting models closer to the real-world deployment. First, we propose a unified evaluation pipeline for forecasting methods with real-world perception inputs, allowing us to compare conventional and end-to-end methods for the first time. Second, our in-depth study uncovers a substantial performance gap when transitioning from curated to perception-based data. In particular, we show that this gap (1) stems not only from differences in precision but also from the nature of imperfect inputs provided by perception modules, and that (2) is not trivially reduced by simply finetuning on perception outputs. Based on extensive experiments, we provide recommendations for critical areas that require improvement and guidance towards more robust motion forecasting in the real world. The evaluation library for benchmarking models under standardized and practical conditions is provided: https://github.com/valeoai/MFEval. Loïck Chambon, Eloi Zablocki, Mickaël Chen, Alexandre Alahi, Matthieu Cord, Patrick Pérez |
ICRA | 4 |
| 2024 | Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealingabstractWe introduce Annealed Multiple Choice Learning (aMCL) which combines simulated annealing with MCL. MCL is a learning framework handling ambiguous tasks by predicting a small set of plausible hypotheses. These hypotheses are trained using the Winner-takes-all (WTA) scheme, which promotes the diversity of the predictions. However, this scheme may converge toward an arbitrarily suboptimal local minimum, due to the greedy nature of WTA. We overcome this limitation using annealing, which enhances the exploration of the hypothesis space during training. We leverage insights from statistical physics and information theory to provide a detailed description of the model training trajectory. Additionally, we validate our algorithm by extensive experiments on synthetic datasets, on the standard UCI benchmark, and on speech separation. David Perera, Victor Letzelter, Théo Mariotte, Adrien Cortés, Mickaël Chen, Slim Essid, Gaël Richard |
NeurIPS | 5 |
| 2023 | OCTET: Object-aware Counterfactual ExplanationsabstractNowadays, deep vision models are being widely deployed in safety-critical applications, e.g., autonomous driving, and explainability of such models is becoming a pressing concern. Among explanation methods, counter-factual explanations aim to find minimal and interpretable changes to the input image that would also change the output of the model to be explained. Such explanations point end-users at the main factors that impact the decision of the model. However, previous methods struggle to explain decision models trained on images with many objects, e.g., urban scenes, which are more difficult to work with but also arguably more critical to explain. In this work, we propose to tackle this issue with an object-centric framework for counterfactual explanation generation. Our method, inspired by recent generative modeling works, encodes the query image into a latent space that is structured in a way to ease object-level manipulations. Doing so, it provides the end-user with control over which search directions (e.g., spatial displacement of objects, style modification, etc.) are to be explored during the counterfactual generation. We conduct a set of experiments on counterfactual explanation benchmarks for driving scenes, and we show that our method can be adapted beyond classification, e.g., to explain semantic segmentation models. To complete our analysis, we design and run a user study that measures the usefulness of counterfactual explanations in understanding a decision model. Code is available at https://github.com/valeoai/OCTET. Mehdi Zemni, Mickaël Chen, Eloi Zablocki, Hédi Ben-Younes, Patrick Pérez, Matthieu Cord |
CVPR | 2 |
| 2023 | Unifying GANs and Score-Based Diffusion as Generative Particle ModelsabstractParticle-based deep generative models, such as gradient flows and score-based diffusion models, have recently gained traction thanks to their striking performance. Their principle of displacing particle distributions using differential equations is conventionally seen as opposed to the previously widespread generative adversarial networks (GANs), which involve training a pushforward generator network. In this paper we challenge this interpretation, and propose a novel framework that unifies particle and adversarial generative models by framing generator training as a generalization of particle models. This suggests that a generator is an optional addition to any such generative model. Consequently, integrating a generator into a score-based diffusion model and training a GAN without a generator naturally emerge from our framework. We empirically test the viability of these original models as proofs of concepts of potential applications of our framework. Jean-Yves Franceschi, Mike Gartrell, Ludovic Dos Santos, Thibaut Issenhuth, Emmanuel de Bézenac, Mickaël Chen, Alain Rakotomamonjy |
NeurIPS | 6 |
| 2023 | Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysisabstractWe introduce Resilient Multiple Choice Learning (rMCL), an extension of the MCL approach for conditional distribution estimation in regression settings where multiple targets may be sampled for each training input.
Multiple Choice Learning is a simple framework to tackle multimodal density estimation, using the Winner-Takes-All (WTA) loss for a set of hypotheses. In regression settings, the existing MCL variants focus on merging the hypotheses, thereby eventually sacrificing the diversity of the predictions. In contrast, our method relies on a novel learned scoring scheme underpinned by a mathematical framework based on Voronoi tessellations of the output space, from which we can derive a probabilistic interpretation.
After empirically validating rMCL with experiments on synthetic data, we further assess its merits on the sound source localization problem, demonstrating its practical usefulness and the relevance of its interpretation. Victor Letzelter, Mathieu Fontaine 0002, Mickaël Chen, Patrick Pérez, Slim Essid, Gaël Richard |
NeurIPS | 3 |
| 2022 | STEEX: Steering Counterfactual Explanations with Semantics
Paul Jacob, Eloi Zablocki, Hédi Ben-Younes, Mickaël Chen, Patrick Pérez, Matthieu Cord |
ECCV (12) | 4 |
| 2022 | A Neural Tangent Kernel Perspective of GANsabstractWe propose a novel theoretical framework of analysis for Generative Adversarial Networks (GANs). We reveal a fundamental flaw of previous analyses which, by incorrectly modeling GANs’ training scheme, are subject to ill-defined discriminator gradients. We overcome this issue which impedes a principled study of GAN training, solving it within our framework by taking into account the discriminator’s architecture. To this end, we leverage the theory of infinite-width neural networks for the discriminator via its Neural Tangent Kernel. We characterize the trained discriminator for a wide range of losses and establish general differentiability properties of the network. From this, we derive new insights about the convergence of the generated distribution, advancing our understanding of GANs’ training dynamics. We empirically corroborate these results via an analysis toolkit based on our framework, unveiling intuitions that are consistent with GAN practice. Jean-Yves Franceschi, Emmanuel de Bézenac, Ibrahim Ayed, Mickaël Chen, Sylvain Lamprier, Patrick Gallinari |
ICML | 4 |
| 2020 | Stochastic Latent Residual Video PredictionabstractDesigning video prediction models that account for the inherent uncertainty of the future is challenging. Most works in the literature are based on stochastic image-autoregressive recurrent networks, which raises several performance and applicability issues. An alternative is to use fully latent temporal models which untie frame synthesis and temporal dynamics. However, no such model for stochastic video prediction has been proposed in the literature yet, due to design and training difficulties. In this paper, we overcome these difficulties by introducing a novel stochastic temporal model whose dynamics are governed in a latent space by a residual update rule. This first-order scheme is motivated by discretization schemes of differential equations. It naturally models video dynamics as it allows our simpler, more interpretable, latent model to outperform prior state-of-the-art methods on challenging datasets. Jean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier, Patrick Gallinari |
ICML | 3 |
| 2020 | Adversarial learning for modeling human motion
Qi Wang 0018, Thierry Artières, Mickaël Chen, Ludovic Denoyer |
Vis. Comput. | 3 |
| 2019 | Unsupervised Object Segmentation by RedrawingabstractObject segmentation is a crucial problem that is usually solved by using supervised learning approaches over very large datasets composed of both images and corresponding object masks. Since the masks have to be provided at pixel level, building such a dataset for any new domain can be very costly. We present ReDO, a new model able to extract objects from images without any annotation in an unsupervised way. It relies on the idea that it should be possible to change the textures or colors of the objects without changing the overall distribution of the dataset. Following this assumption, our approach is based on an adversarial architecture where the generator is guided by an input sample: given an image, it extracts the object mask, then redraws a new object at the same location. The generator is controlled by a discriminator that ensures that the distribution of generated images is aligned to the original one. We experiment with this method on different datasets and demonstrate the good quality of extracted masks. Mickaël Chen, Thierry Artières, Ludovic Denoyer |
NeurIPS | 1 |
| 2018 | transferring style in motion capture sequences with adversarial learning
Qi Wang 0018, Mickaël Chen, Thierry Artières, Ludovic Denoyer |
ESANN | 2 |
| 2018 | Multi-View Data Generation Without View Supervision
Mickaël Chen, Ludovic Denoyer, Thierry Artières |
ICLR (Poster) | 1 |
| 2017 | Multi-view Generative Adversarial Networks
Mickaël Chen, Ludovic Denoyer |
ECML/PKDD (2) | 1 |