EDBT 2026 Demo / reviewers in the wild / expert
François Fleuret
dblp:90/5265
· DBLP profile ↗
111ranked-venue papers
9as first author
27since 2021 · last 2025
0000-0001-9457-7393ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 106 · 9 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pareto Low-Rank Adapters: Efficient Multi-Task Learning with PreferencesabstractMulti-task trade-offs in machine learning can be addressed via Pareto Front Learning (PFL) methods that parameterize the Pareto Front (PF) with a single model. PFL permits to select the desired operational point during inference, contrary to traditional Multi-Task Learning (MTL) that optimizes for a single trade-off decided prior to training. However, recent PFL methodologies suffer from limited scalability, slow convergence, and excessive memory requirements, while exhibiting inconsistent mappings from preference to objective space. We introduce PaLoRA, a novel parameter-efficient method that addresses these limitations in two ways. First, we augment any neural network architecture with task-specific low-rank adapters and continuously parameterize the Pareto Front in their convex hull. Our approach steers the original model and the adapters towards learning general and task-specific features, respectively. Second, we propose a deterministic sampling schedule of preference vectors that reinforces this division of labor, enabling faster convergence and strengthening the validity of the mapping from preference to objective space throughout training. Our experiments show that PaLoRA outperforms state-of-the-art MTL and PFL baselines across various datasets, scales to large networks, reducing the memory overhead $23.8-31.7$ times compared with competing PFL baselines in scene understanding benchmarks. Nikolaos Dimitriadis, Pascal Frossard, François Fleuret |
ICLR | 3 |
| 2025 | LiNeS: Post-training Layer Scaling Prevents Forgetting and Enhances Model MergingabstractFine-tuning pre-trained models has become the standard approach to endow them with specialized knowledge, but it poses fundamental challenges. In particular, (i) fine-tuning often leads to catastrophic forgetting, where improvements on a target domain degrade generalization on other tasks, and (ii) merging fine-tuned checkpoints from disparate tasks can lead to significant performance loss. To address these challenges, we introduce LiNeS, Layer-increasing Network Scaling, a post-training editing technique designed to preserve pre-trained generalization while enhancing fine-tuned task performance. LiNeS scales parameter updates linearly based on their layer depth within the network, maintaining shallow layers close to their pre-trained values to preserve general features while allowing deeper layers to retain task-specific representations. In multi-task model merging scenarios, layer-wise scaling of merged parameters reduces negative task interference. LiNeS demonstrates significant improvements in both single-task and multi-task settings across various benchmarks in vision and natural language processing. It mitigates forgetting, enhances out-of-distribution generalization, integrates seamlessly with existing multi-task model merging baselines improving their performance across benchmarks and model sizes, and can boost generalization when merging LLM policies aligned with different rewards via RLHF. Our method is simple to implement, computationally efficient and complementary to many existing techniques. Our source code is available at github.com/wang-kee/LiNeS. Nikolaos Dimitriadis, Alessandro Favero, Guillermo Ortiz-Jiménez, François Fleuret, Pascal Frossard |
ICLR | 5 |
| 2024 | Efficient World Models with Context-Aware TokenizationabstractScaling up deep Reinforcement Learning (RL) methods presents a significant challenge. Following developments in generative modelling, model-based RL positions itself as a strong contender. Recent advances in sequence modelling have led to effective transformer-based world models, albeit at the price of heavy computations due to the long sequences of tokens required to accurately simulate environments. In this work, we propose $\Delta$-IRIS, a new agent with a world model architecture composed of a discrete autoencoder that encodes stochastic deltas between time steps and an autoregressive transformer that predicts future deltas by summarizing the current state of the world with continuous tokens. In the Crafter benchmark, $\Delta$-IRIS sets a new state of the art at multiple frame budgets, while being an order of magnitude faster to train than previous attention-based approaches. We release our code and models at https://github.com/vmicheli/delta-iris. Vincent Micheli, Eloi Alonso, François Fleuret |
ICML | 3 |
| 2024 | Localizing Task Information for Improved Model Merging and CompressionabstractModel merging and task arithmetic have emerged as promising scalable approaches to merge multiple single-task checkpoints to one multi-task model, but their applicability is reduced by significant performance loss. Previous works have linked these drops to interference in the weight space and erasure of important task-specific features. Instead, in this work we show that the information required to solve each task is still preserved after merging as different tasks mostly use non-overlapping sets of weights. We propose TALL-masks, a method to identify these task supports given a collection of task vectors and show that one can retrieve $>$99% of the single task accuracy by applying our masks to the multi-task vector, effectively compressing the individual checkpoints. We study the statistics of intersections among constructed masks and reveal the existence of selfish and catastrophic weights, i.e., parameters that are important exclusively to one task and irrelevant to all tasks but detrimental to multi-task fusion. For this reason, we propose Consensus Merging, an algorithm that eliminates such weights and improves the general performance of existing model merging approaches. Our experiments in vision and NLP benchmarks with up to 20 tasks, show that Consensus Merging consistently improves existing approaches. Furthermore, our proposed compression scheme reduces storage from 57Gb to 8.2Gb while retaining 99.7% of original performance. Nikolaos Dimitriadis, Guillermo Ortiz-Jiménez, François Fleuret, Pascal Frossard |
ICML | 4 |
| 2024 | DeepEMD: A Transformer-Based Fast Estimation of the Earth Mover's Distance
Atul Kumar Sinha, François Fleuret |
ICPR (4) | 2 |
| 2024 | Diffusion for World Modeling: Visual Details Matter in AtariabstractWorld models constitute a promising approach for training reinforcement learning agents in a safe and sample-efficient manner. Recent world models predominantly operate on sequences of discrete latent variables to model environment dynamics. However, this compression into a compact discrete representation may ignore visual details that are important for reinforcement learning. Concurrently, diffusion models have become a dominant approach for image generation, challenging well-established methods modeling discrete latents. Motivated by this paradigm shift, we introduce DIAMOND (DIffusion As a Model Of eNvironment Dreams), a reinforcement learning agent trained in a diffusion world model. We analyze the key design choices that are required to make diffusion suitable for world modeling, and demonstrate how improved visual details can lead to improved agent performance. DIAMOND achieves a mean human normalized score of 1.46 on the competitive Atari 100k benchmark; a new best for agents trained entirely within a world model. We further demonstrate that DIAMOND's diffusion world model can stand alone as an interactive neural game engine by training on static *Counter-Strike: Global Offensive* gameplay. To foster future research on diffusion for world modeling, we release our code, agents, videos and playable world models at https://diamond-wm.github.io. Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos J. Storkey, Tim Pearce, François Fleuret |
NeurIPS | 7 |
| 2024 | DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingabstractThe transformer architecture by Vaswani et al. (2017) is now ubiquitous across application domains, from natural language processing to speech processing and image understanding. We propose DenseFormer, a simple modification to the standard architecture that improves the perplexity of the model without increasing its size---adding a few thousand parameters for large-scale models in the 100B parameters range. Our approach relies on an additional averaging step after each transformer block, which computes a weighted average of current and past representations---we refer to this operation as Depth-Weighted-Average (DWA). The learned DWA weights exhibit coherent patterns of information flow, revealing the strong and structured reuse of activations from distant layers. Experiments demonstrate that DenseFormer is more data efficient, reaching the same perplexity of much deeper transformer models, and that for the same perplexity, these new models outperform transformer baselines in terms of memory efficiency and inference time. Matteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin Jaggi |
NeurIPS | 3 |
| 2024 | σ-GPTs: A New Approach to Autoregressive Models
Arnaud Pannatier, Evann Courdier, François Fleuret |
ECML/PKDD (7) | 3 |
| 2023 | HyperMixer: An MLP-based Low Cost Alternative to TransformersabstractFlorian Mai, Arnaud Pannatier, Fabio Fehr, Haolin Chen, Francois Marelli, Francois Fleuret, James Henderson. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Florian Mai, Arnaud Pannatier, Fabio Fehr, François Marelli, François Fleuret, James Henderson 0001 |
ACL (1) | 6 |
| 2023 | Human control redressed: Comparing AI and human predictability in a real-effort task
Serhiy Kandul, Vincent Micheli, Juliane Beck, Thomas Burri, François Fleuret, Markus Kneer, Markus Christen |
CogSci | 5 |
| 2023 | ESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance FieldsabstractWe present ESLAM, an efficient implicit neural representation method for Simultaneous Localization and Mapping (SLAM). ESLAM reads RGB-D frames with unknown camera poses in a sequential manner and incrementally reconstructs the scene representation while estimating the current camera position in the scene. We incorporate the latest advances in Neural Radiance Fields (NeRF) into a SLAM system, resulting in an efficient and accurate dense visual SLAM method. Our scene representation consists of multi-scale axis-aligned perpendicular feature planes and shallow decoders that, for each point in the continuous space, decode the interpolated features into Truncated Signed Distance Field (TSDF) and RGB values. Our extensive experiments on three standard datasets, Replica, ScanNet, and TUM RGB-D show that ESLAM improves the accuracy of 3D reconstruction and camera localization of state-of-the-art dense visual SLAM methods by more than 50%, while it runs up to ×10 faster and does not require any pre-training. Project page: https://www.idiap.ch/paper/eslam. Mohammad Mahdi Johari, Camilla Carta, François Fleuret |
CVPR | 3 |
| 2023 | Transformers are Sample-Efficient World Models
Vincent Micheli, Eloi Alonso, François Fleuret |
ICLR | 3 |
| 2023 | Agree to Disagree: Diversity through Disagreement for Better Transferability
Matteo Pagliardini, Martin Jaggi, François Fleuret, Sai Praneeth Karimireddy |
ICLR | 3 |
| 2023 | Pareto Manifold Learning: Tackling multiple tasks via ensembles of single-task modelsabstractIn Multi-Task Learning (MTL), tasks may compete and limit the performance achieved on each other, rather than guiding the optimization to a solution, superior to all its single-task trained counterparts. Since there is often not a unique solution optimal for all tasks, practitioners have to balance tradeoffs between tasks' performance, and resort to optimality in the Pareto sense. Most MTL methodologies either completely neglect this aspect, and instead of aiming at learning a Pareto Front, produce one solution predefined by their optimization schemes, or produce diverse but discrete solutions. Recent approaches parameterize the Pareto Front via neural networks, leading to complex mappings from tradeoff to objective space. In this paper, we conjecture that the Pareto Front admits a linear parameterization in parameter space, which leads us to propose *Pareto Manifold Learning*, an ensembling method in weight space. Our approach produces a continuous Pareto Front in a single training run, that allows to modulate the performance on each task during inference. Experiments on multi-task learning benchmarks, ranging from image classification to tabular datasets and scene understanding, show that *Pareto Manifold Learning* outperforms state-of-the-art single-point algorithms, while learning a better Pareto parameterization than multi-point baselines. Nikolaos Dimitriadis, Pascal Frossard, François Fleuret |
ICML | 3 |
| 2023 | Fast Attention Over Long Sequences With Dynamic Sparse Flash AttentionabstractTransformer-based language models have found many diverse applications requiring them to process sequences of increasing length. For these applications, the causal self-attention---which is the only component scaling quadratically w.r.t. the sequence length---becomes a central concern. While many works have proposed schemes to sparsify the attention patterns and reduce the computational overhead of self-attention, those are often limited by implementation concerns and end up imposing a simple and static structure over the attention matrix. Conversely, implementing more dynamic sparse attention often results in runtimes significantly slower than computing the full attention using the Flash implementation from Dao et al. (2022). We extend FlashAttention to accommodate a large class of attention sparsity patterns that, in particular, encompass key/query dropping and hashing-based attention. This leads to implementations with no computational complexity overhead and a multi-fold runtime speedup on top of FlashAttention. Even with relatively low degrees of sparsity, our method improves visibly upon FlashAttention as the sequence length increases. Without sacrificing perplexity, we increase the training speed of a transformer language model by $2.0\times$ and $3.3\times$ for sequences of respectively $8k$ and $16k$ tokens. Matteo Pagliardini, Daniele Paliotta, Martin Jaggi, François Fleuret |
NeurIPS | 4 |
| 2023 | SUPA: A Lightweight Diagnostic Simulator for Machine Learning in Particle PhysicsabstractDeep learning methods have gained popularity in high energy physics for fast modeling of particle showers in detectors. Detailed simulation frameworks such as the gold standard \textsc{Geant4} are computationally intensive, and current deep generative architectures work on discretized, lower resolution versions of the detailed simulation. The development of models that work at higher spatial resolutions is currently hindered by the complexity of the full simulation data, and by the lack of simpler, more interpretable benchmarks. Our contribution is \textsc{SUPA}, the SUrrogate PArticle propagation simulator, an algorithm and software package for generating data by simulating simplified particle propagation, scattering and shower development in matter. The generation is extremely fast and easy to use compared to \textsc{Geant4}, but still exhibits the key characteristics and challenges of the detailed simulation. The proposed simulator generates thousands of particle showers per second on a desktop machine, a speed up of up to 6 orders of magnitudes over \textsc{Geant4}, and stores detailed geometric information about the shower propagation. \textsc{\textsc{SUPA}} provides much greater flexibility for setting initial conditions and defining multiple benchmarks for the development of models. Moreover, interpreting particle showers as point clouds creates a connection to geometric machine learning and provides challenging and fundamentally new datasets for the field. Atul Kumar Sinha, Daniele Paliotta, Bálint Máté, John A. Raine, Tobias Golling, François Fleuret |
NeurIPS | 6 |
| 2022 | PAUMER: Patch Pausing Transformer for Semantic Segmentation
Evann Courdier, Prabhu Teja Sivaprasad, François Fleuret |
BMVC | 3 |
| 2022 | GeoNeRF: Generalizing NeRF with Geometry PriorsabstractWe present GeoNeRF, a generalizable photorealistic novel view synthesis method based on neural radiance fields. Our approach consists of two main stages: a ge-ometry reasoner and a renderer. To render a novel view, the geometry reasoner first constructs cascaded cost volumes for each nearby source view. Then, using a Transformer- based attention mechanism and the cascaded cost volumes, the renderer infers geometry and appearance, and ren-ders detailed images via classical volume rendering techniques. This architecture, in particular, allows sophis-ticated occlusion reasoning, gathering information from consistent source views. Moreover, our method can eas-ily be fine-tuned on a single scene, and renders com-petitive results with per-scene optimized neural rendering methods with a fraction of computational cost. Ex-periments show that GeoNeRF outperforms state-of-the- art generalizable neural rendering models on various syn-thetic and real datasets. Lastly, with a slight modification to the geometry reasoner, we also propose an alter-native model that adapts to RGBD images. This model di-rectly exploits the depth information often available thanks to depth sensors. The implementation code is available at https://www.idiap.ch/paper/geonerf. Mohammad Mahdi Johari, Yann Lepoittevin, François Fleuret |
CVPR | 3 |
| 2022 | Borrowing from yourself: Faster future video segmentation with partial channel updateabstractSemantic segmentation is a well-addressed topic in the computer vision literature, but the design of fast and accurate video processing networks remains challenging. In addition, to run on embedded hardware, computer vision models often have to make compromises on accuracy to run at the required speed, so that a latency/accuracy trade-off is usually at the heart of these real-time systems’ design. For the specific case of videos, models have the additional possibility to make use of computations made for previous frames to mitigate the accuracy loss while being real-time.In this work, we propose to tackle the task of fast future video segmentation prediction through the use of convolutional layers with time-dependent channel masking. This technique only updates a chosen subset of the feature maps at each timestep, bringing simultaneously less computation and latency, and allowing the network to leverage previously computed features. We apply this technique to several fast architectures and experimentally confirm its benefits for the future prediction subtask. Evann Courdier, François Fleuret |
ICPR | 2 |
| 2022 | Flowification: Everything is a normalizing flowabstractThe two key characteristics of a normalizing flow is that it is invertible (in particular, dimension preserving) and that it monitors the amount by which it changes the likelihood of data points as samples are propagated along the network. Recently, multiple generalizations of normalizing flows have been introduced that relax these two conditions \citep{nielsen2020survae,huang2020augmented}. On the other hand, neural networks only perform a forward pass on the input, there is neither a notion of an inverse of a neural network nor is there one of its likelihood contribution. In this paper we argue that certain neural network architectures can be enriched with a stochastic inverse pass and that their likelihood contribution can be monitored in a way that they fall under the generalized notion of a normalizing flow mentioned above. We term this enrichment \emph{flowification}. We prove that neural networks only containing linear and convolutional layers and invertible activations such as LeakyReLU can be flowified and evaluate them in the generative setting on image datasets. Bálint Máté, Samuel Klein, Tobias Golling, François Fleuret |
NeurIPS | 4 |
| 2022 | Efficient Training of Low-Curvature Neural NetworksabstractStandard deep neural networks often have excess non-linearity, making them susceptible to issues such as low adversarial robustness and gradient instability. Common methods to address these downstream issues, such as adversarial training, are expensive and often sacrifice predictive accuracy. In this work, we address the core issue of excess non-linearity via curvature, and demonstrate low-curvature neural networks (LCNNs) that obtain drastically lower curvature than standard models while exhibiting similar predictive performance. This leads to improved robustness and stable gradients, at a fraction of the cost of standard adversarial training. To achieve this, we decompose overall model curvature in terms of curvatures and slopes of its constituent layers. To enable efficient curvature minimization of constituent layers, we introduce two novel architectural components: first, a non-linearity called centered-softplus that is a stable variant of the softplus non-linearity, and second, a Lipschitz-constrained batch normalization layer.Our experiments show that LCNNs have lower curvature, more stable gradients and increased off-the-shelf adversarial robustness when compared to standard neural networks, all without affecting predictive performance. Our approach is easy to use and can be readily incorporated into existing neural network architectures. Suraj Srinivas, Kyle Matoba, Himabindu Lakkaraju, François Fleuret |
NeurIPS | 4 |
| 2022 | Efficient Wind Speed Nowcasting with GPU-Accelerated Nearest Neighbors Algorithm
Arnaud Pannatier, Ricardo Picatoste, François Fleuret |
SDM | 3 |
| 2021 | Uncertainty Reduction for Model Adaptation in Semantic SegmentationabstractTraditional methods for Unsupervised Domain Adaptation (UDA) targeting semantic segmentation exploit information common to the source and target domains, using both labeled source data and unlabeled target data. In this paper, we investigate a setting where the source data is un-available, but the classifier trained on the source data is; hence named "model adaptation". Such a scenario arises when data sharing is prohibited, for instance, because of privacy, or Intellectual Property (IP) issues.To tackle this problem, we propose a method that reduces the uncertainty of predictions on the target domain data. We accomplish this in two ways: minimizing the entropy of the predicted posterior, and maximizing the noise robustness of the feature representation. We show the efficacy of our method on the transfer of segmentation from computer generated images to real-world driving images, and transfer between data collected in different cities, and surprisingly reach performance comparable with that of the methods that have access to source data. Prabhu Teja Sivaprasad, François Fleuret |
CVPR | 2 |
| 2021 | Language Models are Few-Shot ButlersabstractPretrained language models demonstrate strong performance in most NLP tasks when fine-tuned on small task-specific datasets.Hence, these autoregressive models constitute ideal agents to operate in text-based environments where language understanding and generative capabilities are essential.Nonetheless, collecting expert demonstrations in such environments is a time-consuming endeavour.We introduce a two-stage procedure to learn from a small set of demonstrations and further improve by interacting with an environment.We show that language models fine-tuned with only 1.2% of the expert demonstrations and a simple reinforcement learning algorithm achieve a 51% absolute improvement in success rate over existing methods in the ALFWorld environment.Goal: Rinse the egg to put it in the microwave.Obs: Looking quickly around you, you see a cabinet, a garbagecan, a coffeemachine, [...], a stoveburner, a sinkbasin and a microwave.Action: go to sinkbasin Obs: You arrive at sinkbasin.You see a butterknife, a potato, a spoon Vincent Micheli, François Fleuret |
EMNLP (1) | 2 |
| 2021 | DepthInSpace: Exploitation and Fusion of Multiple Video Frames for Structured-Light Depth EstimationabstractWe present DepthInSpace, a self-supervised deep-learning method for depth estimation using a structured-light camera. The design of this method is motivated by the commercial use case of embedded depth sensors in nowadays smartphones. We first propose to use estimated optical flow from ambient information of multiple video frames as a complementary guide for training a single-frame depth estimation network, helping to preserve edges and reduce over-smoothing issues. Utilizing optical flow, we also propose to fuse the data of multiple video frames to get a more accurate depth map. In particular, fused depth maps are more robust in occluded areas and incur less in flying pixels artifacts. We finally demonstrate that these more precise fused depth maps can be used as self-supervision for fine-tuning a single-frame depth estimation network to improve its performance. Our models’ effectiveness is evaluated and compared with state-of-the-art models on both synthetic and our newly introduced real datasets. The implementation code, training procedure, and both synthetic and captured real datasets are available at https://www.idiap.ch/paper/depthinspace. Mohammad Mahdi Johari, Camilla Carta, François Fleuret |
ICCV | 3 |
| 2021 | Taming GANs with Lookahead-Minmax
Tatjana Chavdarova, Matteo Pagliardini, Sebastian U. Stich, François Fleuret, Martin Jaggi |
ICLR | 4 |
| 2021 | Rethinking the Role of Gradient-based Attribution Methods for Model Interpretability
Suraj Srinivas, François Fleuret |
ICLR | 2 |
| 2020 | Real-Time Segmentation Networks Should be Latency Aware
Evann Courdier, François Fleuret |
ACCV (1) | 2 |
| 2020 | On the importance of pre-training data volume for compact language modelsabstractRecent advances in language modeling have led to computationally intensive and resourcedemanding state-of-the-art models.In an effort towards sustainable practices, we study the impact of pre-training data volume on compact language models.Multiple BERT-based models are trained on gradually increasing amounts of French text.Through fine-tuning on the French Question Answering Dataset (FQuAD), we observe that well-performing models are obtained with as little as 100 MB of text.In addition, we show that past critically low amounts of pre-training data, an intermediate pre-training step on the task-specific corpus does not yield substantial improvements. Vincent Micheli, Martin d'Hoffschmidt, François Fleuret |
EMNLP (1) | 3 |
| 2020 | Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionabstractTransformers achieve remarkable performance in several tasks but due to their quadratic complexity, with respect to the input’s length, they are prohibitively slow for very long sequences. To address this limitation, we express the self-attention as a linear dot-product of kernel feature maps and make use of the associativity property of matrix products to reduce the complexity from $\bigO{N^2}$ to $\bigO{N}$, where $N$ is the sequence length. We show that this formulation permits an iterative implementation that dramatically accelerates autoregressive transformers and reveals their relationship to recurrent neural networks. Our \emph{Linear Transformers} achieve similar performance to vanilla Transformers and they are up to 4000x faster on autoregressive prediction of very long sequences. Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas 0002, François Fleuret |
ICML | 4 |
| 2020 | Optimizer Benchmarking Needs to Account for Hyperparameter TuningabstractThe performance of optimizers, particularly in deep learning, depends considerably on their chosen hyperparameter configuration. The efficacy of optimizers is often studied under near-optimal problem-specific hyperparameters, and finding these settings may be prohibitively costly for practitioners. In this work, we argue that a fair assessment of optimizers’ performance must take the computational cost of hyperparameter tuning into account, i.e., how easy it is to find good hyperparameter configurations using an automatic hyperparameter search. Evaluating a variety of optimizers on an extensive set of standard datasets and architectures, our results indicate that Adam is the most practical solution, particularly in low-budget scenarios. Prabhu Teja Sivaprasad, Florian Mai, Thijs Vogels, Martin Jaggi, François Fleuret |
ICML | 5 |
| 2020 | Fast Transformers with Clustered AttentionabstractTransformers have been proven a successful model for a variety of tasks in sequence modeling. However, computing the attention matrix, which is their key component, has quadratic complexity with respect to the sequence length, thus making them prohibitively expensive for large sequences. To address this, we propose clustered attention, which instead of computing the attention for every query, groups queries into clusters and computes attention just for the centroids. To further improve this approximation, we use the computed clusters to identify the keys with the highest attention per query and compute the exact key/query dot products. This results in a model with linear complexity with respect to the sequence length for a fixed number of clusters. We evaluate our approach on two automatic speech recognition datasets and show that our model consistently outperforms vanilla transformers for a given computational budget. Finally, we demonstrate that our model can approximate arbitrarily complex attention distributions with a minimal number of clusters by approximating a pretrained BERT model on GLUE and SQuAD benchmarks with only 25 clusters and no loss in performance. Apoorv Vyas, Angelos Katharopoulos, François Fleuret |
NeurIPS | 3 |
| 2019 | Learning an Event Sequence Embedding for Dense Event-Based Deep StereoabstractToday, a frame-based camera is the sensor of choice for machine vision applications. However, these cameras, originally developed for acquisition of static images rather than for sensing of dynamic uncontrolled visual environments, suffer from high power consumption, data rate, latency and low dynamic range. An event-based image sensor addresses these drawbacks by mimicking a biological retina. Instead of measuring the intensity of every pixel in a fixed time-interval, it reports events of significant pixel intensity changes. Every such event is represented by its position, sign of change, and timestamp, accurate to the microsecond. Asynchronous event sequences require special handling, since traditional algorithms work only with synchronous, spatially gridded data. To address this problem we introduce a new module for event sequence embedding, for use in difference applications. The module builds a representation of an event sequence by firstly aggregating information locally across time, using a novel fully-connected layer for an irregularly sampled continuous domain, and then across discrete spatial domain. Based on this module, we design a deep learning-based stereo method for event-based cameras. The proposed method is the first learning-based stereo method for an event-based camera and the only method that produces dense results. We show that large performance increases on the Multi Vehicle Stereo Event Camera Dataset (MVSEC), which became the standard set for benchmarking of event-based stereo methods. Stepan Tulyakov, François Fleuret, Martin Kiefel, Peter V. Gehler |
ICCV | 2 |
| 2019 | Processing Megapixel Images with Deep Attention-Sampling ModelsabstractExisting deep architectures cannot operate on very large signals such as megapixel images due to computational and memory constraints. To tackle this limitation, we propose a fully differentiable end-to-end trainable model that samples and processes only a fraction of the full resolution input image. The locations to process are sampled from an attention distribution computed from a low resolution view of the input. We refer to our method as attention sampling and it can process images of several megapixels with a standard single GPU setup. We show that sampling from the attention distribution results in an unbiased estimator of the full model with minimal variance, and we derive an unbiased estimator of the gradient that we use to train our model end-to-end with a normal SGD procedure. This new method is evaluated on three classification tasks, where we show that it allows to reduce computation and memory footprint by an order of magnitude for the same accuracy as classical architectures. We also show the consistency of the sampling that indeed focuses on informative parts of the input images. Angelos Katharopoulos, François Fleuret |
ICML | 2 |
| 2019 | Reducing Noise in GAN Training with Variance Reduced ExtragradientabstractWe study the effect of the stochastic gradient noise on the training of generative adversarial networks (GANs) and show that it can prevent the convergence of standard game optimization methods, while the batch version converges. We address this issue with a novel stochastic variance-reduced extragradient (SVRE) optimization algorithm, which for a large class of games improves upon the previous convergence rates proposed in the literature. We observe empirically that SVRE performs similarly to a batch method on MNIST while being computationally cheaper, and that SVRE yields more stable GAN training on standard datasets. Tatjana Chavdarova, Gauthier Gidel, François Fleuret, Simon Lacoste-Julien |
NeurIPS | 3 |
| 2019 | Full-Gradient Representation for Neural Network VisualizationabstractWe introduce a new tool for interpreting neural nets, namely full-gradients, which decomposes the neural net response into input sensitivity and per-neuron sensitivity components. This is the first proposed representation which satisfies two key properties: completeness and weak dependence, which provably cannot be satisfied by any saliency map-based interpretability method. Using full-gradients, we also propose an approximate saliency map representation for convolutional nets dubbed FullGrad, obtained by aggregating the full-gradient components. We experimentally evaluate the usefulness of FullGrad in explaining model behaviour with two quantitative tests: pixel perturbation and remove-and-retrain. Our experiments reveal that our method explains model behavior correctly, and more comprehensively than other methods in the literature. Visual inspection also reveals that our saliency maps are sharper and more tightly confined to object regions than other methods. Suraj Srinivas, François Fleuret |
NeurIPS | 2 |
| 2018 | WILDTRACK: A Multi-Camera HD Dataset for Dense Unscripted Pedestrian DetectionabstractPeople detection methods are highly sensitive to occlusions between pedestrians, which are extremely frequent in many situations where cameras have to be mounted at a limited height. The reduction of camera prices allows for the generalization of static multi-camera set-ups. Using joint visual information from multiple synchronized cameras gives the opportunity to improve detection performance. In this paper, we present a new large-scale and high-resolution dataset. It has been captured with seven static cameras in a public open area, and unscripted dense groups of pedestrians standing and walking. Together with the camera frames, we provide an accurate joint (extrinsic and intrinsic) calibration, as well as 7 series of 400 annotated frames for detection at a rate of 2 frames per second. This results in over 40 000 bounding boxes delimiting every person present in the area of interest, for a total of more than 300 individuals. We provide a series of benchmark results using baseline algorithms published over the recent months for multi-view detection with deep neural networks, and trajectory estimation using a non-Markovian model. Tatjana Chavdarova, Pierre Baqué, Stéphane Bouquet, Andrii Maksai, Cijo Jose, Timur M. Bagautdinov, Louis Lettry, Pascal Fua, Luc Van Gool, François Fleuret |
CVPR | 10 |
| 2018 | SGAN: An Alternative Training of Generative Adversarial NetworksabstractThe Generative Adversarial Networks (GANs) have demonstrated impressive performance for data synthesis, and are now used in a wide range of computer vision tasks. In spite of this success, they gained a reputation for being difficult to train, what results in a time-consuming and human-involved development process to use them. We consider an alternative training process, named SGAN, in which several adversarial "local" pairs of networks are trained independently so that a "global" supervising pair of networks can be trained against them. The goal is to train the global pair with the corresponding ensemble opponent for improved performances in terms of mode coverage. This approach aims at increasing the chances that learning will not stop for the global pair, preventing both to be trapped in an unsatisfactory local minimum, or to face oscillations often observed in practice. To guarantee the latter, the global pair never affects the local ones. The rules of SGAN training are thus as follows: the global generator and discriminator are trained using the local discriminators and generators, respectively, whereas the local networks are trained with their fixed local opponent. Experimental results on both toy and real-world problems demonstrate that this approach outperforms standard training in terms of better mitigating mode collapse, stability while converging and that it surprisingly, increases the convergence speed as well. Tatjana Chavdarova, François Fleuret |
CVPR | 2 |
| 2018 | Geodesic Convolutional Shape OptimizationabstractAerodynamic shape optimization has many industrial applications. Existing methods, however, are so computationally demanding that typical engineering practices are to either simply try a limited number of hand-designed shapes or restrict oneself to shapes that can be parameterized using only few degrees of freedom. In this work, we introduce a new way to optimize complex shapes fast and accurately. To this end, we train Geodesic Convolutional Neural Networks to emulate a fluidynamics simulator. The key to making this approach practical is remeshing the original shape using a poly-cube map, which makes it possible to perform the computations on GPUs instead of CPUs. The neural net is then used to formulate an objective function that is differentiable with respect to the shape parameters, which can then be optimized using a gradient-based technique. This outperforms state-of-the-art methods by 5 to 20% for standard problems and, even more importantly, our approach applies to cases that previous methods cannot handle. Pierre Baqué, Edoardo Remelli, François Fleuret, Pascal Fua |
ICML | 3 |
| 2018 | Kronecker Recurrent UnitsabstractOur work addresses two important issues with recurrent neural networks: (1) they are over-parametrized, and (2) the recurrent weight matrix is ill-conditioned. The former increases the sample complexity of learning and the training time. The latter causes the vanishing and exploding gradient problem. We present a flexible recurrent neural network model called Kronecker Recurrent Units (KRU). KRU achieves parameter efficiency in RNNs through a Kronecker factored recurrent matrix. It overcomes the ill-conditioning of the recurrent matrix by enforcing soft unitary constraints on the factors. Thanks to the small dimensionality of the factors, maintaining these constraints is computationally efficient. Our experimental results on seven standard data-sets reveal that KRU can reduce the number of parameters by three orders of magnitude in the recurrent weight matrix compared to the existing recurrent models, without trading the statistical performance. These results in particular show that while there are advantages in having a high dimensional recurrent space, the capacity of the recurrent part of the model can be dramatically reduced. Cijo Jose, Moustapha Cissé, François Fleuret |
ICML | 3 |
| 2018 | Not All Samples Are Created Equal: Deep Learning with Importance SamplingabstractDeep Neural Network training spends most of the computation on examples that are properly handled, and could be ignored. We propose to mitigate this phenomenon with a principled importance sampling scheme that focuses computation on "informative" examples, and reduces the variance of the stochastic gradients during training. Our contribution is twofold: first, we derive a tractable upper bound to the per-sample gradient norm, and second we derive an estimator of the variance reduction achieved with importance sampling, which enables us to switch it on when it will result in an actual speedup. The resulting scheme can be used by changing a few lines of code in a standard SGD procedure, and we demonstrate experimentally on image classification, CNN fine-tuning, and RNN training, that for a fixed wall-clock time budget, it provides a reduction of the train losses of up to an order of magnitude and a relative improvement of test errors between 5% and 17%. Angelos Katharopoulos, François Fleuret |
ICML | 2 |
| 2018 | Knowledge Transfer with Jacobian MatchingabstractClassical distillation methods transfer representations from a “teacher” neural network to a “student” network by matching their output activations. Recent methods also match the Jacobians, or the gradient of output activations with the input. However, this involves making some ad hoc decisions, in particular, the choice of the loss function. In this paper, we first establish an equivalence between Jacobian matching and distillation with input noise, from which we derive appropriate loss functions for Jacobian matching. We then rely on this analysis to apply Jacobian matching to transfer learning by establishing equivalence of a recent transfer learning procedure to distillation. We then show experimentally on standard image datasets that Jacobian-based penalties improve distillation, robustness to noisy inputs, and transfer learning. Suraj Srinivas, François Fleuret |
ICML | 2 |
| 2018 | Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matchingabstractEnd-to-end deep-learning networks recently demonstrated extremely good performance for stereo matching. However, existing networks are difficult to use for practical applications since (1) they are memory-hungry and unable to process even modest-size images, (2) they have to be fully re-trained to handle a different disparity range. The Practical Deep Stereo (PDS) network that we propose addresses both issues: First, its architecture relies on novel bottleneck modules that drastically reduce the memory footprint in inference, and additional design choices allow to handle greater image size during training. This results in a model that leverages large image context to resolve matching ambiguities. Second, a novel sub-pixel cross-entropy loss combined with a MAP estimator make this network less sensitive to ambiguous matches, and applicable to any disparity range without re-training. We compare PDS to state-of-the-art methods published over the recent months, and demonstrate its superior performance on FlyingThings3D and KITTI sets. Stepan Tulyakov, Anton Ivanov, François Fleuret |
NeurIPS | 3 |
| 2017 | A Sub-Quadratic Exact Medoid AlgorithmabstractWe present a new algorithm, ‘trimed’ for obtaining the medoid of a set, that is the element of the set which minimises the mean distance to all other elements. The algorithm is shown to have, under certain assumptions, expected run time $O(N^(3/2))$ in $R^d$ where N is the set size, making it the first sub-quadratic exact medoid algorithm for $d > 1$. Experiments show that it performs very well on spatial network data, frequently requiring two orders of magnitude fewer distance calculations than state-of-the-art approximate algorithms. As an application, we show how trimed can be used as a component in an accelerated K-medoids algorithm, and then how it can be relaxed to obtain further computational gains with only a minor loss in cluster quality. James Newling, François Fleuret |
AISTATS | 2 |
| 2017 | Social Scene Understanding: End-to-End Multi-person Action Localization and Collective Activity RecognitionabstractWe present a unified framework for understanding human social behaviors in raw image sequences. Our model jointly detects multiple individuals, infers their social actions, and estimates the collective actions with a single feed-forward pass through a neural network. We propose a single architecture that does not rely on external detection algorithms but rather is trained end-to-end to generate dense proposal maps that are refined via a novel inference scheme. The temporal consistency is handled via a person-level matching Recurrent Neural Network. The complete model takes as input a sequence of frames and outputs detections along with the estimates of individual actions and collective activities. We demonstrate state-of-the-art performance of our algorithm on multiple publicly available benchmarks. Timur M. Bagautdinov, Alexandre Alahi, François Fleuret, Pascal Fua, Silvio Savarese |
CVPR | 3 |
| 2017 | Multi-modal Mean-Fields via Cardinality-Based Clamping
Pierre Baqué, François Fleuret, Pascal Fua |
CVPR | 2 |
| 2017 | Deep Occlusion Reasoning for Multi-camera Multi-target DetectionabstractPeople detection in single 2D images has improved greatly in recent years. However, comparatively little of this progress has percolated into multi-camera multi-people tracking algorithms, whose performance still degrades severely when scenes become very crowded. In this work, we introduce a new architecture that combines Convolutional Neural Nets and Conditional Random Fields to explicitly model those ambiguities. One of its key ingredients are high-order CRF terms that model potential occlusions and give our approach its robustness even when many people are present. Our model is trained end-to-end and we show that it outperforms several state-of-the-art algorithms on challenging scenes. Pierre Baqué, François Fleuret, Pascal Fua |
ICCV | 2 |
| 2017 | Non-Markovian Globally Consistent Multi-object TrackingabstractMany state-of-the-art approaches to multi-object tracking rely on detecting them in each frame independently, grouping detections into short but reliable trajectory segments, and then further grouping them into full trajectories. This grouping typically relies on imposing local smoothness constraints but almost never on enforcing more global ones on the trajectories. In this paper, we propose a non-Markovian approach to imposing global consistency by using behavioral patterns to guide the tracking algorithm. When used in conjunction with state-of-the-art tracking algorithms, this further increases their already good performance on multiple challenging datasets. We show significant improvements both in supervised settings where ground truth is available and behavioral patterns can be learned from it, and in completely unsupervised settings. Andrii Maksai, Xinchao Wang, François Fleuret, Pascal Fua |
ICCV | 3 |
| 2017 | Weakly Supervised Learning of Deep Metrics for Stereo ReconstructionabstractDeep-learning metrics have recently demonstrated extremely good performance to match image patches for stereo reconstruction. However, training such metrics requires large amount of labeled stereo images, which can be difficult or costly to collect for certain applications (consider, for example, satellite stereo imaging). The main contribution of our work is a new weakly supervised method for learning deep metrics from unlabeled stereo images, given coarse information about the scenes and the optical system. Our method alternatively optimizes the metric with a standard stochastic gradient descent, and applies stereo constraints to regularize its prediction. Experiments on reference data-sets show that, for a given network architecture, training with this new method without ground-truth produces a metric with performance as good as state-of-the-art baselines trained with the said ground-truth. This work has three practical implications. Firstly, it helps to overcome limitations of training sets, in particular noisy ground truth. Secondly it allows to use much more training data during learning. Thirdly, it allows to tune deep metric for a particular stereo system, even if ground truth is not available. Stepan Tulyakov, Anton Ivanov, François Fleuret |
ICCV | 3 |
| 2017 | Deep Multi-camera People DetectionabstractThis paper addresses the problem of multi-view people occupancy map estimation. Existing solutions either operate per-view, or rely on a background subtraction preprocessing. Both approaches lessen the detection performance as scenes become more crowded. The former does not exploit joint information, whereas the latter deals with ambiguous input due to the foreground blobs becoming more and more interconnected as the number of targets increases. Although deep learning algorithms have proven to excel on remarkably numerous computer vision tasks, such a method has not been applied yet to this problem. In large part this is due to the lack of large-scale multi-camera data-set. The core of our method is an architecture which makes use of monocular pedestrian data-set, available at larger scale than the multi-view ones, applies parallel processing to the multiple video streams, and jointly utilises it. Our end-to-end deep learning method outperforms existing methods by large margins on the commonly used PETS 2009 data-set. Furthermore, we make publicly available a new three-camera HD data-set. Tatjana Chavdarova, François Fleuret |
ICMLA | 2 |
| 2017 | K-Medoids For K-Means SeedingabstractWe show experimentally that the algorithm CLARANS of Ng and Han (1994) finds better K-medoids solutions than the Voronoi iteration algorithm of Hastie et al. (2001). This finding, along with the similarity between the Voronoi iteration algorithm and Lloyd's K-means algorithm, motivates us to use CLARANS as a K-means initializer. We show that CLARANS outperforms other algorithms on 23/23 datasets with a mean decrease over k-means++ of 30% for initialization mean squared error (MSE) and 3% for final MSE. We introduce algorithmic improvements to CLARANS which improve its complexity and runtime, making it a viable initialization scheme for large datasets. James Newling, François Fleuret |
NIPS | 2 |
| 2017 | Machine learning-based tools to model and to remove the off-target effect
Riwal Lefort, Ludovico Fusco, Olivier Pertz, François Fleuret |
Pattern Anal. Appl. | 4 |
| 2016 | Principled Parallel Mean-Field Inference for Discrete Random FieldsabstractMean-field variational inference is one of the most popular approaches to inference in discrete random fields. Standard mean-field optimization is based on coordinate descent and in many situations can be impractical. Thus, in practice, various parallel techniques are used, which either rely on ad hoc smoothing with heuristically set parameters, or put strong constraints on the type of models. In this paper, we propose a novel proximal gradient-based approach to optimizing the variational objective. It is naturally parallelizable and easy to implement. We prove its convergence, and demonstrate that, in practice, it yields faster convergence and often finds better optima than more traditional mean-field optimization techniques. Moreover, our method is less sensitive to the choice of parameters. Pierre Baqué, Timur M. Bagautdinov, François Fleuret, Pascal Fua |
CVPR | 3 |
| 2016 | Large Scale Hard Sample Mining with Monte Carlo Tree SearchabstractWe investigate an efficient strategy to collect false positives from very large training sets in the context of object detection.Our approach scales up the standard bootstrapping procedure by using a hierarchical decomposition of an image collection which reflects the statistical regularity of the detector's responses.Based on that decomposition, our procedure uses a Monte Carlo Tree Search to prioritize the sampling toward sub-families of images which have been observed to be rich in false positives, while maintaining a fraction of the sampling toward unexplored sub-families of images.The resulting procedure increases substantially the proportion of false positive samples among the visited ones compared to a naive uniform sampling.We apply experimentally this new procedure to face detection with a collection of ∼100,000 background images and to pedestrian detection with ∼32,000 images.We show that for two standard detectors, the proposed strategy cuts the number of images to visit by half to obtain the same amount of false positives and the same final performance. Olivier Canévet, François Fleuret |
CVPR | 2 |
| 2016 | Scalable Metric Learning via Weighted Approximate Rank Component Analysis
Cijo Jose, François Fleuret |
ECCV (5) | 2 |
| 2016 | Importance Sampling Tree for Large-scale Empirical ExpectationabstractWe propose a tree-based procedure inspired by the Monte-Carlo Tree Search that dynamically modulates an importance-based sampling to prioritize computation, while getting unbiased estimates of weighted sums. We apply this generic method to learning on very large training sets, and to the evaluation of large-scale SVMs. The core idea is to reformulate the estimation of a score - whether a loss or a prediction estimate - as an empirical expectation, and to use such a tree whose leaves carry the samples to focus efforts over the problematic "heavy weight" ones. We illustrate the potential of this approach on three problems: to improve Adaboost and a multi-layer perceptron on 2D synthetic tasks with several million points, to train a large-scale convolution network on several millions deformations of the CIFAR data-set, and to compute the response of a SVM with several hundreds of thousands of support vectors. In each case, we show how it either cuts down computation by more than one order of magnitude and/or allows to get better loss estimates. Olivier Canévet, Cijo Jose, François Fleuret |
ICML | 3 |
| 2016 | Fast k-means with accurate boundsabstractWe propose a novel accelerated exact k-means algorithm, which outperforms the current state-of-the-art low-dimensional algorithm in 18 of 22 experiments, running up to 3 times faster. We also propose a general improvement of existing state-of-the-art accelerated exact k-means algorithms through better estimates of the distance bounds used to reduce the number of distance calculations, obtaining speedups in 36 of 44 experiments, of up to 1.8 times. We have conducted experiments with our own implementations of existing methods to ensure homogeneous evaluation of performance, and we show that our implementations perform as well or better than existing available implementations. Finally, we propose simplified variants of standard approaches and show that they are faster than their fully-fledged counterparts in 59 of 62 experiments. James Newling, François Fleuret |
ICML | 2 |
| 2016 | Nested Mini-Batch K-MeansabstractA new algorithm is proposed which accelerates the mini-batch k-means algorithm of Sculley (2010) by using the distance bounding approach of Elkan (2003). We argue that, when incorporating distance bounds into a mini-batch algorithm, already used data should preferentially be reused. To this end we propose using nested mini-batches, whereby data in a mini-batch at iteration t is automatically reused at iteration t+1. Using nested mini-batches presents two difficulties. The first is that unbalanced use of data can bias estimates, which we resolve by ensuring that each data sample contributes exactly once to centroids. The second is in choosing mini-batch sizes, which we address by balancing premature fine-tuning of centroids with redundancy induced slow-down. Experiments show that the resulting nmbatch algorithm is very effective, often arriving within 1\% of the empirical minimum 100 times earlier than the standard mini-batch algorithm. James Newling, François Fleuret |
NIPS | 2 |
| 2016 | Jointly Informative Feature Selection Made Tractable by Gaussian ModelingabstractWe address the problem of selecting groups of jointly informative, continuous, features in the context of classification and propose several novel criteria for performing this selection. The proposed class of methods is based on combining a Gaussian modeling of the feature responses with derived bounds on and approximations to their mutual information with the class label. Furthermore, specific algorithmic implementations of these criteria are presented which reduce the computational complexity of the proposed feature selection algorithms by up to two-orders of magnitude. Consequently we show that feature selection based on the joint mutual information of features and class label is in fact tractable; this runs contrary to prior works that largely depend on marginal quantities. An empirical evaluation using several types of classifiers on multiple data sets show that this class of methods outperforms state-of-the-art baselines, both in terms of speed and classification accuracy. Leonidas Lefakis, François Fleuret |
J. Mach. Learn. Res. | 2 |
| 2016 | Adaptive relevance feedback for large-scale image retrieval
Nicolae Suditu, François Fleuret |
Multim. Tools Appl. | 2 |
| 2016 | Tracking Interacting Objects Using Intertwined FlowsabstractIn this paper, we show that tracking different kinds of interacting objects can be formulated as a network-flow mixed integer program. This is made possible by tracking all objects simultaneously using intertwined flow variables and expressing the fact that one object can appear or disappear at locations where another is in terms of linear flow constraints. Our proposed method is able to track invisible objects whose only evidence is the presence of other objects that contain them. Furthermore, our tracklet-based implementation yields real-time tracking performance. We demonstrate the power of our approach on scenes involving cars and pedestrians, bags being carried and dropped by people, and balls being passed from one player to the next in team sports. In particular, we show that by estimating jointly and globally the trajectories of different types of objects, the presence of the ones which were not initially detected based solely on image evidence can be inferred from the detections of the others. Xinchao Wang, Engin Türetken, François Fleuret, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Probability occupancy maps for occluded depth imagesabstractWe propose a novel approach to computing the probabilities of presence of multiple and potentially occluding objects in a scene from a single depth map. To this end, we use a generative model that predicts the distribution of depth images that would be produced if the probabilities of presence were known and then to optimize them so that this distribution explains observed evidence as closely as possible. This allows us to exploit very effectively the available evidence and outperform state-of-the-art methods without requiring large amounts of data, or without using the RGB signal that modern RGB-D sensors also provide. Timur M. Bagautdinov, François Fleuret, Pascal Fua |
CVPR | 2 |
| 2015 | Kullback-Leibler Proximal Variational InferenceabstractWe propose a new variational inference method based on the Kullback-Leibler (KL) proximal term. We make two contributions towards improving efficiency of variational inference. Firstly, we derive a KL proximal-point algorithm and show its equivalence to gradient descent with natural gradient in stochastic variational inference. Secondly, we use the proximal framework to derive efficient variational algorithms for non-conjugate models. We propose a splitting procedure to separate non-conjugate terms from conjugate ones. We then linearize the non-conjugate terms and show that the resulting subproblem admits a closed-form solution. Overall, our approach converts a non-conjugate model to subproblems that involve inference in well-known conjugate models. We apply our method to many models and derive generalizations for non-conjugate exponential family. Applications to real-world datasets show that our proposed algorithms are easy to implement, fast to converge, perform well, and reduce computations. Mohammad E. Khan, Pierre Baqué, François Fleuret, Pascal Fua |
NIPS | 3 |
| 2014 | LETHA: Learning from High Quality Inputs for 3D Pose Estimation in Low Quality ImagesabstractWe introduce LETHA (Learning on Easy data, Test on Hard), a new learning paradigm consisting of building strong priors from high quality training data, and combining them with discriminative machine learning to deal with low-quality test data. Our main contribution is an implementation of that concept for pose estimation. We first automatically build a 3D model of the object of interest from high-definition images, and devise from it a pose-indexed feature extraction scheme. We then train a single classifier to process these feature vectors. Given a low quality test image, we visit many hypothetical poses, extract features consistently and evaluate the response of the classifier. Since this process uses locations recorded during learning, it does not require matching points anymore. We use a boosting procedure to train this classifier common to all poses, which is able to deal with missing features, due in this context to self-occlusion. Our results demonstrate that the method combines the strengths of global image representations, discriminative even for very tiny images, and the robustness to occlusions of approaches based on local feature point descriptors. Adrián Peñate Sánchez, Francesc Moreno-Noguer, Juan Andrade-Cetto, François Fleuret |
3DV | 4 |
| 2014 | Efficient Sample Mining for Object Detection
Olivier Canévet, François Fleuret |
ACML | 2 |
| 2014 | Sample Distillation for Object Detection and Image Classification
Olivier Canévet, Leonidas Lefakis, François Fleuret |
ACML | 3 |
| 2014 | Jointly Informative Feature SelectionabstractWe propose several novel criteria for the selection of groups of jointly informative continuous features in the context of classification. Our approach is based on combining a Gaussian modeling of the feature responses, with derived upper bounds on their mutual information with the class label and their joint entropy. We further propose specific algorithmic implementations of these criteria which reduce the computational complexity of the algorithms by up to two-orders of magnitude, making these strategies tractable in practice. Experiments on multiple computer-vision data-bases, and using several types of classifiers, show that this class of methods outperforms state-of-the-art baselines, both in terms of speed and classification accuracy. Leonidas Lefakis, François Fleuret |
AISTATS | 2 |
| 2014 | Tracking Interacting Objects Optimally Using Integer Programming
Xinchao Wang, Engin Türetken, François Fleuret, Pascal Fua |
ECCV (1) | 3 |
| 2014 | Dynamic Programming Boosting for Discriminative Macro-Action DiscoveryabstractWe consider the problem of automatic macro-action discovery in imitation learning, which we cast as one of change-point detection. Unlike prior work in change-point detection, the present work leverages discriminative learning algorithms. Our main contribution is a novel supervised learning algorithm which extends the classical Boosting framework by combining it with dynamic programming. The resulting process alternatively improves the performance of individual strong predictors and the estimated change-points in the training sequence. Empirical evaluation is presented for the proposed method on tasks where change-points arise naturally as part of a classification problem. Finally we show the applicability of the algorithm to macro-action discovery in imitation learning and demonstrate it allows us to solve complex image-based goal-planning problems with thousands of features. Leonidas Lefakis, François Fleuret |
ICML | 2 |
| 2014 | Adaptive sampling for large scale boosting
Charles Dubout, François Fleuret |
J. Mach. Learn. Res. | 2 |
| 2014 | Multi-Commodity Network Flow for Tracking Multiple PeopleabstractIn this paper, we show that tracking multiple people whose paths may intersect can be formulated as a multi-commodity network flow problem. Our proposed framework is designed to exploit image appearance cues to prevent identity switches. Our method is effective even when such cues are only available at distant time intervals. This is unlike many current approaches that depend on appearance being exploitable from frame-to-frame. Furthermore, our algorithm lends itself to a real-time implementation. We validate our approach on three publicly available datasets that contain long and complex sequences, the APIDIS basketball dataset, the ISSIA soccer dataset, and the PETS'09 pedestrian dataset. We also demonstrate its performance on a newer basketball dataset that features complete world championship basketball matches. In all cases, our approach preserves identity better than state-of-the-art tracking algorithms. Horesh Ben Shitrit, Jérôme Berclaz, François Fleuret, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Deformable Part Models with Individual Part ScalingabstractCurrent deformable part models such as the ones introduced by Felzenszwalb et al. let the parts deform only at a fixed predetermined scale relative to that of the root of the models (typically at twice the resolution). They do so because it allows them to find the optimal placement of each part efficiently, using a fast 2D distance transform algorithm. We demonstrate in this paper that if one settles for approximately optimal placements, it is possible to efficiently deform the parts across scales as well. Allowing parts to move in 3D increases the expressivity of the models, allowing them to compensate for a wider class of deformations, and might approximate an increase in the scanning resolution. As the number of parameters remains (nearly) constant, overfitting is not a problem. 1 Charles Dubout, François Fleuret |
BMVC | 2 |
| 2013 | Fast Object Detection with Entropy-Driven EvaluationabstractCascade-style approaches to implementing ensemble classifiers can deliver significant speed-ups at test time. While highly effective, they remain challenging to tune and their overall performance depends on the availability of large validation sets to estimate rejection thresholds. These characteristics are often prohibitive and thus limit their applicability. We introduce an alternative approach to speeding-up classifier evaluation which overcomes these limitations. It involves maintaining a probability estimate of the class label at each intermediary response and stopping when the corresponding uncertainty becomes small enough. As a result, the evaluation terminates early based on the sequence of responses observed. Furthermore, it does so independently of the type of ensemble classifier used or the way it was trained. We show through extensive experimentation that our method provides 2 to 10 fold speed-ups, over existing state-of-the-art methods, at almost no loss in accuracy on a number of object classification tasks. Raphael Sznitman, Carlos J. Becker, François Fleuret, Pascal Fua |
CVPR | 3 |
| 2013 | Reservoir Boosting : Between Online and Offline Ensemble LearningabstractWe propose to train an ensemble with the help of a reservoir in which the learning algorithm can store a limited number of samples. This novel approach lies in the area between offline and online ensemble approaches and can be seen either as a restriction of the former or an enhancement of the latter. We identify some basic strategies that can be used to populate this reservoir and present our main contribution, dubbed Greedy Edge Expectation Maximization (GEEM), that maintains the reservoir content in the case of Boosting by viewing the samples through their projections into the weak classifier response space. We propose an efficient algorithmic implementation which makes it tractable in practice, and demonstrate its efficiency experimentally on several compute-vision data-sets, on which it outperforms both online and offline methods in a memory constrained setting. Leonidas Lefakis, François Fleuret |
NIPS | 2 |
| 2013 | treeKL: A distance between high dimension empirical distributions
Riwal Lefort, François Fleuret |
Pattern Recognit. Lett. | 2 |
| 2012 | Iterative relevance feedback with adaptive exploration/exploitation trade-offabstractContent-based image retrieval systems have to cope with two different regimes: understanding broadly the categories of interest to the user, and refining the search in this or these categories to converge to specific images among them. Here, in contrast with other types of retrieval systems, these two regimes are of great importance since the search initialization is hardly optimal (i.e. the page-zero problem) and the relevance feedback must tolerate the semantic gap of the image's visual features. Nicolae Suditu, François Fleuret |
CIKM | 2 |
| 2012 | Exact Acceleration of Linear Object Detectors
Charles Dubout, François Fleuret |
ECCV (3) | 2 |
| 2012 | A tree-based distance between distributions: Application to classification of neuronsabstractThe usual strategy for computing a distance between two distributions consists of modeling the distributions in feature space, and of computing the distance between the models. We propose here to model the distributions of points by using unsupervised trees. Our main contribution is the definition of a tree-based approximation of the Kullback-Leibler divergence for very large feature spaces, from which we derive a symmetric distance. Our tree-based KL divergence consists first of building for each set of samples a balanced tree. Then, for any pair of sets of samples, we effectively compute the KL divergence between the empirical distributions at the leaves for the set used to build the tree, and the empirical distribution at the leaves for the other set. We show experimentally on synthetic data the consistency between this quantity and the exact KL divergence, and demonstrate its efficiency for both unsupervised and supervised classification on multiple standard real-world data-sets. Our main application is the characterization of abnormal neuron development. Riwal Lefort, François Fleuret |
ICASSP | 2 |
| 2012 | Macro-action Discovery Based on Change Point Detection and BoostingabstractWe present a novel approach to automatic macroaction discovery and its application to a complex goal-planning task. The problem of macro-action discovery is framed as one of multiple change point detection and is addressed with the help of the Dynamic Programming Boosting algorithm. The procedure is then employed to solve a complex goal-planning problem which entails an avatar navigating a 3D environment. By using DPBoost to decompose the problem into a number of simpler ones, we are able to successfully address both the complexity and partial observability of the environment. Leonidas Lefakis, François Fleuret |
ICMLA (1) | 2 |
| 2012 | A Real-Time Deformable DetectorabstractWe propose a new learning strategy for object detection. The proposed scheme forgoes the need to train a collection of detectors dedicated to homogeneous families of poses, and instead learns a single classifier that has the inherent ability to deform based on the signal of interest. We train a detector with a standard AdaBoost procedure by using combinations of pose-indexed features and pose estimators. This allows the learning process to select and combine various estimates of the pose with features able to compensate for variations in pose without the need to label data for training or explore the pose space in testing. We validate our framework on three types of data: hand video sequences, aerial images of cars, and face images. We compare our method to a standard boosting framework, with access to the same ground truth, and show a reduction in the false alarm rate of up to an order of magnitude. Where possible, we compare our method to the state of the art, which requires pose annotations of the training data, and demonstrate comparable performance. Karim Ali 0002, François Fleuret, David Hasler, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | FlowBoost - Appearance learning from sparsely annotated videoabstractWe propose a new learning method which exploits temporal consistency to successfully learn a complex appearance model from a sparsely labeled training video. Our approach consists in iteratively improving an appearance-based model built with a Boosting procedure, and the reconstruction of trajectories corresponding to the motion of multiple targets. We demonstrate the efficiency of our procedure on pedestrian detection in videos and cell detection in microscopy image sequences. In both cases, our method is demonstrated to reduce the labeling requirement by one to two orders of magnitude. We show that in some instances, our method trained with sparse labels on a video sequence is able to outperform a standard learning procedure trained with the fully labeled sequence. Karim Ali 0002, David Hasler, François Fleuret |
CVPR | 3 |
| 2011 | Tasting families of features for image classificationabstractUsing multiple families of image features is a very efficient strategy to improve performance in object detection or recognition. However, such a strategy induces multiple challenges for machine learning methods, both from a computational and a statistical perspective. The main contribution of this paper is a novel feature sampling procedure dubbed "Tasting" to improve the efficiency of Boosting in such a context. Instead of sampling features in a uniform manner, Tasting continuously estimates the expected loss reduction for each family from a limited set of features sampled prior to the learning, and biases the sampling accordingly. We evaluate the performance of this procedure with tens of families of features on four image classification and object detection data-sets. We show that Tasting, which does not require the tuning of any meta-parameter, outperforms systematically variants of uniform sampling and state-of- the-art approaches based on bandit strategies. Charles Dubout, François Fleuret |
ICCV | 2 |
| 2011 | Tracking multiple people under global appearance constraintsabstractIn this paper, we show that tracking multiple people whose paths may intersect can be formulated as a convex global optimization problem. Our proposed framework is designed to exploit image appearance cues to prevent identity switches. Our method is effective even when such cues are only available at distant time intervals. This is unlike many current approaches that depend on appearance being exploitable from frame to frame. We validate our approach on three multi-camera sport and pedestrian datasets that contain long and complex sequences. Our algorithm perseveres identities better than state-of-the-art algorithms while keeping similar MOTA scores. Horesh Ben Shitrit, Jérôme Berclaz, François Fleuret, Pascal Fua |
ICCV | 3 |
| 2011 | HEAT: Iterative relevance feedback with one million imagesabstractIt has been shown repeatedly that iterative relevance feedback is a very efficient solution for content-based image retrieval. However, no existing system scales gracefully to hundreds of thousands or millions of images. We present a new approach dubbed Hierarchical and Expandable Adaptive Trace (HEAT) to tackle this problem. Our approach modulates on-the-fly the resolution of the interactive search in different parts of the image collection, by relying on a hierarchical organization of the images computed off-line. Internally, the strategy is to maintain an accurate approximation of the probabilities of relevance of the individual images while fixing an upper bound on the re- quired computation. Our system is compared on the ImageNet database to the state-of-the-art approach it extends, by conducting user evaluations on a sub-collection of 33,000 images. Its scalability is then demonstrated by conducting similar evaluations on 1,000,000 images. Nicolae Suditu, François Fleuret |
ICCV | 2 |
| 2011 | Boosting with Maximum Adaptive SamplingabstractClassical Boosting algorithms, such as AdaBoost, build a strong classifier without concern about the computational cost. Some applications, in particular in computer vision, may involve up to millions of training examples and features. In such contexts, the training time may become prohibitive. Several methods exist to accelerate training, typically either by sampling the features, or the examples, used to train the weak learners. Even if those methods can precisely quantify the speed improvement they deliver, they offer no guarantee of being more efficient than any other, given the same amount of time. This paper aims at shading some light on this problem, i.e. given a fixed amount of time, for a particular problem, which strategy is optimal in order to reduce the training loss the most. We apply this analysis to the design of new algorithms which estimate on the fly at every iteration the optimal trade-off between the number of samples and the number of features to look at in order to maximize the expected loss reduction. Experiments in object recognition with two standard computer vision data-sets show that the adaptive methods we propose outperform basic sampling and state-of-the-art bandit methods. Charles Dubout, François Fleuret |
NIPS | 2 |
| 2011 | The MASH Project
François Fleuret, Philip Abbet, Charles Dubout, Leonidas Lefakis |
ECML/PKDD (3) | 1 |
| 2011 | Multiple Object Tracking Using K-Shortest Paths OptimizationabstractMulti-object tracking can be achieved by detecting objects in individual frames and then linking detections across frames. Such an approach can be made very robust to the occasional detection failure: If an object is not detected in a frame but is in previous and following ones, a correct trajectory will nevertheless be produced. By contrast, a false-positive detection in a few frames will be ignored. However, when dealing with a multiple target problem, the linking step results in a difficult optimization problem in the space of all possible families of trajectories. This is usually dealt with by sampling or greedy search based on variants of Dynamic Programming which can easily miss the global optimum. In this paper, we show that reformulating that step as a constrained flow optimization results in a convex problem. We take advantage of its particular structure to solve it using the k-shortest paths algorithm, which is very fast. This new approach is far simpler formally and algorithmically than existing techniques and lets us demonstrate excellent performance in two very different contexts. Jérôme Berclaz, François Fleuret, Engin Türetken, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Delineating trees in noisy 2D images and 3D image-stacksabstractWe present a novel approach to fully automated delineation of tree structures in noisy 2D images and 3D image stacks. Unlike earlier methods that rely mostly on local evidence, our method builds a set of candidate trees over many different subsets of points likely to belong to the final one and then chooses the best one according to a global objective function. Since we are not systematically trying to span all nodes, our algorithm is able to eliminate noise while retaining the right tree structure. Manually annotated dendrite micrographs and retinal scans are used to evaluate the performance of our method, which is shown to be able to reject noise while retaining the tree structure. Germán González, Engin Türetken, François Fleuret, Pascal Fua |
CVPR | 3 |
| 2010 | Joint Cascade Optimization Using A Product Of Boosted ClassifiersabstractThe standard strategy for efficient object detection consists of building a cascade composed of several binary classifiers. The detection process takes the form of a lazy evaluation of the conjunction of the responses of these classifiers, and concentrates the computation on difficult parts of the image which can not be trivially rejected. We introduce a novel algorithm to construct jointly the classifiers of such a cascade. We interpret the response of a classifier as a probability of a positive prediction, and the overall response of the cascade as the probability that all the predictions are positive. From this noisy-AND model, we derive a consistent loss and a Boosting procedure to optimize that global probability on the training set. Such a joint learning allows the individual predictors to focus on a more restricted modeling problem, and improves the performance compared to a standard cascade. We demonstrate the efficiency of this approach on face and pedestrian detection with standard data-sets and comparisons with reference baselines. Leonidas Lefakis, François Fleuret |
NIPS | 2 |
| 2009 | Learning rotational features for filament detectionabstractState-of-the-art approaches for detecting filament-like structures in noisy images rely on filters optimized for signals of a particular shape, such as an ideal edge or ridge. While these approaches are optimal when the image conforms to these ideal shapes, their performance quickly degrades on many types of real data where the image deviates from the ideal model, and when noise processes violate a Gaussian assumption. In this paper, we show that by learning rotational features, we can outperform state-of-the-art filament detection techniques on many different kinds of imagery. More specifically, we demonstrate superior performance for the detection of blood vessel in retinal scans, neurons in brightfield microscopy imagery, and streets in satellite imagery. Germán González, François Fleuret, Pascal Fua |
CVPR | 2 |
| 2009 | Joint pose estimator and feature learning for object detectionabstractA new learning strategy for object detection is presented. The proposed scheme forgoes the need to train a collection of detectors dedicated to homogeneous families of poses, and instead learns a single classifier that has the inherent ability to deform based on the signal of interest. Specifically, we train a detector with a standard AdaBoost procedure by using combinations of pose-indexed features and pose estimators instead of the usual image features. This allows the learning process to select and combine various estimates of the pose with features able to implicitly compensate for variations in pose. We demonstrate that a detector built in such a manner provides noticeable gains on two hand video sequences and analyze the performance of our detector as these data sets are synthetically enriched in pose while not increased in size. Karim Ali 0002, François Fleuret, David Hasler, Pascal Fua |
ICCV | 2 |
| 2009 | Steerable Features for Statistical 3D Dendrite Detection
Germán González, François Aguet, François Fleuret, Michael Unser, Pascal Fua |
MICCAI (1) | 3 |
| 2009 | Classification-Based Probabilistic Modeling of Texture Transition for Fast Line Search Tracking and DelineationabstractWe introduce a classification-based approach to finding occluding texture boundaries. The classifier is composed of a set of weak learners which operate on image intensity discriminative features which are defined on small patches and fast to compute. A database which is designed to simulate digitized occluding contours of textured objects in natural images is used to train the weak learners. The trained classifier score is then used to obtain a probabilistic model for the presence of texture transitions which can readily be used for line search texture boundary detection in the direction normal to an initial boundary estimate. This method is fast and therefore suitable for real-time and interactive applications. It works as a robust estimator which requires a ribbon like search region and can handle complex texture structures without requiring a large number of observations. We demonstrate results both in the context of interactive 2-D delineation and fast 3-D tracking and compare its performance with other existing methods for line search boundary detection. Ali Shahrokni, Tom Drummond, François Fleuret, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Multi-layer boosting for pattern recognition
François Fleuret |
Pattern Recognit. Lett. | 1 |
| 2008 | Multi-camera Tracking and Atypical Motion Detection with Behavioral Maps
Jérôme Berclaz, François Fleuret, Pascal Fua |
ECCV (3) | 2 |
| 2008 | Automated Delineation of Dendritic Networks in Noisy Image Stacks
Germán González, François Fleuret, Pascal Fua |
ECCV (4) | 2 |
| 2008 | Multicamera People Tracking with a Probabilistic Occupancy MapabstractGiven two to four synchronized video streams taken at eye level and from different angles, we show that we can effectively combine a generative model with dynamic programming to accurately follow up to six individuals across thousands of frames in spite of significant occlusions and lighting changes. In addition, we also derive metrically accurate trajectories for each one of them. Our contribution is twofold. First, we demonstrate that our generative model can effectively handle occlusions in each time frame independently, even when the only data available comes from the output of a simple background subtraction algorithm and when the number of individuals is unknown a priori. Second, we show that multi-person tracking can be reliably achieved by processing individual trajectories separately over long sequences, provided that a reasonable heuristic is used to rank these individuals and avoid confusing them with one another. François Fleuret, Jérôme Berclaz, Richard Lengagne, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Robust Multi-View Change DetectionabstractWe present a multi-view change detection approach aimed at being robust \nwith respect to common “disturbance factors” yielding image changes in realworld \napplications. Disturbance factors causing “slow” or “fast-and-global” \nimage variations, such as light changes and dynamic adjustments of camera \nparameters (e.g. auto-exposure and auto-gain control), are dealt with by a \nproper single-view change detector run independently on each view. The \ncomputed change masks are then fused into a “synergy mask” defined into a \ncommon virtual top-view, so as to detect and filter-out “fast-and-local” image \nchanges due to physical points lying on the ground surface (e.g. shadows cast \nby moving objects and light spots hitting the ground surface). Alessandro Lanza, Luigi Di Stefano, Jérôme Berclaz, François Fleuret, Pascal Fua |
BMVC | 4 |
| 2007 | Occam's Hammer
Gilles Blanchard, François Fleuret |
COLT | 2 |
| 2006 | Robust People Tracking with Global Trajectory OptimizationabstractGiven three or four synchronized videos taken at eye level and from different angles, we show that we can effectively use dynamic programming to accurately follow up to six individuals across thousands of frames in spite of significant occlusions. In addition, we also derive metrically accurate trajectories for each one of them. Our main contribution is to show that multi-person tracking can be reliably achieved by processing individual trajectories separately over long sequences, provided that a reasonable heuristic is used to rank these individuals and avoid confusing them with one another. In this way, we achieve robustness by finding optimal trajectories over many frames while avoiding the combinatorial explosion that would result from simultaneously dealing with all the individuals. Jérôme Berclaz, François Fleuret, Pascal Fua |
CVPR (1) | 2 |
| 2006 | Feature Harvesting for Tracking-by-Detection
Mustafa Özuysal, Vincent Lepetit, François Fleuret, Pascal Fua |
ECCV (3) | 3 |
| 2005 | Classifier-based Contour Tracking for Rigid and Deformable ObjectsabstractThis paper proposes a machine learning approach to the problem of modelbased contour tracking for rigid or deformable objects. The motion of the target is calculated by tracking its contours in a video sequence. We develop a probabilistic representation of contours that allows robust contour tracking in presence of texture and clutter. We use boosting to train a predictor of the conditional probability of texture transition, given the pixel intensities. The most likely connected contours are obtained by maximizing the posterior probability of object model parameters. The probabilistic formulation allows spatial connectivity of the contours to be formulated in a natural manner. For deformable objects, we use a Hidden Markov Model to calculate the joint law of the conditional probabilities of contour points while for rigid objects geometric properties of the model are used in a framework of random sample consensus algorithm to find the optimal model pose. We demonstrate that the proposed method is fast and robust for tracking deformable and rigid objects. We also compare our algorithm to several other contour tracking methods. 1 Ali Shahrokni, François Fleuret, Pascal Fua |
BMVC | 2 |
| 2005 | The GCS Kernel for SVM-Based Image Recognition
Sabri Boughorbel, Jean-Philippe Tarel, François Fleuret, Nozha Boujemaa |
ICANN (2) | 3 |
| 2005 | Fixed Point Probability Field for Complex Occlusion HandlingabstractIn this paper, we show that in a multi-camera context, we can effectively handle occlusions in real-time at each frame independently, even when the only available data comes from the binary output of a simple blob detector, and the number of present individuals is a priori unknown. We start from occupancy probability estimates in a top view and rely on a generative model to yield probability images to be compared with the actual input images. We then refine the estimates so that the probability images match the binary input images as well as possible. We demonstrate the quality of our results on several sequences involving complex occlusions. François Fleuret, Richard Lengagne, Pascal Fua |
ICCV | 1 |
| 2005 | A Bayesian kernel for the prediction of neuron properties from binary gene profilesabstractPredicting cellular properties from molecular or genetic data is a challenge for bioinformatics and machine learning. In brain slices of neuronal tissue, it has become possible to both measure electro-physiological properties of a given neuron and to extract a sample of its cytoplasm so that expressed genes can be amplified. Thus, the presence or absence of genes related to ion channels in the neuronal cell membrane can be correlated with neuronal behavior encoded as a set of electro-physiological parameters. A typical gene amplification process is asymmetric in the sense that false positives are very rare, whereas false negatives (genes expressed but not amplified) are rather common. An analysis of a probabilistic model of that process yields a similarity measure between two strings of amplified genes that takes the asymmetry of the amplification process into account. This similarity measure can be put under the form of a conformal-transformed kernel. We provide experiments with support-vector machines on artificial and neuronal data. François Fleuret, Wulfram Gerstner |
ICMLA | 1 |
| 2005 | Pattern Recognition from One Example by ChoppingabstractWe investigate the learning of the appearance of an object from a single image of it. Instead of using a large number of pictures of the object to recognize, we use a labeled reference database of pictures of other ob- jects to learn invariance to noise and variations in pose and illumination. This acquired knowledge is then used to predict if two pictures of new objects, which do not appear on the training pictures, actually display the same object. We propose a generic scheme called chopping to address this task. It relies on hundreds of random binary splits of the training set chosen to keep together the images of any given object. Those splits are extended to the complete image space with a simple learning algorithm. Given two images, the responses of the split predictors are combined with a Bayesian rule into a posterior probability of similarity. Experiments with the COIL-100 database and with a database of 150 de- graded LATEX symbols compare our method to a classical learning with several examples of the positive class and to a direct learning of the sim- ilarity. François Fleuret, Gilles Blanchard |
NIPS | 1 |
| 2004 | Non-Mercer Kernels for SVM Object RecognitionabstractOn the one hand, Support Vector Machines have met with significant success in solving difficult pattern recognition problems with global features representation. On the other hand, local features in images have shown to be suitable representations for efficient object recognition. Therefore, it is natural to try to combine SVM approach with local features representation to gain advantages on both sides. We study in this paper the Mercer property of matching kernels which mimic classical matching algorithms used in techniques based on points of interest. We introduce a new statistical approach of kernel positiveness. We show that despite the absence of an analytical proof of the Mercer property, we can provide bounds on the probability that the Gram matrix is actually positive definite for kernels in large class of functions, under reasonable assumptions. A few experiments validate those on object recognition tasks. Sabri Boughorbel, Jean-Philippe Tarel, François Fleuret |
BMVC | 3 |
| 2004 | Fast Binary Feature Selection with Conditional Mutual Information
François Fleuret |
J. Mach. Learn. Res. | 1 |
| 2002 | Theoretical properties of functional Multi Layer Perceptrons
Fabrice Rossi, Brieuc Conan-Guez, François Fleuret |
ESANN | 3 |
| 2001 | Coarse-to-Fine Face Detection
François Fleuret, Donald Geman |
Int. J. Comput. Vis. | 1 |
| 2000 | DEA: An Architecture for Goal Planning and ClassificationabstractWe introduce the differential efficiency algorithm, which partitions a perceptive space during unsupervised learning into categories and uses them to solve goal-planning and classification problems. This algorithm is inspired by a biological model of the cortex proposing the cortical column as an elementary unit. We validate the generality of this approach by testing it on four problems with continuous time and no reinforcement signal until the goal is reached (constrained object moves, Hanoi tower problem, animat control, and simple character recognition). François Fleuret, Eric Brunet-Gouet |
Neural Comput. | 1 |