Patrick Gallinari

dblp:g/PatrickGallinari · DBLP profile ↗
← Back
195ranked-venue papers
4as first author
34since 2021 · last 2025
0000-0001-9060-9001ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 133 · 3 first-author · 31 since 2021Databases, data management, data science and information retrieval · 79 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 1 first-authorHuman-computer interaction and ubiquitous computing · 10Applied, interdisciplinary, general and emerging computing · 8Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MEXMA: Token-level objectives improve sentence representations
abstract
Cross-lingual sentence encoders (CLSE) create fixed-size sentence representations with aligned translations.Current pre-trained CLSE approaches use sentence-level objectives only.This can lead to loss of information, especially for tokens, which then degrades the sentence representation.We propose MEXMA, a novel approach that integrates both sentence-level and token-level objectives.The sentence representation in one language is used to predict masked tokens in another language, with both the sentence representation and all tokens directly updating the encoder.We show that adding token-level objectives greatly improves the sentence representation quality across several tasks.Our approach outperforms current pre-trained cross-lingual sentence encoders on bitext mining as well as several downstream tasks.We also analyse the information encoded in our tokens, and how the sentence representation is built from them.
João Maria Janeiro, Benjamin Piwowarski, Patrick Gallinari, Loïc Barrault
ACL (1)3
2025 Mixture of Languages: Improved Multilingual Encoders Through Language Grouping
abstract
João Maria Janeiro, Belen Alastruey, Francisco Massa, Maha Elbayad, Benjamin Piwowarski, Patrick Gallinari, Loic Barrault. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
João Maria Janeiro, Belen Alastruey, Francisco Massa, Maha Elbayad, Benjamin Piwowarski, Patrick Gallinari, Loïc Barrault
EMNLP6
2025 Learning a Neural Solver for Parametric PDEs to Enhance Physics-Informed Methods
abstract
Physics-informed deep learning often faces optimization challenges due to the complexity of solving partial differential equations (PDEs), which involve exploring large solution spaces, require numerous iterations, and can lead to unstable training. These challenges arise particularly from the ill-conditioning of the optimization problem, caused by the differential terms in the loss function. To address these issues, we propose learning a solver, i.e., solving PDEs using a physics-informed iterative algorithm trained on data. Our method learns to condition a gradient descent algorithm that automatically adapts to each PDE instance, significantly accelerating and stabilizing the optimization process and enabling faster convergence of physics-aware models. Furthermore, while traditional physics-informed methods solve for a single PDE instance, our approach addresses parametric PDEs. Specifically, our method integrates the physical loss gradient with the PDE parameters to solve over a distribution of PDE parameters, including coefficients, initial conditions, or boundary conditions. We demonstrate the effectiveness of our method through empirical experiments on multiple datasets, comparing training and test-time optimization performance.
Lise Le Boudec, Emmanuel de Bézenac, Louis Serrano, Ramon Daniel Regueiro-Espino, Patrick Gallinari
ICLR6
2025 SCOPE: A Self-supervised Framework for Improving Faithfulness in Conditional Text Generation
abstract
Large Language Models (LLMs), when used for conditional text generation, often produce hallucinations, i.e., information that is unfaithful or not grounded in the input context. This issue arises in typical conditional text generation tasks, such as text summarization and data-to-text generation, where the goal is to produce fluent text based on contextual input. When fine-tuned on specific domains, LLMs struggle to provide faithful answers to a given context, often adding information or generating errors. One underlying cause of this issue is that LLMs rely on statistical patterns learned from their training data. This reliance can interfere with the model's ability to stay faithful to a provided context, leading to the generation of ungrounded information. We build upon this observation and introduce a novel self-supervised method for generating a training set of unfaithful samples. We then refine the model using a training process that encourages the generation of grounded outputs over unfaithful ones, drawing on preference-based training. Our approach leads to significantly more grounded text generation, outperforming existing self-supervised techniques in faithfulness, as evaluated through automatic metrics, LLM-based assessments, and human evaluations.
Song Duong, Florian Le Bronnec, Alexandre Allauzen, Vincent Guigue, Alberto Lumbreras, Laure Soulier, Patrick Gallinari
ICLR7
2025 Zebra: In-Context Generative Pretraining for Solving Parametric PDEs
abstract
Solving time-dependent parametric partial differential equations (PDEs) is challenging for data-driven methods, as these models must adapt to variations in parameters such as coefficients, forcing terms, and initial conditions. State-of-the-art neural surrogates perform adaptation through gradient-based optimization and meta-learning to implicitly encode the variety of dynamics from observations. This often comes with increased inference complexity. Inspired by the in-context learning capabilities of large language models (LLMs), we introduce Zebra, a novel generative auto-regressive transformer designed to solve parametric PDEs without requiring gradient adaptation at inference. By leveraging in-context information during both pre-training and inference, Zebra dynamically adapts to new tasks by conditioning on input sequences that incorporate context example trajectories. As a generative model, Zebra can be used to generate new trajectories and allows quantifying the uncertainty of the predictions. We evaluate Zebra across a variety of challenging PDE scenarios, demonstrating its adaptability, robustness, and superior performance compared to existing approaches.
Louis Serrano, Armand Kassaï Koupaï, Thomas X. Wang, Pierre Erbacher, Patrick Gallinari
ICML5
2025 ENMA: Tokenwise Autoregression for Continuous Neural PDE Operators
abstract
Solving time-dependent parametric partial differential equations (PDEs) remains a fundamental challenge for neural solvers, particularly when generalizing across a wide range of physical parameters and dynamics. When data is uncertain or incomplete—as is often the case—a natural approach is to turn to generative models. We introduce ENMA, a generative neural operator designed to model spatio-temporal dynamics arising from physical phenomena. ENMA predicts future dynamics in a compressed latent space using a generative masked autoregressive transformer trained with flow matching loss, enabling tokenwise generation. Irregularly sampled spatial observations are encoded into uniform latent representations via attention mechanisms and further compressed through a spatio-temporal convolutional encoder. This allows ENMA to perform in-context learning at inference time by conditioning on either past states of the target trajectory or auxiliary context trajectories with similar dynamics. The result is a robust and adaptable framework that generalizes to new PDE regimes and supports one-shot surrogate modeling of time-dependent parametric PDEs.
Armand Kassaï Koupaï, Lise Le Boudec, Louis Serrano, Patrick Gallinari
NeurIPS4
2024 LOCOST: State-Space Models for Long Document Abstractive Summarization
abstract
Florian Le Bronnec, Song Duong, Mathieu Ravaut, Alexandre Allauzen, Nancy Chen, Vincent Guigue, Alberto Lumbreras, Laure Soulier, Patrick Gallinari. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Florian Le Bronnec, Song Duong, Mathieu Ravaut, Alexandre Allauzen, Nancy F. Chen, Vincent Guigue, Alberto Lumbreras, Laure Soulier, Patrick Gallinari
EACL (1)9
2024 Boosting Generalization in Parametric PDE Neural Solvers through Adaptive Conditioning
abstract
Solving parametric partial differential equations (PDEs) presents significant challenges for data-driven methods due to the sensitivity of spatio-temporal dynamics to variations in PDE parameters. Machine learning approaches often struggle to capture this variability. To address this, data-driven approaches learn parametric PDEs by sampling a very large variety of trajectories with varying PDE parameters. We first show that incorporating conditioning mechanisms for learning parametric PDEs is essential and that among them, \textit{adaptive conditioning}, allows stronger generalization. As existing adaptive conditioning methods do not scale well with respect to the number of parameters to adapt in the neural solver, we propose GEPS, a simple adaptation mechanism to boost GEneralization in Pde Solvers via a first-order optimization and low-rank rapid adaptation of a small set of context parameters. We demonstrate the versatility of our approach for both fully data-driven and for physics-aware neural solvers. Validation performed on a whole range of spatio-temporal forecasting problems demonstrates excellent performance for generalizing to unseen conditions including initial conditions, PDE coefficients, forcing terms and solution domain. *Project page*: https://geps-project.github.io
Armand Kassaï Koupaï, Jorge Mifsut Benet, Jean-Noël Vittaut, Patrick Gallinari
NeurIPS5
2024 AROMA: Preserving Spatial Structure for Latent PDE Modeling with Local Neural Fields
abstract
We present AROMA (Attentive Reduced Order Model with Attention), a framework designed to enhance the modeling of partial differential equations (PDEs) using local neural fields. Our flexible encoder-decoder architecture can obtain smooth latent representations of spatial physical fields from a variety of data types, including irregular-grid inputs and point clouds. This versatility eliminates the need for patching and allows efficient processing of diverse geometries. The sequential nature of our latent representation can be interpreted spatially and permits the use of a conditional transformer for modeling the temporal dynamics of PDEs. By employing a diffusion-based formulation, we achieve greater stability and enable longer rollouts compared to conventional MSE training. AROMA's superior performance in simulating 1D and 2D equations underscores the efficacy of our approach in capturing complex dynamical behaviors.
Louis Serrano, Thomas X. Wang, Etienne Le Naour, Jean-Noël Vittaut, Patrick Gallinari
NeurIPS5
2023 Learning from Multiple Sources for Data-to-Text and Text-to-Data
abstract
Data-to-text (D2T) and text-to-data (T2D) are dual tasks that convert structured data, such as graphs or tables into fluent text, and vice versa. These tasks are usually handled separately and use corpora extracted from a single source. Current systems leverage pre-trained language models fine-tuned on D2T or T2D tasks. This approach has two main limitations: first, a separate system has to be tuned for each task and source; second, learning is limited by the scarcity of available corpora. This paper considers a more general scenario where data are available from multiple heterogeneous sources. Each source, with its specific data format and semantic domain, provides a non-parallel corpus of text and structured data. We introduce a variational auto-encoder model with disentangled style and content variables that allows us to represent the diversity that stems from multiple sources of text and data. Our model is designed to handle the tasks of D2T and T2D jointly. We evaluate our model on several datasets, and show that by learning from multiple sources, our model closes the performance gap with its supervised single-source counterpart and outperforms it in some cases.
Song Duong, Alberto Lumbreras, Mike Gartrell, Patrick Gallinari
AISTATS4
2023 Continuous PDE Dynamics Forecasting with Implicit Neural Representations
Matthieu Kirchmeyer, Jean-Yves Franceschi, Alain Rakotomamonjy, Patrick Gallinari
ICLR5
2023 Enhancing factualness and controllability of Data-to-Text Generation via data Views and constraints
abstract
Neural data-to-text systems lack the control and factual accuracy required to generate useful and insightful summaries of multidimensional data.We propose a solution in the form of data views, where each view describes an entity and its attributes along specific dimensions.A sequence of views can then be used as a high-level schema for document planning, with the neural model handling the complexities of micro-planning and surface realization.We show that our view-based system retains factual accuracy while offering high-level control of output that can be tailored based on user preference or other norms within the domain.
Craig Thomson, Clément Rebuffel, Ehud Reiter, Laure Soulier, Somayajulu Sripada, Patrick Gallinari
INLG6
2023 Module-wise Training of Neural Networks via the Minimizing Movement Scheme
abstract
Greedy layer-wise or module-wise training of neural networks is compelling in constrained and on-device settings where memory is limited, as it circumvents a number of problems of end-to-end back-propagation. However, it suffers from a stagnation problem, whereby early layers overfit and deeper layers stop increasing the test accuracy after a certain depth. We propose to solve this issue by introducing a simple module-wise regularization inspired by the minimizing movement scheme for gradient flows in distribution space. We call the method TRGL for Transport Regularized Greedy Learning and study it theoretically, proving that it leads to greedy modules that are regular and that progressively solve the task. Experimentally, we show improved accuracy of module-wise training of various architectures such as ResNets, Transformers and VGG, when our regularization is added, superior to that of other module-wise training methods and often to end-to-end training, with as much as 60% less memory usage.
Skander Karkar, Ibrahim Ayed, Emmanuel de Bézenac, Patrick Gallinari
NeurIPS4
2023 Operator Learning with Neural Fields: Tackling PDEs on General Geometries
abstract
Machine learning approaches for solving partial differential equations require learning mappings between function spaces. While convolutional or graph neural networks are constrained to discretized functions, neural operators present a promising milestone toward mapping functions directly. Despite impressive results they still face challenges with respect to the domain geometry and typically rely on some form of discretization. In order to alleviate such limitations, we present CORAL, a new method that leverages coordinate-based networks for solving PDEs on general geometries. CORAL is designed to remove constraints on the input mesh, making it applicable to any spatial sampling and geometry. Its ability extends to diverse problem domains, including PDE solving, spatio-temporal forecasting, and inverse problems like geometric design. CORAL demonstrates robust performance across multiple resolutions and performs well in both convex and non-convex domains, surpassing or performing on par with state-of-the-art models.
Louis Serrano, Lise Le Boudec, Armand Kassaï Koupaï, Thomas X. Wang, Jean-Noël Vittaut, Patrick Gallinari
NeurIPS7
2023 Adversarial Sample Detection Through Neural Network Transport Dynamics
Skander Karkar, Patrick Gallinari, Alain Rakotomamonjy
ECML/PKDD (1)2
2022 Constrained Physical-Statistics Models for Dynamical System Identification and Prediction
Jérémie Donà, Marie Déchelle, Patrick Gallinari, Marina Levy
ICLR3
2022 Mapping conditional distributions for domain adaptation under generalized target shift
Matthieu Kirchmeyer, Alain Rakotomamonjy, Emmanuel de Bézenac, Patrick Gallinari
ICLR4
2022 A Neural Tangent Kernel Perspective of GANs
abstract
We propose a novel theoretical framework of analysis for Generative Adversarial Networks (GANs). We reveal a fundamental flaw of previous analyses which, by incorrectly modeling GANs’ training scheme, are subject to ill-defined discriminator gradients. We overcome this issue which impedes a principled study of GAN training, solving it within our framework by taking into account the discriminator’s architecture. To this end, we leverage the theory of infinite-width neural networks for the discriminator via its Neural Tangent Kernel. We characterize the trained discriminator for a wide range of losses and establish general differentiability properties of the network. From this, we derive new insights about the convergence of the generated distribution, advancing our understanding of GANs’ training dynamics. We empirically corroborate these results via an analysis toolkit based on our framework, unveiling intuitions that are consistent with GAN practice.
Jean-Yves Franceschi, Emmanuel de Bézenac, Ibrahim Ayed, Mickaël Chen, Sylvain Lamprier, Patrick Gallinari
ICML6
2022 Generalizing to New Physical Systems via Context-Informed Dynamics Model
abstract
Data-driven approaches to modeling physical systems fail to generalize to unseen systems that share the same general dynamics with the learning domain, but correspond to different physical contexts. We propose a new framework for this key problem, context-informed dynamics adaptation (CoDA), which takes into account the distributional shift across systems for fast and efficient adaptation to new dynamics. CoDA leverages multiple environments, each associated to a different dynamic, and learns to condition the dynamics model on contextual parameters, specific to each environment. The conditioning is performed via a hypernetwork, learned jointly with a context vector from observed data. The proposed formulation constrains the search hypothesis space for fast adaptation and better generalization across environments with few samples. We theoretically motivate our approach and show state-of-the-art generalization results on a set of nonlinear dynamics, representative of a variety of application domains. We also show, on these systems, that new system parameters can be inferred from context vectors with minimal supervision.
Matthieu Kirchmeyer, Jérémie Donà, Nicolas Baskiotis, Alain Rakotomamonjy, Patrick Gallinari
ICML6
2022 AirfRANS: High Fidelity Computational Fluid Dynamics Dataset for Approximating Reynolds-Averaged Navier-Stokes Solutions
abstract
Surrogate models are necessary to optimize meaningful quantities in physical dynamics as their recursive numerical resolutions are often prohibitively expensive. It is mainly the case for fluid dynamics and the resolution of Navier–Stokes equations. However, despite the fast-growing field of data-driven models for physical systems, reference datasets representing real-world phenomena are lacking. In this work, we develop \textsc{AirfRANS}, a dataset for studying the two-dimensional incompressible steady-state Reynolds-Averaged Navier–Stokes equations over airfoils at a subsonic regime and for different angles of attacks. We also introduce metrics on the stress forces at the surface of geometries and visualization of boundary layers to assess the capabilities of models to accurately predict the meaningful information of the problem. Finally, we propose deep learning baselines on four machine learning tasks to study \textsc{AirfRANS} under different constraints for generalization considerations: big and scarce data regime, Reynolds number, and angle of attack extrapolation.
Florent Bonnet, Jocelyn Ahmed Mazari, Paola Cinnella, Patrick Gallinari
NeurIPS4
2022 Diverse Weight Averaging for Out-of-Distribution Generalization
abstract
Standard neural networks struggle to generalize under distribution shifts in computer vision. Fortunately, combining multiple networks can consistently improve out-of-distribution generalization. In particular, weight averaging (WA) strategies were shown to perform best on the competitive DomainBed benchmark; they directly average the weights of multiple networks despite their nonlinearities. In this paper, we propose Diverse Weight Averaging (DiWA), a new WA strategy whose main motivation is to increase the functional diversity across averaged models. To this end, DiWA averages weights obtained from several independent training runs: indeed, models obtained from different runs are more diverse than those collected along a single run thanks to differences in hyperparameters and training procedures. We motivate the need for diversity by a new bias-variance-covariance-locality decomposition of the expected error, exploiting similarities between WA and standard functional ensembling. Moreover, this decomposition highlights that WA succeeds when the variance term dominates, which we show occurs when the marginal distribution changes at test time. Experimentally, DiWA consistently improves the state of the art on DomainBed without inference overhead.
Alexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy, Patrick Gallinari, Matthieu Cord
NeurIPS5
2022 Controlling hallucinations at word level in data-to-text generation
abstract
Abstract Data-to-Text Generation (DTG) is a subfield of Natural Language Generation aiming at transcribing structured data in natural language descriptions. The field has been recently boosted by the use of neural-based generators which exhibit on one side great syntactic skills without the need of hand-crafted pipelines; on the other side, the quality of the generated text reflects the quality of the training data, which in realistic settings only offer imperfectly aligned structure-text pairs. Consequently, state-of-art neural models include misleading statements –usually called hallucinations—in their outputs. The control of this phenomenon is today a major challenge for DTG, and is the problem addressed in the paper. Previous work deal with this issue at the instance level: using an alignment score for each table-reference pair. In contrast, we propose a finer-grained approach, arguing that hallucinations should rather be treated at the word level. Specifically, we propose a Multi-Branch Decoder which is able to leverage word-level labels to learn the relevant parts of each training instance. These labels are obtained following a simple and efficient scoring procedure based on co-occurrence analysis and dependency parsing. Extensive evaluations, via automated metrics and human judgment on the standard WikiBio benchmark, show the accuracy of our alignment labels and the effectiveness of the proposed Multi-Branch Decoder. Our model is able to reduce and control hallucinations, while keeping fluency and coherence in generated texts. Further experiments on a degraded version of ToTTo show that our model could be successfully used on very noisy settings.
Clément Rebuffel, Marco Roberti, Laure Soulier, Geoffrey Scoutheeten, Rossella Cancelliere, Patrick Gallinari
Data Min. Knowl. Discov.6
2022 Modelling spatiotemporal dynamics from Earth observation data with neural differential equations
Ibrahim Ayed, Emmanuel de Bézenac, Arthur Pajot, Patrick Gallinari
Mach. Learn.4
2021 Data-QuestEval: A Referenceless Metric for Data-to-Text Semantic Evaluation
abstract
Clement Rebuffel, Thomas Scialom, Laure Soulier, Benjamin Piwowarski, Sylvain Lamprier, Jacopo Staiano, Geoffrey Scoutheeten, Patrick Gallinari. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Clément Rebuffel, Thomas Scialom, Laure Soulier, Benjamin Piwowarski, Sylvain Lamprier, Jacopo Staiano, Geoffrey Scoutheeten, Patrick Gallinari
EMNLP (1)8
2021 QuestEval: Summarization Asks for Fact-based Evaluation
abstract
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, Patrick Gallinari. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Patrick Gallinari
EMNLP (1)7
2021 Separating Retention from Extraction in the Evaluation of End-to-end Relation Extraction
abstract
State-of-the-art NLP models can adopt shallow heuristics that limit their generalization capability (McCoy et al., 2019).Such heuristics include lexical overlap with the training set in Named-Entity Recognition (Taillé et al., 2020a) and Event or Type heuristics in Relation Extraction (Rosenman et al., 2020).In the more realistic end-to-end RE setting, we can expect yet another heuristic: the mere retention of training relation triples.In this paper we propose several experiments confirming that retention of known facts is a key factor of performance on standard benchmarks.Furthermore, one experiment suggests that a pipeline model able to use intermediate type representations is less prone to over-rely on retention.
Bruno Taillé, Vincent Guigue, Geoffrey Scoutheeten, Patrick Gallinari
EMNLP (1)4
2021 PDE-Driven Spatiotemporal Disentanglement
Jérémie Donà, Jean-Yves Franceschi, Sylvain Lamprier, Patrick Gallinari
ICLR4
2021 Augmenting Physical Models with Deep Networks for Complex Dynamics Forecasting
Vincent Le Guen, Jérémie Donà, Emmanuel de Bézenac, Ibrahim Ayed, Nicolas Thome, Patrick Gallinari
ICLR7
2021 Stochastic sparse adversarial attacks
abstract
This paper introduces stochastic sparse adversarial attacks (SSAA), standing as simple, fast and purely noise-based targeted and untargeted attacks of neural network classifiers (NNC). SSAA offer new examples of sparse (or L0) attacks for which only few methods have been proposed previously. These attacks are devised by exploiting a small-time expansion idea widely used for Markov processes. Experiments on small and large datasets (CIFAR-10 and ImageNet) illustrate several advantages of SSAA in comparison with the-state-of-the-art methods. For instance, in the untargeted case, our method called Voting Folded Gaussian Attack (VFGA) scales efficiently to ImageNet and achieves a significantly lower L0score than SparseFool (up to $\frac{2}{5}$) while being faster. Moreover, VFGA achieves better L0scores on ImageNet than Sparse-RS when both attacks are fully successful on a large number of samples.
Manon Césaire, Lucas Schott, Hatem Hajri, Sylvain Lamprier, Patrick Gallinari
ICTAI5
2021 Exemplars and Counterexemplars Explanations for Image Classifiers, Targeting Skin Lesion Labeling
abstract
Explainable AI consists in developing mechanisms allowing for an interaction between decision systems and humans by making the decisions of the formers understandable. This is particularly important in sensitive contexts like in the medical domain. We propose a use case study, for skin lesion diagnosis, illustrating how it is possible to provide the practitioner with explanations on the decisions of a state of the art deep neural network classifier trained to characterize skin lesions from examples. Our framework consists of a trained classifier onto which an explanation module operates. The latter is able to offer the practitioner exemplars and counterexemplars for the classification diagnosis thus allowing the physician to interact with the automatic diagnosis system. The exemplars are generated via an adversarial autoencoder. We illustrate the behavior of the system on representative examples.
Carlo Metta, Riccardo Guidotti, Patrick Gallinari, Salvatore Rinzivillo
ISCC4
2021 LEADS: Learning Dynamical Systems that Generalize Across Environments
abstract
When modeling dynamical systems from real-world data samples, the distribution of data often changes according to the environment in which they are captured, and the dynamics of the system itself vary from one environment to another. Generalizing across environments thus challenges the conventional frameworks. The classical settings suggest either considering data as i.i.d and learning a single model to cover all situations or learning environment-specific models. Both are sub-optimal: the former disregards the discrepancies between environments leading to biased solutions, while the latter does not exploit their potential commonalities and is prone to scarcity problems. We propose LEADS, a novel framework that leverages the commonalities and discrepancies among known environments to improve model generalization. This is achieved with a tailored training formulation aiming at capturing common dynamics within a shared model while additional terms capture environment-specific dynamics. We ground our approach in theory, exhibiting a decrease in sample complexity w.r.t classical alternatives. We show how theory and practice coincides on the simplified case of linear dynamics. Moreover, we instantiate this framework for neural networks and evaluate it experimentally on representative families of nonlinear dynamics. We show that this new setting can exploit knowledge extracted from environment-dependent data and improves generalization for both known and novel environments.
Ibrahim Ayed, Emmanuel de Bézenac, Nicolas Baskiotis, Patrick Gallinari
NeurIPS5
2021 CycleGAN Through the Lens of (Dynamical) Optimal Transport
Emmanuel de Bézenac, Ibrahim Ayed, Patrick Gallinari
ECML/PKDD (2)3
2021 Differentiable Feature Selection, A Reparameterization Approach
Jérémie Donà, Patrick Gallinari
ECML/PKDD (3)2
2021 Unsupervised domain adaptation with non-stochastic missing data
Matthieu Kirchmeyer, Patrick Gallinari, Alain Rakotomamonjy, Amin Mantrach
Data Min. Knowl. Discov.2
2020 A Hierarchical Model for Data-to-Text Generation
Clément Rebuffel, Laure Soulier, Geoffrey Scoutheeten, Patrick Gallinari
ECIR (1)4
2020 Contextualized Embeddings in Named-Entity Recognition: An Empirical Study on Generalization
abstract
Contextualized embeddings use unsupervised language model pretraining to compute word representations depending on their context. This is intuitively useful for generalization, especially in Named-Entity Recognition where it is crucial to detect mentions never seen during training. However, standard English benchmarks overestimate the importance of lexical over contextual features because of an unrealistic lexical overlap between train and test mentions. In this paper, we perform an empirical analysis of the generalization capabilities of state-of-the-art contextualized embeddings by separating mentions by novelty and with out-of-domain evaluation. We show that they are particularly beneficial for unseen mentions detection, especially out-of-domain. For models trained on CoNLL03, language model contextualization leads to a +1.2% maximal relative micro-F1 score increase in-domain against +13% out-of-domain on the WNUT dataset (The code is available at https://github.com/btaille/contener ).
Bruno Taillé, Vincent Guigue, Patrick Gallinari
ECIR (2)3
2020 Let's Stop Incorrect Comparisons in End-to-end Relation Extraction!
abstract
Despite efforts to distinguish three different evaluation setups (Bekoulis et al., 2018), numerous end-to-end Relation Extraction (RE) articles present unreliable performance comparison to previous work. In this paper, we first identify several patterns of invalid comparisons in published papers and describe them to avoid their propagation. We then propose a small empirical study to quantify the impact of the most common mistake and evaluate it leads to overestimating the final RE performance by around 5% on ACE05. We also seize this opportunity to study the unexplored ablations of two recent developments: the use of language model pretraining (specifically BERT) and span-level NER. This meta-analysis emphasizes the need for rigor in the report of both the evaluation setting and the datasets statistics and we call for unifying the evaluation setting in end-to-end RE.
Bruno Taillé, Vincent Guigue, Geoffrey Scoutheeten, Patrick Gallinari
EMNLP (1)4
2020 Resume: A Robust Framework for Professional Profile Learning & Evaluation
Clara Gainon de Forsan de Gabriac, Constance Scherer, Amina Djelloul, Vincent Guigue, Patrick Gallinari
ESANN5
2020 Learning the Spatio-Temporal Dynamics of Physical Processes from Partial Observations
abstract
We consider the problem of automatically learning the dynamics of physical processes evolving in space and time from incomplete observations. This is a central problem in many fields that remains complicated for large observation spaces and complex dynamics. We propose a data-driven framework, where the system's dynamics are modeled by an unknown time-varying differential equation and the evolution term for the state is estimated from the partially observed data only, using a deep convolutional neural network. Our method yields improvements w.r.t. state-of-the-art recurrent deep network models for the forecast of observations corresponding to complex fluid dynamics. We analyze the latent state representations learned by this model and propose two settings that help interpret the learned states. The model is evaluated on the incompressible Navier Stokes equations.
Ibrahim Ayed, Emmanuel de Bézenac, Arthur Pajot, Patrick Gallinari
ICASSP4
2020 Deep-SST-Eddies: A Deep Learning Framework to Detect Oceanic Eddies in Sea Surface Temperature Images
abstract
Until now, mesoscale oceanic eddies have been automatically detected through physical methods on satellite altimetry. Nevertheless, they often have a visible signature on Sea Surface Temperature (SST) satellite images, which have not been yet sufficiently exploited. We introduce a novel method that employs Deep Learning to detect eddy signatures on such input. We provide the first available dataset for this task, retaining SST images through altimetric-based region proposal. We train a CNN-based classifier which succeeds in accurately detecting eddy signatures in well-defined examples. Our experiments show that the difficulty of classifying a large set of automatically retained images can be tackled by training on a smaller subset of manually labeled data. The difference in performance on the two sets is explained by the noisy automatic labeling and intrinsic complexity of the SST signal. This approach can provide to oceanographers a tool for validation of altimetric eddy detection through SST.
Evangelos Moschos, Olivier Schwander, Alexandre Stegner, Patrick Gallinari
ICASSP4
2020 Stochastic Latent Residual Video Prediction
abstract
Designing video prediction models that account for the inherent uncertainty of the future is challenging. Most works in the literature are based on stochastic image-autoregressive recurrent networks, which raises several performance and applicability issues. An alternative is to use fully latent temporal models which untie frame synthesis and temporal dynamics. However, no such model for stochastic video prediction has been proposed in the literature yet, due to design and training difficulties. In this paper, we overcome these difficulties by introducing a novel stochastic temporal model whose dynamics are governed in a latent space by a residual update rule. This first-order scheme is motivated by discretization schemes of differential equations. It naturally models video dynamics as it allows our simpler, more interpretable, latent model to outperform prior state-of-the-art methods on challenging datasets.
Jean-Yves Franceschi, Edouard Delasalles, Mickaël Chen, Sylvain Lamprier, Patrick Gallinari
ICML5
2020 PARENTing via Model-Agnostic Reinforcement Learning to Correct Pathological Behaviors in Data-to-Text Generation
abstract
In language generation models conditioned by structured data, the classical training via maximum likelihood almost always leads models to pick up on dataset divergence (i.e., hallucinations or omissions), and to incorporate them erroneously in their own generations at inference.In this work, we build ontop of previous Reinforcement Learning based approaches and show that a model-agnostic framework relying on the recently introduced PARENT metric is efficient at reducing both hallucinations and omissions.Evaluations on the widely used WikiBIO and WebNLG benchmarks demonstrate the effectiveness of this framework compared to state-of-the-art models.
Clément Rebuffel, Laure Soulier, Geoffrey Scoutheeten, Patrick Gallinari
INLG4
2020 What BERT Sees: Cross-Modal Transfer for Visual Question Generation
abstract
Pre-trained language models have recently contributed to significant advances in NLP tasks.Recently, multi-modal versions of BERT have been developed, using heavy pretraining relying on vast corpora of aligned textual and image data, primarily applied to classification tasks such as VQA.In this paper, we are interested in evaluating the visual capabilities of BERT out-of-the-box, by avoiding pre-training made on supplementary data.We choose to study Visual Question Generation, a task of great interest for grounded dialog, that enables to study the impact of each modality (as input can be visual and/or textual).Moreover, the generation aspect of the task requires an adaptation since BERT is primarily designed as an encoder.We introduce BERT-gen, a BERT-based architecture for text generation, able to leverage on either monoor multi-modal representations.The results reported under different configurations indicate an innate capacity for BERT-gen to adapt to multi-modal data and text generation, even with few data available, avoiding expensive pre-training.The proposed model obtains substantial improvements over the state-of-the-art on two established VQG datasets.txt1
Thomas Scialom, Patrick Bordes, Paul-Alexis Dray, Jacopo Staiano, Patrick Gallinari
INLG5
2020 Normalizing Kalman Filters for Multivariate Time Series Analysis
abstract
This paper tackles the modelling of large, complex and multivariate time series panels in a probabilistic setting. To this extent, we present a novel approach reconciling classical state space models with deep learning methods. By augmenting state space models with normalizing flows, we mitigate imprecisions stemming from idealized assumptions in state space models. The resulting model is highly flexible while still retaining many of the attractive properties of state space models, e.g., uncertainty and observation errors are properly accounted for, inference is tractable, sampling is efficient, good generalization performance is observed, even in low data regimes. We demonstrate competitiveness against state-of-the-art deep learning methods on the tasks of forecasting real world data and handling varying levels of missing data.
Emmanuel de Bézenac, Syama Sundar Rangapuram, Konstantinos Benidis, Michael Bohlke-Schneider, Richard Kurle, Lorenzo Stella, Hilaf Hasson, Patrick Gallinari, Tim Januschowski
NeurIPS8
2020 A Principle of Least Action for the Training of Neural Networks
Skander Karkar, Ibrahim Ayed, Emmanuel de Bézenac, Patrick Gallinari
ECML/PKDD (2)4
2020 Interpretable time series kernel analytics by pre-image estimation
Thi Phuong Thao Tran, Ahlame Douzal Chouakria, Saeed Varasteh Yazdi, Paul Honeine, Patrick Gallinari
Artif. Intell.5
2020 Knowledge graph entity typing via learning connecting embeddings
Yu Zhao 0019, Anxiang Zhang, Huali Feng, Qing Li 0005, Patrick Gallinari, Fuji Ren
Knowl. Based Syst.5
2019 Incorporating Visual Semantics into Sentence Representations within a Grounded Space
abstract
Patrick Bordes, Eloi Zablocki, Laure Soulier, Benjamin Piwowarski, Patrick Gallinari. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Patrick Bordes, Eloi Zablocki, Laure Soulier, Benjamin Piwowarski, Patrick Gallinari
EMNLP/IJCNLP (1)5
2019 Unsupervised Adversarial Image Reconstruction
Arthur Pajot, Emmanuel de Bézenac, Patrick Gallinari
ICLR (Poster)3
2019 Context-Aware Zero-Shot Learning for Object Recognition
abstract
Zero-Shot Learning (ZSL) aims at classifying unlabeled objects by leveraging auxiliary knowledge, such as semantic representations. A limitation of previous approaches is that only intrinsic properties of objects, e.g. their visual appearance, are taken into account while their context, e.g. the surrounding objects in the image, is ignored. Following the intuitive principle that objects tend to be found in certain contexts but not others, we propose a new and challenging approach, context-aware ZSL, that leverages semantic representations in a new way to model the conditional likelihood of an object to appear in a given context. Finally, through extensive experiments conducted on Visual Genome, we show that contextual information can substantially improve the standard ZSL approach and is robust to unbalanced classes.
Eloi Zablocki, Patrick Bordes, Laure Soulier, Benjamin Piwowarski, Patrick Gallinari
ICML5
2019 Copy Mechanism and Tailored Training for Character-Based Data-to-Text Generation
Marco Roberti, Giovanni Bonetta, Rossella Cancelliere, Patrick Gallinari
ECML/PKDD (2)4
2019 Contextual bandits with hidden contexts: a focused data capture from social media streams
Sylvain Lamprier, Thibault Gisselbrecht, Patrick Gallinari
Data Min. Knowl. Discov.3
2019 Real-time detection of driver distraction: random projections for pseudo-inversion-based neural training
Marco Botta, Rossella Cancelliere, Leo Ghignone, Fabio Tango, Patrick Gallinari, Clara Luison
Knowl. Inf. Syst.5
2019 Spatio-temporal neural networks for space-time data modeling and relation discovery
Edouard Delasalles, Ali Ziat, Ludovic Denoyer, Patrick Gallinari
Knowl. Inf. Syst.4
2018 Learning Multi-Modal Word Representation Grounded in Visual Context
abstract
Representing the semantics of words is a long-standing problem for the natural language processing community. Most methods compute word semantics given their textual context in large corpora. More recently, researchers attempted to integrate perceptual and visual features. Most of these works consider the visual appearance of objects to enhance word representations but they ignore the visual environment and context in which objects appear. We propose to unify text-based techniques with vision-based techniques by simultaneously leveraging textual and visual context to learn multimodal word embeddings. We explore various choices for what can serve as a visual context and present an end-to-end method to integrate visual context elements in a multimodal skip-gram model. We provide experiments and extensive analysis of the obtained results.
Eloi Zablocki, Benjamin Piwowarski, Laure Soulier, Patrick Gallinari
AAAI4
2018 Regularize and explicit collaborative filtering with textual attention
Charles-Emmanuel Dias, Vincent Guigue, Patrick Gallinari
ESANN3
2018 Deep Learning for Physical Processes: Incorporating Prior Scientific Knowledge
Emmanuel de Bézenac, Arthur Pajot, Patrick Gallinari
ICLR (Poster)3
2018 Time Warp Invariant Dictionary Learning for Time Series Clustering: Application to Music Data Stream Analysis
Saeed Varasteh Yazdi, Ahlame Douzal Chouakria, Patrick Gallinari, Manuel Moussallam
ECML/PKDD (1)3
2018 A novel non-Gaussian embedding based model for recommender systems
Sheng Gao 0001, Qinjie Lyu, Jun Guo 0002, Patrick Gallinari
Neurocomputing5
2018 Anomaly detection in smart card logs and distant evaluation with Twitter: a robust framework
Emeric Tonnelier, Nicolas Baskiotis, Vincent Guigue, Patrick Gallinari
Neurocomputing4
2018 Profile-Based Bandit with Unknown Profiles
abstract
Stochastic bandits have been widely studied since decades. A very large panel of settings have been introduced, some of them for the inclusion of some structure between actions. If actions are associated with feature vectors that underlie their usefulness, the discovery of a mapping parameter between such profiles and rewards can help the exploration process of the bandit strategies. This is the setting studied in this paper, but in our case the action profiles (constant feature vectors) are unknown beforehand. Instead, the agent is only given sample vectors, with mean centered on the true profiles, for a subset of actions at each step of the process. In this new bandit instance, policies have thus to deal with a doubled uncertainty, both on the profile estimators and the reward mapping parameters learned so far. We propose a new algorithm, called \textit{SampLinUCB}, specifically designed for this case. Theoretical convergence guarantees are given for this strategy, according to various profile samples delivery scenarios. Finally, experiments are conducted on both artificial data and a task of focused data capture from online social networks. Obtained results demonstrate the relevance of the approach in various settings.
Sylvain Lamprier, Thibault Gisselbrecht, Patrick Gallinari
J. Mach. Learn. Res.3
2018 A distributed Frank-Wolfe framework for learning low-rank matrices with the trace norm
Wenjie Zheng 0001, Aurélien Bellet, Patrick Gallinari
Mach. Learn.3
2018 Representation Learning for Classification in Heterogeneous Graphs with Application to Social Networks
abstract
We address the task of node classification in heterogeneous networks, where the nodes are of different types, each type having its own set of labels, and the relations between nodes may also be of different types. A typical example is provided by social networks where node types may for example be users, content, or films, and relations friendship , like , authorship . Learning and performing inference on such heterogeneous networks is a recent task requiring new models and algorithms. We propose a model, Labeling Heterogeneous Network (LaHNet) , a transductive approach to classification that learns to project the different types of nodes into a common latent space. This embedding is learned so as to reflect different characteristics of the problem such as the correlation between node labels, as well as the graph topology. The application focus is on social graphs, but the algorithm is general and can be used for other domains. The model is evaluated on five datasets representative of different instances of social data.
Ludovic Dos Santos, Benjamin Piwowarski, Ludovic Denoyer, Patrick Gallinari
ACM Trans. Knowl. Discov. Data4
2017 Anomaly detection and characterization in smart card logs using NMF and Tweets
Emeric Tonnelier, Nicolas Baskiotis, Vincent Guigue, Patrick Gallinari
ESANN4
2017 Spatio-Temporal Neural Networks for Space-Time Series Forecasting and Relations Discovery
abstract
We introduce a dynamical spatio-temporal model formalized as a recurrent neural network for forecasting time series of spatial processes, i.e. series of observations sharing temporal and spatial dependencies. The model learns these dependencies through a structured latent dynamical component, while a decoder predicts the observations from the latent representations. We consider several variants of this model, corresponding to different prior hypothesis about the spatial relations between the series. The model is evaluated and compared to state-of-the-art baselines, on a variety of forecasting problems representative of different application areas: epidemiology, geo-spatial statistics and car-traffic prediction. Besides these evaluations, we also describe experiments showing the ability of this approach to extract relevant spatial relations.
Ali Ziat, Edouard Delasalles, Ludovic Denoyer, Patrick Gallinari
ICDM4
2017 Variational Thompson Sampling for Relational Recurrent Bandits
Sylvain Lamprier, Thibault Gisselbrecht, Patrick Gallinari
ECML/PKDD (2)3
2017 Gaussian Embeddings for Collaborative Filtering
abstract
Most collaborative filtering systems, such as matrix factorization, use vector representations for items and users. Those representations are deterministic, and do not allow modeling the uncertainty of the learned representation, which can be useful when a user has a small number of rated items (cold start), or when there is conflicting information about the behavior of a user or the ratings of an item. In this paper, we leverage recent works in learning Gaussian embeddings for the recommendation task. We show that this model performs well on three representative collections (Yahoo, Yelp and MovieLens) and analyze learned representations.
Ludovic Dos Santos, Benjamin Piwowarski, Patrick Gallinari
SIGIR3
2017 Multiple Bayesian discriminant functions for high-dimensional massive data classification
Jianfei Zhang 0002, Shengrui Wang, Lifei Chen, Patrick Gallinari
Data Min. Knowl. Discov.4
2017 A Novel Embedding Method for Information Diffusion Prediction in Social Network Big Data
abstract
With the increase of social networking websites and the interaction frequency among users, the prediction of information diffusion is required to support effective generalization and efficient inference in the context of social big data era. However, the existing models either rely on expensive probabilistic modeling of information diffusion based on partially known network structures, or discover the implicit structures of diffusion from users' behaviors without considering the impacts of different diffused contents. To address the issues, in this paper, we propose a novel information-dependent embedding-based diffusion prediction (IEDP) model to map the users in observed diffusion process into a latent embedding space, then the temporal order of users with the timestamps in the cascade can be preserved by the embedding distance of users. Our proposed model further learns the propagation probability of information in the cascade as a function of the relative positions of information-specific user embeddings in the information-dependent subspace. Then, the problem of temporal propagation prediction can be converted into the task of spatial probability learning in the embedding space. Moreover, we present an efficient margin-based optimization algorithm with a fast computation to make the inference of the information diffusion in the latent embedding space. When applying our proposed method to several social network datasets, the experimental results show the effectiveness of our proposed approach for the information diffusion prediction and the efficiency with respect to the inference speed compared with the state-of-the-art methods.
Sheng Gao 0001, Huacan Pang, Patrick Gallinari, Jun Guo 0002, Nei Kato
IEEE Trans. Ind. Informatics3
2016 Dynamic Data Capture from Social Media Streams: A Contextual Bandit Approach
Thibault Gisselbrecht, Sylvain Lamprier, Patrick Gallinari
ICWSM3
2016 Learning Distributed Representations of Users for Source Detection in Online Social Networks
Simon Bourigault, Sylvain Lamprier, Patrick Gallinari
ECML/PKDD (2)3
2016 Linear Bandits in Unknown Environments
Thibault Gisselbrecht, Sylvain Lamprier, Patrick Gallinari
ECML/PKDD (2)3
2016 Multilabel Classification on Heterogeneous Graphs with Gaussian Embeddings
Ludovic Dos Santos, Benjamin Piwowarski, Patrick Gallinari
ECML/PKDD (2)3
2016 Representation Learning for Information Diffusion through Social Networks: an Embedded Cascade Model
abstract
In this paper, we focus on information diffusion through social networks. Based on the well-known Independent Cascade model, we embed users of the social network in a latent space to extract more robust diffusion probabilities than those defined by classical graphical learning approaches. Better generalization abilities provided by the use of such a projection space allows our approach to present good performances on various real-world datasets, for both diffusion prediction and influence relationships inference tasks. Additionally, the use of a projection space enables our model to deal with larger social networks.
Simon Bourigault, Sylvain Lamprier, Patrick Gallinari
WSDM3
2016 Local search and pseudoinversion: an hybrid approach to neural network training
Luca Rubini, Rossella Cancelliere, Patrick Gallinari, Andrea Grosso
Knowl. Inf. Syst.3
2015 Leveraging Rating Behavior to Predict Negative Social Ties
abstract
User social networks are a useful information for many information access related tasks, such as recommendation or information retrieval. In such tasks, recent papers have exploited the polarity of these links (friend/enemy) by capturing more precisely social patterns. This negative information being relatively scarce, a recent work proposed to infer it in social networks that contain none. However, this work relies on the direct interaction between users. In this paper, we pursue this approach under the assumption that we do not have access to this kind of data neither, thus allowing to cope with most social networks, where users can rate items and have friendship relationships. We exploit the user ratings polarity, i.e the fact that a rating can be positive (like) or negative (dislike), to infer negative ties. Experiments on the Epinions dataset show the potential of our approach.
Luc-Aurélien Gauthier, Benjamin Piwowarski, Patrick Gallinari
ASONAM3
2015 Extracting Diffusion Channels from Real-World Social Data: a Delay-Agnostic Learning of Transmission Probabilities
abstract
Probabilistic cascade models consider information diffusion as an iterative process in which information transits from users to others in a network. The problem of diffusion modeling then comes down to learning transmission probability distributions, depending on hidden influence relationships between users, in order to discover the main diffusion channels of the network. Various learning models have been proposed in the literature, but we argue that the diffusion mechanisms defined in most of these models are too complex for real social networks, where transmissions of content occur between human users. Classical models usually have some difficulties for extracting the main regularities in such real-world settings. In this paper, we propose a relaxed learning process of the well-known Independent Cascade model that, rather than attempting to explain exact timestamps of users' infections, focus on infection probabilities knowing sets of previously infected users. Experiments show the effectiveness of our proposals, by considering the learned models for real-world prediction tasks.
Sylvain Lamprier, Simon Bourigault, Patrick Gallinari
ASONAM3
2015 Policies for Contextual Bandit Problems with Count Payoffs
abstract
The contextual bandit problem has been of major interest in the last few years. This corresponds to a sequential decision process where an agent has to choose at each iteration an action to perform, according to some knowledge about the decision environment and the current available actions, with the aim to maximize a cumulative amount of rewards over time. Many instances of the problem exist, depending on the kind of rewards we collect - real, binary, natural - and various algorithms are known to be efficient for some of these instances, either empirically or theoretically. In this paper we focus on the case of count payoffs, which corresponds to bandit problems where rewards are integer rewards, potentially unbounded. Based on a Bayesian Poisson regression model, we propose two new contextual bandit algorithms for this particular case with several concrete applications in real life: an Upper Confidence Bound algorithm and a Thompson Sampling strategy. Our approaches present the advantage to remain analytically tractable and computationally efficient. We experiment the algorithms on both simulated data and a real world scenario of spread maximization on a social network.
Thibault Gisselbrecht, Sylvain Lamprier, Patrick Gallinari
ICTAI3
2015 WhichStreams: A Dynamic Approach for Focused Data Capture from Large Social Media
Thibault Gisselbrecht, Ludovic Denoyer, Patrick Gallinari, Sylvain Lamprier
ICWSM3
2015 Latent Trajectory Modeling: A Light and Efficient Way to Introduce Time in Recommender Systems
abstract
For recommender systems, time is often an important source of information but it is also a complex dimension to apprehend. We propose here to learn item and user representations such that any timely ordered sequence of items selected by a user will be represented as a trajectory of the user in a representation space. This allows us to rank new items for this user. We then enrich the item and user representations in order to perform rating prediction using a classical matrix factorization scheme. We demonstrate the interest of our approach regarding both item ranking and rating prediction on a series of classical benchmarks.
Élie Guàrdia-Sebaoun, Vincent Guigue, Patrick Gallinari
RecSys3
2015 An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition
abstract
BACKGROUND: This article provides an overview of the first BIOASQ challenge, a competition on large-scale biomedical semantic indexing and question answering (QA), which took place between March and September 2013. BIOASQ assesses the ability of systems to semantically index very large numbers of biomedical scientific articles, and to return concise and user-understandable answers to given natural language questions by combining information from biomedical articles and ontologies. RESULTS: The 2013 BIOASQ competition comprised two tasks, Task 1a and Task 1b. In Task 1a participants were asked to automatically annotate new PUBMED documents with MESH headings. Twelve teams participated in Task 1a, with a total of 46 system runs submitted, and one of the teams performing consistently better than the MTI indexer used by NLM to suggest MESH headings to curators. Task 1b used benchmark datasets containing 29 development and 282 test English questions, along with gold standard (reference) answers, prepared by a team of biomedical experts from around Europe and participants had to automatically produce answers. Three teams participated in Task 1b, with 11 system runs. The BIOASQ infrastructure, including benchmark datasets, evaluation mechanisms, and the results of the participants and baseline methods, is publicly available. CONCLUSIONS: A publicly available evaluation infrastructure for biomedical semantic indexing and QA has been developed, which includes benchmark datasets, and can be used to evaluate systems that: assign MESH headings to published articles or to English questions; retrieve relevant RDF triples from ontologies, relevant articles and snippets from PUBMED Central; produce "exact" and paragraph-sized "ideal" answers (summaries). The results of the systems that participated in the 2013 BIOASQ competition are promising. In Task 1a one of the systems performed consistently better from the NLM's MTI indexer. In Task 1b the systems received high scores in the manual evaluation of the "ideal" answers; hence, they produced high quality summaries as answers. Overall, BIOASQ helped obtain a unified view of how techniques from text classification, semantic indexing, document and passage retrieval, question answering, and text summarization can be combined to allow biomedical experts to obtain concise, user-understandable answers to questions reflecting their real information needs.
George Tsatsaronis 0001, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R. Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, Yannis Almirantis, John Pavlopoulos, Nicolas Baskiotis, Patrick Gallinari, Thierry Artières, Axel-Cyrille Ngonga Ngomo, Norman Heino, Éric Gaussier, Liliana Barrio-Alvers, Michael Schroeder 0001, Ion Androutsopoulos, Georgios Paliouras
BMC Bioinform.14
2015 Knowledge base completion by learning pairwise-interaction differentiated embeddings
Yu Zhao 0019, Sheng Gao 0001, Patrick Gallinari, Jun Guo 0002
Data Min. Knowl. Discov.3
2015 OCReP: An Optimally Conditioned Regularization for pseudoinversion based neural training
Rossella Cancelliere, Mario Gai, Patrick Gallinari, Luca Rubini
Neural Networks3
2014 Graph Anonymization Using Machine Learning
abstract
Data privacy is a major problem that has to be considered before releasing datasets to the public or even to a partner company that would compute statistics or make a deep analysis of these data. This is insured by performing data anonymization as required by legislation. In this context, many different anonymization techniques have been proposed in the literature. These methods are usually specific to a particular de-anonymization procedure-or attack-one wants to avoid, and to a particular known set of characteristics that have to be preserved after the anonymization. They are difficult to use in a general context where attacks can be of different types, and where measures are not known to the anonymizer. The paper proposes a novel approach for automatically finding an anonymization procedure given a set of possible attacks and a set of measures to preserve. The approach is generic and based on machine learning techniques. It allows us to learn directly an anonymization function from a set of training data so as to optimize a trade off between privacy risk and utility loss. The algorithm thus allows one to get a good anonymization procedure for any kind of attacks, and any characteristic in a given set. Experiments made on two datasets show the effectiveness and the genericity of the approach.
Maria Laura Maag, Ludovic Denoyer, Patrick Gallinari
AINA3
2014 Learning social network embeddings for predicting information diffusion
abstract
Analyzing and modeling the temporal diffusion of information on social media has mainly been treated as a diffusion process on known graphs or proximity structures. The underlying phenomenon results however from the interactions of several actors and media and is more complex than what these models can account for and cannot be explained using such limiting assumptions. We introduce here a new approach to this problem whose goal is to learn a mapping of the observed temporal dynamic onto a continuous space. Nodes participating to diffusion cascades are projected in a latent representation space in such a way that information diffusion can be modeled efficiently using a heat diffusion process. This amounts to learning a diffusion kernel for which the proximity of nodes in the projection space reflects the proximity of their infection time in cascades. The proposed approach possesses several unique characteristics compared to existing ones. Since its parameters are directly learned from cascade samples without requiring any additional information, it does not rely on any pre-existing diffusion structure. Because the solution to the diffusion equation can be expressed in a closed form in the projection space, the inference time for predicting the diffusion of a new piece of information is greatly reduced compared to discrete models. Experiments and comparisons with baselines and alternative models have been performed on both synthetic networks and real datasets. They show the effectiveness of the proposed method both in terms of prediction quality and of inference speed.
Simon Bourigault, Cédric Lagnier, Sylvain Lamprier, Ludovic Denoyer, Patrick Gallinari
WSDM5
2014 Learning latent representations of nodes for classifying in heterogeneous social networks
abstract
Social networks are heterogeneous systems composed of different types of nodes (e.g. users, content, groups, etc.) and relations (e.g. social or similarity relations). While learning and performing inference on homogeneous networks have motivated a large amount of research, few work exists on heterogeneous networks and there are open and challenging issues for existing methods that were previously developed for homogeneous networks. We address here the specific problem of nodes classification and tagging in heterogeneous social networks, where different types of nodes are considered, each type with its own label or tag set. We propose a new method for learning node representations onto a latent space, common to all the different node types. Inference is then performed in this latent space. In this framework, two nodes connected in the network will tend to share similar representations regardless of their types. This allows bypassing limitations of the methods based on direct extensions of homogenous frameworks and exploiting the dependencies and correlations between the different node types. The proposed method is tested on two representative datasets and compared to state-of-the-art methods and to baselines.
Yann Jacob, Ludovic Denoyer, Patrick Gallinari
WSDM3
2014 Web-scale classification: web classification in the big data era
abstract
This paper provides an overview of the workshop Web-Scale Classification: Web Classification in the Big Data Era which was held in New York City, on February 28th as a workshop of the seventh International Conference on Web Search and Data Mining. The goal of the workshop was to discuss and assess recent research focusing on classification and mining in Web-scale category systems. The workshop brought together members of several communities such web mining, machine learning, text classification and social media mining.
Ioannis Partalas, Massih-Reza Amini, Ion Androutsopoulos, Thierry Artières, Patrick Gallinari, Éric Gaussier, Georgios Paliouras
WSDM5
2013 Choosing which message to publish on social networks: a contextual bandit approach
abstract
Maximizing the spread and influence of the messages being published is a challenge for many social network users. Selecting the right content according to the information context and the user characteristics is essential for achieving this goal. We propose a model to automatically choose which information to publish on social networks given a set of possible messages. This model will tend to maximize the spread of the published message for a specific audience. The algorithm is based on the use of a contextual bandit model treating each new potential message as an arm to be selected. We conduct experiments on a Twitter dataset, comparing different algorithms and exploring the influence of the content and the characteristics of the messages on the information spread. The results demonstrate the model's ability to maximize the published information flow as well as it's ability to adapt its behavior to each particular audience.
Ricardo Lage, Ludovic Denoyer, Patrick Gallinari, Peter Dolog
ASONAM3
2013 Latent Factor BlockModel for Modelling Relational Data
Sheng Gao 0001, Ludovic Denoyer, Patrick Gallinari, Jun Guo 0002
ECIR3
2013 Predicting Information Diffusion in Social Networks Using Content and User's Profiles
Cédric Lagnier, Ludovic Denoyer, Éric Gaussier, Patrick Gallinari
ECIR4
2013 Multiview semi-supervised ranking for automatic image annotation
abstract
Most photo sharing sites give their users the opportunity to manually label images. The labels collected that way are usually very incomplete due to the size of the image collections: most images are not labeled according to all the categories they belong to, and, conversely, many class have relatively few representative examples. Automated image systems that can deal with small amounts of labeled examples and unbalanced classes are thus necessary to better organize and annotate images. In this work, we propose a multiview semi-supervised bipartite ranking model which allows to leverage the information contained in unlabeled sets of images in order to improve the prediction performance, using multiple descriptions, or views of images. For each topic class, our approach first learns as many view-specific rankers as available views using the labeled data only. These rankers are then improved iteratively by adding pseudo-labeled pairs of examples on which all view-specific rankers agree over the ranking of examples within these pairs. We report on experiments carried out on the NUS-WIDE dataset, which show that the multiview ranking process improves predictive performances when a small number of labeled examples is available specially for unbalanced classes. We show also that our approach achieves significant improvements over a state-of-the art semi-supervised multiview classification model.
Ali Fakeri-Tabrizi, Massih-Reza Amini, Patrick Gallinari
ACM Multimedia3
2013 Robust Bloom Filters for Large MultiLabel Classification Tasks
abstract
This paper presents an approach to multilabel classification (MLC) with a large number of labels. Our approach is a reduction to binary classification in which label sets are represented by low dimensional binary vectors. This representation follows the principle of Bloom filters, a space-efficient data structure originally designed for approximate membership testing. We show that a naive application of Bloom filters in MLC is not robust to individual binary classifiers' errors. We then present an approach that exploits a specific feature of real-world datasets when the number of labels is large: many labels (almost) never appear together. Our approch is provably robust, has sublinear training and inference complexity with respect to the number of labels, and compares favorably to state-of-the-art algorithms on two large scale multilabel datasets.
Moustapha Cissé, Nicolas Usunier, Thierry Artières, Patrick Gallinari
NIPS4
2013 Cross-Domain Recommendation via Cluster-Level Latent Factor Model
Sheng Gao 0001, Shantao Li, Patrick Gallinari, Jun Guo 0002
ECML/PKDD (2)5
2013 A system for the extraction and representation of summary of product characteristics content
Stefania Rubrichi, Silvana Quaglini, Alex Spengler, Paola Russo, Patrick Gallinari
Artif. Intell. Medicine5
2013 Calibration and regret bounds for order-preserving surrogate losses in learning to rank
Clément Calauzènes, Nicolas Usunier, Patrick Gallinari
Mach. Learn.3
2012 Matrix Pseudoinversion for Image Neural Processing
Rossella Cancelliere, Mario Gai, Thierry Artières, Patrick Gallinari
ICONIP (5)4
2012 Coping with the Document Frequency Bias in Sentiment Classification
Abdelhalim Rafrafi, Vincent Guigue, Patrick Gallinari
ICWSM3
2012 "On the (Non-)existence of Convex, Calibrated Surrogate Losses for Ranking"
abstract
We study surrogate losses for learning to rank, in a framework where the rankings are induced by scores and the task is to learn the scoring function. We focus on the calibration of surrogate losses with respect to a ranking evaluation metric, where the calibration is equivalent to the guarantee that near-optimal values of the sur- rogate risk imply near-optimal values of the risk defined by the evaluation metric. We prove that if a surrogate loss is a convex function of the scores, then it is not calibrated with respect to two evaluation metrics widely used for search engine evaluation, namely the Average Precision and the Expected Reciprocal Rank. We also show that such convex surrogate losses cannot be calibrated with respect to the Pairwise Disagreement, an evaluation metric used when learning from pair- wise preferences. Our results cast lights on the intrinsic difficulty of some ranking problems, as well as on the limitations of learning-to-rank algorithms based on the minimization of a convex surrogate risk.
Clément Calauzènes, Nicolas Usunier, Patrick Gallinari
NIPS3
2012 Learning Compact Class Codes for Fast Inference in Large Multi Class Classification
Moustapha Cissé, Thierry Artières, Patrick Gallinari
ECML/PKDD (1)3
2012 Fast Reinforcement Learning with Large Action Sets Using Error-Correcting Output Codes for MDP Factorization
Gabriel Dulac-Arnold, Ludovic Denoyer, Philippe Preux, Patrick Gallinari
ECML/PKDD (2)4
2012 Ranking with non-random missing ratings: influence of popularity and positivity on evaluation metrics
abstract
The evaluation of recommender systems in terms of ranking has recently gained attention, as it seems to better fit the top-k recommendation task than the usual ratings prediction task. In that context, several authors have proposed to consider missing ratings as some form of negative feedback to compensate for the skewed distribution of observed ratings when users choose the items they rate. In this work, we study two major biases of the selection of items: the first one is that some items obtain more ratings than others (popularity effect), and the second one is that positive ratings are observed more frequently than negative ratings (positivity effect). We present a theoretical analysis and experiments on the Yahoo! dataset with randomly selected items, which show that considering missing data as a form of negative feedback during training may improve performances, but also that it can be misleading when testing, favoring models of popularity more than models of user preferences.
Bruno Pradel, Nicolas Usunier, Patrick Gallinari
RecSys3
2012 Sequential approaches for learning datum-wise sparse representations
Gabriel Dulac-Arnold, Ludovic Denoyer, Philippe Preux, Patrick Gallinari
Mach. Learn.4
2012 A Learning to Rank framework applied to text-image retrieval
David Buffoni, Sabrina Tollari, Patrick Gallinari
Multim. Tools Appl.3
2011 Extracting Information from Summary of Product Characteristics for Improving Drugs Prescription Safety
Stefania Rubrichi, Silvana Quaglini, Alex Spengler, Patrick Gallinari
AIME4
2011 Link Pattern Prediction with tensor decomposition in multi-relational networks
abstract
We address the problem of link prediction in collections of objects connected by multiple relation types, where each type may play a distinct role. While traditional link prediction models are limited to single-type link prediction we attempt here to jointly model and predict the multiple relation types, which we refer to as the Link Pattern Prediction (LPP) problem. For that, we propose a tensor decomposition model to solve the LPP problem, which allows to capture the correlations among different relation types and reveal the impact of various relations on prediction performance. The proposed tensor decomposition model is efficiently learned with a conjugate gradient based optimization method. Extensive experiments on real-world datasets demonstrate that this model outperforms the traditional mono-relational model and can achieve better prediction quality.
Sheng Gao 0001, Ludovic Denoyer, Patrick Gallinari
CIDM3
2011 Temporal link prediction by integrating content and structure information
abstract
In this paper we address the problem of temporal link prediction, i.e., predicting the apparition of new links, in time-evolving networks. This problem appears in applications such as recommender systems, social network analysis or citation analysis. Link prediction in time-evolving networks is usually based on the topological structure of the network only. We propose here a model which exploits multiple information sources in the network in order to predict link occurrence probabilities as a function of time. The model integrates three types of information: the global network structure, the content of nodes in the network if any, and the local or proximity information of a given vertex. The proposed model is based on a matrix factorization formulation of the problem with graph regularization. We derive an efficient optimization method to learn the latent factors of this model. Extensive experiments on several real world datasets suggest that our unified framework outperforms state-of-the-art methods for temporal link prediction tasks.
Sheng Gao 0001, Ludovic Denoyer, Patrick Gallinari
CIKM3
2011 Classification and annotation in social corpora using multiple relations
abstract
We consider the problem of learning to annotate documents with concepts or keywords in content information networks, where the documents may share multiple relations. The concepts associated to a document will depend both on its content and on its neighbors in the network through the different relations. We formalize this problem as single and multi-label classification in a multi-graph, the nodes being the documents and the edges representing the different relations. The proposed algorithm learns to weight the different relations according to their importance for the annotation task. We perform experiments on different corpora corresponding to different annotation tasks on scientific articles, emails and Flickr images and show how the model may take advantage of the rich relational information.
Yann Jacob, Ludovic Denoyer, Patrick Gallinari
CIKM3
2011 The Importance of the Depth for Text-Image Selection Strategy in Learning-To-Rank
David Buffoni, Sabrina Tollari, Patrick Gallinari
ECIR3
2011 Text Classification: A Sequential Reading Approach
Gabriel Dulac-Arnold, Ludovic Denoyer, Patrick Gallinari
ECIR3
2011 Learning Scoring Functions with Order-Preserving Losses and Standardized Supervision
David Buffoni, Clément Calauzènes, Patrick Gallinari, Nicolas Usunier
ICML3
2011 Datum-Wise Classification: A Sequential Approach to Sparsity
Gabriel Dulac-Arnold, Ludovic Denoyer, Philippe Preux, Patrick Gallinari
ECML/PKDD (1)4
2010 Iterative Annotation of Multi-relational Social Networks
abstract
We consider here the task of multi-label classification for data organized in a multi-relational graph. We propose the IMMCA model - Iterative Multi-label Multi-Relational Classification Algorithm - a general algorithm for solving the inference and learning problems for this task. Inference is performed iteratively by propagating scores according to the multi-relational structure of the data. We detail two instances of this general model, implementing two different label propagation schemes on the multi-graph. This is the first collective classification method able to handle multiple relations and to perform multi-label classification in multi-graphs. The target application is image annotation in large social media sharing web sites (Flickr). The goal is to assign labels for images when users and images are connected through multiple relations - authorship, friendship, or visual/textual similarities. We show that our model is able to deal with both content and social relations and performs well on real datasets. Additional experiments on artificial data allow us analyzing the behavior of our method in different situations.
Stéphane Peters, Ludovic Denoyer, Patrick Gallinari
ASONAM3
2010 Document structure meets page layout: loopy random fields for web news content extraction
abstract
Web content extraction is concerned with the automatic identification of semantically interesting web page regions. To generalize to pages from unknown sites, it is crucial to exploit not only the local characteristics of a particular web page region, but also the rich interdependencies that exist between the regions and their latent semantics. We therefore propose a loopy conditional random field which combines semantic intra-page dependencies derived from both document structure and page layout, uses a realistic set of local and relational features and is efficiently learnt in the tree-based reparameterization framework. The results of our empirical analysis on a corpus of real-world news web pages from 177 distinct sites with multiple annotations on DOM node level demonstrate that our combination of document structure and layout-driven interdependencies leads to a significant error reduction on the semantically interesting regions of a web page.
Alex Spengler, Patrick Gallinari
ACM Symposium on Document Engineering2
2010 A Ranking Based Model for Automatic Image Annotation in a Social Network
Ludovic Denoyer, Patrick Gallinari
ICWSM2
2010 Multi-view clustering of multilingual documents
abstract
We propose a new multi-view clustering method which uses clustering results obtained on each view as a voting pattern in order to construct a new set of multi-view clusters. Our experiments on a multilingual corpus of documents show that performance increases significantly over simple concatenation and another multi-view clustering technique.
Massih-Reza Amini, Cyril Goutte, Patrick Gallinari
SIGIR4
2010 Improving document clustering in a learned concept space
Jean-François Pessiot, Massih-Reza Amini, Patrick Gallinari
Inf. Process. Manag.4
2010 Erratum: SGDQN is Less Careful than Expected
Antoine Bordes, Léon Bottou, Patrick Gallinari, Jonathan D. Chang, S. Alex Smith
J. Mach. Learn. Res.3
2009 Exploiting Visual Concepts to Improve Text-Based Image Retrieval
Sabrina Tollari, Marcin Detyniecki, Christophe Marsala, Ali Fakeri-Tabrizi, Massih-Reza Amini, Patrick Gallinari
ECIR6
2009 A self-training method for learning to rank with unlabeled data
Tuong-Vinh Truong, Massih-Reza Amini, Patrick Gallinari
ESANN3
2009 Ranking with ordered weighted pairwise classification
abstract
In ranking with the pairwise classification approach, the loss associated to a predicted ranked list is the mean of the pairwise classification losses. This loss is inadequate for tasks like information retrieval where we prefer ranked lists with high precision on the top of the list. We propose to optimize a larger class of loss functions for ranking, based on an ordered weighted average (OWA) (Yager, 1988) of the classification losses. Convex OWA aggregation operators range from the max to the mean depending on their weights, and can be used to focus on the top ranked elements as they give more weight to the largest losses. When aggregating hinge losses, the optimization problem is similar to the SVM for interdependent output spaces. Moreover, we show that OWA aggregates of margin-based classification losses have good generalization properties. Experiments on the Letor 3.0 benchmark dataset for information retrieval validate our approach.
Nicolas Usunier, David Buffoni, Patrick Gallinari
ICML3
2009 Simulated Iterative Classification A New Learning Procedure for Graph Labeling
Francis Maes, Stéphane Peters, Ludovic Denoyer, Patrick Gallinari
ECML/PKDD (2)4
2009 SGD-QN: Careful Quasi-Newton Stochastic Gradient Descent
Antoine Bordes, Léon Bottou, Patrick Gallinari
J. Mach. Learn. Res.3
2009 Structured prediction with reinforcement learning
Francis Maes, Ludovic Denoyer, Patrick Gallinari
Mach. Learn.3
2008 An extension of PLSA for document clustering
abstract
In this paper we propose an extension of the PLSA model in which an extra latent variable allows the model to co-cluster documents and terms simultaneously. We show on three datasets that our extended model produces statistically significant improvements with respect to two clustering measures over the original PLSA and the multinomial mixture MM models.
Jean-François Pessiot, Massih-Reza Amini, Patrick Gallinari
CIKM4
2008 Efficient Data Clustering by Local Density Approximation
abstract
The clustering task is a key part of the data mining process. In today's context of massive data, methods with a computational complexity more than linear are unlikely to be applied practically. In this paper, we begin by a simple assumption: local projections of the data should allow to distinguish local cluster structures. From there, we describe how to obtain “pure” local sub-groupings of points, from projections on randomly chosen lines. The clustering of the data is obtained from the clustering of these sub-groupings. Our method has a linear complexity in the dataset size, and requires only one pass on the original dataset. Being local in essence, it can handle twisted geometries typical of many high-dimensional datasets. We describe the steps of our method and report encouraging results.
Marc-Ismaël Akodjènou-Jeannin, Patrick Gallinari
ECAI2
2008 Experimental Evaluation of the Value of Structure: How to Efficiently Exploit Interdependencies in Sequence Labeling
abstract
Many problems in natural language processing, information extraction or bioinformatics consist in predicting a label for each element of a sequence of observations. The sequence of labels generally presents multiple dependencies that restrict the possible labels the elements can take. Therefore, relations between labels intuitively provide information valuable for the prediction. Several approaches have been proposed to take advantage of this additional information. However, experimental results show that taking relations into account does not always improve prediction performances, while it significantly increases the computational cost of both learning and prediction. In this work, we aim at both explaining these surprising results and proposing a simple but computationally efficient approach for labeling sequences.
Guillaume Wisniewski, Patrick Gallinari
ICDM2
2007 Sequence Labeling with Reinforcement Learning and Ranking Algorithms
Francis Maes, Ludovic Denoyer, Patrick Gallinari
ECML3
2007 Solving multiclass support vector machines with LaRank
abstract
Optimization algorithms for large margin multiclass recognizers are often too costly to handle ambitious problems with structured outputs and exponential numbers of classes. Optimization algorithms that rely on the full gradient are not effective because, unlike the solution, the gradient is not sparse and is very large. The LaRank algorithm sidesteps this difficulty by relying on a randomized exploration inspired by the perceptron algorithm. We show that this approach is competitive with gradient based optimizers on simple multiclass problems. Furthermore, a single LaRank pass over the training examples delivers test error rates that are nearly as good as those of the final solution.
Antoine Bordes, Léon Bottou, Patrick Gallinari, Jason Weston
ICML3
2007 Flexible Grid-Based Clustering
Marc-Ismaël Akodjènou-Jeannin, Kavé Salamatian, Patrick Gallinari
PKDD3
2007 Relaxation Labeling for Selecting and Exploiting Efficiently Non-local Dependencies in Sequence Labeling
Guillaume Wisniewski, Patrick Gallinari
PKDD2
2007 Online Handwritten Shape Recognition Using Segmental Hidden Markov Models
abstract
We investigate a new approach for online handwritten shape recognition. Interesting features of this approach include learning without manual tuning, learning from very few training samples, incremental learning of characters, and adaptation to the user-specific needs. The proposed system can deal with two-dimensional graphical shapes such as Latin and Asian characters, command gestures, symbols, small drawings, and geometric shapes. It can be used as a building block for a series of recognition tasks with many applications.
Thierry Artières, Sanparith Marukatat, Patrick Gallinari
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Precision recall with user modeling (PRUM): Application to structured information retrieval
abstract
Standard Information Retrieval (IR) metrics are not well suited for new paradigms like XML or Web IR in which retrievable information units are document elements and/or sets of related documents. Part of the problem stems from the classical hypotheses on the user models: They do not take into account the structural or logical context of document elements or the possibility of navigation between units. This article proposes an explicit and formal user model that encompasses a large variety of user behaviors. Based on this model, we extend the probabilistic precision-recall metric to deal with the new IR paradigms.
Benjamin Piwowarski, Patrick Gallinari, Georges Dupret
ACM Trans. Inf. Syst.2
2006 Machine Learning Ranking for Structured Information Retrieval
Jean-Noël Vittaut, Patrick Gallinari
ECIR2
2006 A Selective Sampling Strategy for Label Ranking
Massih-Reza Amini, Nicolas Usunier, François Laviolette, Alexandre Lacasse, Patrick Gallinari
ECML5
2006 A Machine Learning based Approach to Evaluating Retrieval Systems
Huyen-Trang Vu, Patrick Gallinari
HLT-NAACL2
2005 Learning to summarise XML documents using content and structure
abstract
Documents formatted in eXtensible Markup Language (XML) are becoming increasingly available in collections of various document types. In this paper, we present an approach for the summarisation of XML documents. The novelty of this approach lies in that it is based on features not only from the content of documents, but also from their logical structure. We follow a machine learning like, sentence extraction-based summarisation technique. To find which features are more effective for producing summaries this approach views sentence extraction as an ordering task. We evaluated our summarisation model using the INEX dataset. The results demonstrate that the inclusion of features from the logical structure of documents increases the effectiveness of the summariser, and that the learnable system is also effective and well-suited to the task of summarisation in the context of XML documents.
Massih-Reza Amini, Anastasios Tombros, Nicolas Usunier, Mounia Lalmas-Roelleke, Patrick Gallinari
CIKM5
2005 Using RankBoost to compare retrieval systems
abstract
This paper presents a new pooling method for constructing the assessment sets used in the evaluation of retrieval systems. Our proposal is based on RankBoost, a machine learning voting algorithm. It leads to smaller pools than classical pooling and thus reduces the manual assessment workload for building test collections. Experimental results obtained on an XML document collection demonstrate the effectiveness of the approach according to different evaluation criteria.
Huyen-Trang Vu, Patrick Gallinari
CIKM2
2005 Automatic Text Summarization Based on Word-Clusters and Ranking Algorithms
Massih-Reza Amini, Nicolas Usunier, Patrick Gallinari
ECIR3
2005 Automatic learning of domain model for personalized hypermedia applications
Hermine Njike Fotzo, Thierry Artières, Patrick Gallinari, Julien Blanchard 0003, Guillaume Letellier
IJCAI3
2005 Generalization error bounds for classifiers trained with interdependent data
abstract
In this paper we propose a general framework to study the generalization properties of binary classifiers trained with data which may be depen- dent, but are deterministically generated upon a sample of independent examples. It provides generalization bounds for binary classification and some cases of ranking problems, and clarifies the relationship between these learning tasks.
Nicolas Usunier, Massih-Reza Amini, Patrick Gallinari
NIPS3
2005 A Bayesian Framework for XML Information Retrieval: Searching and Learning with the INEX Collection
Benjamin Piwowarski, Patrick Gallinari
Inf. Retr.2
2005 Semi-supervised learning with an imperfect supervisor
Massih-Reza Amini, Patrick Gallinari
Knowl. Inf. Syst.2
2004 A Model-Based Approach to Sequence Clustering
Henri Binsztok, Thierry Artières, Patrick Gallinari
ECAI3
2004 Bayesian network model for semi-structured document classification
Ludovic Denoyer, Patrick Gallinari
Inf. Process. Manag.2
2003 Structured multimedia document classification
abstract
International audience
Ludovic Denoyer, Jean-Noël Vittaut, Patrick Gallinari, Sylvie Brunessaux, Stephan Brunessaux
ACM Symposium on Document Engineering3
2003 Classification and Tracking of Hypermedia Navigation Patterns
Patrick Gallinari, Sylvain Bidel, Laurent Lemoine, Frédéric Piat, Thierry Artières
ICANN1
2003 A Flexible Recognition Engine for Complex On-line Handwritten Character Recognition
abstract
A major feature of new mobiles terminals using penbased interfaces, such as personal assistants or e-book, is their personal character, implying that a good interface should be easily customizable in order to meet various users ’ needs. We proposed recently a new recognition engine with strong adaptation abilities that allows learning a user’s writing style or new symbols easily. It dealt with characters with simple shape only; we describe here an extension of this recognition system that deals with more complex graphical symbols. We propose an adaptive learning scheme that can learn more and more as the writer gives new samples. 1.
Sanparith Marukatat, Rudy Sicard, Thierry Artières, Patrick Gallinari
ICDAR4
2003 Semi-Supervised Learning with Explicit Misclassification Modeling
Massih-Reza Amini, Patrick Gallinari
IJCAI2
2003 Using Belief Networks and Fisher Kernels for Structured Document Classification
Ludovic Denoyer, Patrick Gallinari
PKDD2
2003 A Hidden Markov Models combination framework for handwriting recognition
Thierry Artières, Nadji Gauthier, Patrick Gallinari, Bernadette Dorizzi
Int. J. Document Anal. Recognit.3
2003 An MLP-SVM combination architecture for offline handwritten digit recognition: Reduction of recognition errors by Support Vector Machines rejection mechanisms
Abdel Bellili, Michel Gilloux, Patrick Gallinari
Int. J. Document Anal. Recognit.3
2002 Semi Supervised Logistic Regression
Massih-Reza Amini, Patrick Gallinari
ECAI2
2002 Learning Classification with Both Labeled and Unlabeled Data
Jean-Noël Vittaut, Massih-Reza Amini, Patrick Gallinari
ECML3
2002 Data Driven Design of an ANN/HMM System for On-line Unconstrained Handwritten Character Recognition
abstract
This paper is dedicated to a data driven design method for a hybrid ANN/HMM based handwriting recognition system. On one hand, a data driven designed neural modelling of handwriting primitives is proposed. ANNs are firstly used as state models in a HMM primitive divider that associates each signal frame with an ANN by minimizing the accumulated prediction error. Then, the neural modelling is realized by training each network on its own frame set. Organizing these two steps in an EM algorithm, precise primitive models are obtained. On the other hand, a data driven systematic method is proposed for the HMM topology inference task. All possible prototypes of a pattern class are firstly merged into several clusters by a tabu search aided clustering algorithm. Then a multiple parallel-path HMM is constructed for the pattern class. Experiments prove an 8% recognition improvement with a saving of 50% of system resources, compared to an intuitively designed referential ANN/HMM system.
Thierry Artières, Patrick Gallinari
ICMI3
2002 State Sharing in a Hybrid Neuro-Markovian On-Line Handwriting Recognition System through a Simple Hierarchical Clustering Algorithm
abstract
HMM has been largely applied in many fields with great success. To achieve a better performance, an easy way is using more states or more free parameters for a better signal modelling. Thus, state sharing and state clipping methods have been proposed to reduce parameter redundancy and to limit the explosive consummation of system resources. We focus on a simple state sharing method for a hybrid neuro-Markovian on-line handwriting recognition system. At first, a likelihood-based distance is proposed for measuring the similarity between two HMM state models. Afterwards, a minimum quantification error aimed hierarchical clustering algorithm is also proposed to select the most representative models. Here, models are shared to the most under the constraint of the minimum system performance loss. As the result, we maintain about 98% of the system performance while about 60% of the parameters are reduced.
Thierry Artières, Patrick Gallinari
ICMI3
2002 The use of unlabeled data to improve supervised learning for text summarization
abstract
With the huge amount of information available electronically, there is an increasing demand for automatic text summarization systems. The use of machine learning techniques for this task allows one to adapt summaries to the user needs and to the corpus characteristics. These desirable properties have motivated an increasing amount of work in this field over the last few years. Most approaches attempt to generate summaries by extracting sentence segments and adopt the supervised learning paradigm which requires to label documents at the text span level. This is a costly process, which puts strong limitations on the applicability of these methods. We investigate here the use of semi-supervised algorithms for summarization. These techniques make use of few labeled data together with a larger amount of unlabeled data. We propose new semi-supervised algorithms for training classification models for text summarization. We analyze their performances on two data sets - the Reuters news-wire corpus and the Computation and Language (cmp_lg) collection of TIPSTER SUMMAC. We perform comparisons with a baseline - non learning - system, and a reference trainable summarizer system.
Massih-Reza Amini, Patrick Gallinari
SIGIR2
2001 Learning for Text Summarization Using Labeled and Unlabeled Sentences
Massih-Reza Amini, Patrick Gallinari
ICANN2
2001 An Hybrid MLP-SVM Handwritten Digit Recognizer
abstract
This paper presents an original hybrid MLP-SVM method for unconstrained handwritten digits recognition. Specialized support vector machines (SVMs) are introduced to improve significantly the multilayer perceptron (MLP) performances in local areas around the separation surfaces between each pair of digit classes, in the input pattern space. This hybrid architecture is based on the idea that the correct digit class almost systematically belongs to the two maximum MLP outputs and that some pairs of digit classes constitute the majority of MLP substitutions (errors). Specialized local SVMs are introduced to detect the correct class among these two classification hypotheses. The hybrid MLP-SVM recognizer achieves a recognition rate of 98.01%, for real mail zip code digits recognition task, a performance better than several classifiers reported in recent researches.
Abdel Bellili, Michel Gilloux, Patrick Gallinari
ICDAR3
2001 Strategies for Combining On-line and Off-line Information in an On-line Handwriting Recognition System
abstract
This paper investigates the cooperation of online and off-line handwriting word recognition systems. Our goal is to improve a mature online recognition system by, exploiting the complementary information present in the off-line representation built from online signal. After describing the online and off-line HMM based handwriting recognition systems, we propose a formal framework, which allows describing different strategies for combining the two HMM. These schemes are then evaluated on the UNIPEN database, both for isolated character and word recognition tasks.
Nadji Gauthier, Thierry Artières, Patrick Gallinari, Bernadette Dorizzi
ICDAR3
2001 Sentence Recognition through Hybrid Neuro-Markovian Modeling
abstract
This paper focuses on designing a handwriting recognition system dealing with on-line signal, i.e. temporal handwriting signal captured through an electronic pen or a digitalized tablet. We present here some new results concerning a hybrid on-line handwriting recognition system based on Hidden Markov Models (HMMs) and Neural Networks (NNs), which has already been presented in several contributions. In our approach, a letter-model is a Left-Right HMM, whose emission probability densities are approximated with mixtures of predictive multilayer perceptrons. The basic letter models are cascaded in order to build models for words and sentences. At the word level, recognition is performed thanks to a dictionary organized with a tree-structure. At the sentence level, a word-predecessor conditioned frame synchronous beam search algorithm allows to perform simultaneously segmentation into words and word recognition. It processes through the building of a word graph from which a set of candidate sentences may be extracted. Word and sentence recognition performances are evaluated on parts of the UNIPEN international database.
Sanparith Marukatat, Thierry Artières, Patrick Gallinari, Bernadette Dorizzi
ICDAR3
2001 Automatic Text Summarization Using Unsupervised and Semi-supervised Learning
Massih-Reza Amini, Patrick Gallinari
PKDD2
2001 Maximum mutual information training for an online neural predictive handwritten word recognition system
Sonia Garcia-Salicetti, Bernadette Dorizzi, Patrick Gallinari, Zsolt Wimmer
Int. J. Document Anal. Recognit.3
2000 Multi-Modal Segmental Models for On-Line Handwriting Recognition
abstract
Hidden Markov models (HMMs) have become within a few years the main technology for online handwritten word recognition (HWR). We consider segment models which generalize HMMs, these models aim at modeling the signal at a global level rather than at the frame level and have been shown to overcome standard HMMs in their modeling ability. We propose a segment model which allows us to automatically handle different writing styles. We compare our system on the isolated character set of the UNIPEN database with a reference system and a baseline segment model.
Thierry Artières, J.-M. Marchand, Patrick Gallinari, Bernadette Dorizzi
ICPR3
1999 Dictionary Preselection in a Neuro-Markovian Word Recognition System
abstract
Previously we have introduced a neural predictive system for on-line and off-line word recognition (Garcia-Salicetti et al., 1995; 1996; 1997). Words are recognized thanks to a dictionary which is used in a postprocessing stage. We focus on this lexical part of the system and more precisely on the preselection technique that is used to reduce the computational complexity. We define and compare two edit distances and show how the results can be improved through the use of the confusion matrix.
Zsolt Wimmer, Bernadette Dorizzi, Patrick Gallinari
ICDAR3
1999 Improved performance in protein secondary structure prediction by inhomogeneous score combination
abstract
MOTIVATION: In many fields of pattern recognition, combination has proved efficient to increase the generalization performance of individual prediction methods. Numerous systems have been developed for protein secondary structure prediction, based on different principles. Finding better ensemble methods for this task may thus become crucial. Furthermore, efforts need to be made to help the biologist in the post-processing of the outputs. RESULTS: An ensemble method has been designed to post-process the outputs of discriminant models, in order to obtain an improvement in prediction accuracy while generating class posterior probability estimates. Experimental results establish that it can increase the recognition rate of protein secondary structure prediction methods that provide inhomogeneous scores, even though their individual prediction successes are largely different. This combination thus constitutes a help for the biologist, who can use it confidently on top of any set of prediction methods. Moreover, the resulting estimates can be used in various ways, for instance to determine which areas in the sequence are predicted with a given level of reliability. AVAILABILITY: The prediction is freely available over the Internet on the Network Protein Sequence Analysis (NPS@) WWW server at http://pbil.ibcp.fr/NPSA/npsa_server.ht ml. The source code of the combiner can be obtained on request for academic use.
Yann Guermeur, Christophe Geourjon, Patrick Gallinari, Gilbert Deléage
Bioinform.3
1999 Practical complexity control in multilayer perceptrons
Patrick Gallinari, Tautvydas Cibas
Signal Process.1
1998 Twelve Numerical, Symbolic and Hybrid Supervised Classification Methods
abstract
Supervised classification has already been the subject of numerous studies in the fields of Statistics, Pattern Recognition and Artificial Intelligence under various appellations which include discriminant analysis, discrimination and concept learning. Many practical applications relating to this field have been developed. New methods have appeared in recent years, due to developments concerning Neural Networks and Machine Learning. These "hybrid" approaches share one common factor in that they combine symbolic and numerical aspects. The former are characterized by the representation of knowledge, the latter by the introduction of frequencies and probabilistic criteria. In the present study, we shall present a certain number of hybrid methods, conceived (or improved) by members of the SYMENU research group. These methods issue mainly from Machine Learning and from research on Classification Trees done in Statistics, and they may also be qualified as "rule-based". They shall be compared with other more classical approaches. This comparison will be based on a detailed description of each of the twelve methods envisaged, and on the results obtained concerning the "Waveform Recognition Problem" proposed by Breiman et al.,4 which is difficult for rule based approaches.
Olivier Gascuel, Bernadette Bouchon-Meunier, Gilles Caraux, Patrick Gallinari, Alain Guénoche, Yann Guermeur, Yves Lechevallier, Christophe Marsala, Laurent Miclet, Jacques Nicolas, Richard Nock, Mohammed Ramdani 0003, Michèle Sebag, Basavanneppa Tallur, Gilles Venturini, Patrick Vitte
Int. J. Pattern Recognit. Artif. Intell.4
1998 Multiple multivariate regression and global sequence optimization: : An application to large-scale models of radiation intensity
Hugo Zaragoza, Patrick Gallinari, R. Curtelin, F. Leglaye
Signal Process.2
1997 Optimal Linear Regression on Classifier Outputs
Yann Guermeur, Florence d'Alché-Buc, Patrick Gallinari
ICANN3
1997 Multiple Multivariate Regression and Global Optimization in a Large Scale Thermodynamical Application
Hugo Zaragoza, Patrick Gallinari
ICANN2
1997 Neural and adaptive controllers for a non-minimum phase varying time-delay system
Abdelmoumène Toudeft, Patrick Gallinari
Artif. Intell. Eng.2
1996 Combining Statistical Models for Protein Secondary Structure Prediction
Yann Guermeur, Patrick Gallinari
ICANN2
1996 Diagnosis Tools for Telecommunication Network Traffic Management
Philippe Leray 0001, Patrick Gallinari, Elisabeth Didelet
ICANN2
1996 From characters to words: dynamical segmentation and predictive neural networks
abstract
We present the extension of a neural predictive system primitively designed for on-line character recognition to words. Feature extraction is performed after resampling the pen trajectory information, recorded by a digitizing tablet. Each word is modeled by the natural concatenation of letter-models corresponding to the letters composing it. Successive parts of a word trajectory are this way modeled by different neural networks and only transitions from each one to itself or to its right neighbors are permitted. A holistic and dynamical segmentation allows one to adjust letter-models to the great variability of handwriting encountered in the words. Our system combines multilayer neural networks and dynamic programming with an underlying left-right hidden Markov model (HMM). Training was performed on 7000 words from 9 writers, leading to good results in the letter-labelling process, without using any language model.
Sonia Garcia-Salicetti, Patrick Gallinari, Bernadette Dorizzi, Zsolt Wimmer, Stéphane Gentric
ICASSP2
1996 Adaptive discrimination in an HMM-based neural predictive system for on-line word recognition
abstract
We have introduced previously (1996) a neural predictive system for on-line word recognition. Our approach implements a hidden Markov model (HMM)-based cooperation of several neural networks. The task of the HMM is to guide the training procedure of neural networks on successive parts of a word. Each word is modeled by the concatenation of letter-models corresponding to the letters composing it. In this article, we present the discriminative training procedures introduced in order to improve the results of our first model. Discriminative training is described at the local level, that is of each extracted parameter vector, and at the global level, that is the level of sequences of labels. We relate this type of training in both cases to the maximum mutual information formalism. Discriminative training was performed on 7000 words from 9 writers, leading to improved results at the character level. Moreover, the use of a neural lexical post-processor (NLPP) gives very good word recognition rates.
Sonia Garcia-Salicetti, Bernadette Dorizzi, Patrick Gallinari, Zsolt Wimmer
ICPR3
1996 Variable selection with neural networks
Tautvydas Cibas, Françoise Fogelman-Soulié, Patrick Gallinari, Sarunas Raudys
Neurocomputing3
1995 A neural predictive approach for on-line cursive script recognition
abstract
We present a neural prediction system for on-line writer-independent character recognition as a first step towards a word recognition system. The input feature vectors contain the pen trajectory information, recorded by a digitizing tablet. Each letter is modeled by a variable number of predictive neural networks, depending on its length. Successive parts of a letter are modeled by different multilayer neural networks, only transitions from each one to itself or to its right neighbors being permitted. To deal with the great variability of cursive handwriting, we introduce a holistic approach for both learning and recognition, combining neural networks and dynamic programming techniques. Our system is able to recognize strongly distorted and truncated letters, obtained by automatic segmentation of 10000 words from 10 different writers. Even on such databases, inappropriate to character recognition (letters in it were not recorded as handwritten isolated characters), quite good recognition rates are obtained.
Sonia Garcia-Salicetti, Patrick Gallinari, Bernadette Dorizzi, Abdelhamid Mellouk, D. Fanchon
ICASSP2
1995 Global discrimination for neural predictive systems based on N-best algorithm
abstract
We describe a general formalism for training neural predictive systems. We then introduce discrimination at the frame level and show how it relates to maximum mutual information training. Finally, we propose an approach for performing discrimination in predictive systems at the sequence level, it makes use of N-best sequence selection. The performance for acoustic-phonetic decoding showed a 77.4% phone accuracy on the 1988 version of the TIMIT database.
Abdelhamid Mellouk, Patrick Gallinari
ICASSP2
1995 A hidden Markov model extension of a neural predictive system for on-line character recognition
abstract
The authors present a neural predictive system for on-line writer-independent character recognition. The data collection of each letter contains the pen trajectory information recorded by a digitizing tablet. Each letter is modeled by a fixed number of predictive neural networks (NN), so that a different multilayer NN models successive parts of a letter. The topology of each letter-model only permits transitions from each NN to itself or to its neighbors. In order to deal with the great variability proper to cursive handwriting in the omni-scriptor framework, they implement a holistic approach during both learning and recognition by performing adaptive segmentation. Also, the recognition step implements interactive recognition and segmentation. The approach compares neural techniques combined with dynamic programming to its extension to the hidden Markov model (HMM) framework. The first system gives quite good recognition rates on letter databases obtained from 10 different writers, and results improve considerably when one considers the extension of the first system to the durational HMM framework.
Sonia Garcia-Salicetti, Bernadette Dorizzi, Patrick Gallinari, Abdelhamid Mellouk, D. Fanchon
ICDAR3
1995 Multi-state predictive neural networks for text-independent speaker recognition
abstract
Both Hidden Markov Models and Neural Networks have already been used as production systems for speaker identification or verification. Recently [9] has shown that ergodic multi-state hidden Markov Models do not outperform one-state "hidden" Markov Models, i.e. Gaussian Mixture Models, for speaker recognition. She put in evidence that the important characteristic of these models is the total number of mixtures and not the number of states. These HMMs are thus unable to make use of temporal information for performing speaker recognition. On the other hand, recent experiments have shown that, for neural predictive systems, modelization of non stationarity allowed to significantly improve the performances [6]. We are interested here in the development of such models which will be refereed to as multi-state predictive neural networks (MSPNNs). We study the ability of these systems for speaker identification and discuss the superiority of multi-state upon one-state models. We provide results...
Thierry Artières, Patrick Gallinari
EUROSPEECH2
1995 Integrating heuristic preferences into a neural understanding system
Adelaide Stevenur, Patrick Gallinari
EUROSPEECH2
1995 Neural networks for discrimination and modelization of speakers
Younès Bennani, Patrick Gallinari
Speech Commun.2
1994 Discriminative training for improved neural prediction systems
abstract
Presents improvements to neural predictive systems for acoustic-phonetic decoding. They allow to raise the performances of these systems close to the state of the art. Important increases have been obtained through carefully selected discriminant criteria.>
Abdelhamid Mellouk, Patrick Gallinari
ICASSP (1)2
1993 A discriminative neural prediction system for speech recognition
Abdelhamid Mellouk, Patrick Gallinari
ICASSP (1)2
1993 Neural models for extracting speaker characteristics in speech modelization systems
abstract
We analyze in this paper inherent limitations of prediction systems for speaker identification. We introduce different approaches for enhancing these systems and present results of tests on TIMIT with neural net predictive systems. Keywords : speaker recognition, predictive neural networks, sequence classification. 1. INTRODUCTION Over the last 15 years, several approaches have been tested for speaker recognition problems. Many of them use static characteristics of speaker utterances computed from long term statistics or short term spectra. More recently dynamical approaches have been used to modelize more accurately speaker characteristics. Most of them have been inspired from speech recognition techniques. Several systems have thus been proposed based on Hidden Markov Models (HMM) [1], Vectorial Autoregressive Models (VAM) [2,3] or Neural Networks (NN) [4]. We will focus here on free text speaker identification by neural networks. Whereas long term statistics may contain enough infor...
Thierry Artières, Patrick Gallinari
EUROSPEECH2
1993 Prediction and discrimination in neural networks for continuous speech recognition
Abdelhamid Mellouk, Patrick Gallinari, F. Rauscher
EUROSPEECH2
1992 A speech recognizer optimally combining learning vector quantization, dynamic programming and multi-layer perceptron
abstract
The authors give a detailed description of a new hybrid system for acoustic decoding. The system features cooperation between a multilayer perceptron (MLP) and an adaptive dynamic programming (DP) module. They show how to train the whole system in an optimal way using an adaptive gradient technique. The DP module optimizes cost functions inspired from k-means and learning vector quantization (LVQ). This module allows the training of synthetic references which incorporate discriminant information and improves the performance and/or speed of usual dynamic programming systems. The authors analyze and provide solutions to some problems which may occur when training the whole hybrid system and show that they are common to many modular architectures. These theoretical issues are illustrated through experiments on an isolated-word database.>
Xavier Driancourt, Patrick Gallinari
ICASSP2
1991 Validation of neural net architectures on speech recognition tasks
abstract
Using two speech recognition tasks, the authors compared the performance and behavior of time delay neural networks (NN), learning vector quantization, and a modular architecture. This set of experiments makes it possible to investigate the capabilities of the models and demonstrate some of their weaknesses. Good performance was obtained through the use of sophisticated architectures which encompass the limitations of more basic NN models. This is particularly clear for a phoneme experiment where it was possible to increase the performances until they were far better than those of traditional classifiers. This improvement was obtained in successive steps by using modified cost functions or algorithms and building a combined architecture. These results illustrate that current NN algorithms can be greatly improved. Modular architectures like the one used are a promising way to do this.>
Younès Bennani, Nasser Chaourar, Patrick Gallinari, Abdelhamid Mellouk
ICASSP3
1991 On the use of TDNN-extracted features information in talker identification
abstract
The authors propose a novel model for text-independent talker identification which uses TDNN (time-delay neural network) extracted feature information. This model has been tested on 20 speakers (10 male and 10 female) from the TIMIT database using an LPC (linear predictive coding) parameterization. An average identification of 98% was observed.>
Younès Bennani, Patrick Gallinari
ICASSP2
1991 On the relations between discriminant analysis and multilayer perceptrons
Patrick Gallinari, Sylvie Thiria, Fouad Badran, Françoise Fogelman-Soulié
Neural Networks1
1990 A connectionist approach for automatic speaker identification
abstract
A connectionist approach to automatic speaker identification based on the learning vector quantization (VQ) algorithm is presented. For each adherent to the identification system, a number of references is fixed. The algorithm is based on a nearest-neighbor principle, with adaptation through learning. The identification is realized by comparing to a given threshold the distance of the unknown utterance to the nearest reference. Preliminary tests run on a ten-speaker set show an identification rate of 97% for MFC coefficients. The identification system and database used and the results obtained for different combinations of parameters are given. The system is evaluated by comparing its performances with a Bayesian system.>
Younès Bennani, Françoise Fogelman-Soulié, Patrick Gallinari
ICASSP3
1990 Cooperation of neural nets for robust classification
abstract
The problem of building robust neural network classification algorithms that perform well on a wide variety of problems is addressed. Evidence of the fact that current procedures are only accurate on a restricted range of tasks is presented. The authors propose multimodule architectures based on the cooperation of two or more neural net techniques as a solution. This idea is illustrated by the cooperation of a multilayer perceptron and a learning vector quantization algorithm. It is shown that this method combines the advantages of its individual components and is much more robust. It is accurate for a large range of problems and is easy to tune. An algorithm that allows the direct training of this multimodule architecture is described. The use of this technique enhances performance when dealing with classification problems
Michel de Bollivier, Patrick Gallinari, Sylvie Thiria
IJCNN2
1990 A Framework for the Cooperation of Learning Algorithms
Léon Bottou, Patrick Gallinari
NIPS2
1990 Speech processing and recognition using integrated neurocomputing techniques (Esprit Basic Research Action 3228: SPRINT)
Khalid Choukri, S. Soudoplatoff, A. Wallyn, Frédéric Bimbot, H. Valbret, Patrick Gallinari, Younès Bennani, A. Varga, Manfred Immendörfer, T. Michaux
Neurocomputing6
1988 Comparing neural networks and data analysis
Patrick Gallinari, Sylvie Thiria, Françoise Fogelman-Soulié
Neural Networks1