Raphaël Féraud

dblp:87/2264 · DBLP profile ↗
← Back
36ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0001-9986-8878ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 8 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 5 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Multi-Armed Bandits Meet Large Language Models
abstract
Bandit algorithms and Large Language Models (LLMs) have emerged as powerful tools in artificial intelligence, each addressing distinct yet complementary challenges in decision-making and natural language processing. This survey explores the synergistic potential between these two fields, highlighting how bandit algorithms can enhance the performance of LLMs and how LLMs, in turn, can provide novel insights for improving bandit-based decision-making. We first examine the role of bandit algorithms in optimizing LLM fine-tuning, prompt engineering, and adaptive response generation, focusing on their ability to balance exploration and exploitation in large-scale learning tasks. Subsequently, we explore how LLMs can augment bandit algorithms through advanced contextual understanding, dynamic adaptation, and improved policy selection using natural language reasoning. By providing a comprehensive review of existing research and identifying key challenges and opportunities, this survey aims to bridge the gap between bandit algorithms and LLMs, paving the way for innovative applications and interdisciplinary research in AI.
Djallel Bouneffouf 0001, Raphaël Féraud
AAAI2
2026 Decentralized multi-agent multi-armed bandits for smart electric vehicles charging
Sharyal Zafar, Raphaël Féraud, Anne Blavette, Guy Camilleri, Hamid Ben Ahmed
Eng. Appl. Artif. Intell.2
2024 A Tutorial on Multi-Armed Bandit Applications for Large Language Models
abstract
This tutorial offers a comprehensive guide on using multi-armed bandit (MAB) algorithms to improve Large Language Models (LLMs). As Natural Language Processing (NLP) tasks grow, efficient and adaptive language generation systems are increasingly needed. MAB algorithms, which balance exploration and exploitation under uncertainty, are promising for enhancing LLMs.
Djallel Bouneffouf 0001, Raphaël Féraud
KDD2
2023 Question Answering System with Sparse and Noisy Feedback
abstract
The rise of personal assistants has made question answering a very popular mechanism for user-system interaction. In Question Answering System, implicit feedbacks can be easily observed (user clicking in the link given by the QA system), but they are noisy. However, receiving an explicit feedback on the quality of the response just given is rare but more valuable. Motivated by a practical need in Question Answering System of processing these two types of rewards, this paper investigates and proposes a new stochastic multi-armed bandit model in which each action has a noisy reward and a sparse reward. We studied this problem in the contextual bandit settings, and proposed and analyzed efficient algorithms that are based on the LINUCB frameworks. Our algorithms are verified by empirical studies on various reward distributions and a real-world dataset and application.
Djallel Bouneffouf 0001, Oznur Alkan, Raphaël Féraud, Baihan Lin
ICASSP3
2023 Massive multi-player multi-armed bandits for IoT networks: An application on LoRa networks
Hiba Dakdouk, Raphaël Féraud, Nadège Varsier, Patrick Maillé, Romain Laroche
Ad Hoc Networks2
2021 Horizontal Scaling in Cloud Using Contextual Bandits
David Delande, Patricia Stolf, Raphaël Féraud, Jean-Marc Pierson, André Bottaro
Euro-Par3
2021 Toward Skills Dialog Orchestration with Online Learning
abstract
Building multi-domain AI agents is a challenging task and an open problem in the area of AI. Within the domain of dialog, the ability to orchestrate multiple independently trained dialog agents, or skills, to create a unified system is of particular significance. In this work, we study the task of online posterior dialog orchestration, where we define posterior orchestration as the task of selecting a subset of skills which most appropriately answer a user input using features extracted from both the user input and the individual skills. To account for the various costs associated with extracting skill features, we consider online posterior orchestration under a skill execution budget. We formalize this setting as Context Attentive Bandit with Observations (CABO), a variant of context attentive bandits, and evaluate it on proprietary conversational datasets.
Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Mayank Agarwal, Yasaman Khazaeni, Irina Rish
ICASSP2
2021 Double-Linear Thompson Sampling for Context-Attentive Bandits
abstract
In this paper, we analyze and extend an online learning frame-work known as Context-Attentive Bandit, motivated by various practical applications, from medical diagnosis to dialog systems, where due to observation costs only a small subset of a potentially large number of context variables can be observed at each iteration; however, the agent has a freedom to choose which variables to observe. We derive a novel algorithm, called Context-Attentive Thompson Sampling (CATS), which builds upon the Linear Thompson Sampling approach, adapting it to Context-Attentive Bandit setting. We provide a theoretical regret analysis and an extensive empirical evaluation demonstrating advantages of the proposed approach over several baseline methods on a variety of real-life datasets.
Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Yasaman Khazaeni, Irina Rish
ICASSP2
2021 Toward Optimal Solution for the Context-Attentive Bandit Problem
abstract
In various recommender system applications, from medical diagnosis to dialog systems, due to observation costs only a small subset of a potentially large number of context variables can be observed at each iteration; however, the agent has a freedom to choose which variables to observe. In this paper, we analyze and extend an online learning framework known as Context-Attentive Bandit, We derive a novel algorithm, called Context-Attentive Thompson Sampling (CATS), which builds upon the Linear Thompson Sampling approach, adapting it to Context-Attentive Bandit setting. We provide a theoretical regret analysis and an extensive empirical evaluation demonstrating advantages of the proposed approach over several baseline methods on a variety of real-life datasets.
Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Irina Rish, Yasaman Khazaeni
IJCAI2
2020 Collaborative Exploration in Stochastic Multi-Player Bandits
abstract
Internet of Things (IoT) faces multiple challenges to achieve high reliability, low-latency and low power consumption. Its performance is affected by many factors such as external interference coming from other coexisting wireless communication technologies that are sharing the same spectrum. To address this problem, we introduce a general approach for the identification of poor-link quality channels. We formulate our problem as a multi-player multi-armed bandit problem, where the devices in an IoT network are the players, and the arms are the radio channels. For a realistic formulation, we do not assume that sensing information is available or that the number of players is below the number of arms. We develop and analyze a collaborative decentralized algorithm that aims to find a set of $m$ $(\epsilon,m)$-optimal arms using an Explore-$m$ algorithm (as denoted by Kalyanakrishnan and Stone (2010)) as a subroutine, and hence blacklisting the suboptimal arms in order to improve the QoS of IoT networks while reducing their energy consumption. We prove analytically and experimentally that our algorithm outperforms selfish algorithms in terms of sample complexity with a low communication cost, and that although playing a smaller set of arms increases the collision rate, playing the optimal arms only improves the QoS of the network.
Hiba Dakdouk, Raphaël Féraud, Nadège Varsier, Patrick Maillé
ACML2
2020 Restarted Bayesian Online Change-point Detector achieves Optimal Detection Delay
abstract
we consider the problem of sequential change-point detection where both the change-points and the distributions before and after the change are assumed to be unknown. For this problem of primary importance in statistical and sequential learning theory, we derive a variant of the Bayesian Online Change Point Detector proposed by \cite{fearnhead2007line} which is easier to analyze than the original version while keeping its powerful message-passing algorithm. We provide a non-asymptotic analysis of the false-alarm rate and the detection delay that matches the existing lower-bound. We further provide the first explicit high-probability control of the detection delay for such approach. Experiments on synthetic and real-world data show that this proposal outperforms the state-of-art change-point detection strategy, namely the Improved Generalized Likelihood Ratio (Improved GLR) while compares favorably with the original Bayesian Online Change Point Detection strategy.
Réda Alami, Odalric-Ambrym Maillard, Raphaël Féraud
ICML3
2019 Decentralized Exploration in Multi-Armed Bandits
abstract
We consider the decentralized exploration problem: a set of players collaborate to identify the best arm by asynchronously interacting with the same stochastic environment. The objective is to insure privacy in the best arm identification problem between asynchronous, collaborative, and thrifty players. In the context of a digital service, we advocate that this decentralized approach allows a good balance between conflicting interests: the providers optimize their services, while protecting privacy of users and saving resources. We define the privacy level as the amount of information an adversary could infer by intercepting all the messages concerning a single user. We provide a generic algorithm DECENTRALIZED ELIMINATION, which uses any best arm identification algorithm as a subroutine. We prove that this algorithm insures privacy, with a low communication cost, and that in comparison to the lower bound of the best arm identification problem, its sample complexity suffers from a penalty depending on the inverse of the probability of the most frequent players. Then, thanks to the genericity of the approach, we extend the proposed algorithm to the non-stationary bandits. Finally, experiments illustrate and complete the analysis.
Raphaël Féraud, Réda Alami, Romain Laroche
ICML1
2018 Reinforcement Learning Algorithm Selection
Romain Laroche, Raphaël Féraud
ICLR (Poster)2
2018 Reinforcement Learning Techniques for Optimized Channel Hopping in IEEE 802.15.4-TSCH Networks
abstract
The Industrial Internet of Things (IIoT) faces multiple challenges to achieve high reliability, low-latency and low power consumption. The IEEE 802.15.4 Time-Slotted Channel Hopping (TSCH) protocol aims to address these issues by using frequency hopping to improve the transmission quality when coping with low-quality channels. However, an optimized transmission system should also try to favor the use of high-quality channels, which are unknown a priori. Hence reinforcement learning algorithms could be useful.
Hiba Dakdouk, Erika Tarazona, Réda Alami, Raphaël Féraud, Georgios Z. Papadopoulos, Patrick Maillé
MSWiM4
2017 Context Attentive Bandits: Contextual Bandit with Restricted Context
abstract
We consider a novel formulation of the multi-armed bandit model, which we call the contextual bandit with restricted context, where only a limited number of features can be accessed by the learner at every iteration. This novel formulation is motivated by different online problems arising in clinical trials, recommender systems and attention modeling.Herein, we adapt the standard multi-armed bandit algorithm known as Thompson Sampling to take advantage of our restricted context setting, and propose two novel algorithms, called the Thompson Sampling with Restricted Context (TSRC) and the Windows Thompson Sampling with Restricted Context (WTSRC), for handling stationary and nonstationary environments, respectively. Our empirical results demonstrate advantages of the proposed approaches on several real-life datasets.
Djallel Bouneffouf 0001, Irina Rish, Guillermo A. Cecchi, Raphaël Féraud
IJCAI4
2017 Selection of learning experts
abstract
The contextual bandits can be viewed as a generalization of online classification models, where only the chosen class is observed. The selection of learning experts allows to find the best parametrization of an expert during its learning, within a set of predefined parameters, and reduces the bias of the hypothesis space, and hence improves the performances. As the contextual bandits learn, their performances tend to increase during time, and hence the choice of the best one implies to solve a non-stationary problem. We provide a theoretical framework to solve this difficult problem. A first approach to handle the selection of learning experts problem is to reduce it to stochastic problems: the experts are learned in parallel during a first phase, then they are explored and exploited in a second phase. A second approach models the selection of learning experts as an adversarial problem with a finite budget of contaminated rewards. We call this setting adversarial bandits with budget. Here, experts are learned, selected and exploited at the same time. When the budget of contaminated rewards is known, we propose and analyze a randomized variant of the algorithm Successive Elimination. When this budget is unknown, we analyze the algorithm EXP 3 for the proposed setting. We illustrate both approaches in the case where the learning experts are based on Bandit Forests initialized with different sets of parameters.
Robin Allesiardo, Raphaël Féraud
IJCNN2
2016 Random Forest for the Contextual Bandit Problem
abstract
To address the contextual bandit problem, we propose an online random forest algorithm. The analysis of the proposed algorithm is based on the sample complexity needed to find the optimal decision stump. Then, the decision stumps are recursively stacked in a random collection of decision trees, BANDIT FOREST. We show that the proposed algorithm is optimal up to logarithmic factors. The dependence of the sample complexity upon the number of contextual variables is logarithmic. The computational cost of the proposed algorithm with respect to the time horizon is linear. These analytical results allow the proposed algorithm to be efficient in real applications , where the number of events to process is huge, and where we expect that some contextual variables, chosen from a large set, have potentially non-linear dependencies with the rewards. In the experiments done to illustrate the theoretical analysis, BANDIT FOREST obtain promising results in comparison with state-of-the-art algorithms.
Raphaël Féraud, Robin Allesiardo, Tanguy Urvoy, Fabrice Clérot
AISTATS1
2016 Multi-armed bandit problem with known trend
Djallel Bouneffouf 0001, Raphaël Féraud
Neurocomputing2
2015 EXP3 with drift detection for the switching bandit problem
abstract
The multi-armed bandit is a model of exploration and exploitation, where one must select, within a finite set of arms, the one which maximizes the cumulative reward up to the time horizon T. For the adversarial multi-armed bandit problem, where the sequence of rewards is chosen by an oblivious adversary, the notion of best arm during the time horizon is too restrictive for applications such as ad-serving, where the best ad could change during time range. In this paper, we consider a variant of the adversarial multi-armed bandit problem, where the time horizon is divided into unknown time periods within which rewards are drawn from stochastic distributions. During each time period, there is an optimal arm which may be different from the optimal arm at the previous time period. We present an algorithm taking advantage of the constant exploration of EXP3 to detect when the best arm changes. Its analysis shows that on a run divided into N periods where the best arm changes, the proposed algorithms achieves a regret in O(N √T log T).
Robin Allesiardo, Raphaël Féraud
DSAA2
2014 A Neural Networks Committee for the Contextual Bandit Problem
Robin Allesiardo, Raphaël Féraud, Djallel Bouneffouf 0001
ICONIP (1)2
2014 Contextual Bandit for Active Learning: Active Thompson Sampling
Djallel Bouneffouf 0001, Romain Laroche, Tanguy Urvoy, Raphaël Féraud, Robin Allesiardo
ICONIP (1)4
2013 Generic Exploration and K-armed Voting Bandits
abstract
We study a stochastic online learning scheme with partial feedback where the utility of decisions is only observable through an estimation of the environment parameters. We propose a generic pure-exploration algorithm, able to cope with various utility functions from multi-armed bandits settings to dueling bandits. The primary application of this setting is to offer a natural generalization of dueling bandits for situations where the environment parameters reflect the idiosyncratic preferences of a mixed crowd.
Tanguy Urvoy, Fabrice Clérot, Raphaël Féraud, Sami Naamane
ICML (2)3
2013 Exploration and exploitation of scratch games
Raphaël Féraud, Tanguy Urvoy
Mach. Learn.1
2010 Modelling Complex Data by Learning Which Variable to Construct
Françoise Fessant, Aurélie Le Cam, Marc Boullé, Raphaël Féraud
DaWak4
2008 Contact personalization using a score understanding method
abstract
This paper presents a method to interpret the output of a classification (or regression) model. The interpretation is based on two concepts: the variable importance and the value importance of the variable. Unlike most of the state of art interpretation methods, our approach allows the interpretation of the model output for every instance. Understanding the score given by a model for one instance can for example lead to an immediate decision in a customer relational management (CRM) system. Moreover the proposed method does not depend on a particular model and is therefore usable for any model or software used to produce the scores.
Vincent Lemaire 0001, Raphaël Féraud, Nicolas Voisine
IJCNN2
2006 Driven Forward Features Selection: A Comparative Study on Neural Networks
Vincent Lemaire 0001, Raphaël Féraud
ICONIP (2)2
2002 A methodology to explain neural network classification
Raphaël Féraud, Fabrice Clérot
Neural Networks1
2001 A Fast and Accurate Face Detector Based on Neural Networks
abstract
Detecting faces in images with complex backgrounds is a difficult task. Our approach, which obtains state of the art results, is based on a neural network model: the constrained generative model (CGM). Generative, since the goal of the learning process is to evaluate the probability that the model has generated the input data, and constrained since some counter-examples are used to increase the quality of the estimation performed by the model. To detect side view faces and to decrease the number of false alarms, a conditional mixture of networks is used. To decrease the computational time cost, a fast search algorithm is proposed. The level of performance reached, in terms of detection accuracy and processing time, allows us to apply this detector to a real world application: the indexing of images and videos.
Raphaël Féraud, Olivier Bernier, Jean-Emmanuel Viallet, Michel Collobert
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 A Fast and Accurate Face Detector for Indexation of Face Images
abstract
Detecting faces in images with complex backgrounds is a difficult task. Our approach, which obtains state-of-the-art results, is based on a generative neural network model: the constrained generative model (CGM). To detect side-view faces and to decrease the number of false alarms, a conditional mixture of networks is used. To decrease the computational time cost, a fast search algorithm is proposed. The level of performance reached, in terms of detection accuracy and processing time, allows us to apply this detector to a real-world application: the indexation of face images on the WWW.
Raphaël Féraud, Olivier Bernier, Jean-Emmanuel Viallet, Michel Collobert
FG1
2000 Kalman and Neural Network Approaches for the Control of a VP Bandwidth in an ATM Network
Raphaël Féraud, Fabrice Clérot, Jean-Louis Simon, Daniel Pallou, Cyril Labbé, Serge Martin
NETWORKING1
1998 Panorama: A What I See Is What I Want Contactless Visual Interface
Jean-Emmanuel Viallet, Michel Collobert, Raphaël Féraud, Olivier Bernier
FG3
1998 MULTRAK: A System for Automatic Multiperson Localization and Tracking in Real-Time
abstract
A real-time system is described for automatic detection and tracking of multiple persons, in the context of video-conferencing systems. This system, called MULTRAK (multiperson locating and tracking automatic kernel) is able to continuously detect and track the position of faces in its field of view. The heart of the system as a modular neural network based face detector, giving fast and accurate face detection.
Olivier Bernier, Michel Collobert, Raphaël Féraud, Vincent Lemaire 0001, Jean-Emmanuel Viallet, Daniel Collobert
ICIP (1)3
1997 A Conditional Mixture of Neural Networks for Face Detection, Applied to Locating and Tracking an Individual Speaker
Raphaël Féraud, Olivier Bernier, Jean-Emmanuel Viallet, Michel Collobert, Daniel Collobert
CAIP1
1997 Ensemble and Modular Approaches for Face Detection: A Comparison
Raphaël Féraud, Olivier Bernier
NIPS1
1997 A Constrained Generative Model Applied to Face Detection
Raphaël Féraud, Olivier Bernier, Daniel Collobert
Neural Process. Lett.1
1996 LISTEN: A System for Locating and Tracking Individual Speakers
abstract
Both visual and acoustical informations provide effective means of telecommunication between persons. In this context, the face is the most important part of the person both visually and acoustically. We describe how the cooperation of image and audio processing allows to track a person's face and to collect the audio information it produces. We present detection techniques of regions of interest (e.g. Moving regions of skin color), coupled with a neural network based face detector with a low false alarm rate, to locate and track faces. The system is connected to a nine microphone array adaptive beam forming which performs immediate beam forming. Visual and acoustical informations from the speaker face are thus obtained in real time.
Michel Collobert, Raphaël Féraud, G. Le Tourneur, Olivier Bernier, Jean-Emmanuel Viallet, Yannick Mahieux, Daniel Collobert
FG2