VLDB 2026 Research / reviewers in the wild / expert
Djallel Bouneffouf 0001
dblp:45/11240-1
· DBLP profile ↗
50ranked-venue papers
28as first author
24since 2021 · last 2026
0000-0003-3342-7513ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 21 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Armed Bandits Meet Large Language ModelsabstractBandit algorithms and Large Language Models (LLMs) have emerged as powerful tools in artificial intelligence, each addressing distinct yet complementary challenges in decision-making and natural language processing. This survey explores the synergistic potential between these two fields, highlighting how bandit algorithms can enhance the performance of LLMs and how LLMs, in turn, can provide novel insights for improving bandit-based decision-making. We first examine the role of bandit algorithms in optimizing LLM fine-tuning, prompt engineering, and adaptive response generation, focusing on their ability to balance exploration and exploitation in large-scale learning tasks. Subsequently, we explore how LLMs can augment bandit algorithms through advanced contextual understanding, dynamic adaptation, and improved policy selection using natural language reasoning. By providing a comprehensive review of existing research and identifying key challenges and opportunities, this survey aims to bridge the gap between bandit algorithms and LLMs, paving the way for innovative applications and interdisciplinary research in AI. Djallel Bouneffouf 0001, Raphaël Féraud |
AAAI | 1 |
| 2026 | The Shepherd Test: How Will Super Intelligent Agents Balance Care and Control in Asymmetric Relationships?abstractThis paper introduces the Shepherd Test, a new conceptual test for assessing the moral and relational dimensions of superintelligent artificial agents. The test is inspired by human interactions with animals, where ethical considerations about care, manipulation, and consumption arise in contexts of asymmetric power and self-preservation. We argue that AI crosses an important, and potentially dangerous, threshold of intelligence when it exhibits the ability to manipulate, nurture, and instrumentally use less intelligent agents, while also managing its own survival and expansion goals. This includes the ability to weigh moral trade-offs between self-interest and the well-being of subordinate agents. The Shepherd Test thus challenges traditional AI evaluation paradigms by emphasizing moral agency, hierarchical behavior, and complex decision-making under existential stakes. We argue that this shift is critical for advancing AI governance, particularly as AI systems become increasingly integrated into multi-agent environments. We conclude by identifying key research directions, including the development of simulation environments for testing moral behavior in AI, and the formalization of ethical manipulation within multi-agent systems. Djallel Bouneffouf 0001, Matthew Riemer, Kush R. Varshney |
AAAI | 1 |
| 2025 | Contextual Value AlignmentabstractDeveloping value-aligned agents is a complex undertaking and an ongoing challenge in the field of AI. Indeed, designing Large Language Models (LLMs) that can balance multiple possibly conflicting moral values based on the context is a problem of paramount importance. In this paper, we propose a system that performs contextual value alignment based on contextual aggregation of possible responses. This aggregation is achieved by integrating a subset of possible LLM responses that are best suited to a user's input while taking into account features extracted about the user's moral preferences. The proposed system trained using the Moral Integrity Corpus displays better alignment to human values than state-of-the-art baselines. Pierre L. Dognin, Jesus Rios, Ronny Luss, Prasanna Sattigeri, Miao Liu 0001, Inkit Padhi, Matthew Riemer, Manish Nagireddy, Kush R. Varshney, Djallel Bouneffouf 0001 |
ICASSP | 10 |
| 2025 | Evaluating the Prompt Steerability of Large Language ModelsabstractErik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy, Elizabeth M. Daly, Kush R. Varshney, Eitan Farchi, Pierre Dognin, Jesus Rios, Djallel Bouneffouf, Miao Liu, Prasanna Sattigeri. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy, Elizabeth Daly, Kush R. Varshney, Eitan Farchi, Pierre L. Dognin, Jesus Rios, Djallel Bouneffouf 0001, Miao Liu 0001, Prasanna Sattigeri |
NAACL (Long Papers) | 9 |
| 2025 | Sequential uncertainty quantification with contextual tensors for social targeting
Tsuyoshi Idé, Keerthiram Murugesan, Djallel Bouneffouf 0001, Naoki Abe |
Knowl. Inf. Syst. | 3 |
| 2024 | ComVas: Contextual Moral Values Alignment System
Inkit Padhi, Pierre L. Dognin, Jesus Rios, Ronny Luss, Swapnaja Achintalwar, Matthew Riemer, Miao Liu 0001, Prasanna Sattigeri, Manish Nagireddy, Kush R. Varshney, Djallel Bouneffouf 0001 |
IJCAI | 11 |
| 2024 | A Tutorial on Multi-Armed Bandit Applications for Large Language ModelsabstractThis tutorial offers a comprehensive guide on using multi-armed bandit (MAB) algorithms to improve Large Language Models (LLMs). As Natural Language Processing (NLP) tasks grow, efficient and adaptive language generation systems are increasingly needed. MAB algorithms, which balance exploration and exploitation under uncertainty, are promising for enhancing LLMs. Djallel Bouneffouf 0001, Raphaël Féraud |
KDD | 1 |
| 2024 | Interpolating Item and User Fairness in Multi-Sided RecommendationsabstractToday's online platforms heavily lean on algorithmic recommendations for bolstering user engagement and driving revenue. However, these recommendations can impact multiple stakeholders simultaneously---the platform, items (sellers), and users (customers)---each with their unique objectives, making it difficult to find the right middle ground that accommodates all stakeholders. To address this, we introduce a novel fair recommendation framework, Problem (FAIR), that flexibly balances multi-stakeholder interests via a constrained optimization formulation. We next explore Problem (FAIR) in a dynamic online setting where data uncertainty further adds complexity, and propose a low-regret algorithm FORM that concurrently performs real-time learning and fair recommendations, two tasks that are often at odds. Via both theoretical analysis and a numerical case study on real-world data, we demonstrate the efficacy of our framework and method in maintaining platform revenue while ensuring desired levels of fairness for both items and users. Qinyi Chen, Jason Cheuk Nam Liang, Negin Golrezaei, Djallel Bouneffouf 0001 |
NeurIPS | 4 |
| 2023 | Question Answering System with Sparse and Noisy FeedbackabstractThe rise of personal assistants has made question answering a very popular mechanism for user-system interaction. In Question Answering System, implicit feedbacks can be easily observed (user clicking in the link given by the QA system), but they are noisy. However, receiving an explicit feedback on the quality of the response just given is rare but more valuable. Motivated by a practical need in Question Answering System of processing these two types of rewards, this paper investigates and proposes a new stochastic multi-armed bandit model in which each action has a noisy reward and a sparse reward. We studied this problem in the contextual bandit settings, and proposed and analyzed efficient algorithms that are based on the LINUCB frameworks. Our algorithms are verified by empirical studies on various reward distributions and a real-world dataset and application. Djallel Bouneffouf 0001, Oznur Alkan, Raphaël Féraud, Baihan Lin |
ICASSP | 1 |
| 2023 | Dialogue System with Missing ObservationabstractWithin the domain of dialogue, the ability to orchestrate multiple independently trained dialogue agents to create a unified system is of particular importance. Where we define orchestration as the task of selecting a subset of skills which most appropriately answer a user input using features extracted from both the user input and the individual skills. In this work, we study the task of online dialogue orchestration where the user feedback associated with the dialogue agent may not always be observed. In order to address the missing feedback setting, we propose to combine the attentive contextual bandit approach with an unsupervised learning mechanism such as clustering. By leveraging clustering to estimate missing reward, we are able to learn from each incoming event, even those with missing rewards. Promising empirical results are obtained on proprietary conversational datasets. Djallel Bouneffouf 0001, Mayank Agarwal, Irina Rish |
ICASSP | 1 |
| 2023 | SupervisorBot: NLP-Annotated Real-Time Recommendations of Psychotherapy Treatment Strategies with Deep Reinforcement LearningabstractWe present a novel recommendation system designed to provide real-time treatment strategies to therapists during psychotherapy sessions. Our system utilizes a turn-level rating mechanism that forecasts the therapeutic outcome by calculating a similarity score between the profound representation of a scoring inventory and the patient's current spoken sentence. By transcribing and segmenting the continuous audio stream into patient and therapist turns, our system conducts immediate evaluation of their therapeutic working alliance. The resulting dialogue pairs, along with their computed working alliance ratings, are then utilized in a deep reinforcement learning recommendation system. In this system, the sessions are treated as users, while the topics are treated as items. To showcase the system's effectiveness, we not only evaluate its performance using an existing dataset of psychotherapy sessions but also demonstrate its practicality through a web app. Through this demo, we aim to provide a tangible and engaging experience of our recommendation system in action. Baihan Lin, Guillermo A. Cecchi, Djallel Bouneffouf 0001 |
IJCAI | 3 |
| 2023 | Non-Stationary Bandits with Auto-Regressive Temporal DependencyabstractTraditional multi-armed bandit (MAB) frameworks, predominantly examined under stochastic or adversarial settings, often overlook the temporal dynamics inherent in many real-world applications such as recommendation systems and online advertising. This paper introduces a novel non-stationary MAB framework that captures the temporal structure of these real-world dynamics through an auto-regressive (AR) reward structure. We propose an algorithm that integrates two key mechanisms: (i) an alternation mechanism adept at leveraging temporal dependencies to dynamically balance exploration and exploitation, and (ii) a restarting mechanism designed to discard out-of-date information. Our algorithm achieves a regret upper bound that nearly matches the lower bound, with regret measured against a robust dynamic benchmark. Finally, via a real-world case study on tourism demand prediction, we demonstrate both the efficacy of our algorithm and the broader applicability of our techniques to more complex, rapidly evolving time series. Qinyi Chen, Negin Golrezaei, Djallel Bouneffouf 0001 |
NeurIPS | 3 |
| 2022 | Bandit Limited Discrepancy Search and Application to Machine Learning Pipeline OptimizationabstractOptimizing a machine learning (ML) pipeline has been an important topic of AI and ML. Despite recent progress, pipeline optimization remains a challenging problem, due to potentially many combinations to consider as well as slow training and validation. We present the BLDS algorithm for optimized algorithm selection (ML operations) in a fixed ML pipeline structure. BLDS performs multi-fidelity optimization for selecting ML algorithms trained with smaller computational overhead, while controlling its pipeline search based on multi-armed bandit and limited discrepancy search. Our experiments on well-known classification benchmarks show that BLDS is superior to competing algorithms. We also combine BLDS with hyperparameter optimization, empirically showing the advantage of BLDS. Akihiro Kishimoto, Djallel Bouneffouf 0001, Radu Marinescu 0002, Parikshit Ram, Ambrish Rawat, Martin Wistuba, Paulito P. Palmes, Adi Botea |
AAAI | 2 |
| 2022 | Optimal Epidemic Control as a Contextual Combinatorial Bandit with BudgetabstractIn light of the COVID-19 pandemic, it is an open challenge and critical practical problem to find a optimal way to dynamically prescribe the best policies that balance both the governmental resources and epidemic control in different countries and regions. To solve this multi-dimensional tradeoff of exploitation and exploration, we formulate this technical challenge as a contextual combinatorial bandit problem that jointly optimizes a multi-criteria reward function. Given the historical daily cases in a region and the past intervention plans in place, the agent should generate useful intervention plans that policy makers can implement in real time to minimizing both the number of daily COVID-19 cases and the stringency of the recommended interventions. We prove this concept with simulations of multiple realistic policy making scenarios and demonstrate a clear advantage in providing a pareto optimal solution in the epidemic intervention problem.1 Baihan Lin, Djallel Bouneffouf 0001 |
FUZZ-IEEE | 2 |
| 2022 | Learning to Generate Image Source-Agnostic Universal Adversarial PerturbationsabstractAdversarial perturbations are critical for certifying the robustness of deep learning models. A ``universal adversarial perturbation'' (UAP) can simultaneously attack multiple images, and thus offers a more unified threat model, obviating an image-wise attack algorithm. However, the existing UAP generator is underdeveloped when images are drawn from different image sources (e.g., with different image resolutions). Towards an authentic universality across image sources, we take a novel view of UAP generation as a customized instance of ``few-shot learning'', which leverages bilevel optimization and learning-to-optimize (L2O) techniques for UAP generation with improved attack success rate (ASR). We begin by considering the popular model agnostic meta-learning (MAML) framework to meta-learn a UAP generator. However, we see that the MAML framework does not directly offer the universal attack across image sources, requiring us to integrate it with another meta-learning framework of L2O. The resulting scheme for meta-learning a UAP generator (i) has better performance (50% higher ASR) than baselines such as Projected Gradient Descent, (ii) has better performance (37% faster) than the vanilla L2O and MAML frameworks (when applicable), and (iii) is able to simultaneously handle UAP generation for different victim models and data sources. Pu Zhao 0001, Parikshit Ram, Songtao Lu, Yuguang Yao, Djallel Bouneffouf 0001, Xue Lin 0001, Sijia Liu 0001 |
IJCAI | 5 |
| 2022 | Linear Upper Confident Bound with Missing Reward: Online Learning with Less DataabstractWe consider a novel variant of the contextual bandit problem (i.e., the multi-armed bandit with side-information, or context, available to a decision-maker) where the reward associated with each context-based decision may not always be observed (“missing rewards”). This new problem is motivated by certain online settings including clinical trial and ad recommendation applications. In order to address the missing rewards setting, we propose to combine the standard contextual bandit approach with an unsupervised learning mechanism such as clustering. Unlike standard contextual bandit methods, by leveraging clustering to estimate missing reward, we are able to learn from each incoming event, even those with missing rewards. Promising empirical results are obtained on several real-life datasets. Djallel Bouneffouf 0001, Sohini Upadhyay, Yasaman Khazaeni |
IJCNN | 1 |
| 2022 | Predicting Human Decision Making with LSTMabstractUnlike traditional time series, the action sequences of human decision making usually involve many cognitive processes such as beliefs, desires, intentions and theory of mind, i.e. what others are thinking. This makes predicting human decision making challenging to be treated agnostically to the underlying psychological mechanisms. We propose to use a recurrent neural network architecture based on long short-term memory networks (LSTM) to predict the time series of the actions taken by the human subjects at each step of their decision making, the first application of such methods in this research domain. In this study, we collate the human data from 8 published literature of the Iterated Prisoner's Dilemma comprising 168,386 individual decisions and postprocess them into 8,257 behavioral trajectories of 9 actions each for both players. Similarly, we collate 617 trajectories of 95 actions from 10 different published studies of Iowa Gambling Task experiments with healthy human subjects. We train our prediction networks on the behavioral data from these published psychological experiments of human decision making, and demonstrate a clear advantage over the state-of-the-art methods in predicting human decision making trajectories in both single-agent scenarios such as the Iowa Gambling Task and multi-agent scenarios such as the Iterated Prisoner's Dilemma. In the prediction, we observe that the weights of the top performers tends to have a wider distribution, and a bigger bias in the LSTM networks, which suggests possible interpretations for the distribution of strategies adopted by each group.11The data and codes to reproduce the empirical results can be accessed and reproduced at https://github.com/doerlbh/HumanLST M. A extended version of this work is to appear in [1]. Baihan Lin, Djallel Bouneffouf 0001, Guillermo A. Cecchi |
IJCNN | 2 |
| 2022 | Online Learning in Iterated Prisoner's Dilemma to Mimic Human Behavior
Baihan Lin, Djallel Bouneffouf 0001, Guillermo A. Cecchi |
PRICAI (3) | 2 |
| 2022 | Linearizing contextual bandits with latent state dynamicsabstractIn many real-world applications of multi-armed bandit problems, both rewards and contexts are often influenced by confounding latent variables which evolve stochastically over time. While the observed contexts and rewards are nonlinearly related, we show that prior knowledge of latent causal structure can be used to reduce the problem to the linear bandit setting. We develop two algorithms, Latent Linear Thompson Sampling (L2TS) and Latent Linear UCB (L2UCB), which use online EM algorithms for hidden Markov models to learn the latent transition model and maintain a posterior belief over the latent state, and then use the resulting posteriors as context features in a linear bandit problem. We upper bound the error in reward estimation in the presence of a dynamical latent state, and derive a novel problem-dependent regret bound for linear Thompson sampling with non-stationarity and unconstrained reward distributions, which we apply to L2TS under certain conditions. Finally, we demonstrate the superiority of our algorithms over related bandit algorithms through experiments. Elliot Nelson, Debarun Bhattacharjya, Tian Gao 0007, Miao Liu 0001, Djallel Bouneffouf 0001, Pascal Poupart |
UAI | 5 |
| 2021 | Corrupted Contextual Bandits: Online Learning with Corrupted ContextabstractWe consider a novel variant of the contextual bandit problem (i.e., the multi-armed bandit with side-information, or context, available to a decision-maker) where the context used at each decision may be corrupted ("useless context"). This new problem is motivated by certain on-line settings including clinical trial and ad recommendation applications. In order to address the corrupted-context setting, we propose to combine the standard contextual bandit approach with a classical multi-armed bandit mechanism. Unlike standard contextual bandit methods, we are able to learn from all iteration, even those with corrupted context, by improving the computing of the expectation for each arm. Promising empirical results are obtained on several real-life datasets. Djallel Bouneffouf 0001 |
ICASSP | 1 |
| 2021 | Online Hyper-Parameter Tuning for the Contextual BanditabstractWe study here the problem of learning the exploration exploitation trad-off in the contextual bandit problem with linear reward function setting. In the traditional algorithms that solve the contextual bandit problem, the exploration is a parameter that is tuned by the user. However, our proposed algorithm learn to choose the right exploration parameters in an online manner based on the observed context, and the immediate reward received for the chosen action. We have presented here two algorithms that uses a bandit to find the optimal exploration of the contextual bandit algorithm, which we hope is the first step toward the automation of the multi-armed bandit algorithm. Djallel Bouneffouf 0001, Emmanuelle Claeys |
ICASSP | 1 |
| 2021 | Toward Skills Dialog Orchestration with Online LearningabstractBuilding multi-domain AI agents is a challenging task and an open problem in the area of AI. Within the domain of dialog, the ability to orchestrate multiple independently trained dialog agents, or skills, to create a unified system is of particular significance. In this work, we study the task of online posterior dialog orchestration, where we define posterior orchestration as the task of selecting a subset of skills which most appropriately answer a user input using features extracted from both the user input and the individual skills. To account for the various costs associated with extracting skill features, we consider online posterior orchestration under a skill execution budget. We formalize this setting as Context Attentive Bandit with Observations (CABO), a variant of context attentive bandits, and evaluate it on proprietary conversational datasets. Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Mayank Agarwal, Yasaman Khazaeni, Irina Rish |
ICASSP | 1 |
| 2021 | Double-Linear Thompson Sampling for Context-Attentive BanditsabstractIn this paper, we analyze and extend an online learning frame-work known as Context-Attentive Bandit, motivated by various practical applications, from medical diagnosis to dialog systems, where due to observation costs only a small subset of a potentially large number of context variables can be observed at each iteration; however, the agent has a freedom to choose which variables to observe. We derive a novel algorithm, called Context-Attentive Thompson Sampling (CATS), which builds upon the Linear Thompson Sampling approach, adapting it to Context-Attentive Bandit setting. We provide a theoretical regret analysis and an extensive empirical evaluation demonstrating advantages of the proposed approach over several baseline methods on a variety of real-life datasets. Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Yasaman Khazaeni, Irina Rish |
ICASSP | 1 |
| 2021 | Toward Optimal Solution for the Context-Attentive Bandit ProblemabstractIn various recommender system applications, from medical diagnosis to dialog systems, due to observation costs only a small subset of a potentially large number of context variables can be observed at each iteration; however, the agent has a freedom to choose which variables to observe. In this paper, we analyze and extend an online learning framework known as Context-Attentive Bandit, We derive a novel algorithm, called Context-Attentive Thompson Sampling (CATS), which builds upon the Linear Thompson Sampling approach, adapting it to Context-Attentive Bandit setting. We provide a theoretical regret analysis and an extensive empirical evaluation demonstrating advantages of the proposed approach over several baseline methods on a variety of real-life datasets. Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Irina Rish, Yasaman Khazaeni |
IJCAI | 1 |
| 2020 | An ADMM Based Framework for AutoML Pipeline ConfigurationabstractWe study the AutoML problem of automatically configuring machine learning pipelines by jointly selecting algorithms and their appropriate hyper-parameters for all steps in supervised learning pipelines. This black-box (gradient-free) optimization with mixed integer & continuous variables is a challenging problem. We propose a novel AutoML scheme by leveraging the alternating direction method of multipliers (ADMM). The proposed framework is able to (i) decompose the optimization problem into easier sub-problems that have a reduced number of variables and circumvent the challenge of mixed variable categories, and (ii) incorporate black-box constraints alongside the black-box optimization objective. We empirically evaluate the flexibility (in utilizing existing AutoML techniques), effectiveness (against open source AutoML toolkits), and unique capability (of executing AutoML with practically motivated black-box constraints) of our proposed scheme on a collection of binary classification data sets from UCI ML & OpenML repositories. We observe that on an average our framework provides significant gains in comparison to other AutoML frameworks (Auto-sklearn & TPOT), highlighting the practical advantages of this framework. Sijia Liu 0001, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf 0001, Gregory Bramble, Horst Samulowitz, Dakuo Wang, Andrew Conn 0001, Alexander G. Gray |
AAAI | 4 |
| 2020 | Data Augmentation for Discrimination Prevention and Bias DisambiguationabstractMachine learning models are prone to biased decisions due to biases in the datasets they are trained on. In this paper, we introduce a novel data augmentation technique to create a fairer dataset for model training that could also lend itself to understanding the type of bias existing in the dataset i.e. if bias arises from a lack of representation for a particular group (sampling bias) or if it arises because of human bias reflected in the labels (prejudice based bias). Given a dataset involving a protected attribute with a privileged and unprivileged group, we create an "ideal world'' dataset: for every data sample, we create a new sample having the same features (except the protected attribute(s)) and label as the original sample but with the opposite protected attribute value. The synthetic data points are sorted in order of their proximity to the original training distribution and added successively to the real dataset to create intermediate datasets. We theoretically show that two different notions of fairness: statistical parity difference (independence) and average odds difference (separation) always change in the same direction using such an augmentation. We also show submodularity of the proposed fairness-aware augmentation approach that enables an efficient greedy algorithm. We empirically study the effect of training models on the intermediate datasets and show that this technique reduces the two bias measures while keeping the accuracy nearly constant for three datasets. We then discuss the implications of this study on the disambiguation of sample bias and prejudice based bias and discuss how pre-processing techniques should be evaluated in general. The proposed method can be used by policy makers who want to use unbiased datasets to train machine learning models for their applications to add a subset of synthetic points to an extent that they are comfortable with to mitigate unwanted bias. Shubham Sharma 0002, Jesús M. Ríos Aliaga, Djallel Bouneffouf 0001, Vinod Muthusamy, Kush R. Varshney |
AIES | 4 |
| 2020 | Survey on Applications of Multi-Armed and Contextual BanditsabstractIn recent years, the multi-armed bandit (MAB) framework has attracted a lot of attention in various applications, from recommender systems and information retrieval to healthcare and finance. This success is due to its stellar performance combined with attractive properties, such as learning from less feedback. The multiarmed bandit field is currently experiencing a renaissance, as novel problem settings and algorithms motivated by various practical applications are being introduced, building on top of the classical bandit problem. This article aims to provide a comprehensive review of top recent developments in multiple real-life applications of the multi-armed bandit. Specifically, we introduce a taxonomy of common MAB-based applications and summarize the state-of-the-art for each of those domains. Furthermore, we identify important current trends and provide new perspectives pertaining to the future of this burgeoning field. Djallel Bouneffouf 0001, Irina Rish, Charu C. Aggarwal |
CEC | 1 |
| 2020 | Survey on Automated End-to-End Data Science?abstractData science is labor-intensive and human experts are scarce but heavily involved in every aspect of it. This makes data science time consuming and restricted to experts with the resulting quality heavily dependent on their experience and skills. To make data science more accessible and scalable, we need its democratization. Automated Data Science (AutoDS) is aimed towards that goal and is emerging as an important research and business topic. We introduce and define the AutoDS challenge, followed by a proposal of a general AutoDS framework that covers existing approaches but also provides guidance for the development of new methods. We categorize and review the existing literature from multiple aspects of the problem setup and employed techniques. Then we provide several views on how AI could succeed in automating end-to-end AutoDS. We hope this survey can serve as insightful guideline for the AutoDS field and provide inspiration for future research. Djallel Bouneffouf 0001, Charu C. Aggarwal, Thanh Hoang, Udayan Khurana, Horst Samulowitz, Beat Buesser, Sijia Liu 0001, Tejaswini Pedapati, Parikshit Ram, Ambrish Rawat, Martin Wistuba, Alexander G. Gray |
IJCNN | 1 |
| 2019 | Incorporating Behavioral Constraints in Online AI SystemsabstractAI systems that learn through reward feedback about the actions they take are increasingly deployed in domains that have significant impact on our daily life. However, in many cases the online rewards should not be the only guiding criteria, as there are additional constraints and/or priorities imposed by regulations, values, preferences, or ethical principles. We detail a novel online agent that learns a set of behavioral constraints by observation and uses these learned constraints as a guide when making decisions in an online setting while still being reactive to reward feedback. To define this agent, we propose to adopt a novel extension to the classical contextual multi-armed bandit setting and we provide a new algorithm called Behavior Constrained Thompson Sampling (BCTS) that allows for online learning while obeying exogenous constraints. Our agent learns a constrained policy that implements the observed behavioral constraints demonstrated by a teacher agent, and then uses this constrained policy to guide the reward-based online exploration and exploitation. We characterize the upper bound on the expected regret of the contextual bandit algorithm that underlies our agent and provide a case study with real world data in two application domains. Our experiments show that the designed agent is able to act within the set of behavior constraints without significantly degrading its overall reward performance. Avinash Balakrishnan, Djallel Bouneffouf 0001, Nicholas Mattei, Francesca Rossi 0001 |
AAAI | 2 |
| 2019 | Scalable Recollections for Continual Lifelong LearningabstractGiven the recent success of Deep Learning applied to a variety of single tasks, it is natural to consider more human-realistic settings. Perhaps the most difficult of these settings is that of continual lifelong learning, where the model must learn online over a continuous stream of non-stationary data. A successful continual lifelong learning system must have three key capabilities: it must learn and adapt over time, it must not forget what it has learned, and it must be efficient in both training time and memory. Recent techniques have focused their efforts primarily on the first two capabilities while questions of efficiency remain largely unexplored. In this paper, we consider the problem of efficient and effective storage of experiences over very large time-frames. In particular we consider the case where typical experiences are O(n) bits and memories are limited to O(k) bits for k Matthew Riemer, Tim Klinger, Djallel Bouneffouf 0001, Michele Franceschini |
AAAI | 3 |
| 2019 | Beyond Backprop: Online Alternating Minimization with Auxiliary VariablesabstractDespite significant recent advances in deep neural networks, training them remains a challenge due to the highly non-convex nature of the objective function. State-of-the-art methods rely on error backpropagation, which suffers from several well-known issues, such as vanishing and exploding gradients, inability to handle non-differentiable nonlinearities and to parallelize weight-updates across layers, and biological implausibility. These limitations continue to motivate exploration of alternative training algorithms, including several recently proposed auxiliary-variable methods which break the complex nested objective function into local subproblems. However, those techniques are mainly offline (batch), which limits their applicability to extremely large datasets, as well as to online, continual or reinforcement learning. The main contribution of our work is a novel online (stochastic/mini-batch) alternating minimization (AM) approach for training deep neural networks, together with the first theoretical convergence guarantees for AM in stochastic settings and promising empirical results on a variety of architectures and datasets. Anna Choromanska, Benjamin Cowen, Sadhana Kumaravel, Ronny Luss, Mattia Rigotti, Irina Rish, Paolo Diachille, Viatcheslav Gurev, Brian Kingsbury, Ravi Tejwani, Djallel Bouneffouf 0001 |
ICML | 11 |
| 2019 | Optimal Exploitation of Clustering and History Information in Multi-armed BanditabstractWe consider the stochastic multi-armed bandit problem and the contextual bandit problem with historical observations and pre-clustered arms. The historical observations can contain any number of instances for each arm, and the pre-clustering information is a fixed clustering of arms provided as part of the input. We develop a variety of algorithms which incorporate this offline information effectively during the online exploration phase and derive their regret bounds. In particular, we develop the META algorithm which effectively hedges between two other algorithms: one which uses both historical observations and clustering, and another which uses only the historical observations. The former outperforms the latter when the clustering quality is good, and vice-versa. Extensive experiments on synthetic and real world datasets on Warafin drug dosage and web server selectionfor latency minimization validate our theoretical insights and demonstrate that META is a robust strategy for optimally exploiting the pre-clustering information. Djallel Bouneffouf 0001, Srinivasan Parthasarathy 0002, Horst Samulowitz, Martin Wistuba |
IJCAI | 1 |
| 2019 | Split Q Learning: Reinforcement Learning with Two-Stream RewardsabstractDrawing an inspiration from behavioral studies of human decision making, we propose here a general parametric framework for a reinforcement learning problem, which extends the standard Q-learning approach to incorporate a two-stream framework of reward processing with biases biologically associated with several neurological and psychiatric conditions, including Parkinson's and Alzheimer's diseases, attention-deficit/hyperactivity disorder (ADHD), addiction, and chronic pain. For AI community, the development of agents that react differently to different types of rewards can enable us to understand a wide spectrum of multi-agent interactions in complex real-world socioeconomic systems. Moreover, from the behavioral modeling perspective, our parametric framework can be viewed as a first step towards a unifying computational model capturing reward processing abnormalities across multiple mental conditions and user preferences in long-term recommendation systems. Baihan Lin, Djallel Bouneffouf 0001, Guillermo A. Cecchi |
IJCAI | 2 |
| 2019 | Teaching AI Agents Ethical Values Using Reinforcement Learning and Policy OrchestrationabstractAutonomous cyber-physical agents play an increasingly large role in our lives. To ensure that they behave in ways aligned with the values of society, we must develop techniques that allow these agents to not only maximize their reward in an environment, but also to learn and follow the implicit constraints of society. We detail a novel approach that uses inverse reinforcement learning to learn a set of unspecified constraints from demonstrations and reinforcement learning to learn to maximize environmental rewards. A contextual bandit-based orchestrator then picks between the two policies: constraint-based and environment reward-based. The contextual bandit orchestrator allows the agent to mix policies in novel ways, taking the best actions from either a reward-maximizing or constrained policy. In addition, the orchestrator is transparent on which policy is being employed at each time step. We test our algorithms using Pac-Man and show that the agent is able to learn to act optimally, act within the demonstrated constraints, and mix these two functions in complex ways. Ritesh Noothigattu, Djallel Bouneffouf 0001, Nicholas Mattei, Rachita Chandra, Piyush Madan, Kush R. Varshney, Murray Campbell, Moninder Singh, Francesca Rossi 0001 |
IJCAI | 2 |
| 2018 | Using Contextual Bandits with Behavioral Constraints for Constrained Online Movie RecommendationabstractAI systems that learn through reward feedback about the actions they take are increasingly deployed in domains that have significant impact on our daily life. In many cases the rewards should not be the only guiding criteria, as there are additional constraints and/or priorities imposed by regulations, values, preferences, or ethical principles. We detail a novel online system, based on an extension of the contextual bandits framework, that learns a set of behavioral constraints by observation and uses these constraints as a guide when making decisions in an online setting while still being reactive to reward feedback. In addition, our system can highlight features of the context which are more predicted to be more rewarding and/or are in line with the behavioral constraints. We demonstrate the system by building an interactive interface for an online movie recommendation agent and show that our system is able to act within a set of behavior constraints without significantly degrading overall performance. Avinash Balakrishnan, Djallel Bouneffouf 0001, Nicholas Mattei, Francesca Rossi 0001 |
IJCAI | 2 |
| 2018 | Eigenspectrum Shape Based Nyström SamplingabstractSpectral clustering has shown a superior performance in analyzing the cluster structure. However, its computational complexity limits its application in analyzing large-scale data. To address this problem, many low-rank matrix approximating algorithms are proposed, including the Nyström method - an approach with proven approximate error bounds. There are several algorithms that provide recipes to construct Nyström approximations with variable accuracies and computing times. This paper proposes a scalable Nyström-based clustering algorithm with a new sampling procedure, Centroid Minimum Sum of Squared Similarities (CMS3), and a heuristic on when to use it. Our heuristic depends on the eigenspectrum shape of the dataset, and yields competitive low-rank approximations in test datasets compared to the other state-of-the-art methods. Djallel Bouneffouf 0001 |
IJCNN | 1 |
| 2017 | Context Attentive Bandits: Contextual Bandit with Restricted ContextabstractWe consider a novel formulation of the multi-armed bandit model, which we call the contextual bandit with restricted context, where only a limited number of features can be accessed by the learner at every iteration. This novel formulation is motivated by different online problems arising in clinical trials, recommender systems and attention modeling.Herein, we adapt the standard multi-armed bandit algorithm known as Thompson Sampling to take advantage of our restricted context setting, and propose two novel algorithms, called the Thompson Sampling with Restricted Context (TSRC) and the Windows Thompson Sampling with Restricted Context (WTSRC), for handling stationary and nonstationary environments, respectively. Our empirical results demonstrate advantages of the proposed approaches on several real-life datasets. Djallel Bouneffouf 0001, Irina Rish, Guillermo A. Cecchi, Raphaël Féraud |
IJCAI | 1 |
| 2016 | Finite-time analysis of the multi-armed bandit problem with known trendabstractWe study here an a problem called “the multi-armed bandit problem with known trend”, where an agent knows the shape of the reward function of each arm but not its distribution. This problem is motivated by different real world tasks, where when an arm is sampled by the model the received reward change according to a known trend. By adapting the standard multi-armed bandit algorithms, we propose to study the regret upper bounds of three algorithms: the two first one assumes a stochastic model; and the last one is based on a Bayesian approach. Djallel Bouneffouf 0001 |
CEC | 1 |
| 2016 | Contextual bandit algorithm for risk-aware recommender systemsabstractContext-Aware Recommender Systems can be naturally modelled as an exploration/exploitation trade-off (exr/exp) problem, where the system has to choose between maximizing its expected rewards dealing with its current knowledge (exploitation) and learning more about the unknown user's preferences to improve its knowledge (exploration). This problem has been addressed by the reinforcement learning community but they do not consider the risk level of the current user's situation, where it may be dangerous to recommend items the user may not desire in her current situation if the risk level is high. We introduce in this paper an algorithm named R-UCB that considers the risk level of the user's situation to adaptively balance between exr and exp. The detailed analysis of the experimental results reveals several important discoveries in the exr/exp behaviour. Djallel Bouneffouf 0001 |
CEC | 1 |
| 2016 | Ensemble Minimum Sum of Squared Similarities sampling for Nyström-based spectral clusteringabstractSpectral clustering is a powerful approach for clustering, with applications across multiple disciplines, including bioinformatics. However, the way its computational complexity scales limits its application in analyzing large datasets. This complexity can be reduced using the Nyström method, which subsamples the input data in a way that preserves its representational diversity. There are different established strategies for subsampling, yet they may have performance limitations for certain complex datasets. This paper we propose an alternative to those methods, introducing a new sampling procedure called Ensemble Minimum Sum of Squared Similarities (EMS3). We further improve on this method by using weight mixtures in subsample selection, yielding more accurate low-rank approximations than existing ensemble Nyström methods. We also provide a theoretical analysis of the upper error bound of the EMS3 algorithm, and demonstrate its performance in comparison to the leading spectral clustering methods that use Nyström sampling. Djallel Bouneffouf 0001, Inanç Birol |
IJCNN | 1 |
| 2016 | Theoretical analysis of the Minimum Sum of Squared Similarities sampling for Nyström-based spectral clusteringabstractSpectral clustering has shown a superior performance in analyzing the cluster structure. However, the exponentially computational complexity limits its application in analyzing large-scale data. To tackle this problem, many low-rank matrix approximating algorithms are proposed, of which the Nyström method is an approach with proved lower approximate errors. The algorithms commonly combine two powerful techniques in machine learning: spectral clustering algorithms and Nyström methods commonly used to obtain good quality low rank approximations of large matrices. This paper proposes to analyze a scalable Nyström-based clustering algorithm with a Minimum Sum of Squared Similarities (MSSS) sampling procedure. We provide theoretical analysis of the performance of the algorithm MSSS and demonstrate its theoretical performance in comparison to the leading spectral clustering methods that use Nyström sampling. Djallel Bouneffouf 0001, Inanç Birol |
IJCNN | 1 |
| 2016 | Multi-armed bandit problem with known trend
Djallel Bouneffouf 0001, Raphaël Féraud |
Neurocomputing | 1 |
| 2015 | Sampling with Minimum Sum of Squared Similarities for Nystrom-Based Large Scale Spectral Clustering
Djallel Bouneffouf 0001, Inanç Birol |
IJCAI | 1 |
| 2014 | A Neural Networks Committee for the Contextual Bandit Problem
Robin Allesiardo, Raphaël Féraud, Djallel Bouneffouf 0001 |
ICONIP (1) | 3 |
| 2014 | Freshness-Aware Thompson Sampling
Djallel Bouneffouf 0001 |
ICONIP (3) | 1 |
| 2014 | Contextual Bandit for Active Learning: Active Thompson Sampling
Djallel Bouneffouf 0001, Romain Laroche, Tanguy Urvoy, Raphaël Féraud, Robin Allesiardo |
ICONIP (1) | 1 |
| 2013 | Risk-Aware Recommender Systems
Djallel Bouneffouf 0001, Amel Bouzeghoub, Alda Lopes Gançarski |
ICONIP (1) | 1 |
| 2013 | Contextual Bandits for Context-Based Information Retrieval
Djallel Bouneffouf 0001, Amel Bouzeghoub, Alda Lopes Gançarski |
ICONIP (2) | 1 |
| 2012 | A Contextual-Bandit Algorithm for Mobile Context-Aware Recommender System
Djallel Bouneffouf 0001, Amel Bouzeghoub, Alda Lopes Gançarski |
ICONIP (3) | 1 |
| 2012 | Hybrid-ε-greedy for Mobile Context-Aware Recommender System
Djallel Bouneffouf 0001, Amel Bouzeghoub, Alda Lopes Gançarski |
PAKDD (1) | 1 |