VLDB 2026 Research / reviewers in the wild / expert
Caroline Ponzoni Carvalho Chanel
dblp:40/9513 · also Caroline P. C. Chanel, Caroline P. Carvalho Chanel
· DBLP profile ↗
20ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-3578-4186ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An offline risk-aware policy selection method for Bayesian Markov decision processesabstractIn Offline Model Learning for Planning and in Offline Reinforcement Learning, the limited data set hinders the estimate of the Value function of the relative Markov Decision Process (MDP). Consequently, the performance of the obtained policy in the real world is bounded and possibly risky, especially when the deployment of a wrong policy can lead to catastrophic consequences. For this reason, several pathways are being followed with the scope of reducing the model error (or the distributional shift between the learned model and the true one) and, more broadly, obtaining risk-aware solutions with respect to model uncertainty. But when it comes to the final application which baseline should a practitioner choose? In an offline context where computational time is not an issue and robustness is the priority we propose Exploitation vs Caution (EvC), a paradigm that (1) elegantly incorporates model uncertainty abiding by the Bayesian formalism, and (2) selects the policy that maximizes a risk-aware objective over the Bayesian posterior between a fixed set of candidate policies provided, for instance, by the current baselines. We validate EvC with state-of-the-art approaches in different discrete, yet simple, environments offering a fair variety of MDP classes. In the tested scenarios EvC manages to select robust policies and hence stands out as a useful tool for practitioners that aim to apply offline planning and reinforcement learning solvers in the real world. Giorgio Angelotti, Nicolas Drougard, Caroline Ponzoni Carvalho Chanel |
Artif. Intell. | 3 |
| 2026 | An online user-centric brain-computer interface based on code-modulated visually evoked potentials and partially observable Markov decision processabstract: Reactive brain–computer interfaces (rBCIs) can deliver fast and reliable performance, yet the decision step — determining when to act based on neural evidence — remains an often overlooked component. In prior work, we proposed a decision-making framework based on partially observable Markov decision processes (POMDP). The present study moves this approach from offline validation to a real-time setting, integrating it into a code-modulated visual evoked potential (c-VEP) rBCI. Twelve healthy participants performed a five-class c-VEP control task across two sessions, each comprising a cued and a self-paced (Pinpad) task with a semi-dry electroencephalography (EEG) system. One session used a conventional accumulation-based decision strategy; the other employed the POMDP-based approach. In the POMDP condition, calibration data were collected during an engaging cued task, allowing the policy to be computed without interrupting user interaction. Across all tasks and conditions, participants achieved mean accuracies above 97%. In the self-paced task, the POMDP significantly reduced mean decoding time (1.55 s) compared to the accumulation baseline (1.97 s), while maintaining equivalent accuracy. This study provides the first online demonstration of a POMDP-based decision framework for rBCI, balancing speed and accuracy under the constrains of real-time operation . By removing the need for individual thresholds and integrating calibration into active use, the approach offers a flexible, user-centric pathway toward more user-centered and practical BCIs. Juan Jesús Torre Tresols, Caroline Ponzoni Carvalho Chanel, Frédéric Dehais |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Online Metrics to Enhance Human-Artificial Agent Collaboration Efficiency: A Narrative Literature ReviewabstractMachines have traditionally served as tools to fulfill human requirements; however, the rapid advancement of artificial intelligence has enabled the development of autonomous systems capable of functioning as fully integrated teammates. These agents can share information, assume roles, and execute tasks within collaborative environments. Effective Human–Artificial Agent collaboration, achieved through the integration of complementary cognitive and operational capabilities, has demonstrated improvements in overall team performance across multiple domains, including industrial robotics, healthcare, and augmented reality. Nevertheless, achieving both optimal performance and interaction fluency remains a significant challenge. Real-time monitoring of tasks, intentions, and constraints of human and artificial partners is still limited, and the application of quantifiable online metrics for this purpose is underexplored. This narrative review systematically examines online metrics derived from behavioral, physiological, and interaction-based approaches, discussing their potential to enhance adaptive mechanisms and optimize team fluency in H–AA collaboration. Adam H. M. Pinto, Christophe Antony Lounis, Mickaël Causse, Caroline Ponzoni Carvalho Chanel |
Int. J. Hum. Comput. Interact. | 4 |
| 2024 | SKATE : Successive Rank-based Task Assignment for Proactive Online PlanningabstractThe development of online applications for services such as package delivery, crowdsourcing, or taxi dispatching has caught the attention of the research community to the domain of online multi-agent multi-task allocation. In online service applications, tasks (or requests) to be performed arrive over time and need to be dynamically assigned to agents. Such planning problems are challenging because: (i) few or almost no information about future tasks is available for long-term reasoning; (ii) agent number, as well as, task number can be impressively high; and (iii) an efficient solution has to be reached in a limited amount of time. In this paper, we propose SKATE, a successive rank-based task assignment algorithm for online multi-agent planning. SKATE can be seen as a meta-heuristic approach which successively assigns a task to the best-ranked agent until all tasks have been assigned. We assessed the complexity of SKATE and showed it is cubic in the number of agents and tasks. To investigate how multi-agent multi-task assignment algorithms perform under a high number of agents and tasks, we compare three multi-task assignment methods in synthetic and real data benchmark environments: Integer Linear Programming (ILP), Genetic Algorithm (GA), and SKATE. In addition, a proactive approach is nested to all methods to determine near-future available agents (resources) using a receding-horizon. Based on the results obtained, we can argue that the classical ILP offers the better quality solutions when treating a low number of agents and tasks, i.e. low load despite the receding-horizon size, while it struggles to respect the time constraint for high load. SKATE performs better than the other methods in high load conditions, and even better when a variable receding-horizon is used. Déborah Conforto Nedelmann, Jérôme Lacan, Caroline Ponzoni Carvalho Chanel |
ICAPS | 3 |
| 2024 | Causal Reinforcement Learning in Iterated Prisoner's DilemmaabstractThe iterated prisoner’s dilemma (IPD) is an archetypal paradigm to model cooperation and has guided studies on social dilemmas. In this work, we develop a causal reinforcement learning (CRL) strategy in a PD game. An agent is designed to have an explicit causal representation of other agents playing strategies from the Axelrod tournament. The collection of policies is assembled in an ensemble RL to choose the best strategy. The agent is then tested against selected Axelrod tournament strategies as well as an adaptive agent trained using traditional RL. Results show that our agent is able to play against all other players and score higher while being adaptive in situations where the strategy of the other players’ changes. Furthermore, the decision taken by the agent can be explained in terms of the causal representation of the interactions. Based on the decision made by the agent, a human observer can understand the chosen strategy. Yosra Kazemi, Caroline Ponzoni Carvalho Chanel, Sidney Givigi |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Data Augmentation Through Expert-Guided Symmetry Detection to Improve Performance in Offline Reinforcement LearningabstractOffline estimation of the dynamical model of a Markov Decision Process (MDP) is a non-trivial task that greatly depends on the data available in the learning phase. Sometimes the dynamics of the model is invariant with respect to some transformations of the current state and action. Recent works showed that an expert-guided pipeline relying on Density Estimation methods as Deep Neural Network based Normalizing Flows effectively detects this structure in deterministic environments, both categorical and continuous-valued. The acquired knowledge can be exploited to augment the original data set, leading eventually to a reduction in the distributional shift between the true and the learned model. Such data augmentation technique can be exploited as a preliminary process to be executed before adopting an Offline Reinforcement Learning architecture, increasing its performance. In this work we extend the paradigm to also tackle non-deterministic MDPs, in particular, 1) we propose a detection threshold in categorical environments based on statistical distances, and 2) we show that the former results lead to a performance improvement when solving the learned MDP and then applying the optimized policy in the real environment. Giorgio Angelotti, Nicolas Drougard, Caroline Ponzoni Carvalho Chanel |
ICAART (2) | 3 |
| 2023 | Modular zk-rollup on-demand
Thomas Lavaur, Jonathan Detchart, Jérôme Lacan, Caroline Ponzoni Carvalho Chanel |
J. Netw. Comput. Appl. | 4 |
| 2022 | Expert-guided Symmetry Detection in Markov Decision ProcessesabstractLearning a Markov Decision Process (MDP) from a fixed batch of trajectories is a non-trivial task whose outcome's quality depends on both the amount and the diversity of the sampled regions of the state-action space. Yet, many MDPs are endowed with invariant reward and transition functions with respect to some transformations of the current state and action. Being able to detect and exploit these structures could benefit not only the learning of the MDP but also the computation of its subsequent optimal control policy. In this work we propose a paradigm, based on Density Estimation methods, that aims to detect the presence of some already supposed transformations of the state-action space for which the MDP dynamics is invariant. We tested the proposed approach in a discrete toroidal grid environment and in two notorious environments of OpenAI's Gym Learning Suite. The results demonstrate that the model distributional shift is reduced when the dataset is augmented with the data obtained by using the detected symmetries, allowing for a more thorough and data-efficient learning of the transition functions. Giorgio Angelotti, Nicolas Drougard, Caroline Ponzoni Carvalho Chanel |
ICAART (2) | 3 |
| 2022 | Towards a POMDP-based Control in Hybrid Brain-Computer InterfacesabstractBrain-Computer Interfaces (BCI) provide a unique communication channel between the brain and computer systems. After extensive research and implementation on ample fields of application, numerous challenges to assure reliable and quick data processing have resulted in the hybrid BCI (hBCI) paradigm, consisting on the combination of two BCI systems. However, not all challenges have been properly addressed (e.g. re-calibration, idle-state modelling, adaptive thresholds, etc) to allow hBCI implementation outside of the lab. In this paper, we review electroencephalography based hBCI studies and state potential limitations. We propose a sequential decision-making framework based on Partially Observable Markov Decision Process (POMDP) to design and to control hBCI systems. The POMDP framework is an excellent candidate to deal with the limitations raised above. To illustrate our opinion, an example of architecture using a POMDP-based hBCI control system is provided, and future directions are discussed. We believe this framework will encourage research efforts to provide relevant means to combine information from BCI systems and push BCI out of the laboratory. Juan Jesús Torre Tresols, Caroline Ponzoni Carvalho Chanel, Frédéric Dehais |
SMC | 2 |
| 2021 | AI can fool us humans, but not at the psycho-physiological level: a hyperscanning and physiological synchrony studyabstractThis study aims at investigating the neural and physiological correlates of human-human and human-AI interactions under ecological settings. We designed a scenario in which a ground controller had to guide his/her pilot to reach a location. We also implemented a Controller-Bot and a Pilot-Bot using AI techniques to behave like real human operators. The cooperation between controllers and pilots were either genuine (‘Coop scenarios’ – four missions), explicitly notified as pilot-Bot and controller-Bot interactions (‘No coop scenarios’ – two missions), or with no notification that they were actually collaborating with their AI counterparts (‘fake coop scenarios’ – two missions). Sixteen participants (8 dyads) equipped with EEG and ECG took part in this experiment. Our findings disclosed that Human-Human dyads exhibited similar performance to Human-Bots dyads whether the human participants were aware that they were playing with a bot or not. Our participants declared that they did not realize they were playing with an AI in the fake cooperation condition. These findings indicate that 1) humans can be fooled by AI, and that 2) humans can behave in a natural way with AI. Interestingly enough, our analyses revealed that the cardiac activity of controllers and pilots was more synchronized when they were collaborating together than when they were playing with AI (being aware or not). Similarly, EEG analyses disclosed a higher cerebral efficiency and connectivity between the two brains when teammates were interacting together than when cooperating with AI. Frédéric Dehais, G. Vergotte, Nicolas Drougard, G. Ferraro, Bertille Somon, Caroline Ponzoni Carvalho Chanel, Raphaëlle N. Roy |
SMC | 6 |
| 2020 | Entropy-Based Adaptive Exploit-Explore Coefficient for Monte-Carlo Path PlanningabstractEfficient path planning for autonomous vehicles in cluttered environments is a challenging sequential decision-making problem under uncertainty. In this context, this paper implements a partially observable stochastic shortest path (PO-SSP) planning problem for autonomous urban navigation of Unmanned Aerial Vehicles (UAVs). To solve this planning problem, the POMCP-GO algorithm is used, which is goal oriented variant of POMCP, one of the fastest online state-of-the-art solvers for partially observable environments based on Monte Carlo Planning. This algorithm relies on the Upper Confidence Bounds (UCB1) algorithm as action selection strategy. UCB1 depends on an exploration constant typically adjusted empirically. Its best value varies significantly between planning problems, and hence, an exhaustive search to find the most suitable value is required. This exhaustive search applied to a complex path planning problem may be extremely time consuming. Moreover, considering real applications where online planning is needed, this extensive search is not suitable. Thereby this paper explores the use of an adaptive exploration coefficient for action selection during planning. Monte-Carlo value backup approximation is also applied which empirically demonstrates to accelerate the policy value convergence. Simulation results show that the use of the adaptive exploration co- \nefficient within a user-defined interval achieves better convergence and success rates when compared with most hand-tuned fixed coefficients in said interval, although never achieving the same results as the best fixed coefficient. Therefore, a compromise must be made between the desired quality of the results and the time one is willing to spend on the exhaustive search for the best coefficient value before planning. Ana Raquel Carmo, Jean-Alexis Delamer, Yoko Watanabe, Rodrigo M. M. Ventura, Caroline Ponzoni Carvalho Chanel |
ECAI | 5 |
| 2020 | Predicting Human Operator's Decisions Based on Prospect TheoryabstractAbstract The aim of this work is to predict human operator’s (HO) decisions in a specific operational context, such as a cooperative human-robot mission, by approximating his/her utility function based on prospect theory (PT). To this aim, a within-subject experiment was designed in which the HO has to decide with limited time and incomplete information. This experiment also involved a framing effect paradigm, a typical cognitive bias causing people to react differently depending on the context. Such an experiment allowed to acquire data concerning the HO’s decisions in two different mission scenarios: search and rescue and Mars rock sampling. The framing was manipulated (e.g. positive vs. negative) and the probability of the outcomes causing people to react differently depending on the context. Statistical results observed for this experiment supported the hypothesis that the way the problem was presented (positively or negatively framed) and the emotional commitment affected the HO’s decisions. Thus, based on the collected data, the present work is willed to propose: (i) a formal approximation of the HO’s utility function founded on the prospect theory and (ii) a model used to predict the HO’s decisions based on the economics approach of multi-dimensional consumption bundle and PT. The obtained results, in terms of utility function fit and prediction accuracy, are promising and show that similar modeling and prediction method should be taken into account when an intelligent cybernetic system drives human–robot interaction. The advantage of predicting the HO’s decision, in this operational context, is to anticipate his/her decision, given the way a question is framed to the HO. Such a predictor lays the foundation for the development of a decision-making system capable of choosing how to present the information to the operator while expecting to align his/her decision with the given operational guideline. Paulo E. U. de Souza, Caroline Ponzoni Carvalho Chanel, Melody Mailliez, Frédéric Dehais |
Interact. Comput. | 2 |
| 2019 | Online ECG-based Features for Cognitive Load AssessmentabstractThis study was concerned with the development and testing of online cognitive-load monitoring methods by means of a working-memory experiment using electrocardiogram (ECG) analyses for future applications in mixed-initiative human-machine interaction (HMI). To this end, we first identified potentially reliable cognitive-workload-related cardiac metrics and algorithms for online processing. We then compared our online results to those conventionally obtained with state-of-the-art offline methods. Finally, we evaluated the possibility of classifying low versus high working-memory load using different classification algorithms. Our results show that both offline and online methods reliably estimate the workload associated with a multi-level working-memory task at the group level, whether it is with the heart rhythm or the heart rate variation (standard deviation of the RR interval). Moreover, we found significant working-memory load classification accuracy using both two-dimensional linear discriminant analyses (LDA) or a support vector machine (SVM). We hence argue that our online algorithm is reliable enough to provide online electrocardiographic metrics as a tool for real-life workload evaluation and can be a valuable feature for mixed-initiative systems. Caroline Ponzoni Carvalho Chanel, Matthew D. Wilson, Sebastien Scannella |
SMC | 1 |
| 2018 | Human-Agent Interaction Model Learning based on CrowdsourcingabstractMissions involving humans interacting with automated systems become increasingly common. Due to the non-deterministic behavior of the human and possibly high risk of failing due to human factors, such an integrated system should react smartly by adapting its behavior when necessary. A promise avenue to design an efficient interaction-driven system is the mixed-initiative paradigm. In this context, this paper proposes a method to learn the model of a mixed-initiative human-robot mission. The first step to set up a reliable model is to acquire enough data. For this aim a crowdsourcing campaign was conducted and learning algorithms were trained on the collected data in order to model the human-robot mission and to optimize a supervision policy with a Markov Decision Process (MDP). This model takes into account the actions of the human operator during the interaction as well as the state of the robot and the mission. Once such a model has been learned, the supervision strategy can be optimized according to a criterion representing the goal of the mission. In this paper, the supervision strategy concerns the robot's operating mode. Simulations based on the MDP model show that planning under uncertainty solvers can be used to adapt robot's mode according to the state of the human-robot system. The optimization of the robot's operation mode seems to be able to improve the team's performance. The dataset that comes from crowdsourcing is therefore a material that can be useful for research in human-machine interaction, that is why it has been made available on our web site. Jack-Antoine Charles, Caroline Ponzoni Carvalho Chanel, Corentin Chauffaut, Pascal Chauvin, Nicolas Drougard |
HAI | 2 |
| 2017 | Pre-stimulus antero-posterior EEG connectivity predicts performance in a UAV monitoring taskabstractLong monitoring tasks without regular actions, are becoming increasingly common from aircraft pilots to train conductors as these systems grow more automated. These task contexts are challenging for the human operator because they require inputs at irregular and highly interspaced moments even though these actions are often critical. It has been shown that such conditions lead to divided and distracted attentional states which in turn reduce the processing of external stimuli (e.g. alarms) and may lead to miss critical events. In this study we explored to which extent it is possible to predict an operator's behavioural performance in a Unmanned Aerial Vehicle (UAV) monitoring task using electroencephalographic (EEG) activity. More specifically we investigated the relevance of large-scale EEG connectivity for performance prediction by correlating relative coherence with reaction times (RT). We show that long-range EEG relative coherence, i.e. between occipital and frontal electrodes, is significantly correlated with RT and that different frequency bands exhibit opposite effects. More specifically we observed that coherence between occipital and frontal electrodes was: negatively correlated with RT at 6Hz (θ band), more coherence leading to better performance, and positively correlated with RT at 8Hz (lower α band), more coherence leading to worse performance. Our results suggest that EEG connectivity measures could be useful in predicting an operator's attentional state and her/his performances in ecological settings. Hence these features could potentially be used in a neuro-adaptive interface to improve operator-system interaction and safety in critical systems. Mehdi Senoussi, Kevin J. Verdiere, Angela Bovo, Caroline Ponzoni Carvalho Chanel, Frédéric Dehais, Raphaëlle N. Roy |
SMC | 4 |
| 2016 | Considering human's non-deterministic behavior and his availability state when designing a collaborative human-robots systemabstractThe objective of this study is to design a human-robots system that takes into account the non-deterministic nature of the human operator's behavior. Such a system is implemented in a proof of concept scenario relying on a (MO)MDP decision framework that takes advantage of an eye-tracker device to estimate the cognitive availability of the human operator, and, some human operator's inputs to deduce where he is focusing his attention. An experiment was conducted with ten participants interacting with a team of autonomous vehicles in a Search & Rescue scenario. Our results demonstrate the advantages of considering the cognitive availability of a human operator in such a complex context and also the interest of using such a decisional framework that can formally integrate the non-deterministic outcomes which model the human behavior. Thibault Gateau, Caroline Ponzoni Carvalho Chanel, Mai-Huy Le, Frédéric Dehais |
IROS | 2 |
| 2016 | Towards human-robot interaction: A framing effect experimentabstractDecision making is a critical issue for humans operating unmanned vehicles. However, it is well admitted that many cognitive biases affect human judgments, leading to suboptimal or irrational decisions. The framing effect is a typical cognitive bias causing people to react differently depending on the context, the probability of the outcomes and how the problem is presented (loss vs. gain). There is a need to better understand the effects of these biases in operational contexts to optimize human-robot interactions. We therefore conducted an experiment involving a framing paradigm in a search and rescue mission (earthquake) and in a Mars rock sampling mission. We manipulated the framing (positive vs. negative) and the probability of the outcomes. Our findings revealed that the way the problem was presented (positively or negatively framed) and the emotional commitment (saving lives vs. collecting the good rock) statistically affected the choices made by the human operators. Paulo E. U. de Souza, Caroline Ponzoni Carvalho Chanel, Frédéric Dehais, Sidney Givigi |
SMC | 2 |
| 2015 | MOMDP-Based Target Search Mission Taking into Account the Human Operator's Cognitive StateabstractThis study discusses the application of sequential decision making under uncertainty and mixed observability in a mixed-initiative robotic target search application. In such a robotic mission, two agents, a ground robot and a human operator, must collaborate to reach a common goal using, each in turn, their recognized skills. The originality of the work relies in considering that the human operator is not a providential agent when the robot fails. Using the data from previous experiments, a Mixed Observability Markov Decision Process (MOMDP) model was designed, which allows to consider aleatory failure events and the partial observable human operator's state while planning for a long-term horizon. Results show that the collaborative system was in general able to successfully complete or terminate the mission, even when many simultaneous sensors, devices and operator failures happened. So, the mixed-initiative framework highlighted in this study shows the relevancy of taking into account the cognitive state of the operator, which permits to compute a policy for the sequential decision problem which prevents to re-planning when unexpected (but known) events occurs. Paulo E. U. de Souza, Caroline Ponzoni Carvalho Chanel, Frédéric Dehais |
ICTAI | 2 |
| 2013 | Multi-Target Detection and Recognition by UAVs Using Online POMDPsabstractThis paper tackles high-level decision-making techniques for robotic missions, which involve both active sensing and symbolic goal reaching, under uncertain probabilistic environments and strong time constraints. Our case study is a POMDP model of an online multi-target detection and recognition mission by an autonomous UAV. The POMDP model of the multi-target detection and recognition problem is generated online from a list of areas of interest, which are automatically extracted at the beginning of the flight from a coarse-grained high altitude observation of the scene. The POMDP observation model relies on a statistical abstraction of an image processing algorithm's output used to detect targets. As the POMDP problem cannot be known and thus optimized before the beginning of the flight, our main contribution is an "optimize-while-execute" algorithmic framework: it drives a POMDP sub-planner to optimize and execute the POMDP policy in parallel under action duration constraints. We present new results from real outdoor flights and SAIL simulations, which highlight both the benefits of using POMDPs in multi-target detection and recognition missions, and of our "optimize-while-execute" paradigm. Caroline Ponzoni Carvalho Chanel, Florent Teichteil-Königsbuch, Charles Lesire |
AAAI | 1 |
| 2013 | Properly Acting under Partial Observability with Action Feasibility Constraints
Caroline Ponzoni Carvalho Chanel, Florent Teichteil-Königsbuch |
ECML/PKDD (1) | 1 |