EDBT 2026 Demo / reviewers in the wild / expert
Sidney Givigi
dblp:03/6019 · also Sidney N. Givigi, Sidney Nascimento Givigi
· DBLP profile ↗
32ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0002-3829-3545ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 16 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 4 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reinforcement Learning Task Assignment for Multi-Pursuer Multi-Evader Reach-Avoid GamesabstractMulti-pursuer multi-evader (MPME) pursuit-evasion (PE) scenarios are common in real-world applications but pose significant computational challenges as the number of agents grows and dynamics become more complex. Hierarchical decomposition simplifies the problem by splitting it into high-level task assignment (TA) and low-level execution. TA, an NP-hard problem with an exponentially expanding solution space, is addressed here using a novel Reinforcement Learning (RL) approach with both static and dynamic TA in a $2\mathbb{D}$ MPME reach-avoid game. The RL method matches the performance of a combinatorial optimization approach with homogeneous pursuers and outperforms it with heterogeneous ones, demonstrating its effectiveness in complex settings. Asif Mahfuz, Kleber M. Cabral, Sidney Givigi |
SMC | 3 |
| 2025 | Can AlphaZero Master Human Concepts in Tic-Tac-Toe?abstractAlphaZero’s success in games like chess, Go, and shogi has raised questions about its learning behaviour and limitations. Its play often mirrors human reasoning, applying familiar and novel concepts alike. In this work, we apply AlphaZero to Tic-Tac-Toe to examine its ability to learn a simple rules-based hierarchy that models human understanding. Despite the game’s simplicity, AlphaZero exhibits counter-intuitive, suboptimal behaviour not typical of human players. However, we show that modifying the reward structure can enforce rule hierarchies in pure MCTS and supervised settings, and improves AlphaZero’s rule learning while reducing sub-optimal play. Anthony J. Marasco, Sidney Givigi |
SMC | 2 |
| 2025 | IPRPAS: A Dataset of Physical Adversarial Samples for Assessing Object Detection in Intelligent Vehicles
Mahdieh Safarzadehvahed, Mohammad Zulkernine, Paulo Ricardo Marques de Araujo, Sidney Givigi |
IEEE Internet Things J. | 4 |
| 2025 | Optimizing a continuous action learning automata (CALA) optimizer for training artificial neural networks
James Lindsay, Sidney Givigi |
Neural Comput. Appl. | 2 |
| 2025 | Altruism in Fuzzy Reinforcement LearningabstractWe propose using a genetic algorithm to select hyperparameters in multiagent reinforcement learning (MARL) settings. In particular, we look at this in the context of cooperation and altruism. We show through the use of three continuous space games, that certain algorithmic hyperparameters are better suited to allow to agents learn altruistic behaviors. The agents learn using fuzzy actor critic learning algorithms in either a hierarchical structure or a single actor critic policy. The genetic algorithm selects the discount factors, the reward weights, and the standard deviation of noise applied to actor during learning. The genetic algorithm uses a fitness function based on the ratio of successful tests the group of agents can pass after training. This automated selection of these specific hyperparameters show that they are important for cooperation and also not trivial to select. Rachel Haighton, Howard M. Schwartz, Sidney Givigi |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Combining Dense and Sparse Rewards to Improve Deep Reinforcement Learning Policies in Reach-Avoid Games with Faster Evaders in Two vs. One ScenariosabstractThis paper investigates a variation of the reach-avoid game, a multi-agent pursuit and evasion scenario applicable to aerial defense, with faster evaders. Using Deep Reinforcement Learning techniques, the study proposes a different reward function that combines dense (distance-based) and sparse (outcome-based) rewards. Focused on the defender’s perspective in aerial defense, this new reward function resulted in effective learned policies against faster evaders, outperforming traditional differential game and DRL strategies with dense-only rewards. Moreover, the learned policy demonstrated versatility across different instances of the problem, including changes in pursuer speeds and winning radii, illustrating its versatility in unseen situations during training. Jefferson Silveira, Kalena McCloskey, Camille Alain Rabbath, Craig Williams, Sidney Givigi |
CoDIT | 5 |
| 2024 | An adaptable fuzzy reinforcement learning method for non-stationary environmentsabstractHow do we know when a reinforcement learning policy needs to adapt? In non-stationary environments, agents must adapt and learn in environments that change dynamically. We propose a finite-horizon model-free solution using a hierarchical learning structure with fuzzy systems. The higher-level learning policy advises the lower-level policy when to start and stop learning based on the temporal differences calculated within the lower-level. Major differences in the temporal difference of each action produced by an agent may indicate environment change. This structure is tested with multi-agent differential games in both the cooperative and competitive aspect. Results show that this method is quick to notice and adapt the policy within relatively few learning episodes. Rachel Haighton, Amirhossein Asgharnia, Howard M. Schwartz, Sidney Givigi |
Neurocomputing | 4 |
| 2024 | Causal Reinforcement Learning in Iterated Prisoner's DilemmaabstractThe iterated prisoner’s dilemma (IPD) is an archetypal paradigm to model cooperation and has guided studies on social dilemmas. In this work, we develop a causal reinforcement learning (CRL) strategy in a PD game. An agent is designed to have an explicit causal representation of other agents playing strategies from the Axelrod tournament. The collection of policies is assembled in an ensemble RL to choose the best strategy. The agent is then tested against selected Axelrod tournament strategies as well as an adaptive agent trained using traditional RL. Results show that our agent is able to play against all other players and score higher while being adaptive in situations where the strategy of the other players’ changes. Furthermore, the decision taken by the agent can be explained in terms of the causal representation of the interactions. Based on the decision made by the agent, a human observer can understand the chosen strategy. Yosra Kazemi, Caroline Ponzoni Carvalho Chanel, Sidney Givigi |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | Cost-efficient blockchain application to secure data transmission in heterogeneous FANETsabstractTo guarantee the success of vehicular networks, it is essential to ensure that the communication process is reliable and safe from malicious actions and that the solution has low computational complexity and energy consumption. Therefore, the present work proposes a proof-of-concept solution to ensure crash fault-tolerant communication in emulated heterogeneous Flying Ad-Hoc Networks (FANETs) using the Proof of Elapsed Time (PoET) consensus algorithm through the Hyperledger Sawtooth blockchain framework. Daniel Paiva Fernandes, Jeremias B. Machado, Sidney Givigi |
CCNC | 3 |
| 2023 | Real-Time Fast Marching Tree for Mobile Robot Motion Planning in Dynamic EnvironmentsabstractThis paper proposes the Real-Time Fast Marching Tree (RT-FMT), a real-time planning algorithm that features local and global path generation, multiple-query planning, and dynamic obstacle avoidance. During the search, RT-FMT quickly looks for the global solution and, in the meantime, generates local paths that can be used by the robot to start execution faster. In addition, our algorithm constantly rewires the tree to keep branches from forming inside the dynamic obstacles and to maintain the tree root near the robot, which allows the tree to be reused multiple times for different goals. Our algorithm is based on the planners Fast Marching Tree (FMT*) and Real-time Rapidly-Exploring Random Tree (RT-RRT*). We show via simulations that RT-FMT outperforms RT- RRT* in both execution cost and arrival time, in most cases. Moreover, we also demonstrate via simulation that it is worthwhile taking the local path before the global path is available in order to reduce arrival time, even though there is a small possibility of taking an inferior path. Jefferson Silveira, Kleber M. Cabral, Sidney Givigi, Joshua A. Marshall |
ICRA | 3 |
| 2022 | Reinforcement Learning-based Dynamic Resource Allocation For Grant-Free AccessabstractCellular networks have evolved to deliver high-speed broadband services to support the requirements of IoT applications, which demand high speed, low latency, and massive capacity. A primary market goal is to provide support for ultra-reliable low latency communication (URLLC). URLLC requires sub-milliseconds-level latencies as defined by the third generation partnership project (3GPP). One of the promising technologies to achieve the aforementioned specifications is grant-free (GF) access for uplink resources. The GF scheme enables the user equipment (UE) to transmit data over pre-allocated resources which reduces communication latency. This paper proposes an intelligent Reinforcement Learning (RL) based allocator of grants trained via Deep Q-Learning. The experimental results show effect of the number of UEs in the network, and the percentage of unstable UEs on the speed of the RL agent's convergence. Mariam Elsayem, Hatem Abou-Zeid, Ali Afana, Sidney Givigi |
GLOBECOM | 4 |
| 2021 | Evolutionary Inherited Neuromodulated Neurocontrollers with Objective Weighted RankingabstractIn the physical world, individuals compete against others within their own population or separately evolving populations. Robotic agents will soon face this coevolutionary and adversarial reality. Many challenges not encountered in the evolution of single populations are encountered in adversarial coevolution. These challenges can render single evolutionary approaches either less effective or ineffective. Most problems have a fundamental goal, but also feature secondary desirable objectives. Here, evader agents in the pursuit-evasion game that do not elude capture are effectively useless. Secondly, when applied to evolve real robots online in a coevolutionary context, non-optimal robots can be erratic and cause damage. Objective hierarchy can be used to define the importance of each objective, and promote quicker optimization of primary objectives such as evasion. The Evolutionary Inherited Neuromodulated Neurocontroller (EINN) method incorporates objective weighted ranking (OWR), a novel objective hierarchy method that promotes optimization of the primary objective in simultaneous multi-objective optimization. EINN is compared to the previously demonstrated Lamarckian-inherited Neuromodulated MultiObjective Evolutionary Neurocontroller (LNMOEN), and shown to be effective in a single evolutionary context. Ian Showalter, Howard M. Schwartz, Sidney Givigi |
CEC | 3 |
| 2021 | The Behavioural and Topological Effects of Measurement Noise on Evolutionary NeurocontrollersabstractThe disparity in performance between simulated and real systems is a major problem in robotics and other fields. Simulating measurement noise is one method of reducing these performance differences. Here, we examine the effect of measurement noise on the behaviour and topology of evolved neuromodulated neurocontrollers applied to control evader agents in a pursuit-evasion game. Measurement noise in the form of a zero-mean, normally distributed random signal is applied to the evader’s radar range and angle signals. The results indicate that increasing the levels of measurement noise increases the number of generations required to evolve fit agents. Noise at the neurocontroller outputs is of lesser amplitude than that at the inputs, suggesting a low-pass filtering operation. When levels of measurement noise different to those with which they were evolved were applied to the neurocontrollers, greater amplitude in the measurement noise signal increased the average length of time required to capture the evader. When the level of measurement noise was changed during evolution, after a few generations of further evolution, the neurocontrollers were able to adapt to both increases and decreases in the amount of noise. The evolutionary neurocontrollers are robust to high levels of measurement noise and can adapt to large changes in noise amplitude. This suggests that the neurocontrollers will be robust when used in the field on real robots, and that they may be a good solution to bridging the gap between simulation and reality. Ian Showalter, Howard M. Schwartz, Sidney Givigi |
SMC | 3 |
| 2020 | Trust in Multi-Vehicle Systems Using MDP Control StrategiesabstractThis paper proposes a protocol that ensures trust between two vehicles in a multi-vehicle system. Trust is the implicit assessment that another vehicle will follow a predetermined strategy. The communication is done through a channel and the quantity of information transferred is guaranteed to be small. For privacy, the channel can be encrypted, but the message can only be decoded if the vehicles know the control strategy being followed. The protocol is implemented for a problem of two Unmanned Aerial Vehicles (UAVs) trying to find a target in a maze. The control strategy is implemented using Markov Decision Processes (MDPs). Simulations of the protocol demonstrate that communication is received and decoded by the teammates without explicitly revealing the tactics being used. Jean-Alexis Delamer, Sidney Givigi |
SMC | 2 |
| 2020 | FPGA-Based Design for Real-Time Crack Detection Based on Particle FilterabstractDue to the related hazards, costly down-time, and detection inconsistencies associated with manual visual inspection for cracks in structures, there has been an emergence of real-time systems capable of conducting inspections. Advanced robotic systems have been used for scanning structures located in remote areas or that pose significant hazards to personnel. However, due to their inherent resource limitations, the current solution is to transfer all applicable sensor data to a ground station where detection will occur at a later time, thereby preventing real-time decision based on the results. To allow on-board decision making, in this article, a crack detection particle filter is optimized for parallel computation and implemented onto an Field-programmable gate array (FPGA). This article shows that an FPGA holds distinct tradeoffs between computational speed, energy consumption, and physical footprint compared to that of traditional CPU designs, allowing for it to be an ideal system for autonomous applications. Tim Chisholm, Romulo Gonçalves Lins, Sidney Givigi |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Convolutional Neural Networks as Asymmetric Volterra Models Based on Generalized Orthonormal Basis FunctionsabstractThis paper introduces a convolutional neural network (CNN) approach to derive Volterra models of dynamical systems based on generalized orthonormal basis function (GOBF)-Volterra. The approach derives the parameters of the model through a CNN and the neural network's learned weights represent the poles of a system. Simulation results show that the parameters of the system can be exactly recovered when no noise is applied. Furthermore, when noise is present, the errors in the parameters are very small for both the linear and nonlinear cases. Finally, the approach is used to identify the model of a quadcopter using data from actual flight tests. Comparisons with previous works demonstrate that CNNs can be satisfactorily used for the identification of dynamical systems. Jeremias B. Machado, Sidney Givigi |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Impulse Response Modeling of Dynamical Systemswith Convolutional Neural NetworksabstractThis paper introduces a convolutional neural network approach to derive the impulse response model of a dynamical system. The structure of the neural network is such that the learned weights are the parameters of the impulse response model. For a linear approximation, a finite impulse response representation is used. The size of the network is defined by the number of regressors in the model. For a nonlinear approximation, a second-order Volterra series model is used. Simulation results show that the parameters of the system can be perfectly recovered when no noise is applied to the system. Furthermore, when noise is applied to the system, the errors in the parameters are very small for both the linear and nonlinear cases. Finally, the approach is used to identify the model of a quadcopter using data from actual flight tests. The mean squared error for the output of the model as well as for the model parameters are shown to be small demonstrating the efficiency of the approach. Jeremias B. Machado, Sidney Givigi |
IJCNN | 2 |
| 2018 | Solving Home Robotics Challenges with Game Theory and Machine LearningabstractThe home environment introduces many challenges for robots to operate, for the home is much more unstructured than the environments in which industrial or commercial robots are found. This paper looks at a path planning problem and how to dynamically tune PID controllers that are to be used in the home environment. The method used for tuning is Learning Automata, specifically Finite Action Learning Automata for the prediction of the presence of people and a game of Continuous Action Learning Automata for the derivation of PID controllers. Results show that the proposed method efficiently derives better controllers for path planning when compared to a PID controller derived with a classical method. Furthermore, the method used to find acceptable waypoints shows that the robot is able to approximate the location of people in a home-care application. James Lindsay, Sidney Givigi |
SMC | 2 |
| 2017 | A Q-Learning Approach to Flocking With UAVs in a Stochastic EnvironmentabstractIn the past two decades, unmanned aerial vehicles (UAVs) have demonstrated their efficacy in supporting both military and civilian applications, where tasks can be dull, dirty, dangerous, or simply too costly with conventional methods. Many of the applications contain tasks that can be executed in parallel, hence the natural progression is to deploy multiple UAVs working together as a force multiplier. However, to do so requires autonomous coordination among the UAVs, similar to swarming behaviors seen in animals and insects. This paper looks at flocking with small fixed-wing UAVs in the context of a model-free reinforcement learning problem. In particular, Peng's Q(λ) with a variable learning rate is employed by the followers to learn a control policy that facilitates flocking in a leader-follower topology. The problem is structured as a Markov decision process, where the agents are modeled as small fixed-wing UAVs that experience stochasticity due to disturbances such as winds and control noises, as well as weight and balance issues. Learned policies are compared to ones solved using stochastic optimal control (i.e., dynamic programming) by evaluating the average cost incurred during flight according to a cost function. Simulation results demonstrate the feasibility of the proposed learning approach at enabling agents to learn how to flock in a leader-follower topology, while operating in a nonstationary stochastic environment. Shao-Ming Hung, Sidney Givigi |
IEEE Trans. Cybern. | 2 |
| 2016 | Towards human-robot interaction: A framing effect experimentabstractDecision making is a critical issue for humans operating unmanned vehicles. However, it is well admitted that many cognitive biases affect human judgments, leading to suboptimal or irrational decisions. The framing effect is a typical cognitive bias causing people to react differently depending on the context, the probability of the outcomes and how the problem is presented (loss vs. gain). There is a need to better understand the effects of these biases in operational contexts to optimize human-robot interactions. We therefore conducted an experiment involving a framing paradigm in a search and rescue mission (earthquake) and in a Mars rock sampling mission. We manipulated the framing (positive vs. negative) and the probability of the outcomes. Our findings revealed that the way the problem was presented (positively or negatively framed) and the emotional commitment (saving lives vs. collecting the good rock) statistically affected the choices made by the human operators. Paulo E. U. de Souza, Caroline Ponzoni Carvalho Chanel, Frédéric Dehais, Sidney Givigi |
SMC | 4 |
| 2015 | A Dyna-Q (Lambda) Approach to Flocking with Fixed-Wing UAVs in a Stochastic EnvironmentabstractUnmanned Aerial Vehicles (UAVs) have demonstrated their efficacy in supporting both military and civilian applications, many of which contain tasks that are parallel in nature, and can benefit from cooperation in terms of effectiveness. One of the fundamental challenges of multi-UAV systems is autonomous team coordination. This paper looks at flocking with small fixed-wing UAVs in the context of a model-free reinforcement learning problem. Dyna-Q ( ) with a variable learning rate is employed by the agents to learn a control policy that facilitates flocking in a leader-follower topology while operating in a stochastic environment. Simulation results demonstrate the followers learning and adapting their policies to non-stationary stochastic environments. Shao-Ming Hung, Sidney Givigi, Aboelmagd Noureldin |
SMC | 2 |
| 2015 | A Novel Machine Vision Approach Applied for Autonomous Robotics NavigationabstractMachine vision is widely used in many applications of engineering. In this paper, a new approach of machine vision is proposed for autonomous robotics navigation. The proposed methodology uses a sequence of images to locate homologous points and rebuild the objects in the robot surroundings. Using this approach a robot is able to measure an obstacle in order to avoid a collision as well as measure its velocity in maneuver using only information provided by cameras. Experimental results validate the application of the proposed method for autonomous robots applications. Romulo Gonçalves Lins, Sidney Givigi, Shao-Ming Hung, Aboelmagd Noureldin |
SMC | 2 |
| 2014 | Multiple-model Q-learning for stochastic reinforcement delaysabstractThe main contribution of this work is a novel machine reinforcement learning algorithm for problems where a Poissonian stochastic time delay is present in the agent's reinforcement signal. Despite the presence of the reinforcement noise, the algorithm can craft a suitable control policy for the agent's environment. The novel approach can deal with reinforcements which may be received out of order in time or may even overlap, which was not previously considered in the literature. The proposed algorithm is simulated and its performance is compared to a standard Q-learning algorithm. Through simulation, the proposed method is found to improve the performance of a learning agent in an environment with Poissonian-type stochastically delayed rewards. Jeffrey S. Campbell, Sidney Givigi, Howard M. Schwartz |
SMC | 2 |
| 2014 | Encirclement of moving target using linear model predictive control via feedback linearizationabstractA team of Unmanned Aerial Vehicles (UAVs) is used for the dynamic encirclement of a moving target in simulation. The encirclement tactic is defined for the situation in which a target is isolated and surrounded by a group of UAVs. It may be employed by a team of UAVs to neutralize the target and restrict its movement. A combination of decentralized Linear Model Predictive Control (LMPC) and Feedback Linearization (FL) is implemented on the team of UAVs in order to accomplish dynamic encirclement around a moving target. The main contribution of this paper lays in the application of LMPC and FL to solve the problem of encirclement of a moving target using an autonomous team of UAVs in simulation. Ahmed T. Hafez, Mohamad Iskandarani, Sidney Givigi, Shahram Yousefi, Aboelmagd Noureldin, Alain Beaulieu |
SMC | 3 |
| 2014 | Robust and efficient multi-robot 3D mapping with octree based occupancy gridsabstractA technique for merging 3D octree based occupancy grid maps robust to error in transformation between map reference frames is proposed and implemented. Recent robotics applications require 3D representations of the environments in which robots are to operate. In many cases, such as Simultaneous Localization and Mapping (SLAM) and when a large environment is required to be mapped within a reasonable time constraint, it is not feasible for a single robot to map the entire portion of the environment required to be mapped. In these cases it is necessary for a team of robots to build maps independently and merge them into a single global map. The contribution of this work lies in the introduction of methods which use map data from commonly mapped portions of the environment with registration techniques such that maps may still be merged coherently despite erroneous relative transformations between maps. The results from this paper demonstrate that not only are octree occupancy grids a suitable representation for multi-robot 3D mapping, but that the proposed techniques for improving erroneous transformation estimates between map frames are valid. James P. Jessup, Sidney Givigi, Alain Beaulieu |
SMC | 2 |
| 2013 | Cooperative Exploration Using Potential GamesabstractIn this paper, we improve on a previously proposed algorithm for exploring a 2-D environment with multiple robots using Potential Games. A potential game is a type of game that results in coordinated behaviours amongst players. This is done by enforcing strict rules for each player in selecting an action from its action set. As part of this game, we define a potential function for the game that is meaningful in terms of achieving the greater objective of exploring a space. Furthermore, an objective function is assigned for each player from this potential function. We then create an algorithm for the exploration of an obstacle-filled bounded space, and demonstrate through simulation how it outperforms an uncoordinated algorithm by reducing the time needed to uncover the space. This algorithm is then improved by having robots predict the future positions of all other robots. George Philip, Howard M. Schwartz, Sidney Givigi |
SMC | 3 |
| 2012 | Map merging of Multi-Robot SLAM using Reinforcement LearningabstractUsing `Simultaneous Localization and Mapping' (SLAM), mobile robots can become truly autonomous in the exploration of their environment. However, once these environments becomes too large, Multi-Robot SLAM becomes a requirement. This paper will outline how a mobile robot should decide when best to merge its maps with another robot's upon rendezvous, as opposed to doing so immediately. This decision will be based on the current status of the mapping particle filters and the current status of the environment. Using Reinforcement Learning, a model can be established and then trained upon to determine a policy capable of deciding when best to merge. This will allow the robot to incur less error during a merge compared to simply merging immediately. This policy is trained and validated using simulated mobile robot datasets. Pierre Dinnissen, Sidney Givigi, Howard M. Schwartz |
SMC | 2 |
| 2012 | A game theoretical representation for the rendezvous problemabstractIn this paper, we explore the application of Transferable Utility games, which is one of the many different types of games that make up the class of cooperative games, to the rendezvous problem. We consider an environment with multiple robots, wherein the objective is for them to agree in a decentralized manner on a rendezvous point in a non holonomic environment. The utility functions to be used by the controllers is such that they are shown to converge to a solution that leads all the robots to the objective. The paper also offers a mathematical proof of convergence for the proposed algorithm. Finally, simulations and experiments confirm our results. James Lindsay, Sidney Givigi |
SMC | 2 |
| 2012 | An experimental validation of reinforcement learning applied to the position control of UAVsabstractIn this paper, we explore the application of Reinforcement Learning (RL) to the derivation of control laws for the flight control of an unmanned aerial vehicle (UAV). The controllers are derived off-line with a simulation and the solutions are ported to an actual aircraft. Experimental results showed that the controllers stabilize the quad-rotor during the path tracking as has been learned in the simulation. Sergio Ronaldo Barros dos Santos, Sidney Givigi, Cairo L. Nascimento Jr. |
SMC | 2 |
| 2011 | Policy Invariance under Reward Transformations for General-Sum Stochastic GamesabstractWe extend the potential-based shaping method from Markov decision processes to multi-player general-sum stochastic games. We prove that the Nash equilibria in a stochastic game remains unchanged after potential-based shaping is applied to the environment. The property of policy invariance provides a possible way of speeding convergence when learning to play a stochastic game. Xiaosong Lu, Howard M. Schwartz, Sidney Givigi |
J. Artif. Intell. Res. | 3 |
| 2009 | An Experimental Adaptive Fuzzy Controller for Differential GamesabstractIn this paper a reinforcement fuzzy learning scheme for robots playing a differential game is derived. A differential game may be considered a Markov decision process in continuous time, with continuous states and actions. The robots receive reinforcements from the environment after they take an action; and this reinforcement is then used to adapt a fuzzy controller that stores the experience accumulated by the robot. Every calculation is done in a physical system based on microcontrollers to control the movement of the robots and sensors to measure their position and angle in a 2D-plane. Filters are also implemented to approximate the derivatives of the states. Experiments of a pursuer-evader game are provided in order to show the feasibility of the technique. It should be noted, though, that the technique may also be used in a multi-game environment. Sidney Givigi, Howard M. Schwartz, Xiaosong Lu |
SMC | 1 |
| 2005 | A New Strategy for Designing Bidirectional Associative Memories
Gengsheng Zheng, Sidney Givigi, Weiyu Zheng |
ISNN (1) | 2 |