VLDB 2026 Research / reviewers in the wild / expert
Toshiharu Sugawara
dblp:78/1202
· DBLP profile ↗
98ranked-venue papers
10as first author
36since 2021 · last 2026
0000-0002-9271-4507ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 10 first-author · 31 since 2021Computer networks · 7Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph Orientation for Efficient Multi-Agent Pickup and Delivery Using Genetic Algorithm
Yuki Miyashita, Toshiharu Sugawara |
COMPSAC | 2 |
| 2026 | Language-Grounded Strategy-Following Multi-Agent Deep Reinforcement Learning for Controllability of Real-World Applications
Yoshinari Motokawa, Toshiharu Sugawara |
COMPSAC | 2 |
| 2026 | Transformer-Based Action Prediction for Efficient Mutual Modeling in Multi-Agent Reinforcement Learning
Kohei Suzuki, Toshiharu Sugawara |
ICAART (1) | 2 |
| 2026 | Block Stacking Problem by Group of Agents Using Particle Swarm Optimization
Kotaro Yamada, Stephen Raharja, Toshiharu Sugawara |
ICAART (1) | 3 |
| 2025 | Multi-Agent Path Finding Using Provisionally Booking Nodes for Pickup and Delivery Problems
Daiki Shimada, Yuki Miyashita, Toshiharu Sugawara |
ICAART (3) | 3 |
| 2025 | Enhanced Interpretability for Conditional Coordinated Behavior in Multi-Agent Reinforcement LearningabstractIn this study, we propose an enhanced distributed attention actor architecture after conditional attention (eDA6-X), a framework for model-free multi-agent reinforcement learning (MADRL), to improve the performance and interpretability of cooperative behaviors. Interpreting behaviors from MADRL is crucial for validating their use in real-world applications, given the diverse coordination structures in multi-agent systems (MAS). By incorporating a conditional module, eDA6-X allows agents to consider contextual and conditional states to determine their behaviors in dynamic MADRL environments. We evaluated the framework in a sequential object collection game, requiring agents to execute multi-step tasks and establish coordination based on role differences. Our results show that eDA6-X agents effectively use attention mechanisms to develop efficient coordination strategies, outperforming baseline methods. The analysis of cooperative behavior through standard, class, and positional attention heatmaps highlights eDA6-X’s enhanced capability to deliver clearer insights into coordination and decision-making processes in complex MAS. Yoshinari Motokawa, Toshiharu Sugawara |
IJCNN | 2 |
| 2025 | MARL-HE: An Improved Multi-agent Reinforcement Learning-based Pathfinding Method for Fire Evacuation GuidanceabstractIn the multi-agent pathfinding problem, an agent may have multiple possible destinations, each of which must be appropriately selected in terms of conflict avoidance and shortest path. For example, in an evacuation system when faced with a fire hazard in a building with multiple exits for evacuation, agents need to consider not only the shortest path but also the conflicts of other agents and the fire spread to respond appropriately. However, recent reinforcement-learning-based evacuation pathfinding methods only consider the scenarios of independent individual agents and ignore the influence of other agents and the spread of fire, limiting their application to evacuation. To address these issues, we propose a multi-agent reinforcement learning for pathfinding during hazard evacuation (MARL-HE), which is integrated into reinforcement learning with two novel components: an efficient artificial potential field adapter (APF-adapter) and policy-sharing-enhanced multi-agent proximity policy optimization (PS-MAPPO). The APF adapter utilizes an artificial potential field to eliminate the dynamic state transitions during hazard evacuation. PS-MAPPO leverages policy sharing by determining which other agents each agent has to share part of their policy model to accelerate training efficiency. To evaluate our method, we built a multi-agent fire evacuation guidance system (MAFEGS), a Unity-based fire hazard environment. The experimental results show that our MARL-HE method achieved a 7.02% higher evacuation success rate and 25.12% lower average evacuation number than MAPPO. MAFEGS can plan a safer and faster evacuation guidance path to keep individual safe and increase evacuation efficiency. The simulation environment and code can be downloaded from: https://github.com/ColaZhang22/PS_MARL_PF. Sun Cherry, Yosuke Fujisawa, Toshiharu Sugawara |
IJCNN | 6 |
| 2025 | 3PS: Periodical Partial Policy Sharing for Distributed Cooperative Multi-agent Reinforcement LearningabstractWe propose a framework for multi-agent reinforcement learning called periodic partial policy sharing (3PS) to balance the efficiency of learning and privacy concerns. Because of the negative impacts of instability on team members’ learning, some correct actions are later considered incorrect. To address this problem, current methods employ centralized training and decentralized execution (CTDE) using a shared experience buffer to train consistent policies for individual agents. However, in sensitive applications, agents are not allowed to share their experiences because of privacy concerns. In the proposed 3PS, agents periodically share only the parameters in their policies to mitigate training inefficiencies and prevent privacy breaches from sharing the experience datasets. Depending on the parameters with the associated importance weights of the individual policies, we introduce three sharing methods in 3PS: average 3PS (A-3PS), reward-scalability 3PS (R-3PS), and personalized 3PS (P-3PS). We evaluated these 3PSs using three backbones, QMIX, VDN, and IPPO, for two multi-agent cooperative tasks: SMAC and multi-agent path planning. The results show that 3PS outperforms the baseline methods in terms of the convergence rate, average reward, and win rate, while ofering better data privacy than centralized training. 3PS code is available at https://github.com/ColaZhang22/3PS . Toshiharu Sugawara |
KES | 4 |
| 2025 | Conflict-Free Multi-Agent Path Generation Using Ant Colony Optimization with Load BalancingabstractWe propose a method using ant colony optimization (ACO) with load-balancing techniques for multi-agent path finding problems. In automated warehouses and manufacturing facilities, multiple automated guided vehicles (AGVs) must efficiently perform transportation tasks without colliding with obstacles or other agents. Although several metaheuristic approaches, including ACO-based methods, address this challenge, conventional methods do not account for scenarios with many agents, and/or the planning efficiency and quality significantly decrease as the number of agents increases. The proposed method is inspired by conventional ACO-based approaches but differs in three aspects. First, we carefully define the collisions by considering the agent size and travel time. Second, we propose a tailored collision avoidance method for numerous agents by determining whether they should wait for other agents or generate detours. Finally, a directional load-balancing method that considers the movement direction is proposed for more effective ACO path generation. Our experiments demonstrate that the proposed method generates superior coordinated paths without increasing CPU time, even in high-density environments. Ryusuke Kazama, Toshiharu Sugawara |
SMC | 2 |
| 2024 | Scheduling and Negotiation Method for Double Synchronized Multi-Agent Pickup and Delivery Problem
Yuki Miyashita, Toshiharu Sugawara |
ICAART (1) | 2 |
| 2024 | Efficient Retraining for Continuous Operability Through Strategic DirectivesabstractWe introduce a method to ensure the controllability of agents in pretrained networks, anticipating changes in managerial requirements and environmental conditions. Advances in multi-agent deep reinforcement learning (MADRL) have fostered sophisticated cooperative behaviors in multi-agent systems in which agents share and execute various complex tasks. However, retraining MADRL systems to adapt to new conditions is costly, because many agents are involved. Our approach introduces several types of directives, termed destination channels (DCs), which allow agents to experience diverse coordination patterns during training without detailed instructions. When changes occur, the system manager assigns appropriate DCs to each agent to facilitate adaptation and maintain a continuous operation. We conducted experiments using object-collection games to evaluate our proposed method by comparing the number of objects recovered and the learning speed of our method with those of existing methods, an implicit quantile network (IQN), and a traditional deep Q-network (DQN). The experimental results demonstrate that agents using this method adapt more swiftly to environmental changes during retraining than baseline methods, sustaining performance without significant degradation. Gentoku Nakasone, Yoshinari Motokawa, Yuki Miyashita, Toshiharu Sugawara |
ICMLA | 4 |
| 2024 | Action Selection in Reinforcement Learning with Upgoing Policy UpdateabstractThis paper proposes the boosted generalized advantage estimation (BGAE), which extends the generalized advantage estimation (GAE) based on proximal policy optimization (PPO) algorithms with a general advantage estimator used in the Atari 2600 game environment. We utilize the preferable feature of the upgoing policy update (UPGO), which is a crucial algorithm behind the well-known multiagent reinforcement learning model, AlphaStar. The proposed method uses UPGO as an action selector to select good actions and amplify their importance. In addition, we discuss the overestimation of the UPGO algorithm in a standard environment. Unlike other approaches that increase model complexity through overestimation, our novel method effectively addresses the overestimation problem in the UPGO method, offering a streamlined solution in a standard learning environment. Implementing UPGO algorithms can improve the convergence speed of conventional PPO learning in the commonly used OpenAI-Gym Atari environment. Toshiharu Sugawara |
ICMLA | 2 |
| 2024 | Performance Improvement for UAV-Assisted Mobile Edge Computing with Multi-Agent Deep Reinforcement LearningabstractIn this paper, we propose a method for mobile edge computing (MEC) using unmanned aerial vehicles (UAVs) to enhance wireless connectivity in areas afflicted by natural disasters or obstructed by tall buildings. Despite the active research on deep reinforcement learning in recent years, challenges persist in its application to the optimization of cooperative and coordinated behavior among multi-agents. MEC aims to improve efficiency and fairness in offloading from user terminals (UTs) to UAVs while minimizing energy consumption. Therefore, UAV groups must operate independently within designated areas to facilitate connections between UTs and servers. We introduce multi-agent deep deterministic policy gradient to MEC and improve the reward design to achieve high-performance MEC with efficient cooperative behavior only using limited local data. Experimental results demonstrate that our approach significantly enhances MEC efficiency through effective cooperative behavior. Specifically, it offloads more tasks/applications to UAVs compared to baseline methods while reducing energy consumption per offloading. Kohei Suzuki, Toshiharu Sugawara |
INISTA | 2 |
| 2024 | Multi-Agent Reinforcement Learning with Clustering and Experience SharingabstractWe propose a training method for a heterogeneous multi-agent system to improve the learning efficiency in sparse-reward environments. Although extensive research on multi-agent deep reinforcement learning are conducted actively, these studies often assume that all agents are homogeneous to share/utilize learning parameters in their networks. Unfortunately, this is not always the case in real-world applications where heterogeneous autonomous agents, i.e., those with different capabilities and perspectives, must properly cooperate and coordinate with each other. In our learning method, which is an extension of the shared experience actor-critic (SEAC) for a heterogeneous agent environment, agents are classified depending on their features (such as trajectories of the observations, actions and received rewards) using variational autoencoder, and share their experience among agents within each cluster to train their individual agents for improving the learning efficiency in a sparse-reward environment. Our experimental evaluation shows that the proposed method is capable of more efficient cooperative/coordinated behaviors than the baselines while remaining the advantages of SEAC. Kaname Inokuchi, Toshiharu Sugawara |
KES | 2 |
| 2024 | Overlap-Free Pattern Formation Method Using Particle Swarm OptimizationabstractWe propose a method based on particle swarm optimization (PSO) for the multi-agent pattern formation problem (MAPFP), which allows agents to form formation patterns with a high completion rate without overlaps. MAPFP is a problem in which a large number of agents autonomously move to appropriate locations to fill in a target pattern, such as line drawings, letters, and geometric figure, in a cooperative manner. They are used in several applications such as route planning and object exploration. Although some studies consider the pattern-formation problem using swarm intelligence (SI), including PSO, they do not often consider conflicts such as collisions and overlaps of particles, and thus cannot be used in real-world applications. Our proposed method, based on PSO, an SI algorithm, allows agents to act independently, while avoiding overlaps, to form a pattern by introducing a local field of view to each particle and secreting pheromones to their neighbor points to attract other agents that have not yet found places to stay. Through experimental evaluation, we show that the proposed method not only eliminates overlaps during movement and unwanted overlaps on a target pattern but also improves the pattern completion rate compared to the baseline. Kotaro Yamada, Toshiharu Sugawara |
KES | 2 |
| 2024 | Fair Path Generation for Formation Control Combining Ant Colony and Particle Swarm Optimizations
Yoshie Suzuki, Toshiharu Sugawara |
KES-AMSTA | 2 |
| 2024 | Path Finding with Flexible Provisional Booking in Multi-agent Pickup and Delivery Problems
Daiki Shimada, Yuki Miyashita, Toshiharu Sugawara |
PRIMA | 3 |
| 2023 | User's Position-Dependent Strategies in Consumer-Generated Media with Monetary RewardsabstractNumerous forms of consumer-generated media (CGM), such as social networking services (SNS), are widely used. Their success relies on users' voluntary participation, often driven by psychological rewards like recognition and connection from reactions by other users. Furthermore, a few CGM platforms offer monetary rewards to users, serving as incentives for sharing items such as articles, images, and videos. However, users have varying preferences for monetary and psychological rewards, and the impact of monetary rewards on user behaviors and the quality of the content they post remains unclear. Hence, we propose a model that integrates some monetary reward schemes into the SNS-norms game, which is an abstraction of CGM. Subsequently, we investigate the effect of each monetary reward scheme on individual agents (users), particularly in terms of their proactivity in posting items and their quality, depending on agents' positions in a CGM network. Our experimental results suggest that these factors distinctly affect the number of postings and their quality. We believe that our findings will help CGM platformers in designing better monetary reward schemes. Shintaro Ueki, Fujio Toriumi, Toshiharu Sugawara |
ASONAM | 3 |
| 2023 | Interpretation Using Classified Gradient-Based Saliency Maps for Two-Player Board GamesabstractIn this study, we propose a gradient-based saliency map generation method to improve the explainability of saliency maps in two-player board games, such as chess and shogi (Japanese chess). In the proposed method, our saliency maps are generated by masking gradients of unnecessary squares to reduce ambiguity and classify the gradient values to express the positive and negative aspects of each piece on a board. The saliency maps can clearly illustrate the roles of each piece in the game. We experimentally evaluate the saliency maps by comparing them against those obtained through conventional methods using chess and shogi board configurations, demonstrating their capacity to effectively identify vital pieces and squares. Thus, we can understand which pieces are important from the visualized saliency maps. This also provides useful insights for human players to learn the games. Gentoku Nakasone, Toshiharu Sugawara |
CoG | 2 |
| 2023 | Modeling Others as a Player in Non-cooperative Game for Multi-agent Coordination
Junjie Zhong, Toshiharu Sugawara |
EANN | 2 |
| 2023 | Autonomous Energy-Saving Behaviors with Fulfilling Requirements for Multi-Agent Cooperative Patrolling Problem
Kohei Matsumoto, Keisuke Yoneda, Toshiharu Sugawara |
ICAART (1) | 3 |
| 2023 | Interpretability for Conditional Coordinated Behavior in Multi-Agent Reinforcement LearningabstractWe propose a model-free reinforcement learning architecture, called distributed attentional actor architecture after conditional attention (DA6-X), to provide better interpretability of conditional coordinated behaviors. The underlying principle involves reusing the saliency vector, which represents the conditional states of the environment, such as the global position of agents. Hence, agents with DA6-X flexibility built into their policy exhibit superior performance by considering the additional information in the conditional states during the decision-making process. The effectiveness of the proposed method was experimentally evaluated by comparing it with conventional methods in an objects collection game. By visualizing the attention weights from DA6-X, we confirmed that agents successfully learn situation-dependent coordinated behaviors by correctly identifying various conditional states, leading to improved interpretability of agents along with superior performance. Yoshinari Motokawa, Toshiharu Sugawara |
IJCNN | 2 |
| 2023 | Strategy-Following Multi-Agent Deep Reinforcement Learning through External High-Level InstructionabstractIn this paper, we propose the strategy-following distributed attentional actor architecture after conditional attention (sfDA6-X) for multi-agent deep reinforcement learning (MADRL). The architecture is designed to provide controllability of coordinated behaviors in MADRL through the use of a saliency vector that captures conditions from the environment and rough, high-level instructions given by external experts such as the system designers or users. We introduce destination channels, which enable agents to be instructed to change their behaviors only by specifying the regions where individual agents should work. To validate the effectiveness of sfDA6-X, we conducted experiments in the object collection game and analyzed how agents change their coordinated and cooperative behaviors based on the given instructions. Our findings suggest that our approach provides a preliminary but promising solution for controlling coordinated behaviors in MADRL via arbitrary instructions represented in destination channels. Yoshinari Motokawa, Toshiharu Sugawara |
KES | 2 |
| 2022 | Distributed and Asynchronous Planning and Execution for Multi-agent Systems through Short-Sighted Conflict ResolutionabstractWe propose a distributed method for a multi-agent pick-up and delivery problem with fluctuations in agent movement speeds while agents perform planning, detect and resolve conflicts (collisions) between the plans, and execute actions in the plans in a distributed manner. Our study assumes that the robot's movement speed can fluctuate, owing to various factors, thus delaying their scheduled tasks. Such delays can rapidly cause other agent conflicts to cascade and render long-term plans useless. Our proposed method allows each agent's plans to be executed and modified using an advanced short-sighted conflict resolution mechanism. Hence, although an agent attempts to follow its given sequence of actions, it performs each one after carefully checking for any conflict in the next few steps. Our method is fully distributed and works effectively, even when the number of task endpoints, which are the pick-up and delivery locations, is small and the agents are concentrated. We experimentally confirm that our method works efficiently without collisions in environments having agent speed fluctuations and deadlocks using example problems from robot movement in a construction site. Further, we compare the performance of our method with that of the baseline method. Yuki Miyashita, Tomoki Yamauchi, Toshiharu Sugawara |
COMPSAC | 3 |
| 2022 | Task Handover Negotiation Protocol for Planned Suspension based on Estimated Chances of Negotiations in Multi-agent Patrolling
Sota Tsuiki, Keisuke Yoneda, Toshiharu Sugawara |
ICAART (1) | 3 |
| 2022 | Flexible Exploration Strategies in Multi-Agent Reinforcement Learning for Instability by Mutual LearningabstractA fundamental challenge in multi-agent reinforcement learning is an effective exploration of state-action spaces because agents must learn their policies in a non-stationary environment due to changing policies of other learning agents. As the agent’s learning progresses, different undesired situations may appear one after another and agents have to learn again to adapt them. Therefore, agents must learn again with a high probability of exploration to find the appropriate actions for the exposed situation. However, existing algorithms can suffer from inability to learn behavior again on the lack of exploration for these situations because agents usually become exploitation-oriented by using simple exploration strategies, such as ε-greedy strategy. Therefore, we propose two types of simple exploration strategies, where each agent monitors the trend of performance and controls the exploration probability, ε, based on the transition of performance. By introducing a coordinated problem called the PushBlock problem, which includes the above issue, we show that the proposed method could improve the overall performance relative to conventional ε-greedy strategies and analyze their effects on the generated behavior. Yuki Miyashita, Toshiharu Sugawara |
ICMLA | 2 |
| 2022 | Imbalanced Equilibrium: Emergence of Social Asymmetric Coordinated Behavior in Multi-agent Games
Yidong Bai, Toshiharu Sugawara |
ICONIP (2) | 2 |
| 2022 | Distributed Multi-Agent Deep Reinforcement Learning for Robust Coordination against NoiseabstractIn multi-agent systems, noise reduction techniques are considerable for improving the overall system reliability as agents are required to rely on limited environmental information to develop cooperative and coordinated behaviors with the surrounding agents. However, previous studies have often applied centralized noise reduction methods to build robust and versatile coordination in noisy multi-agent environments, while distributed and decentralized autonomous agents are more plausible for real-world application. In this paper, we introduce a distributed attentional actor architecture model for a multi-agent system (DA3-X), using which we demonstrate that agents with DA3-X can selectively learn the noisy environment and behave cooperatively. We experimentally evaluate the effectiveness of DA3-X by comparing learning methods with and without DA3-X and show that agents with DA3-X can achieve better performance than baseline agents. Furthermore, we visualize heatmaps of attentional weights from the DA3-X to analyze how the decision-making process and coordinated behavior are influenced by noise. Yoshinari Motokawa, Toshiharu Sugawara |
IJCNN | 2 |
| 2022 | Deadlock-Free Method for Multi-Agent Pickup and Delivery Problem Using Priority Inheritance with Temporary PriorityabstractThis paper proposes a control method for the multi-agent pickup and delivery problem (MAPD problem) by extending the priority inheritance with backtracking (PIBT) method to make it applicable to more general environments. PIBT is an effective algorithm that introduces a priority to each agent, and at each timestep, the agents, in descending order of priority, decide their next neighboring locations in the next timestep through communications only with the local agents. Unfortunately, PIBT is only applicable to environments that are modeled as a bi-connected area, and if it contains dead-ends, such as tree-shaped paths, PIBT may cause deadlocks. However, in the real-world environment, there are many dead-end paths to locations such as the shelves where materials are stored as well as loading/unloading locations to transportation trucks. Our proposed method enables MAPD tasks to be performed in environments with some tree-shaped paths without deadlock while preserving the PIBT feature; it does this by allowing the agents to have temporary priorities and restricting agents’ movements in the trees. First, we demonstrate that agents can always reach their delivery location without deadlock. Our experiments indicate that the proposed method is very efficient, even in environments where PIBT is not applicable, by comparing them with those obtained using the well-known token passing method as a baseline. Yukita Fujitani, Tomoki Yamauchi, Yuki Miyashita, Toshiharu Sugawara |
KES | 4 |
| 2022 | Task Selection Algorithm for Multi-Agent Pickup and Delivery with Time Synchronization
Tomoki Yamauchi, Yuki Miyashita, Toshiharu Sugawara |
PRIMA | 3 |
| 2021 | Path and Action Planning in Non-uniform Environments for Multi-agent Pickup and Delivery Tasks
Tomoki Yamauchi, Yuki Miyashita, Toshiharu Sugawara |
EUMAS | 3 |
| 2021 | Effective Area Partitioning in a Multi-Agent Patrolling Domain for Better Efficiency
Katsuya Hattori, Toshiharu Sugawara |
ICAART (1) | 2 |
| 2021 | Distributed Service Area Control for Ride Sharing by using Multi-Agent Deep Reinforcement Learning
Naoki Yoshida, Itsuki Noda, Toshiharu Sugawara |
ICAART (1) | 3 |
| 2021 | MAT-DQN: Toward Interpretable Multi-agent Deep Reinforcement Learning for Coordinated Activities
Yoshinari Motokawa, Toshiharu Sugawara |
ICANN (4) | 2 |
| 2021 | Multi-Agent Task Allocation Based on Reciprocal Trust in Distributed Environments
Koki Sato, Toshiharu Sugawara |
KES-AMSTA | 2 |
| 2021 | Analysis of coordinated behavior structures with multi-agent deep reinforcement learningabstractAbstract Cooperation and coordination are major issues in studies on multi-agent systems because the entire performance of such systems is greatly affected by these activities. The issues are challenging however, because appropriate coordinated behaviors depend on not only environmental characteristics but also other agents’ strategies. On the other hand, advances in multi-agent deep reinforcement learning (MADRL) have recently attracted attention, because MADRL can considerably improve the entire performance of multi-agent systems in certain domains. The characteristics of learned coordination structures and agent’s resulting behaviors, however, have not been clarified sufficiently. Therefore, we focus here on MADRL in which agents have their own deep Q-networks (DQNs), and we analyze their coordinated behaviors and structures for the pickup and floor laying problem , which is an abstraction of our target application. In particular, we analyze the behaviors around scarce resources and long narrow passages in which conflicts such as collisions are likely to occur. We then indicated that different types of inputs to the networks exhibit similar performance but generate various coordination structures with associated behaviors, such as division of labor and a shared social norm, with no direct communication. Yuki Miyashita, Toshiharu Sugawara |
Appl. Intell. | 2 |
| 2020 | Multi-Agent Pattern Formation with Deep Reinforcement Learning (Student Abstract)abstractWe propose a decentralized multi-agent deep reinforcement learning architecture to investigate pattern formation under the local information provided by the agents' sensors. It consists of tasking a large number of homogeneous agents to move to a set of specified goal locations, addressing both the assignment and trajectory planning sub-problems concurrently. We then show that agents trained on random patterns can organize themselves into very complex shapes. Elhadji Amadou Oury Diallo, Toshiharu Sugawara |
AAAI | 2 |
| 2020 | Learning Efficient Coordination Strategy for Multi-step Tasks in Multi-agent Systems using Deep Reinforcement Learning
Zean Zhu, Elhadji Amadou Oury Diallo, Toshiharu Sugawara |
ICAART (1) | 3 |
| 2020 | Coordinated Behavior for Sequential Cooperative Task Using Two-Stage Reward Assignment with Decay
Yuki Miyashita, Toshiharu Sugawara |
ICONIP (2) | 2 |
| 2020 | Multi-Agent Pattern Formation: a Distributed Model-Free Deep Reinforcement Learning ApproachabstractIn this paper, we investigate how a large-scale system of independently learning agents can collectively form acceptable two-dimensional patterns (pattern formation) from any initial configuration. We propose a decentralized multi-agent deep reinforcement learning architecture MAPF-DQN (Multi-Agent Pattern Formation DQN) in which a set of independent and distributed agents capture their local visual field and learn how to act so as to collectively form target shapes. Agents exploit their individual networks with a central replay memory and target networks that are used to store and update the representation of the environment as well as learning the dynamics of the other agents. We then show that agents trained on random patterns using MAPF-DQN can organize themselves into very complex shapes in large-scale environments. Our results suggest that the proposed framework achieves zero-shot generalization on most of the environments independently of the depth of view of agents. Elhadji Amadou Oury Diallo, Toshiharu Sugawara |
IJCNN | 2 |
| 2020 | Meta-Reward Model Based on Trajectory Data with k-Nearest Neighbors MethodabstractReward shaping is a crucial method to speed up the process of reinforcement learning (RL). However, designing reward shaping functions usually requires many expert demonstrations and much hand-engineering. Moreover, by using the potential function to shape the training rewards, an RL agent can perform Q-learning well to converge the associated Q-table faster without using the expert data, but in deep reinforcement learning (DRL), which is RL using neural networks, Q-learning is sometimes slow to learn the parameters of networks, especially in a long horizon and sparse reward environment. In this paper, we propose a reward model to shape the training rewards for DRL in real time to learn the agent's motions with a discrete action space. This model and reward shaping method use a combination of agent self-demonstrations and a potential-based reward shaping method to make the neural networks converge faster in every task and can be used in both deep Q-learning and actor-critic methods. We experimentally showed that our proposed method could speed up the DRL in the classic control problems of an agent in various environments. Toshiharu Sugawara |
IJCNN | 2 |
| 2020 | Multi-Agent Task Allocation Based on the Learning of Managers and Local Preference SelectionsabstractThis paper discusses an adaptive distributed allocation method in which agents individually learn strategies for preferences to decide on the rank of tasks which they want to be allocated by a manager. In a distributed edge-computing environment, multiple managers that control the provision of a variety of services requested from different locations have to allocate the corresponding tasks to appropriate agents, which are usually programs developed by different companies. In our proposed method, each agent learns which manager will allocate tasks it performs well and how to declare its preferred tasks. We experimentally evaluated the proposed learning method and showed that agents using the proposed method could effectively execute requested tasks and could adapt to changes in patterns of the requested tasks. Yuka Ishihara, Toshiharu Sugawara |
KES | 2 |
| 2020 | Policy Advisory Module for Exploration Hindrance Problem in Multi-agent Deep Reinforcement Learning
Toshiharu Sugawara |
PRIMA | 2 |
| 2020 | Analysis of Coordination Structures of Partially Observing Cooperative Agents by Multi-agent Deep Q-Learning
Ken Smith, Yuki Miyashita, Toshiharu Sugawara |
PRIMA | 3 |
| 2020 | Coordinated behavior of cooperative agents using deep reinforcement learning
Elhadji Amadou Oury Diallo, Ayumi Sugiyama, Toshiharu Sugawara |
Neurocomputing | 3 |
| 2019 | Learning of Activity Cycle Length based on Battery Limitation in Multi-agent Continuous Cooperative Patrol ProblemsabstractWe propose a learning method that decides the appropriate activity cycle length (ACL) according to environmental characteristics and other agents’ behavior in the (multi-agent) continuous cooperative patrol problem. With recent advances in computer and sensor technologies, agents, which are intelligent control programs running on computers and robots, obtain high autonomy so that they can operate in various fields without pre-defined knowledge. However, cooperation/coordination between agents is sophisticated and complicated to implement. We focus on the ACL which is time length from starting patrol to returning to charging base for cooperative patrol when agents like robots have batteries with limited capacity. Long ACL enable agent to visit distant location, but it requires long rest. The basic idea of our method is that if agents have long-life batteries, they can appropriately shorten the ACL by frequently recharging. Appropriate ACL depends on many elements such as environmental size, the number of agents, and workload in an environment. Therefore, we propose a method in which agents autonomously learn the appropriate ACL on the basis of the number of events detected per cycle. We experimentally indicate that our agents are able to learn appropriate ACL depending on established spatial divisional cooperation. Ayumi Sugiyama, Lingying Wu, Toshiharu Sugawara |
ICAART (1) | 3 |
| 2019 | Cooperation and Coordination Regimes by Deep Q-Learning in Multi-agent Task Executions
Yuki Miyashita, Toshiharu Sugawara |
ICANN (1) | 2 |
| 2019 | Coordination in Adversarial Multi-Agent with Deep Reinforcement Learning Under Partial ObservabilityabstractWe propose a method using several variants of deep Q-network for learning strategic formations in large-scale adversarial multi-agent systems. The goal is to learn how to maximize the probability of acting jointly as coordinated as possible. Our method is called the centralized training and decentralized testing (CTDT) framework that is based on the POMDP during training and dec-POMDP during testing. During the training phase, the centralized neural network's inputs are the collections of local observations of agents of the same team. Although agents only know their action, the centralized network decides the joint action and subsequently distributes these actions to the individual agents. During the test, however, each agent uses a copy of the centralized network and independently decides its action based on its policy and local view. We show that deep reinforcement learning techniques using the CTDT framework can converge and generate several strategic group formations in large-scale multi-agent systems. We also compare the results using the CTDT with those derived from a centralized shared DQN and then we investigate the characteristics of the learned behaviors. Elhadji Amadou Oury Diallo, Toshiharu Sugawara |
ICTAI | 2 |
| 2019 | Energy-Efficient Strategies for Multi-Agent Continuous Cooperative Patrolling ProblemsabstractWhereas research of the multi-agent patrolling problem has been widely conducted from different aspects, the issue of energy minimization has not been sufficiently studied. When considering real-world applications with a trade-off between energy efficiency and level of perfection, it is usually more desirable to minimize the energy cost and carry out the tasks to the required level of quality instead of fulfilling tasks perfectly by ignoring energy efficiency. This paper proposes a series of coordinated behavioral strategies and an autonomous learning method of target decision strategies to reduce of energy consumption on the premise of satisfying quality requirements in continuous patrolling problems by multiple cooperative agents. We extended our previous method of target decision strategy learning by incorporating a number of behavioral strategies, with which agents individually estimate whether the requirement is reached and then modify their action plans to reduce energy consumption. It is experimentally shown that agents with the proposed methods learn to decide the appropriate strategies based on energy cost and performance efficiency and are able to reduce energy consumption while cooperatively meeting the given requirements of quality. Lingying Wu, Ayumi Sugiyama, Toshiharu Sugawara |
KES | 3 |
| 2019 | Fair and Effective Elevator Car Dispatching Method in Elevator Group Control System using CamerasabstractWe propose a control method for an elevator group control system to allocate elevator cars for all types of passengers, including general passengers and special passengers who are likely to be unfairly treated (e.g., with strollers, wheelchairs, or bulky luggage), in order to achieve fair waiting times as well as efficient transportation. Elevators are necessary for people to move vertically within high-rise buildings. Since the number of elevator cars is fixed, they have to be carefully controlled for effective dispatch. Furthermore, due to the limited capacities of elevator cars, some special passengers who require more space are often forced to wait much longer than general passengers for cars with sufficient empty space to arrive. These days, as cameras and other sensors that monitor the environment have become more common, and thanks to the recent advances in computer vision technologies, we can estimate the number of waiting passengers and the size of their belongings in elevator halls. By using such information gathered from muliple agents that monitor a specific elevator car or elevator hall, the proposed control enables effective dispatch for shorter and fairer waiting times. Experimental results using the simulated elevator control showed that our method could make waiting times fairer and achieved total efficiency to carry passengers. We discuss the reasons for the improvement as well as the limitation of our method. Tomoki Yamauchi, Rina Ide, Toshiharu Sugawara |
KES | 3 |
| 2019 | Coordination in Collaborative Work by Deep Reinforcement Learning with Various State Descriptions
Yuki Miyashita, Toshiharu Sugawara |
PRIMA | 2 |
| 2019 | Strategies for Energy-Aware Multi-agent Continuous Cooperative Patrolling Problems Subject to Requirements
Lingying Wu, Toshiharu Sugawara |
PRIMA | 2 |
| 2019 | Multiple-World Genetic Algorithm to Identify Locally Reasonable Behaviors in Complex Social NetworksabstractWe propose a novel method for evolutionary network analysis that uses the genetic algorithm (GA), called the multiple world genetic algorithm, to coevolve appropriate in-dividual behaviors of many agents on complex networks without sacrificing diversity. The GA is the powerful way, and thus, used in many domains, such as economics, biology, and social science as well as computer science, to find the interaction strategies on networks of agents. In evolutionary network analysis using GA, parents for reproduction of offspring are often selected among their neighbors under the assumption that neighbors' better strategies are useful. However, if they are on complex networks, agents exist in distinctive and diverse situations. Therefore, agents have their own appropriate interaction strategies that may be affected by a large number of neighboring agents. Here, we propose the evolutionary computation method that uses a GA on fixed networks to coevolve diverse strategies for individual agents. We conducted the experiments using simulated games of social networking services to evaluate the proposed method. The results indicate that it could effectively evolve the diverse strategy for each agent and the resulting fitness values were almost always larger than those derived through evolution using the conventional evolutionary network analysis using the GA. Yutaro Miura, Fujio Toriumi, Toshiharu Sugawara |
SMC | 3 |
| 2019 | Emergence of divisional cooperation with negotiation and re-learning and evaluation of flexibility in continuous cooperative patrol problem
Ayumi Sugiyama, Vourchteang Sea, Toshiharu Sugawara |
Knowl. Inf. Syst. | 3 |
| 2018 | Frequency-Based Multi-agent Patrolling Model and Its Area Partitioning Solution Method for Balanced Workload
Vourchteang Sea, Ayumi Sugiyama, Toshiharu Sugawara |
CPAIOR | 3 |
| 2018 | Learning Strategic Group Formation for Coordinated Behavior in Adversarial Multi-Agent with Double DQN
Elhadji Amadou Oury Diallo, Toshiharu Sugawara |
PRIMA | 2 |
| 2017 | Learning to Coordinate with Deep Reinforcement Learning in Doubles Pong GameabstractThis paper discusses the emergence of cooperative and coordinated behaviors between joint and concurrent learning agents using deep Q-learning. Multi-agent systems (MAS) arise in a variety of domains. The collective effort is one of the main building blocks of many fundamental systems that exist in the world, and thus, sequential decision making under uncertainty for collaborative work is one of the important and challenging issues for intelligent cooperative multiple agents. However, the decisions for cooperation are highly sophisticated and complicated because agents may have a certain shared goal or individual goals to achieve and their behavior is inevitably influenced by each other. Therefore, we attempt to explore whether agents using deep Q-networks (DQN) can learn cooperative behavior. We use doubles pong game as an example and we investigate how they learn to divide their works through iterated game executions. In our approach, agents jointly learn to divide their area of responsibility and each agent uses its own DQN to modify its behavior. We also investigate how learned behavior changes according to environmental characteristics including reward schemes and learning techniques. Our experiments indicate that effective cooperative behaviors with balanced division of workload emerge. These results help us to better understand how agents behave and interact with each other in complex environments and how they coherently choose their individual actions such that the resulting joint actions are optimal. Elhadji Amadou Oury Diallo, Ayumi Sugiyama, Toshiharu Sugawara |
ICMLA | 3 |
| 2017 | Adaptive Task Allocation Based on Social Utility and Individual Preference in Distributed EnvironmentsabstractRecent advances in computer and network technologies enable the provision of many services combining multiple types of information and different computational capabilities. The tasks for these services are executed by allocating them to appropriate collaborative agents, which are computational entities with specific functionality. However, the number of these tasks is huge, and these tasks appear simultaneously, and appropriate allocation strongly depends on the agent’s capability and the resource patterns required to complete tasks. Thus, we first propose a task allocation method in which, although the social utility for the shared and required performance is attempted to be maximized, agents also give weight to individual preferences based on their own specifications and capabilities. We also propose a learning method in which collaborative agents autonomously decide the preference adaptively in the dynamic environment. We experimentally demonstrate that the appropriate strategy to decide the preference depends on the type of task and the features of the task reward. We then show that agents using the proposed learning method adaptively decided their preference and could maintain excellent performance in a changing environment. Naoki Iijima, Ayumi Sugiyama, Masashi Hayano, Toshiharu Sugawara |
KES | 4 |
| 2017 | Robust spread of cooperation by expectation-of-cooperation strategy with simple labeling methodabstractThis paper proposes an interaction strategy called the extended expectation-of-cooperation (EEoC) that is intended to spread cooperative activities in prisoner's dilemma situations over an entire agent network. Recently developed computer and communications applications run on the network and interact with each other as delegates of the owners, so they often encounter social dilemma situations. To improve social efficiency, they are required to cooperate, but one-sided cooperation is meaningless and loses some payoff due to a rip-off by defecting agents. The concept of EEoC is that when agents encounter mutual cooperation, they continue to cooperate a few times with the desire to see the emergence of cooperative behavior in their neighbors. EEoC is easy to implement in computer systems. We experimentally show that EEoC can effectively spread cooperative activities in dilemma situations in complete, Erdös-Rényi, and regular networks. We also clarify the robustness against defecting agents and the limitation of the EEoC strategy. Tomoaki Otsuka, Toshiharu Sugawara |
WI | 2 |
| 2016 | Switching Behavioral Strategies for Effective Team Formation by Autonomous Agent OrganizationabstractIn this work, we propose agents that switch their behavioral strategy between rationality and reciprocity depending on their internal states to achieve efficient team formation. With the recent advances in computer science, mechanics, and electronics, there are an increasing number of applications with services/goals that are achieved by teams of different agents. To efficiently provide these services, the tasks to achieve a service must be allocated to agents that have the required capabilities and the agents must not be overloaded. Conventional distributed allocation methods often lead to conflicts in large and busy environments because high-capability agents are likely to be identified as the best team member by many agents, resulting in inefficiency of the entire
system due to concentration of task allocation. Our proposed agents switch their strategies in accordance with their local evaluation to avoid conflicts occurring in busy environments. They also establish an organization in which a number of groups are autonomously generated in a bottom-up manner on the basis of dependability in order to avoid the conflict in advance while ignoring tasks allocated by undependable/unreliable agents. We experimentally evaluate our proposed method and analyze the structure of the organization that the agents established. Masashi Hayano, Yuki Miyashita, Toshiharu Sugawara |
ICAART (1) | 3 |
| 2016 | Effective Task Allocation by Enhancing Divisional Cooperation in Multi-Agent Continuous Patrolling TasksabstractThis paper proposes an effective autonomous task allocation method that can achieve efficient cooperative work by divisional cooperation in multi-agent contexts. Computer and network technology has enabled agents/robots to behave autonomously and to be used in a variety of applications such as cleaning and security patrolling. However, to cover large environments, cooperation and collaboration among several agents are mandatory for efficiency and for the required task quality. However, how agents cooperate is a challenging issue because actual environments are usually complicated and because their own (very uncommon) characteristics. Thus, we first define the continuous cooperative patrolling problem, in which agents split up and move around the environments with the required frequencies that are defined for every location. Then, we extend the previous cooperation method to prompt autonomous and effective division of labor by introducing the negotiation for task (re) allocations. We experimentally show that agents with our method enable effective division and fair allocation by identifying their own responsible locations in a bottom-up manner and that they could achieve considerably improved results compared with those of the previous method. We also investigated the structure of the resulting regime for cooperation and analyzed why our method could achieve the effective task allocation. Ayumi Sugiyama, Vourchteang Sea, Toshiharu Sugawara |
ICTAI | 3 |
| 2016 | Assignment Problem with Preference and an Efficient Solution Method Without Dissatisfaction
Kengo Saito, Toshiharu Sugawara |
KES-AMSTA | 2 |
| 2016 | Spread of Cooperation in Complex Agent Networks Based on Expectation of Cooperation
Ryosuke Shibusawa, Tomoaki Otsuka, Toshiharu Sugawara |
PRIMA | 3 |
| 2015 | Single-object resource allocation in multiple bid declaration with preferential orderabstractThis paper proposes solutions to a problem called single-object resource allocation with preferential orders and efficient algorithms for semi-optimal allocations. Formalizations of resource allocation problems are widely used in many applications. Although many studies on allocation methods have focused on maximizing social welfare or total revenues, they have rarely taken into account agents' individual preferential orders that may have interfered with one another. Our proposed framework allocates one unit of resources to individual users but allows them to declare multiple resources with their own preferential orders. It then tries to allocate a resource to each agent by not only maximizing the total values but also considering the agent's preferences, at least, by ensuring that no or few dissatisfactions are reported. This is obviously a combinatorial problem to find optimal solutions. Thus, we propose efficient methods for semi-optimal solutions (allocations) that satisfy as many user preferences as possible. Finally, we analyze the quality of solutions and computation time by comparing them with the solutions obtained by CPLEX. Then, we experimentally demonstrate that the proposed methods are extremely efficient, while the reduced quality of their solutions is quite small. Kengo Saito, Toshiharu Sugawara |
ICIS | 2 |
| 2015 | Fair Assessment of Group Work by Mutual Evaluation with Irresponsible and Collusive Students Using Trust Networks
Yumeno Shiba, Haruna Umegaki, Toshiharu Sugawara |
PRIMA | 3 |
| 2015 | Learning and relearning of target decision strategies in continuous coordinated cleaning tasks with shallow coordinationabstractWe propose a method of autonomous learning of target decision strategies for coordination in the continuous cleaning domain. With ongoing advances in computer and sensor technologies, we can expect robot applications for covering large areas that often require coordinated/cooperative activities by multiple robots. We focus on the cleaning tasks by multiple robots or by agents which are programs to control the robots in this paper. We assumed situations where agents did not directly exchange deep and complicated internal information and reasoning results such as plans, strategies and long-term targets for their sophisticated coordinated activities, but rather exchanged superficial information such as the locations of other agents (using the equipment deployed) for their shallow coordination and individually learned appropriate strategies by observing how much dirt/dust had been vacuumed up in multi-agent system environments. We will first discuss the preliminary method of improving the coordinated activities by autonomously learning to select cleaning strategies to determine which targets to move to clear them. Although we could have improved the efficiency of cleaning, we observed a phenomenon where performance degraded if agents continued to learn strategies. This is because so many agents overly selected the same strategy (over-selection) by using autonomous learning. In addition, the preliminary method assumed information given about which regions in the environment easily became dirty. Thus, we propose a method that was extended by incorporating the preliminary method with (1) environmental learning to identify which places were likely to be dirty and (2) autonomous relearning through self-monitoring the amount of vacuumed dirt to avoid strategies from being over-selected. We experimentally evaluated the proposed method by comparing its performance with those obtained by the regimes of agents with a single strategy and obtained with the preliminary method. The experimental results revealed that the proposed method enabled agents to select target decision strategies and, if necessary, to abandon the current strategies from their own perspectives, resulting in appropriate combinations of multiple strategies. We also found that environmental learning on dirt accumulation was effectively learned. Keisuke Yoneda, Ayumi Sugiyama, Chihiro Kato, Toshiharu Sugawara |
Web Intell. | 4 |
| 2014 | Fair assessment of group work by mutual evaluation based on trust networkabstractWe propose a method for fair and accurate assessment of group work based on trust networks generated by mutual evaluations. Group work is often used for educational activities in universities since it is an effective way to acquire useful knowledge in a number of practical subjects. One drawback is the difficulty of deciding on final marks. Some students may work quite hard whereas others may rarely participate in the group work, but it is almost impossible for professors/instructors to identify contributions of individual students in detail. In contrast, students in the same group are obvious choices for appropriate evaluators of other members since they have first-hand knowledge of the collaborative work. However, some students may be irresponsible for their ratings and submit disputable evaluations, resulting in inaccurate marks. We introduce a simple mutual evaluation method and generate trust networks expressing the distances between evaluations in this paper. After that, disputable evaluations are excluded and students are marked again. We also examine a grouping strategy to detect irresponsible students more accurately. We demonstrate the effectiveness and limitations of our method using multi-agent simulation. Results show that our method can help with the marking of individual students in a group work. Yumeno Shiba, Toshiharu Sugawara |
FIE | 2 |
| 2014 | Autonomous Strategy Determination with Learning of Environments in Multi-agent Continuous Cleaning
Ayumi Sugiyama, Toshiharu Sugawara |
PRIMA | 2 |
| 2014 | Role and member selection in team formation using resource estimation for large-scale multi-agent systems
Masashi Hayano, Dai Hamada, Toshiharu Sugawara |
Neurocomputing | 3 |
| 2013 | Role and Member Selection in Team Formation Using Resource EstimationabstractWe propose an efficient team formation method for multi-agent systems consisting of self-interested agents in task-oriented domains. Services computing on computer networks have been rapidly increasing. Efficient team formation for service tasks is considered to be a way to improve performance. Our method is based on our previous parameter learning method enabling agents to efficiently form teams but requiring prior knowledge about all others' resources. We extended that method by adding a resource estimation method so as to increase its applicability to actual application systems. We experimentally evaluated our method by comparing it with the previous method and the task allocation using contract net protocol (CNP). The results demonstrated that the proposed method outperformed other methods even though it did not require prior knowledge about resources in other agents. We discuss the reason for this improvement. Masashi Hayano, Dai Hamano, Toshiharu Sugawara |
KES-AMSTA | 3 |
| 2013 | Decentralized Area Partitioning for a Cooperative Cleaning Task
Chihiro Kato, Toshiharu Sugawara |
PRIMA | 2 |
| 2013 | ADMIRE: Anomaly detection method using entropy-based PCA with three-step sketches
Yoshiki Kanda, Romain Fontugne, Kensuke Fukuda, Toshiharu Sugawara |
Comput. Commun. | 4 |
| 2012 | Deciding Roles for Efficient Team Formation by Parameter Learning
Dai Hamada, Toshiharu Sugawara |
KES-AMSTA | 2 |
| 2012 | Two-Sided Parameter Learning of Role Selections for Efficient Team Formation
Dai Hamada, Toshiharu Sugawara |
PRIMA | 2 |
| 2011 | Adaptive Routing Point Control in Virtualized Local Area Networks Using Particle Swarm OptimizationsabstractThis paper describes methods for controlling routing points of VLAN domains using binary particle swarm optimization (BPSO) and angle modulated particle swarm optimization (AMPSO). Virtual LAN (VLAN) is a technique for virtualizing data link layer (or L2) and can construct arbitrary logical networks on top of a physical network. However, VLAN often causes much redundant traffic due to inappropriate deployments of network-layer (L3) routing capabilities in VLAN networks. We propose two methods using BPSO and AMPSO, and show that they can adaptively select the routing points dynamically in accordance with the observed traffic patterns and thus reduce the redundant traffic. The convergence features are compared with those of the conventional method on the basis of a statistical method. Then we also show that the scalability of the algorithm using AMPOS is high and thus we can expect that it is applicable to practical large VLAN environments. Kensuke Takahashi, Toshio Hirotsu, Toshiharu Sugawara |
ICTAI | 3 |
| 2011 | Emergence and Stability of Social Conventions in Conflict SituationsabstractWe investigate the emergence and stability of social conventions for efficiently resolving conflicts through reinforcement learning. Facilitation of coordination and conflict resolution is an important issue in multi-agent systems. However, exhibiting coordinated and negotiation activities is computationally expensive. In this paper, we first describe a conflict situation using a Markov game which is iterated if the agents fail to resolve their conflicts, where the repeated failures result in an inefficient society. Using this game, we show that social conventions for resolving conflicts emerge, but their stability and social efficiency depend on the payoff matrices that characterize the agents. We also examine how unbalanced populations and small heterogeneous agents affect efficiency and stability of the resulting conventions. Our results show that (a) a type of indecisive agent that is generous for adverse results leads to unstable societies, and (b) selfish agents that have an explicit order of benefits make societies stable and efficient. Toshiharu Sugawara |
IJCAI | 1 |
| 2010 | A PCA Analysis of Daily Unwanted TrafficabstractThis paper investigates the macroscopic behavior of unwanted traffic (e.g., virus, worm, backscatter of (D)DoS or misconfiguration) passing through the Internet. The data set we used are unwanted packets measured at /18 darknet in Japan from Oct. 2006 to Apr. 2009 that included the recent Conficker outbreak. The traffic behavior is quantified by the entropy of ten packet features (e.g., 5-tuple). Then, we apply PCA (principal component analysis) to a ten dimensional entropy time series matrix to obtain a suitable representation of unwanted traffic. PCA is a well-known and studied method for finding out normal and anomalous behaviors in Internet backbone traffic, however, few studies applied it to darknet traffic. We first demonstrate the high variability nature of the entropy time series for ten packet features. Next, we show that the top four principal components are sufficiently enough to describe the original traffic behavior. In particular, the first component can be interpreted as the type of unwanted traffic (i. e., worm/virus or scanning), and the second one as the difference in communication patterns (e. g., one-to-many or many-to-one). Those two components account for 63.8\% of the original data set in terms of the total variance. On the other hand, the outliers in the higher components indicate the presence of specific anomalies although most of mapped data to the components have less variability. Furthermore, we show that the scatter plot of the first and second principal component scores provides us with a better view of the macroscopic unwanted traffic behavior. Kensuke Fukuda, Toshio Hirotsu, Osamu Akashi, Toshiharu Sugawara |
AINA | 4 |
| 2010 | Adaptive probabilistic task allocation in large-scale multi-agent systems and its evaluationabstractIn this paper, we introduce the probabilistic awardee selection strategy, under which awardee is selected with a fixed probability, into the award phase of contract net protocol. We then point out that, by changing the probabilities in this strategy according the local workload, the overall performance can be considerably improved. Toshiharu Sugawara, Kensuke Fukuda, Toshio Hirotsu, Satoshi Kurihara |
GECCO | 1 |
| 2010 | Evaluation of Anomaly Detection Based on Sketch and PCAabstractUsing traffic random projections (sketches) and Principal Component Analysis (PCA) for Internet traffic anomaly detection has become popular topics in the anomaly detection fields, but few studies have been undertaken on the subjective and quantitative comparison of multiple methods using the data traces open to the community. In this paper, we propose a new method that combines sketches and PCA to detect and identify the source IP addresses associated with the traffic anomalies in the backbone traces measured at a single link. We compare the results with those of a method incorporating sketches and multi-resolution gamma modeling using the trans-Pacific link traces. The comparison indicates that each method has its own advantages and disadvantages. Our method is good at detecting worm activities with many packets, whereas the gamma method is good at detecting scan activities for peer hosts with only a few packets, but it reports many false positives for traces of worm outbreaks. Therefore, their use in combination would be effective. We also examined the impact of adaptive decision making on a parameter (the number of normal subspaces in PCA) on the basis of the cumulative proportion of each sketched traffic and conclude that it performs at a higher level than the previous method deciding only on one specific value of the parameter for every divided traffics. Yoshiki Kanda, Kensuke Fukuda, Toshiharu Sugawara |
GLOBECOM | 3 |
| 2010 | Probabilistic Award Strategy for Contract Net Protocol in Massively Multi-agent Systems
Toshiharu Sugawara, Toshio Hirotsu, Kensuke Fukuda |
ICAART (2) | 1 |
| 2010 | A Flow Analysis for Mining Traffic AnomaliesabstractAlthough analyzing anomalous network traffic behavior is a popular research topic, few studies have been undertaken on the analysis of communication pattern per host based on their flows to characterize the anomalous Internet traffic. This paper discusses the possibility of using a flow-based communication pattern per host as a metric to identify anomalies. The key idea underlining our method is that scanning worm-infected hosts reveal the intrinsic characteristics of host's communication pattern and such patterns are distinguishable from those of other hosts. In particular, we found that scanning of worm-infected hosts that generated a lot of flows revealed the intrinsic communication pattern and the pattern could be classified from those of other hosts by k-means clustering. We also found that our flow-based metric could isolate the anomalies that have little influence upon the volumetric information of traffic and flow as "lines", which is remarkable in that the hosts that caused the hidden anomalies were mined out. Yoshiki Kanda, Kensuke Fukuda, Toshiharu Sugawara |
ICC | 3 |
| 2010 | Dynamic and distributed routing control for virtualized local area networksabstractAdvanced Layer-3 (L3) switches achieve high-speed IP packet forwarding by storing parts of the header information from transmitted packets into the flow cache in the switch fabric when relaying the IP packets between subnets. When IP traffic is overloaded on an L3 switch, the flow cache is easily exhausted, decreasing the IP packet forwarding performance. Virtual LAN (VLAN), a virtualization technology at the data-link layer, is widely used for the internal networks of many organizations because it allows network configurations to be changed easily and provides design flexibility. In a VLAN-based local area network with multiple L3 switches, the relaying point for each VLAN can be placed on any of the L3 switches. We developed a new network control scheme called distributed virtual routing, which dynamically controls the packet exchange points for each VLAN to suppress the consumption of flow caches. We describe the basic concept and then evaluate the reduction of the relaying flows through the simulation using the real network data. Toshio Hirotsu, Satoshi Kurihara, Kensuke Fukuda, Osamu Akashi, Hirotake Abe, Toshiharu Sugawara |
LCN | 6 |
| 2010 | Effect of Alternative Distributed Task Allocation Strategy Based on Local Observations in Contract Net Protocol
Toshiharu Sugawara, Kensuke Fukuda, Toshio Hirotsu, Satoshi Kurihara |
PRIMA | 1 |
| 2010 | Fluctuated peer selection policy and its performance in large-scale multi-agent systemsabstractThis paper describes how, in large-scale multi-agent systems, each agent's adaptive selection of peer agents for collaborative tasks affects the overall performance and how this performance varies with the workload of the system and with fluctuations Toshiharu Sugawara, Kensuke Fukuda, Toshio Hirotsu, Shin-ya Sato, Osamu Akashi, Satoshi Kurihara |
Web Intell. Agent Syst. | 1 |
| 2009 | Estimating Relevance of Items on Basis of Proximity of User Groups on BlogspaceabstractWe describe a new method to estimate the relevance of two items (such as products and works of art) on the basis of the relationship between the corresponding user (blogger) groups on a blogspace, where a user group refers to a collection of users interested in an item. We estimated the strength of the relationship between user groups on the basis of their proximity on the blogspace. We validated our approach through experimental studies using actual data. In developing the method for estimating relevance among items, we introduced a new technique for measuring the proximity of two groups of vertices on a network, which can be thought of as an extension of conventional co-occurrence analysis. Shin-ya Sato, Kensuke Fukuda, Toshio Hirotsu, Satoshi Kurihara, Toshiharu Sugawara |
Web Intelligence | 5 |
| 2008 | Correlation Among Piecewise Unwanted Traffic Time SeriesabstractIn this paper, we investigate temporal and spatial correlations of time series of unwanted traffic (i.e., darknet or network telescope traffic) in order to estimate statistical behavior of unwanted activities from a small size of darknet address block. First, from the analysis of long-range dependency, we point out that TCP time series has a weak temporal correlation though UDP time series without huge flooding is well-modeled using a Poisson process. Next, we analyze the spatial correlation between two traffic time series divided by different sized darknet address blocks. We confirm that a TCP SYN traffic time series (e.g, virus or worm) has a clear spatial correlation in the arrival of packets between two neighboring address blocks. Indeed, this spatial correlation remains in traffic time series 1,000 addresses far from the target time series, even if a darknet address block is small (e.g., /26). On the other hand, TCP SYNACK traffic (e.g., backscatter) and UDP traffic (e.g., virus or worm) have less spatial correlation between two adjacent large address blocks. Finally, we estimate the average propagation delay of global unwanted activities appearing in TCP SYN traffic by using the generalized inter-correlation coefficient. Kensuke Fukuda, Toshio Hirotsu, Osamu Akashi, Toshiharu Sugawara |
GLOBECOM | 4 |
| 2008 | Estimation of sensor-network topology from time-series sensor data using ant colony optimization methodabstractWe propose a method for estimating sensor network topology from only with time-series sensor data and without prior knowledge about the locations of sensors. The proposed method is based on ant colony optimization (ACO) but is further improved, compared with previous work[s], to construct a more accurate topology through an examination of the reliability of the acquired sensor data for the adjacency estimation. This reliability value is used to control the amount of pheromones deposited. We evaluate our method using actual sensor data and show that it can estimate adjacencies, in which the error rate is approximately 87% less than that of the previous method. Kensuke Takahashi, Toshiharu Sugawara |
SIS | 2 |
| 2008 | Co-occurrence Analysis Focused on Blogger CommunitiesabstractWe studied the problem of finding a subspace of Web pages that is contextually consistent for co-occurrence analysis. We looked at blogs and proposed blogger-based co-occurrence analysis, which assumes that two items are relevant to each other if they appear in any of the blog entries posted by the same blogger. We show that (1) blogger-based analysis outperforms conventional page-based analysis in solving context-sensitive problems and that (2) analysis focused on bloggers forming a community yields better performance compared with that focused on isolated bloggers. Shin-ya Sato, Kensuke Fukuda, Toshio Hirotsu, Satoshi Kurihara, Toshiharu Sugawara |
Web Intelligence | 5 |
| 2008 | Policy-based BGP-control architecture for inter-AS routing adjustment
Osamu Akashi, Kensuke Fukuda, Toshio Hirotsu, Toshiharu Sugawara |
Comput. Commun. | 4 |
| 2007 | Performance variation due to interference among a large number of self-interested agentsabstractThe performance features of a massively multiagent system (MMAS) when applying the contract net protocol (CNP) are examined. The recent growth in the volume of e-commerce on the Internet is increasing the opportunities for coordinated transactions by agents, concurrently occurring everywhere. Because of limited CPU and network resources, running many interactive tasks among agents can lower the quality or efficiency of MMASs. Although CNP is a widely used negotiation protocol that can allocate tasks and resources to appropriate agents, it is unclear how effectively CNP works in an MMAS where thousands of agents work together and interfere with each other. The performance of CNP in such an MMAS, especially the overall efficiency and the reliability of promised completion times, is investigated by using an MAS simulation environment. The results show that only managerside control of CNP can improve performance in an MMAS. Toshiharu Sugawara, Toshio Hirotsu, Satoshi Kurihara, Kensuke Fukuda |
IEEE Congress on Evolutionary Computation | 1 |
| 2007 | Generating Extensional Definitions of Concepts from Ostensive Definitions by Using Web
Shin-ya Sato, Kensuke Fukuda, Satoshi Kurihara, Toshio Hirotsu, Toshiharu Sugawara |
WISE | 5 |
| 2006 | Adaptive Agent Selection in Large-Scale Multi-Agent Systems
Toshiharu Sugawara, Kensuke Fukuda, Toshio Hirotsu, Shin-ya Sato, Satoshi Kurihara |
PRICAI | 1 |
| 2005 | ARTISTE: Agent Organization Management System for Multi-Agent Systems
Atsushi Terauchi, Osamu Akashi, Mitsuru Maruyama, Kensuke Fukuda, Toshiharu Sugawara, Toshio Hirotsu, Satoshi Kurihara |
PRIMA | 5 |
| 2003 | Secure and Manageable Virtual Private Networks for End-usersabstractThis paper presents personal networks, which integrate a VPN and the per-VPN execution environments of the hosts included in the VPN. The key point is that each execution environment called a portspace is bound to only one VPN, i.e., single-homed. Using this feature of portspaces, personal networks address several problems at multi-homed hosts that use multiple VPNs. Information flow is separated by personal networks so that it is not mixed at multi-homed hosts. IP addressing in a personal network is independent of the other personal networks, even the base network, and therefore does not conflict with those of other networks at multi-homed hosts. In addition, personal networks provide facilities for easy bootstrapping so that the end-users can construct such isolated networks easily. Inheritance of portspaces supports the creation of new portspaces based on existing portspaces. Self-construction of personal networks enables end-users to construct personal networks without help from the base network. Kenichi Kourai, Toshio Hirotsu, Koji Sato, Osamu Akashi, Kensuke Fukuda, Toshiharu Sugawara, Shigeru Chiba |
LCN | 6 |
| 2003 | Extraction of Implicit Resource Relationships in Multi-agent Systems
Toshiharu Sugawara, Satoshi Kurihara, Osamu Akashi |
PRIMA | 1 |
| 2001 | A multi-agent monitoring and diagnostic system for TCP/IP-based network and its coordination
Toshiharu Sugawara, Ken-ichiro Murakami, Shigeki Goto |
Knowl. Based Syst. | 1 |
| 2000 | Multiagent-based cooperative inter-AS diagnosis in ENCOREabstractIf the Internet is to operate reliably, it is important to verify whether the routing information about an autonomous system (AS) is correctly spread throughout the Internet as the originating AS intends or not. Because the routing information changes as it spreads, it is necessary for the inter-AS diagnostic system to observe the routing information from outside the AS. Each AS is controlled by a single administrative authority based on that AS's own policy, so a cooperative distributed solution is desirable. To cope with these requirements, we have proposed a multi-agent-based inter-AS diagnostic system called ENCORE, where a collection of intelligent agents are located in multiple AS and perform collective observation and analysis. This paper describes its diagnostic functions. It focuses on the agents' autonomous observation actions and their ability to perform cooperative and efficient analysis of routing anomalies. We also discuss the effectiveness and limitations of ENCORE based on analysis of a real AS. Osamu Akashi, Toshiharu Sugawara, Ken-ichiro Murakami, Mitsuru Maruyama, Naohisa Takahashi |
NOMS | 2 |
| 1998 | Learning to Improve Coordinated Actions in Cooperative Distributed Problem-Solving Environments
Toshiharu Sugawara, Victor R. Lesser |
Mach. Learn. | 1 |