Yuki Miyashita

dblp:170/3541 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-1676-9346ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Graph Orientation for Efficient Multi-Agent Pickup and Delivery Using Genetic Algorithm
Yuki Miyashita, Toshiharu Sugawara
COMPSAC1
2025 Multi-Agent Path Finding Using Provisionally Booking Nodes for Pickup and Delivery Problems
Daiki Shimada, Yuki Miyashita, Toshiharu Sugawara
ICAART (3)2
2024 Scheduling and Negotiation Method for Double Synchronized Multi-Agent Pickup and Delivery Problem
Yuki Miyashita, Toshiharu Sugawara
ICAART (1)1
2024 Efficient Retraining for Continuous Operability Through Strategic Directives
abstract
We introduce a method to ensure the controllability of agents in pretrained networks, anticipating changes in managerial requirements and environmental conditions. Advances in multi-agent deep reinforcement learning (MADRL) have fostered sophisticated cooperative behaviors in multi-agent systems in which agents share and execute various complex tasks. However, retraining MADRL systems to adapt to new conditions is costly, because many agents are involved. Our approach introduces several types of directives, termed destination channels (DCs), which allow agents to experience diverse coordination patterns during training without detailed instructions. When changes occur, the system manager assigns appropriate DCs to each agent to facilitate adaptation and maintain a continuous operation. We conducted experiments using object-collection games to evaluate our proposed method by comparing the number of objects recovered and the learning speed of our method with those of existing methods, an implicit quantile network (IQN), and a traditional deep Q-network (DQN). The experimental results demonstrate that agents using this method adapt more swiftly to environmental changes during retraining than baseline methods, sustaining performance without significant degradation.
Gentoku Nakasone, Yoshinari Motokawa, Yuki Miyashita, Toshiharu Sugawara
ICMLA3
2024 Path Finding with Flexible Provisional Booking in Multi-agent Pickup and Delivery Problems
Daiki Shimada, Yuki Miyashita, Toshiharu Sugawara
PRIMA2
2022 Distributed and Asynchronous Planning and Execution for Multi-agent Systems through Short-Sighted Conflict Resolution
abstract
We propose a distributed method for a multi-agent pick-up and delivery problem with fluctuations in agent movement speeds while agents perform planning, detect and resolve conflicts (collisions) between the plans, and execute actions in the plans in a distributed manner. Our study assumes that the robot's movement speed can fluctuate, owing to various factors, thus delaying their scheduled tasks. Such delays can rapidly cause other agent conflicts to cascade and render long-term plans useless. Our proposed method allows each agent's plans to be executed and modified using an advanced short-sighted conflict resolution mechanism. Hence, although an agent attempts to follow its given sequence of actions, it performs each one after carefully checking for any conflict in the next few steps. Our method is fully distributed and works effectively, even when the number of task endpoints, which are the pick-up and delivery locations, is small and the agents are concentrated. We experimentally confirm that our method works efficiently without collisions in environments having agent speed fluctuations and deadlocks using example problems from robot movement in a construction site. Further, we compare the performance of our method with that of the baseline method.
Yuki Miyashita, Tomoki Yamauchi, Toshiharu Sugawara
COMPSAC1
2022 Flexible Exploration Strategies in Multi-Agent Reinforcement Learning for Instability by Mutual Learning
abstract
A fundamental challenge in multi-agent reinforcement learning is an effective exploration of state-action spaces because agents must learn their policies in a non-stationary environment due to changing policies of other learning agents. As the agent’s learning progresses, different undesired situations may appear one after another and agents have to learn again to adapt them. Therefore, agents must learn again with a high probability of exploration to find the appropriate actions for the exposed situation. However, existing algorithms can suffer from inability to learn behavior again on the lack of exploration for these situations because agents usually become exploitation-oriented by using simple exploration strategies, such as ε-greedy strategy. Therefore, we propose two types of simple exploration strategies, where each agent monitors the trend of performance and controls the exploration probability, ε, based on the transition of performance. By introducing a coordinated problem called the PushBlock problem, which includes the above issue, we show that the proposed method could improve the overall performance relative to conventional ε-greedy strategies and analyze their effects on the generated behavior.
Yuki Miyashita, Toshiharu Sugawara
ICMLA1
2022 Deadlock-Free Method for Multi-Agent Pickup and Delivery Problem Using Priority Inheritance with Temporary Priority
abstract
This paper proposes a control method for the multi-agent pickup and delivery problem (MAPD problem) by extending the priority inheritance with backtracking (PIBT) method to make it applicable to more general environments. PIBT is an effective algorithm that introduces a priority to each agent, and at each timestep, the agents, in descending order of priority, decide their next neighboring locations in the next timestep through communications only with the local agents. Unfortunately, PIBT is only applicable to environments that are modeled as a bi-connected area, and if it contains dead-ends, such as tree-shaped paths, PIBT may cause deadlocks. However, in the real-world environment, there are many dead-end paths to locations such as the shelves where materials are stored as well as loading/unloading locations to transportation trucks. Our proposed method enables MAPD tasks to be performed in environments with some tree-shaped paths without deadlock while preserving the PIBT feature; it does this by allowing the agents to have temporary priorities and restricting agents’ movements in the trees. First, we demonstrate that agents can always reach their delivery location without deadlock. Our experiments indicate that the proposed method is very efficient, even in environments where PIBT is not applicable, by comparing them with those obtained using the well-known token passing method as a baseline.
Yukita Fujitani, Tomoki Yamauchi, Yuki Miyashita, Toshiharu Sugawara
KES3
2022 Task Selection Algorithm for Multi-Agent Pickup and Delivery with Time Synchronization
Tomoki Yamauchi, Yuki Miyashita, Toshiharu Sugawara
PRIMA2
2021 Path and Action Planning in Non-uniform Environments for Multi-agent Pickup and Delivery Tasks
Tomoki Yamauchi, Yuki Miyashita, Toshiharu Sugawara
EUMAS2
2021 Analysis of coordinated behavior structures with multi-agent deep reinforcement learning
abstract
Abstract Cooperation and coordination are major issues in studies on multi-agent systems because the entire performance of such systems is greatly affected by these activities. The issues are challenging however, because appropriate coordinated behaviors depend on not only environmental characteristics but also other agents’ strategies. On the other hand, advances in multi-agent deep reinforcement learning (MADRL) have recently attracted attention, because MADRL can considerably improve the entire performance of multi-agent systems in certain domains. The characteristics of learned coordination structures and agent’s resulting behaviors, however, have not been clarified sufficiently. Therefore, we focus here on MADRL in which agents have their own deep Q-networks (DQNs), and we analyze their coordinated behaviors and structures for the pickup and floor laying problem , which is an abstraction of our target application. In particular, we analyze the behaviors around scarce resources and long narrow passages in which conflicts such as collisions are likely to occur. We then indicated that different types of inputs to the networks exhibit similar performance but generate various coordination structures with associated behaviors, such as division of labor and a shared social norm, with no direct communication.
Yuki Miyashita, Toshiharu Sugawara
Appl. Intell.1
2020 Coordinated Behavior for Sequential Cooperative Task Using Two-Stage Reward Assignment with Decay
Yuki Miyashita, Toshiharu Sugawara
ICONIP (2)1
2020 Analysis of Coordination Structures of Partially Observing Cooperative Agents by Multi-agent Deep Q-Learning
Ken Smith, Yuki Miyashita, Toshiharu Sugawara
PRIMA2
2019 Cooperation and Coordination Regimes by Deep Q-Learning in Multi-agent Task Executions
Yuki Miyashita, Toshiharu Sugawara
ICANN (1)1
2019 Coordination in Collaborative Work by Deep Reinforcement Learning with Various State Descriptions
Yuki Miyashita, Toshiharu Sugawara
PRIMA1
2016 Switching Behavioral Strategies for Effective Team Formation by Autonomous Agent Organization
abstract
In this work, we propose agents that switch their behavioral strategy between rationality and reciprocity depending on their internal states to achieve efficient team formation. With the recent advances in computer science, mechanics, and electronics, there are an increasing number of applications with services/goals that are achieved by teams of different agents. To efficiently provide these services, the tasks to achieve a service must be allocated to agents that have the required capabilities and the agents must not be overloaded. Conventional distributed allocation methods often lead to conflicts in large and busy environments because high-capability agents are likely to be identified as the best team member by many agents, resulting in inefficiency of the entire system due to concentration of task allocation. Our proposed agents switch their strategies in accordance with their local evaluation to avoid conflicts occurring in busy environments. They also establish an organization in which a number of groups are autonomously generated in a bottom-up manner on the basis of dependability in order to avoid the conflict in advance while ignoring tasks allocated by undependable/unreliable agents. We experimentally evaluate our proposed method and analyze the structure of the organization that the agents established.
Masashi Hayano, Yuki Miyashita, Toshiharu Sugawara
ICAART (1)2