Gyu Seon Kim

dblp:333/1068 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-5559-9749ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Hybrid Large Language Models and Reinforcement Learning for Energy-Efficient Multisatellite Scheduling: Boosting the Performance From Scratch
abstract
Low Earth Orbit (LEO) satellite constellations are crucial for global connectivity by providing extensive coverage and reduced delays. However, scheduling data transmission in these dynamic networks is challenging due to rapidly changing satellite positions. This research introduces BoostRL (boosted reinforcement learning), a novel framework integrating large language models (LLMs) with reinforcement learning (RL) for efficient scheduling in LEO satellite constellations. BoostRL leverages LLM-generated initial policies to accelerate convergence, thereby guiding the early stages of policy learning while adapting swiftly to dynamic network conditions. It employs a hybrid policy approach which transitions smoothly from LLM recommendations to autonomous RL policies. A tailored initialization of Q-value parameters and an enhanced loss function further optimize learning efficiency, by aligning the initial learning phase with LLM-generated insights. Simulations using two-line element (TLE) orbital data demonstrate that BoostRL achieves rapid convergence and improved efficiency, thereby validating its potential as a scalable, adaptive solution for managing satellite communication networks.
Hyojun Ahn, Gyu Seon Kim, In-Sop Cho, Soyi Jung, Joongheon Kim
IEEE Internet Things J.2
2026 Learning-Based Resilient Dynamic Routing for Fault-Tolerant Robust LEO Satellite Networks
abstract
Providing seamless global Internet connectivity remains a significant challenge due to geographical barriers, economic constraints, and vulnerabilities to catastrophic events, including natural disasters and warfare. Non-terrestrial Networks (NTN), particularly satellite communication systems, are crucial for overcoming these limitations, but traditional geostationary Earth orbit (GEO) satellites experience inherent latency issues due to their high-altitude positioning. Low Earth orbit (LEO) satellites have emerged as superior alternatives, significantly reducing communication delays and enhancing network performance. Despite their advantages, LEO satellite networks face significant challenges arising from rapidly changing network topologies driven by high orbital speeds, resulting in frequent shifts in inter-satellite link (ISL) availability. Conventional routing algorithms struggle to adapt effectively to these dynamic conditions, resulting in suboptimal routing decisions and increased latency due to congestion-related issues. To address these complexities, this paper proposes an adaptive reinforcement learning-based NTN routing (ARL-NTNR) algorithm designed explicitly for dynamic LEO environments. The proposed ARL-NTNR algorithm autonomously learns and adapts optimal routing policies in real-time by continually interacting with the evolving satellite network. It effectively balances shortest-path routing with real-time queue backlog congestion management, significantly reducing packet delay and minimizing packet loss.
Gyu Seon Kim, Chaemoon Im, Yeong Goo Kim, Jaekyoung Ha, Joongheon Kim
IEEE Internet Things J.1
2026 Quantum Multi-Agent Reinforcement Learning for Cooperative Mobile Access in Space-Air-Ground Integrated Networks
abstract
Achieving global space-air-ground integrated network (SAGIN) access only with CubeSats presents significant challenges such as the access sustainability limitations in specific regions (e.g.,polar regions) and the energy efficiency limitations in CubeSats. To tackle these problems, high-altitude long-endurance unmanned aerial vehicles (HALE-UAVs) can complement these CubeSat shortcomings for providing cooperatively global access sustainability and energy efficiency. However, as the number of CubeSats and HALE-UAVs, increases, the scheduling dimension of each ground station (GS) increases. As a result, each GS can fall into the curse of dimensionality, and this challenge becomes one major hurdle for efficient global access. Therefore, this paper provides a quantum multi-agent reinforcement Learning (QMARL)-based method for scheduling between GSs and CubeSats/HALE-UAVs in order to improve global access availability and energy efficiency. The main reason why the QMARL-based scheduler can be beneficial is that the algorithm facilitates a logarithmic-scale reduction in scheduling action dimensions, which is one critical feature as the number of CubeSats and HALE-UAVs expands. Additionally, individual GSs have different traffic demands depending on their locations and characteristics, thus it is essential to provide differentiated access services. The superiority of the proposed scheduler is validated through data-intensive experiments in realistic CubeSat/HALE-UAV settings.
Gyu Seon Kim, Yeryeong Cho, Jaehyun Chung, SooHyun Park, Soyi Jung, Zhu Han 0001, Joongheon Kim
IEEE Trans. Mob. Comput.1
2026 Joint Sustainable Control and Quantum Reinforcement Learning for Energy-Efficient Cube-Satellite Networks
abstract
Satellites have been envisioned as primary non-terrestrial networks capable of seamless global network and surveillance services. Among various satellite types, Cube Satellites (CubeSats) have been actively researched because multiple CubeSats can be conveniently positioned in a target orbit simultaneously and in proximity to Earth. However, CubeSats are small-scale, and thus, they are not able to accommodate a sizable battery, imposing constraints on the duration of their mission. Considering this energy limitation, in order to realize global network services using multiple CubeSats, this paper proposes a novel two-stage Reinforcement Learning (RL) algorithm for energy-efficient CubeSats where RL is utilized for dynamic control under uncertainty. Firstly, sustainable control for single-CubeSat orbital maneuver is considered using deep deterministic policy gradient for vertical position adjustment over a continuous action domain. Secondly, a novel quantum multi-agent RL algorithm for multi-CubeSat cooperative scheduling is designed to realize action dimension reduction into a logarithmic scale based on our proposed Projection-Valued Measure (PVM) over the quantum domain. It is highlighted that our considering two single- and multi-CubeSat problems cannot be separately considered for extreme energy management. The performance evaluation results demonstrate that the proposed algorithm outperforms other benchmarks with 1.51× higher performance in orbital control, 2.71× higher converged reward in enormous action dimensions, and 2.27× higher average network performance.
SooHyun Park, Gyu Seon Kim, Soyi Jung, Zhu Han 0001, Joongheon Kim
IEEE Trans. Mob. Comput.2
2025 Hallucination-Aware Generative Pretrained Transformer for Cooperative Aerial Mobility Control
abstract
This paper proposes SafeGPT, a two-tiered framework that integrates generative pretrained transformers (GPTs) with reinforcement learning (RL) for efficient and reliable un-manned aerial vehicle (UAV) last-mile deliveries. In the proposed design, a Global GPT module assigns high-level tasks such as sector allocation, while an On-Device GPT manages real-time local route planning. An RL-based safety filter monitors each GPT decision and overrides unsafe actions that could lead to battery depletion or duplicate visits, effectively mitigating hallucinations. Furthermore, a dual replay buffer mechanism helps both the GPT modules and the RL agent refine their strategies over time. Simulation results demonstrate that SafeGPT achieves higher delivery success rates compared to a GPT-only baseline, while substantially reducing battery consumption and travel distance. These findings validate the efficacy of combining GPT-based semantic reasoning with formal safety guarantees, contributing a viable solution for robust and energy-efficient UAV logistics.
Hyojun Ahn, Seungcheol Oh, Gyu Seon Kim, Soyi Jung, SooHyun Park, Joongheon Kim
GLOBECOM3
2025 Quantum Reinforcement Learning for Coordinated Satellite Systems
abstract
Reinforcement learning (RL) using conventional neural networks (NN) has significantly progressed in various applications. However, conventional RL needs help training in environments with large-scale action dimensions, such as coordinated mobility/satellite systems. Quantum reinforcement learning (QRL) with quantum NN (QNN) can address this problem through superposition and entanglement, one of the great features of quantum mechanics. Based on its ‘i) fast convergence’ and ‘ii) high scalability’, unique advantages of QRL that distinguish it from conventional RL, this paper highlights the potential for QRL utilization in coordinated mobility and satellite systems.
Gyu Seon Kim, Samuel Yen-Chi Chen, SooHyun Park, Joongheon Kim
ICASSP1
2025 Stabilized Robust Control for Lightweight Autonomous Aircraft Mobility: A Quantum Reinforcement Learning Approach
abstract
The stability of aircraft remains vulnerable to sudden external disturbances and unpredictable vortices. The aircraft's attitude angles undergo rapid changes due to random turbulence. Consequently, to ensure safety, it is essential to control the aircraft's control surfaces, i.e., ailerons, elevators, and rudder angles, to maintain its static stability. Although classical closed-loop control methods have been widely adopted, their limited adaptability to changing dynamics calls for more robust solutions. Reinforcement learning (RL) offers adaptive capabilities but often demands a large number of training parameters and substantial computational resources, which may be impractical for real-time lightweight aircraft applications. To overcome these limitations, this paper introduces a quantum aircraft with the quantum actorcritic networks-based aircraft control (QACN-AC) algorithm. By utilizing quantum neural networks (QNN), QACN-AC significantly reduces the number of parameters required for training, thus mitigating computational overhead while preserving robust control performance. The QACN-AC's effectiveness is validated through realistic simulations leveraging Boeing's B777 specifications. The results highlight QACN-AC's superiority over conventional RL, evidenced by a$1.25 \times$higher control performance and a$760 \times$reduction in the number of required parameters.
Gyu Seon Kim, Jaehyun Chung, Trung Quang Duong, SooHyun Park, Joongheon Kim
WiOpt1
2025 Quantum Reinforcement Learning for Lightweight LEO Satellite Routing
abstract
Low Earth orbit (LEO) satellite networks have emerged as a promising solution, offering advantages such as lower propagation delay, broader coverage, and rapid deployment capabilities. However, the dynamic topology and frequent handovers inherent in LEO satellite systems, coupled with limited onboard computational resources, necessitate the development of efficient and lightweight routing algorithms. Therefore, this paper proposes quantum reinforcement learning-based satellite routing (QRL-SR) tailored for LEO satellite networks. The QRL-SR algorithm addresses three critical considerations: (i) adapting to the dynamic and time-varying environment of LEO satellite networks; (ii) incorporating LEO satellite geometry by transforming celestial coordinate data, specifically two-line element, into orbital coordinate systems for accurate LEO satellite positioning over time; and (iii) being designed to be lightweight by leveraging QRL to reduce the number of training parameters. The proposed QRL-SR efficiently trains routing policies with fewer parameters, aligning with LEO satellites’ small-size, weight, and power (SWaP) constraints. The primary purpose of the QRL-SR-based LEO satellites is to reduce free space path loss, delay time, and the number of hops needed for routing through the inter-satellite links. Finally, experimental results demonstrate that the QRL-SR achieves routing performance comparable to or outperforms conventional algorithms while significantly reducing computational resources.
Gyu Seon Kim, Sungjoon Lee, In-Sop Cho, SooHyun Park, Joongheon Kim
IEEE Internet Things J.1
2025 Joint Quantum Reinforcement Learning and Neural Myerson Auction for High-Quality Digital-Twin Services in Multitier Networks
abstract
In order to build realistic digital-twin systems, this article proposes a novel two-stage algorithm for high-quality digital-twin services in cloud-assisted multitier networks. In our proposed algorithm, the first stage is quantum multiagent reinforcement learning (QMARL)-based scheduling for differentiated quality control of individual segments of digital-twin virtual objects in our cloud. As the number of segments selected by each edge increases, the edge’s action dimension expands exponentially, posing significant challenges to learning with conventional MARL. To solve this problem, the quantum-inspired MARL-based scheduler is considered in order to reduce the scheduling action dimensions into a logarithmic-scale. For the scheduling formulation, age-of-information (AoI) is also considered for low-latency high-quality digital-twin services. Additionally, the second stage is for the fast and seamless distribution of differentiated quality-controlled segments of virtual objects. For this objective, each user requests its desired segments and one of nearby edges is selected. Among various approaches, this second stage considers second price auction for truthful and distributed computation. Furthermore, low-complexity computation can be realized by avoiding integer-programming-based computation which is NP-hard. The proposed two-stage algorithm achieves performance levels that are 8.33 and 1.18 times higher in terms of reward value in high dimensions and revenue, respectively, compared to other benchmarks.
SooHyun Park, Gyu Seon Kim, Joongheon Kim
IEEE Internet Things J.2
2025 Slimmable Federated Reinforcement Learning for Energy-Efficient Proactive Caching
abstract
Recent advances in deep learning have successfully replaced classical algorithms with machine learning models based on neural networks (NNs). This is particularly prevalent in proactive caching. As NNs grows more capable as their size in terms of storage and computation increases, NN-based proactive caching achieves performance improvement. Nonetheless, there remain challenges in implementing NN-based proactive caching in realistic environments with dynamic user movement. These are due to the fixed structure of NNs that should expand the input size to match the dimensions of the input with the dimensions of the NN’s input units. To address these challenges, this paper proposes a scalable proactive caching framework, named slimmable federated reinforcement learning (SlimFRL). By adopting slimmable neural networks (SNNs) in FRL, our SlimFRL easily adjusts the widths of the SNNs during training according to the number of users. Moreover, due to the scalability of SNNs, our SlimFRL can set the appropriate input dimension while not using imputation, leading to performance improvement. This paper also validates the performance and advantages of SlimFRL in terms of reward and additional cost functions. Additionally, this paper proposes several training algorithms for SlimFRL and corroborates their superiority with convergence analysis and various experiments.
Hankyul Baek, Gyu Seon Kim, SooHyun Park, Andreas F. Molisch, Joongheon Kim
IEEE Trans. Netw.2
2024 Advanced Taxiing Path Guidance Using Multi-Agent Reinforcement Learning for Air Traffic Management
Sungjoon Lee, Gyu Seon Kim, SooHyun Park, Joongheon Kim
WiOpt2
2024 Markov Decision Policies for Distributed Angular Routing in LEO Mobile Satellite Constellation Networks
abstract
This article proposes a distributed angular routing algorithm in time-varying dynamic low Earth orbit (LEO) satellite constellation networks. For designing satellite routing algorithms, it is essential to consider 1) distributed operation due to the difficulty in global centralized computation and 2) angle-based computation under the consideration of orbit coordinate systems. Therefore, our proposed routing algorithm is based on distributed angular computation. Moreover, the proposed algorithm is designed by the Markov decision process (MDP) for discrete-time sequential decision making in time-varying LEO satellite networks. As a result, this article proposes an MDP-based distributed angular routing (MDAR) algorithm for seamless LEO routing. Based on the reward formulation in terms of angular differences in MDP formulation, our proposed distributed angular routing algorithm pursues orbit-geometrically straight-line data delivery from the source to its associated destination. Finally, our proposed routing algorithm is evaluated in the realistic environment with real-world satellite data, i.e., two line elements (TLEs), and the results confirm that our proposed algorithm outperforms the others in terms of routing success rate, reward convergence, and successful throughput.
SooHyun Park, Gyu Seon Kim, Soyi Jung, Joongheon Kim
IEEE Internet Things J.2
2023 Multi-Agent Deep Reinforcement Learning for Efficient Passenger Delivery in Urban Air Mobility
abstract
It has been considered that urban air mobility (UAM), also known as drone-taxi or electrical vertical takeoff and landing (eVTOL), will play a key role in future transportation. By putting UAM into practical future transportation, several benefits can be realized, i.e., (i) the total travel time of passengers can be reduced compared to traditional transportation and (ii) there is no environmental pollution and no special labor costs to operate the system because electric batteries will be used in UAM system. However, there are various dynamic and uncertain factors in the flight environment, i.e., passenger sudden service requests, battery discharge, and collision among UAMs. Therefore, this paper proposes a novel cooperative multiagent deep reinforcement learning (MADRL) algorithm based on centralized training and distributed execution (CTDE) concepts for reliable and efficient passenger delivery in UAM networks. According to the performance evaluation results, we confirm that the proposed algorithm outperforms other existing algorithms in terms of the number of serviced passengers increase (30%) and the waiting time per serviced passenger decrease (26% ).
Chanyoung Park 0002, SooHyun Park, Gyu Seon Kim, Soyi Jung, Joongheon Kim
ICC3