Song Han 0001

dblp:80/806-1 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-9072-420XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Computer networks · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
abstract
The emergence of Large Language Models (LLMs) with strong reasoning capabilities marks a significant milestone, unlocking new frontiers in complex problem-solving. However, training these reasoning models, typically using Reinforcement Learning (RL), encounters critical efficiency bottlenecks: response generation during RL training exhibits a persistent long-tail distribution, where a few very long responses dominate execution time, wasting resources and inflating costs. To address this, we propose TLT, a system that accelerates reasoning RL training losslessly by integrating adaptive speculative decoding. Applying speculative decoding in RL is challenging due to the dynamic workloads, evolving target model, and draft model training overhead. TLT overcomes these obstacles with two synergistic components: (1) Adaptive Drafter, a lightweight draft model trained continuously on idle GPUs during long-tail generation to maintain alignment with the target model at no extra cost; and (2) Adaptive Rollout Engine, which maintains a memory-efficient pool of pre-captured CUDAGraphs and adaptively select suitable SD strategies for each input batch. Evaluations demonstrate that TLT achieves over 1.7x end-to-end RL training speedup over state-of-the-art systems, preserves the model accuracy, and yields a high-quality draft model as a free byproduct suitable for efficient deployment. Code is released at https://github.com/mit-han-lab/fastrl.
Qinghao Hu 0004, Shang Yang, Junxian Guo, Xiaozhe Yao, Yujun Lin 0001, Yuxian Gu, Han Cai, Chuang Gan 0001, Ana Klimovic, Song Han 0001
ASPLOS (2)10
2026 Cross-layer feature consistency and dual-transformer residual framework for underwater image enhancement
Xinbin Li, Song Han 0001, Hui Dang, Muge Li
Eng. Appl. Artif. Intell.3
2026 A heterogeneous reinforcement learning approach for joint relay selection and power allocation in time-varying UASNs with energy harvesting
Song Han 0001, Yuming He, Aijia Li, Xinbin Li, Zhixin Liu 0001, Lei Yan 0010, Tongwei Zhang, Huimin Kang
Inf. Sci.1
2025 Hierarchical-Learning-Based Task Assignment for Heterogeneous Multi-AUV-UG Collaborative System to Collect Data From Underwater Sensors
abstract
In this study, the task assignment problem for heterogeneous underwater vehicle collaborative system, which involves autonomous underwater vehicles (AUVs) and underwater gliders (UGs), is studied for high-efficiency data collection. UGs and AUVs show different motion modes. The advantages of different motion modes can be mutually complemented to achieve the preference-matched task assignment results, which show greatly promising prospect to enhance the data collection efficiency. Most of existing underwater task assignment algorithms focus on the single-type vehicles, which can not be applied to the heterogeneous system. To address this issue, a hierarchical learning algorithm is proposed. Firstly, based on the evaluated emergency degree of tasks, the preliminary-task-assignment hierarchy is proposed to assign the emergency tasks to AUVs and assign the non-emergency tasks to UGs, thereby achieving the preference-matched task assignment. Therefore, the collaborative efficiency of heterogeneous system can be enhanced. Then, in the UG-task-assignment hierarchy, the adaptive serial cluster mechanism is proposed to extract the high utility task-connectivity regions for UGs, thereby fully leveraging the UG advantages in region data collection. Furthermore, in the AUV-task-assignment hierarchy, the extended self-organizing mapping neural network is constructed to eliminate the disorganization of neuronal loops. As a result, the crossed paths of AUVs can be excluded to reduce the energy consumption. Finally, the superior performance is verified by numerical results.
Jiaao Zhao, Song Han 0001, Xinbin Li, Junzhi Yu 0001, Zhixin Liu 0001, Tongwei Zhang
IEEE Trans. Intell. Transp. Syst.2
2025 An Extended Bandit-Based Game Scheme for Distributed Joint Resource Allocation in Underwater Acoustic Communication Networks
abstract
This paper investigates a joint discrete-channel and continuous-power allocation problem for multi-user underwater acoustic communication networks. The unknown underwater acoustic Channel State Information (CSI) and the distributed optimization requirement make the proposed hybrid discrete-continuous optimization problem full of challenges. Firstly, an adversarial multi-player bandit game model is formulated, which enables each user to independently optimize its own strategy, thereby achieving the distributed decision. In the strategic game, the Multi-armed Bandit (MAB) learning theory is exploited to achieve the best response strategy of independent user without prior CSI. Secondly, an evolutive finite discrete strategy pool learning structure is proposed to achieve an efficient search for the hybrid discrete-continuous space. The constant evolvement of strategy pool endows the proposed MAB-based algorithm with the ability to search the whole continuous power space, thereby avoiding missing the superior strategy caused by the discretization of continuous space. Thirdly, a selection probability setting rule is proposed, which promotes the exploration-exploitation balance for the dynamic strategy pool, thereby improving the learning efficiency. Finally, simulation results demonstrate the superiority of the proposed algorithm.
Xinbin Li, Song Han 0001, Junzhi Yu 0001, Zhixin Liu 0001, Tongwei Zhang
IEEE Trans. Netw. Serv. Manag.3
2024 Multi-hop relay selection for underwater acoustic sensor networks: A dynamic combinatorial multi-armed bandit learning approach
Xinbin Li, Song Han 0001, Zhixin Liu 0001, Haihong Zhao, Lei Yan 0010
Comput. Networks3
2024 Joint Multiple Resources Allocation for Underwater Acoustic Cooperative Communication in Time-Varying IoUT Systems: A Double Closed-Loop Adversarial Bandit Approach
abstract
This article deals with a joint multiple resources (relay, channel, and power) allocation problem for underwater acoustic (UWA) cooperative communication in time-varying Internet of Underwater Things scenarios. The strong coupling of multiple resources and the unknown time-varying characteristic of UWA communication scenes make the joint optimization problem full of challenges. To address this issue, the adversarial multiarmed bandit online learning model without any prior channel information and statistic assumptions is employed. Furthermore, a double closed-loop learning structure with multiple intelligent experts assistance is proposed. Multiple experts embedded in inner loop can intelligently learn the derived inferential information to provide more efficient advice for the player in outer loop, thereby enriching learning information and enhancing learning ability. In addition, the expert diversity learning mechanism is proposed to fully reflect the characteristics of seeking advantages and avoiding disadvantages in the double closed-loop learning structure. As a result, the learning speed and performance of the proposed algorithms are significantly improved. The superiorities of the proposed algorithms are demonstrated through numerical results.
Song Han 0001, Xinbin Li, Junzhi Yu 0001, Zhixin Liu 0001, Lei Yan 0010, Tongwei Zhang
IEEE Internet Things J.1
2024 The Unified Task Assignment for Underwater Data Collection With Multi-AUV System: A Reinforced Self-Organizing Mapping Approach
abstract
This article deals with the task assignment problem for multiple autonomous underwater vehicles to efficiently collect underwater data from sensors. We formulate a unified framework to consistently address the heterogeneous task assignment problem (nonemergency and emergency cases) without strictly distinguishing the mixed cases. First, a unified problem, which bridges the gap between different constraints and optimization objectives of different cases, is constructed. Then, the proposed reinforced self-organizing mapping algorithm is reinforced in three aspects: the regional learning rate, the self-configuring neuron (SCN) strategy, and the workload balance mechanism. Specifically, the proposed regional learning rate comprehensively considers the individual worth of tasks and the topology to generate the regional learning rate of dynamic task regions, which consists of dynamic remaining tasks and the reconstructed topology. Based on this idea, the constructed unified problem can be solved consistently. Furthermore, the proposed SCN strategy optimizes the neuron population both in quality and quantity, and guides the update of neurons with enriched historical information to improve the mapping ability. This strategy greatly improves learning efficiency and applicability in a wide range of scenarios. Meanwhile, the proposed workload balance mechanism takes into consideration of both the work capability and consumed energy to extend the continuous working capability. The numerical results validate the effectiveness and adaptability of the proposed unified task assignment framework.
Song Han 0001, Xinbin Li, Junzhi Yu 0001, Tongwei Zhang, Zhixin Liu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Joint Resource Allocation for Time-Varying Underwater Acoustic Communication System: A Self-Reflection Adversarial Bandit Approach
abstract
This study deals with a joint channel selection and power allocation problem for time-varying underwater acoustic communication system. Without any prior channel information, designing a highly adaptable resource allocation algorithm to cope with the fast time-varying environment is a very challenging issue. To address this issue, a hierarchical learning approach, which is combined with adversarial multiarmed bandit theory and outdated pilot-based feedback information, is proposed. The proposed learning approach can online optimize joint resource allocate strategy without any prior channel state information. Specifically, a hierarchical self-reflection learning structure is proposed to offer different learning manners and spaces for the actual played information and outdated feedback information, thereby balancing the exploitation and exploration to cope with the time-varying environment effectively. Further, an integration learning structure is proposed to alleviate the solving difficulty and policy explosion of joint multiple substrategies problem. The user can rapidly achieve a few superior strategies in low-dimension space, then efficiently search the expected optimal strategy in high-dimension space, as a result, the learning efficiency is significantly improved. The proposed algorithms show strong tolerance for delay and noncomplete information due to the elaborate learning structures. The superiority of the proposed algorithms is demonstrated through numerical results.
Song Han 0001, Xinbin Li, Junzhi Yu 0001, Haihong Zhao, Zhixin Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2023 Multiple attentional path aggregation network for marine object detection
Xinbin Li, Yankai Feng, Song Han 0001
Appl. Intell.4
2023 Underwater vision enhancement based on GAN with dehazing evaluation
Xinbin Li, Yankai Feng, Song Han 0001
Appl. Intell.4
2023 Collaboration-Aware Relay Selection for AUV in Internet of Underwater Network: Evolving Contextual Bandit Learning Approach
abstract
In Internet of Underwater Things, data collection is assisted by autonomous underwater vehicle (AUV) to enhance the reliable transmission. AUV acts as a mobile collector and transmits the collected data to the station via relay nodes. However, the highly mobile nature of AUV needs an adaptive and efficient relay selection scheme for achieving good capacity performance. In this article, we propose a new contextual multiarmed bandit with evolving relay set (CMAB-ERS) learning framework, which successfully addresses crucial issues, including dynamic environment conditions and evolving relay set. To deal with the evolving relay set, CMAB-ERS incorporates collaborative effects into inference as well as learning processes, the new relays will acquire prior knowledge by having experienced nodes sharing observations, reducing the learning time significantly. To overcome the uncertainty of environmental information, we exploit the contextual environment factors to assist relay reward estimation and execute time-sensitive parameter update after every transmit–receive cycle, aiming for minimizing potential loss due to the time-varying channel. Correspondingly, the collaboration-aware online contextual bandit learning (COCBL) algorithm is designed that enables AUV to switch optimal relay adaptively and promises high-capacity transmission. Further, we rigorously prove the convergence of the COCBL algorithm by considering the evolving relay set and give its upper bound on the cumulative regret. Finally, extensive simulation results elucidate the effectiveness of the proposed COCBL.
Haihong Zhao, Xinbin Li, Song Han 0001, Lei Yan 0010, Junzhi Yu 0001
IEEE Internet Things J.3
2023 Adaptive Relay Selection Strategy in Underwater Acoustic Cooperative Networks: A Hierarchical Adversarial Bandit Learning Approach
abstract
Relay selection solutions for underwater acoustic cooperative networks suffer significant performance degradation as they fail to adapt to incomplete information, noisy interference and overwhelming dynamics. To address this challenge, a hierarchical adversarial multi-armed bandit learning framework by proposing an online reward estimation layer is designed to improve adaptive relay decision control. In online reward estimation layer, adaptive Kalman filter estimator is developed to properly handle noisy observation to support accurate reward. Meanwhile, an online predict mechanism is projected for all relays to enrich learning information. Furthermore, based on estimate error variance, an adaptive exploration structure is developed to accelerate the balance between exploration and exploitation. All gathered information are exploited to learn relay quality for the decision-making. Accordingly, we present a Hierarchical Adversarial Bandit Learning (HABL) algorithm to fully exploit the heuristic interaction between the hierarchical framework. HABL integrates reward estimation, information prediction, adaptive exploration and decision making carefully in a holistic algorithm to maximize the learning efficiency. Thereby, the HABL-based relay selection algorithm has higher system throughput and lower communication cost. Further, we rigorously analyze the convergence of HABL algorithm and give its upper bound on the cumulative regret. Finally, extensive simulations elucidate the effectiveness of the HABL.
Haihong Zhao, Xinbin Li, Song Han 0001, Lei Yan 0010, Junzhi Yu 0001
IEEE Trans. Mob. Comput.3
2022 An adaptive multi-zone geographic routing protocol for underwater acoustic sensor networks
Xinbin Li, Haihong Zhao, Song Han 0001, Lei Yan 0010
Wirel. Networks4
2021 A modified nature-inspired meta-heuristic methodology for heterogeneous unmanned aerial vehicle system task assignment problem
Song Han 0001, Xinbin Li, Yi Yuan 0002
Soft Comput.2
2020 Adaptive OFDM underwater acoustic transmission: An adversarial bandit approach
Haihong Zhao, Xinbin Li, Song Han 0001, Lei Yan 0010, Xin-Ping Guan
Neurocomputing3
2020 A Robust Game-Based Algorithm for Downlink Joint Resource Allocation in Hierarchical OFDMA Femtocell Network System
abstract
Femtocell is a promising technology for wireless service networks to facilitate sustainable and efficient services for users. This paper deals with a downlink joint channel assignment and power allocation problem with multiple channels, users, constraints, and uncertainties in the orthogonal frequency division of a multiple-access hierarchical femtocell network system. Specifically, a hierarchical robust Stackelberg game, which aims to achieve robust equilibrium, is first proposed for resource allocation with uncertainties. Then, a low-complexity, low-interference, high-efficiency, and high-performance algorithm is presented to handle the complex robust joint allocation problem. Considering the demand capacity of macro-base stations, an efficient fitness function in conjunction with particle swarm optimization-constriction factor, is utilized to yield the best response in the upper game to satisfy multiple constraints. Meanwhile, an iterative waterfilling algorithm is exploited to achieve the best response of femto-base stations in the lower game. Lastly, a stop protocol is established for two-tier users to accomplish an efficient robust Stackelberg equilibrium. Comparative results demonstrated that the proposed algorithm is superior to the existing game-based algorithms.
Junzhi Yu 0001, Song Han 0001, Xinbin Li
IEEE Trans. Syst. Man Cybern. Syst.2
2019 MAB-based two-tier learning algorithms for joint channel and power allocation in stochastic underwater acoustic communication networks
Song Han 0001, Xinbin Li, Lei Yan 0010, Zhixin Liu 0001, Xin-Ping Guan
Soft Comput.1
2018 Game-based hierarchical multi-armed bandit learning algorithm for joint channel and power allocation in underwater acoustic communication networks
Song Han 0001, Xinbin Li, Lei Yan 0010, Zhixin Liu 0001, Xin-Ping Guan
Neurocomputing1
2018 Joint resource allocation in underwater acoustic communication networks: A game-based hierarchical adversarial multiplayer multiarmed bandit algorithm
Song Han 0001, Xinbin Li, Lei Yan 0010, Jiajie Xu 0003, Zhixin Liu 0001, Xin-Ping Guan
Inf. Sci.1
2016 Joint Relay Selection and Power Allocation in Underwater Cognitive Acoustic Cooperative System with Limited Feedback
abstract
We study the problem of joint relay selection and power allocation in a underwater cooperative system with multiple users assisted by multiple relays. Due to the harsh underwater environments, the channel state information (CSI) at the transmitter is imperfect, which leads to the performance degrading in the underwater cooperative acoustic system. Therefore, we analyze the cooperative underwater acoustic channel with limited feedback to increase the sum-rate of the system. Meanwhile, different from other researches, we do not only focus on the single system scenario, but also consider the presence of nearby acoustic activities and the problem of joint relay selection and power allocation is solved in a cognitive acoustic (CA) scenario. Thus the codebook of interference CSI and the codebook of quantized relay selection and power allocation strategy are designed, respectively. Simulation results show that a few bits feedback can significantly improve the performance of the CA cooperative acoustic system.
Lei Yan 0010, Xinbin Li, Kai Ma 0001, Jing Yan 0001, Song Han 0001
VTC Spring5
2016 Distributed hierarchical game-based algorithm for downlink power allocation in OFDMA femtocell networks
Song Han 0001, Xinbin Li, Zhixin Liu 0001, Xin-Ping Guan
Comput. Networks1
2016 Hierarchical-game-based algorithm for downlink joint subchannel and power allocation in OFDMA femtocell networks
Song Han 0001, Xinbin Li, Zhixin Liu 0001, Xin-Ping Guan
J. Netw. Comput. Appl.1