VLDB 2026 Research / reviewers in the wild / expert
Jikun Kang
dblp:299/0233
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM PersonalizationabstractLinfeng Du, Ye Yuan, Zichen Zhao, Fuyuan Lyu, Emiliano Penaloza, Xiuying Chen, Zipeng Sun, Jikun Kang, Laurent Charlin, Xue Liu, Haolun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Linfeng Du, Ye Yuan 0017, Fuyuan Lyu, Emiliano Penaloza, Xiuying Chen, Zipeng Sun, Jikun Kang, Laurent Charlin, Xue (Steve) Liu, Haolun Wu |
ACL (1) | 8 |
| 2026 | Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable PersonalizationabstractWeixu Zhang, Ye Yuan, Changjiang Han, Yuxing Tian, Zipeng Sun, Linfeng Du, Jikun Kang, Hong Kang, Xue Liu, Haolun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weixu Zhang, Ye Yuan 0017, Changjiang Han, Yuxing Tian, Zipeng Sun, Linfeng Du, Jikun Kang, Hong Kang, Xue Liu 0004, Haolun Wu |
ACL (1) | 7 |
| 2024 | Think Before You Act: Decision Transformers with Working MemoryabstractDecision Transformer-based decision-making agents have shown the ability to generalize across multiple tasks. However, their performance relies on massive data and computation. We argue that this inefficiency stems from the forgetting phenomenon, in which a model memorizes its behaviors in parameters throughout training. As a result, training on a new task may deteriorate the model’s performance on previous tasks. In contrast to LLMs’ implicit memory mechanism, the human brain utilizes distributed memory storage, which helps manage and organize multiple skills efficiently, mitigating the forgetting phenomenon. Inspired by this, we propose a working memory module to store, blend, and retrieve information for different downstream tasks. Evaluation results show that the proposed method improves training efficiency and generalization in Atari games and Meta-World object manipulation tasks. Moreover, we demonstrate that memory fine-tuning further enhances the adaptability of the proposed architecture. Jikun Kang, Romain Laroche, Xingdi Yuan, Adam Trischler, Xue (Steve) Liu, Jie Fu 0001 |
ICML | 1 |
| 2024 | Incentive Temperature Control for Green Colocation Data Centers via Reinforcement LearningabstractIncreasing supply air temperatures is a rule-of-thumb approach to reduce cooling energy usage of data centers (DCs). However, colocation DCs are short of incentive programs to move tenants from the current over-cooling strategy despite the expanding allowable temperature ranges of the computing equipment. This paper considers an essential incentive mechanism, in which the DC operator offers monetary incentives to offset tenants’ electricity payments. We propose an encoder-embedded multi-agent reinforcement learning solution to let the operator agent and tenant agents collaboratively find their policies for deciding the incentives and supply air temperatures, respectively, which are coupled in determining the DC’s total cooling power usage. The solution does not require the cooling power model, which is complex and in general unavailable in practice. Moreover, as each tenant agent learns in the other tenants’ latent state spaces defined by their pre-trained variational autoencoders, only encoded tenants’ states are exchanged, thereby mitigating information leakage concerns. Extensive trace-driven evaluation and comparison with three baselines show that our solution effectively incentivizes tenants to move from the over-cooling strategy and achieves substantial cooling power savings. Duc Van Le, Jikun Kang, Rui Tan 0001, Xue (Steve) Liu |
IWQoS | 3 |
| 2023 | Multi-Agent Attention Actor-Critic Algorithm for Load Balancing in Cellular NetworksabstractIn cellular networks, User Equipment (UE) handoff from one Base Station (BS) to another, giving rise to the load balancing problem among the BSs. To address this problem, BSs can work collaboratively to deliver a smooth migration (or handoff) and satisfy the UEs' service requirements. This paper formulates the load balancing problem as a Markov game and proposes a Robust Multi-agent Attention Actor-Critic (Robust-MA3C) algorithm that can facilitate collaboration among the BSs (i.e., agents). In particular, to solve the Markov game and find a Nash equilibrium policy, we embrace the idea of adopting a nature agent to model the system uncertainty. Moreover, we utilize the self-attention mechanism, which encourages high-performance BSs to assist low-performance BSs. In addition, we consider two types of schemes, which can facilitate load balancing for both active UEs and idle UEs. We carry out extensive evaluations by simulations, and simulation results illustrate that, compared to the state-of-the-art MARL methods, Robust-MA3C scheme can improve the overall performance by up to 45%. Jikun Kang, Di Wu 0044, Ju Wang 0003, Ekram Hossain 0001, Xue Liu 0004, Gregory Dudek |
ICC | 1 |
| 2022 | A Generalized Load Balancing Policy With Multi-Teacher Reinforcement LearningabstractAlthough reinforcement learning (RL) shows advantages in cellular network load balancing, it suffers from a low generalization ability, preventing it from real-world applications. Specifically, if network traffic pattern changes, the learned RL policy cannot adapt accordingly, resulting in system performance degradation. To address this issue, we propose a Multi-teacher MOdel BAsed Reinforcement Learning algorithm (MOBA), which leverages multi-teacher knowledge distillation theory to learn a generalized load balancing policy for adapting the real-world traffic pattern changes. The key is that different teachers represent different traffic patterns, and can learn various system models. By distilling and transferring the teacher knowledge, the student network is able to learn a generalized system model that covers different traffic patterns and unseen situations. Moreover, to improve the robustness of multi-teacher knowledge transfer, we learn a set of student models and use an ensemble method to jointly predict system dynamics. Results show that, compared with state-of-the-art RL methods, MOBA improves the minimal throughput and total throughput of a cellular network by up to 28.6% and 23.2%. Results also show that MOBA improves the training efficiency by up to 64%. Jikun Kang, Ju Wang 0003, Chengming Hu, Xue Liu 0004, Gregory Dudek |
GLOBECOM | 1 |
| 2021 | AFB: Improving Communication Load Forecasting Accuracy with Adaptive Feature BoostingabstractPrediction of key system characteristics, such as the communication load, is required to overcome the delays in wireless communication systems. State-of-The-Art (SOTA) approaches mostly apply existing Neural Network (NN) structures, and extract latent features purely based on their sensitivity to the forecasting accuracy. This way of feature extraction may neglect some non-obvious yet informative dimensions in the model input, leading to inaccurate forecasting results. In this paper, we present an Adaptive Feature Boosting (AFB) approach, which integrates multiple AutoEncoders (AEs) to automatically extract robust and comprehensive latent features for communication load forecasting. The recurrent and residual connections among the AEs make sure that the extracted latent features are representative for all input dimensions. With more comprehensive information extracted from the history, the forecasting accuracy is thus improved. We evaluate AFB against existing approaches on a real-world dataset that contains Call Detail Records (CDRs) of the Milan city over a period of two months. The evaluation shows that our AFB-based approach achieves 35.2% more accurate load forecasting results than the SOTA deep approaches. Chengming Hu, Xi Chen 0009, Ju Wang 0003, Jikun Kang, Yi Tian Xu, Xue Liu 0004, Di Wu 0044, Seowoo Jang, Intaik Park, Gregory Dudek |
GLOBECOM | 5 |
| 2021 | Load Balancing for Communication Networks via Data-Efficient Deep Reinforcement LearningabstractWithin a cellular network, load balancing between different cells is of critical importance to network performance and quality of service. Most existing load balancing algorithms are manually designed and tuned rule-based methods where near-optimality is almost impossible to achieve. These rule-based meth-ods are difficult to adapt quickly to traffic changes in real-world environments. Given the success of Reinforcement Learning (RL) algorithms in many application domains, there have been a number of efforts to tackle load balancing for communication systems using RL-based methods. To our knowledge, none of these efforts have addressed the need for data efficiency within the RL framework, which is one of the main obstacles in applying RL to wireless network load balancing. In this paper, we formulate the communication load balancing problem as a Markov Decision Process and propose a data-efficient transfer deep reinforcement learning algorithm to address it. Experimental results show that the proposed method can significantly improve the system performance over other baselines and is more robust to environmental changes. Di Wu 0044, Jikun Kang, Yi Tian Xu, Jimmy Li 0001, Xi Chen 0009, Dmitriy Rivkin, Michael R. M. Jenkin, Taeseop Lee, Intaik Park, Xue Liu 0004, Gregory Dudek |
GLOBECOM | 2 |
| 2021 | Hierarchical Policy Learning for Hybrid Communication Load BalancingabstractDue to the uneven demographic distribution and people’s daily activities, communication systems usually experience highly imbalanced load across different cells. This imbalance leads to unsatisfied users in the congested cells and under-utilized resources in the less-loaded cells. To deal with this issue, existing work migrates the load from heavily loaded cells to lightly loaded cells, by either handing over active mode User Equipment (UEs) to other serving cells, or re-selecting the camping cells for idle mode UEs. In this paper, we further advance the research on Load Balancing (LB) with a hybrid control of both active and idle UEs. This task is challenging, due to the conflicts between Active-UE LB (AULB) and Idle-UE LB (IULB) policies. To overcome this challenge, we propose a Hierarchical Policy Learning (HPL) framework, which coordinates the actions between LB policies with a two-level learning structure. In this way, HPL produces AULB and IULB policies that are better aligned with each other. Extensive simulation results illustrate the efficiency and efficacy of the proposed HPL. Jikun Kang, Xi Chen 0009, Di Wu 0044, Yi Tian Xu, Xue Liu 0004, Gregory Dudek, Taeseop Lee, Intaik Park |
ICC | 1 |