EDBT 2026 Demo / reviewers in the wild / expert
Shaokang Dong
dblp:213/0943
· DBLP profile ↗
22ranked-venue papers
8as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic StudyabstractSpeech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12× faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency. Xiaoran Fan, Yangfan Gao, Jingfei Xiong, Hang Yan 0001, Yifei Cao, Zhihao Zhang 0002, Zhiheng Xi, Yuhao Zhou 0005, Senjie Jin, Changhao Jiang, Junjie Ye 0005, Ming Zhang 0030, Zhenhua Han, Yunke Zhang, Demei Yan, Shaokang Dong, Tao Gui |
AAAI | 20 |
| 2026 | LMFENet: A hybrid local-global and multi-scale feature extraction network for oil spill type classification using sentinel-1 imagery
Shaokang Dong, Jiangfan Feng |
Expert Syst. Appl. | 1 |
| 2026 | MMKLTrack: Illumination-robust UAV tracking with progressive localization
Kuan Yin, Jiangfan Feng, Shaokang Dong, Ying Long |
Expert Syst. Appl. | 3 |
| 2026 | A unified and efficient training framework for open-ended non-transitive games
Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001 |
Neural Networks | 1 |
| 2025 | Beyond Mandatory Federations: Balancing Egoism, Utilitarianism and Egalitarianism in Mixed-Motive GamesabstractIn the field of mixed-motive games, extensive multi-agent learning studies have explored the balance between egoism (individual interest), utilitarianism (collective interest), and egalitarianism (fairness). Traditional approaches often rely on manually designed reward functions, social norms, and alliance/federation mechanisms to transition agents from individualistic behaviors toward cooperative strategies. However, these methods typically require all agents to share private local information or to mandatorily participate in federations, which is impractical in real-world applications. To address these issues, this paper proposes a Flexible-Participation Federation (FPF) framework that allows agents to participate in the federation voluntarily. Furthermore, we extend the federation from a global to a Local Multi-Federation (LMF) framework, enabling agents to form multiple localized federations, thereby promoting more efficient and adaptive cooperation. Theoretical evidence demonstrates that the global FPF model, along with the discrepancy between decentralized egoistic policies and federated utilitarian policies, achieves an O(1/T) convergence rate. Agents in the LMF framework also reach consensus within a sublinear gap. Extensive experiments show that agents opting out of federation participation experience a reduction in egoism, and our approach outperforms multiple baselines in terms of both utilitarianism and egalitarianism. Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001 |
AAAI | 1 |
| 2025 | TASO: Task-Aligned Sparse Optimization for Parameter-Efficient Model AdaptationabstractDaiye Miao, Yufang Liu, Jie Wang, Changzhi Sun, Yunke Zhang, Demei Yan, Shaokang Dong, Qi Zhang, Yuanbin Wu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Daiye Miao, Yufang Liu, Jie Wang 0126, Changzhi Sun, Yunke Zhang, Demei Yan, Shaokang Dong, Qi Zhang 0001, Yuanbin Wu |
EMNLP | 7 |
| 2025 | Towards Empowerment Gain through Causal Structure Learning in Model-Based Reinforcement LearningabstractIn Model-Based Reinforcement Learning (MBRL), incorporating causal structures into dynamics models provides agents with a structured understanding of the environments, enabling efficient decision.
Empowerment as an intrinsic motivation enhances the ability of agents to actively control their environments by maximizing the mutual information between future states and actions.
We posit that empowerment coupled with causal understanding can improve controllability, while enhanced empowerment gain can further facilitate causal reasoning in MBRL.
To improve learning efficiency and controllability, we propose a novel framework, Empowerment through Causal Learning (ECL), where an agent with the awareness of causal dynamics models achieves empowerment-driven exploration and optimizes its causal structure for task learning.
Specifically, ECL operates by first training a causal dynamics model of the environment based on collected data. We then maximize empowerment under the causal structure for exploration, simultaneously using data gathered through exploration to update causal dynamics model to be more controllable than dense dynamics model without causal structure. In downstream task learning, an intrinsic curiosity reward is included to balance the causality, mitigating overfitting.
Importantly, ECL is method-agnostic and is capable of integrating various causal discovery methods.
We evaluate ECL combined with $3$ causal discovery methods across $6$ environments including pixel-based tasks, demonstrating its superior performance compared to other causal MBRL methods, in terms of causal discovery, sample efficiency, and asymptotic performance. Hongye Cao, Shaokang Dong, Tianpei Yang, Jing Huo, Yang Gao 0001 |
ICLR | 4 |
| 2025 | Coordinating Multi-Agent Reinforcement Learning via Dual Collaborative Constraints
Shaokang Dong, Shangdong Yang, Yujing Hu, Wenbin Li 0006, Yang Gao 0001 |
Neural Networks | 2 |
| 2025 | Multi-Task Multi-Agent Reinforcement Learning With Interaction and Task RepresentationsabstractMulti-task multi-agent reinforcement learning (MT-MARL) is capable of leveraging useful knowledge across multiple related tasks to improve performance on any single task. While recent studies have tentatively achieved this by learning independent policies on a shared representation space, we pinpoint that further advancements can be realized by explicitly characterizing agent interactions within these multi-agent tasks and identifying task relations for selective reuse. To this end, this article proposes Representing Interactions and Tasks (RIT), a novel MT-MARL algorithm that characterizes both intra-task agent interactions and inter-task task relations. Specifically, for characterizing agent interactions, RIT presents the interactive value decomposition to explicitly take the dependency among agents into policy learning. Theoretical analysis demonstrates that the learned utility value of each agent approximates its Shapley value, thus representing agent interactions. Moreover, we learn task representations based on per-agent local trajectories, which assess task similarities and accordingly identify task relations. As a result, RIT facilitates the effective transfer of interaction knowledge across similar multi-agent tasks. Structurally, RIT develops universal policy structure for scalable multi-task policy learning. We evaluate RIT against multiple state-of-the-art baselines in various cooperative tasks, and its significant performance under both multi-task and zero-shot settings demonstrates its effectiveness. Shaokang Dong, Shangdong Yang, Yujing Hu, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Optimistic Value Instructors for Cooperative Multi-Agent Reinforcement LearningabstractIn cooperative multi-agent reinforcement learning, decentralized agents hold the promise of overcoming the combinatorial explosion of joint action space and enabling greater scalability. However, they are susceptible to a game-theoretic pathology called relative overgeneralization that shadows the optimal joint action. Although recent value-decomposition algorithms guide decentralized agents by learning a factored global action value function, the representational limitation and the inaccurate sampling of optimal joint actions during the learning process make this problem still. To address this limitation, this paper proposes a novel algorithm called Optimistic Value Instructors (OVI). The main idea behind OVI is to introduce multiple optimistic instructors into the value-decomposition paradigm, which are capable of suggesting potentially optimal joint actions and rectifying the factored global action value function to recover these optimal actions. Specifically, the instructors maintain optimistic value estimations of per-agent local actions and thus eliminate the negative effects caused by other agents' exploratory or sub-optimal non-cooperation, enabling accurate identification and suggestion of optimal joint actions. Based on the instructors' suggestions, the paper further presents two instructive constraints to rectify the factored global action value function to recover these optimal joint actions, thus overcoming the RO problem. Experimental evaluation of OVI on various cooperative multi-agent tasks demonstrates its superior performance against multiple baselines, highlighting its effectiveness. Jianqi Wang, Yujing Hu, Shaokang Dong, Wenbin Li 0012, Tangjie Lv, Changjie Fan, Yang Gao 0001 |
AAAI | 5 |
| 2024 | Multi-Agent Exploration via Self-Learning and Social LearningabstractSelf-learning and social learning stand as two pivotal constituents in multi-agent exploration. Inspired by the fact that animals and humans explore unfamiliar environments to learn survival skills by training themselves using unlabeled data and replicating others’ successful experiences, we propose a multi-agent reinforcement learning method, named Self-Learning and Social Learning (S2L), which aims to address the complex tasks caused by sparse rewards and intricate sequential structures. Specifically, in Self-Learning, we incorporate both task-specific and task-agnostic intrinsic rewards. These incentives steer individual agents towards exploration and comprehension of the environment. Furthermore, in Social Learning, different independent agents can implicitly share the successful experience by observing others in view and without additional communication or parameter-sharing overhead. Finally, experimental evaluation of S2L on the complex task characterized by sparse rewards and intricate sequential structures demonstrates its superior performance against other competing exploration baselines. Shaokang Dong, Wubing Chen, Hongye Cao, Yang Gao 0001 |
ICASSP | 1 |
| 2024 | Multi-Agent Sparse Interaction Modeling is an Anomaly Detection ProblemabstractMost real-world multi-agent tasks exhibit the characteristic of sparse interaction, wherein agents interact with each other in a limited number of crucial states while largely acting independently. Effectively modeling the sparse interaction and leveraging the learned interaction structure to instruct agents’ learning processes can enhance the efficiency of multi-agent reinforcement learning algorithms. However, it remains unclear how to identify these specific interactive states solely through trials and errors within current multi-agent tasks. To address this challenge, this paper introduces a novel algorithm called Sparse Interaction as Anomaly (SIA), which innovatively casts the sparse interaction modeling into an anomaly detection problem. The underlying intuition is that interactive states appear rarely in agents’ trajectories and exhibit distinct dynamics compared to other commonplace states. Building upon this insight, SIA first employs variational inference to model the latent dynamics of agents’ trajectories. It then designates states with anomalous dynamics as the elusive interactive states and subsequently instructs agents to explore these states more extensively. This facilitates the emergence of interactive behaviors and promotes the learning of multi-agent policies. Experimental evaluation of SIA across various multi-agent tasks demonstrates its superior performance against multiple baselines, highlighting its effectiveness. Shaokang Dong, Shangdong Yang, Hongye Cao, Yang Gao 0001 |
ICASSP | 2 |
| 2024 | Decentralized Counterfactual Value with Threat Detection for Multi-Agent Reinforcement Learning in mixed cooperative and competitive environments
Shaokang Dong, Shangdong Yang, Yang Gao 0001 |
Expert Syst. Appl. | 1 |
| 2024 | Egoism, utilitarianism and egalitarianism in multi-agent reinforcement learning
Shaokang Dong, Shangdong Yang, Bo An 0001, Wenbin Li 0006, Yang Gao 0001 |
Neural Networks | 1 |
| 2024 | WToE: Learning When to Explore in Multiagent Reinforcement LearningabstractExisting multiagent exploration works focus on how to explore in the fully cooperative task, which is insufficient in the environment with nonstationarity induced by agent interactions. To tackle this issue, we propose When to Explore (WToE), a simple yet effective variational exploration method to learn WToE under nonstationary environments. WToE employs an interaction-oriented adaptive exploration mechanism to adapt to environmental changes. We first propose a novel graphical model that uses a latent random variable to model the step-level environmental change resulting from interaction effects. Leveraging this graphical model, we employ the supervised variational auto-encoder (VAE) framework to derive a short-term inferred policy from historical trajectories to deal with the nonstationarity. Finally, agents engage in exploration when the short-term inferred policy diverges from the current actor policy. The proposed approach theoretically guarantees the convergence of the Q -value function. In our experiments, we validate our exploration mechanism in grid examples, multiagent particle environments and the battle game of MAgent environments. The results demonstrate the superiority of WToE over multiple baselines and existing exploration methods, such as MAEXQ, NoisyNets, EITI, and PR2. Shaokang Dong, Hangyu Mao, Shangdong Yang, Shengyu Zhu 0001, Wenbin Li 0006, Jianye Hao, Yang Gao 0001 |
IEEE Trans. Cybern. | 1 |
| 2023 | Leveraging transition exploratory bonus for efficient exploration in Hard-Transiting reinforcement learning problems
Shangdong Yang, Shaokang Dong, Xingguo Chen |
Future Gener. Comput. Syst. | 3 |
| 2023 | Online attentive kernel-based temporal difference learning
Xingguo Chen, Guang Yang 0066, Shangdong Yang, Shaokang Dong, Yang Gao 0001 |
Knowl. Based Syst. | 5 |
| 2022 | DDMA: Discrepancy-Driven Multi-agent Reinforcement Learning
Yujing Hu, Pinzhuo Tian, Shaokang Dong, Yang Gao 0001 |
PRICAI (3) | 4 |
| 2022 | Application of Artificial Intelligence in an Unsupervised Algorithm for Trajectory Segmentation Based on Multiple Motion FeaturesabstractWith the development of the wireless network, location‐based services (e.g., the place of interest recommendation) play a crucial role in daily life. However, the data acquired is noisy, massive, it is difficult to mine it by artificial intelligence algorithm. One of the fundamental problems of trajectory knowledge discovery is trajectory segmentation. Reasonable segmentation can reduce computing resources and improvement of storage effectiveness. In this work, we propose an unsupervised algorithm for trajectory segmentation based on multiple motion features (TS‐MF). The proposed algorithm consists of two steps: segmentation and mergence. The segmentation part uses the Pearson coefficient to measure the similarity of adjacent trajectory points and extract the segmentation points from a global perspective. The merging part optimizes the minimum description length (MDL) value by merging local sub‐trajectories, which can avoid excessive segmentation and improve the accuracy of trajectory segmentation. To demonstrate the effectiveness of the proposed algorithm, experiments are conducted on two real datasets. Evaluations of the algorithm’s performance in comparison with the state‐of‐the‐art indicate the proposed method achieves the highest harmonic average of purity and coverage. Wenjin Xu, Shaokang Dong |
Wirel. Commun. Mob. Comput. | 2 |
| 2020 | Consistent MetaReg: Alleviating Intra-task Discrepancy for Better Meta-knowledgeabstractIn the few-shot learning scenario, the data-distribution discrepancy between training data and test data in a task usually exists due to the limited data. However, most existing meta-learning approaches seldom consider this intra-task discrepancy in the meta-training phase which might deteriorate the performance. To overcome this limitation, we develop a new consistent meta-regularization method to reduce the intra-task data-distribution discrepancy. Moreover, the proposed meta-regularization method could be readily inserted into existing optimization-based meta-learning models to learn better meta-knowledge. Particularly, we provide the theoretical analysis to prove that using the proposed meta-regularization, the conventional gradient-based meta-learning method can reach the lower regret bound. The extensive experiments also demonstrate the effectiveness of our method, which indeed improves the performances of the state-of-the-art gradient-based meta-learning models in the few-shot classification task. Pinzhuo Tian, Lei Qi 0001, Shaokang Dong, Yinghuan Shi, Yang Gao 0001 |
IJCAI | 3 |
| 2019 | Measuring Structural Similarities in Finite MDPsabstractIn this paper, we investigate the structural similarities within a finite Markov decision process (MDP). We view a finite MDP as a heterogeneous directed bipartite graph and propose novel measures for state similarity and action similarity in a mutual reinforcement manner. We prove that the state similarity is a metric and the action similarity is a pseudometric. We also establish the connection between the proposed similarity measures and the optimal values of the MDP. Extensive experiments show that the proposed measures are effective. Hao Wang 0013, Shaokang Dong, Ling Shao 0001 |
IJCAI | 2 |
| 2017 | Job and Candidate Recommendation with Big Data Support: A Contextual Online Learning ApproachabstractTo make every user conveniently have access to his or her most interested jobs and candidates (recommendation items) in the current employment market, the recruitment networks need to meet the demand of fast and accurate recommendation. But a key challenge is that the total items may have a large quantity in the big data scenarios. And another problem is that the personalization of different users is diverse. In order to handle these challenges, this paper proposes a mining and prediction system for job and candidate recommendation with contextual online learning. It predicts a proper item by utilizing the feedback reward of previous users in the nearby context region. Besides that, we introduce a Monte-Carlo Tree Search (MCTS) method in which the similar items can be amalgamated into a cluster to reduce the computing load. Our algorithm can achieve sublinear regret and space complexity. Finally, some experiments are conducted to test our algorithm based on a large database from \emph{Work4} (the global leader in social and mobile recruiting), which can show the outstanding performance of our algorithm when compared with other existing algorithms. Shaokang Dong, Zijian Lei, Pan Zhou 0001, Kaigui Bian, Guanghui Liu 0001 |
GLOBECOM | 1 |