EDBT 2026 Demo / reviewers in the wild / expert
Kaiyan Zhao
dblp:328/6324
· DBLP profile ↗
16ranked-venue papers
2as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot ManipulationabstractHumanoid robots exhibit significant potential in executing diverse human-level skills. However, current research predominantly relies on data-driven approaches that necessitate extensive training datasets to achieve robust multimodal decision-making capabilities and generalizable visuomotor control. These methods raise concerns due to the neglect of geometric reasoning in unseen scenarios and the inefficient modeling of robot-target relationships within the training data, resulting in a significant waste of training resources. To address these limitations, we present the Recurrent Geometric-prior Multimodal Policy (RGMP), an end-to-end framework that unifies geometric-semantic skill reasoning with data-efficient visuomotor control. For perception capabilities, we propose the Geometric-prior Skill Selector, which infuses geometric inductive biases into a vision language model, producing adaptive skill sequences for unseen scenes with minimal spatial common sense tuning. To achieve data-efficient robotic motion synthesis, we introduce the Adaptive Recursive Gaussian Network, which parameterizes robot-object interactions as a compact hierarchy of Gaussian processes that recursively encode multi-scale spatial relationships, yielding dexterous, data-efficient motion synthesis even from sparse demonstrations. Evaluated on both our humanoid robot and desktop robot, the RGMP framework achieves 87% task success in generalization tests and exhibits 5× greater data efficiency than the state-of-the-art model. This performance underscores its superior cross-domain generalization, paving the way for more versatile and data-efficient robotic systems. Xuetao Li, Wenke Huang 0003, Nengyuan Pan, Kaiyan Zhao, Songhua Yang, Mengde Li, Mang Ye, Jifeng Xuan, Miao Li 0002 |
AAAI | 4 |
| 2026 | Explore to Learn: Latent Exploration Through Disentangled Synergy Patterns for Reinforcement Learning in Overactuated ControlabstractControl in high-dimensional action spaces remains a fundamental challenge in reinforcement learning (RL), primarily due to inefficient exploration of the action space. While recent methods attempt to guide exploration, they often fall short of achieving the agility and coordination exhibited in biological motor control. Inspired by how organisms exploit muscle synergies for efficient movement, we propose Explore to Learn (ETL), a two-stage framework that first discovers fundamental synergy patterns and then leverages them for task-specific policy learning. In the first stage, ETL discovers underlying synergy patterns by deploying a targeted exploration policy. These patterns are modeled as latent directions in a low-dimensional space, along which the agent is guided to collect diverse and structured muscle activation trajectories. A variational autoencoder (VAE) is then trained to encode high-dimensional actions into a latent space whose dimensions correspond to the synergy patterns. In the second stage, the policy is trained entirely in this synergy-aware latent space, producing synergy coefficients that the decoder maps back to full-dimensional muscle actions. This structured representation significantly reduces the complexity of learning, while the decoder is further fine-tuned to enhance expressiveness and generalization across downstream tasks. Extensive experiments across musculoskeletal environments and the DMControl suite demonstrate that ETL consistently outperforms prior methods in both exploration efficiency and control performance, achieving superior scalability and generalization in overactuated control tasks. Kaiyan Zhao, Xu Li 0039, Yan Li 0122, Steven Morad, Leong Hou U |
AAAI | 2 |
| 2026 | DSAP: Enhancing Generalization in Goal-Conditioned Reinforcement LearningabstractGoal-conditioned Reinforcement Learning (RL) is a promising direction for training agents capable of tackling a variety of tasks. However, generalizing to new goals in different environments remains a central challenge for goal-conditioned RL agents. Existing methods often rely on state abstraction, which involves learning abstracted state representations by excluding irrelevant features, to improve generalization. Despite their success in simplified settings, these methods often fail to generalize effectively to realistic environments with varied goals. In this work, we propose to enhance generalization through state abstraction from the perspective of causal inference. We hypothesize that the generalization gap arises in part due to unobserved confounders: latent variables that simultaneously influence both the global and goal states. To address this, we introduce Deconfounded State Abstraction for Policy learning (DSAP), a novel framework that mitigates backdoor confounding by employing a learned causal graph as a *proxy* for the hidden confounders. We provide theoretical analysis demonstrating that DSAP improves both the learning process and the generalization capability of goal-conditioned policies. Extensive experiments across different settings of multiple benchmarks show that our method significantly outperforms existing methods. Kaiyan Zhao, Yan Li 0122, Furui Liu, Leong Hou U |
AAAI | 2 |
| 2026 | Latent State-Predictive Exploration for Deep Reinforcement LearningabstractReinforcement learning (RL) has achieved promising results in continuous control tasks, where efficient exploration of the state space is crucial for success. However, many recent RL approaches still struggle with sample inefficiency and insufficient exploration for long-horizon tasks, particularly in environments characterized by high-dimensional and complex state spaces. To address these challenges, we propose a novel exploration framework, Latent State Predictive Exploration (LSPE). The core idea behind LSPE is to endow the agent with a form of ``foresight" to enhance exploration in long-horizon settings. Specifically, LSPE employs a state encoder to learn compact latent representations from high-dimensional visual observations, effectively filtering out irrelevant or noisy information. To further enrich and stabilize these representations, we incorporate a diffusion-based self-predictive module that enforces temporal consistency by predicting future states, thereby improving both exploration and downstream predictive control. Additionally, we introduce an Exploration Reward Function (ERF) that explicitly encourages the agent to visit novel latent states. This reward signal promotes more efficient and scalable exploration in complex environments. We evaluate LSPE across a diverse set of challenging long-horizon navigation and manipulation tasks, spanning simulation environments such as Habitat and Robosuite, as well as deployment on a real robot in a **physical indoor environment**. Experimental results show that LSPE substantially enhances exploration efficiency and scales effectively to complex, high-dimensional tasks. Kaiyan Zhao, Borong Zhang, Yan Li 0122, Leong Hou U |
AAAI | 2 |
| 2026 | NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement LearningabstractNeologism-aware machine translation 1 aims to translate source sentences containing neologisms into target languages.This field remains underexplored compared with general machine translation (MT).In this paper, we propose an agentic framework, NeoAMT, for neologism-aware machine translation equipped with a Wiktionary-based search toolkit.Specifically, we first construct a dedicated dataset for neologism-aware machine translation and build a search toolkit grounded in Wiktionary.The dataset covers 16 languages and 75 translation directions in total, derived from approximately 10 million records of an English Wiktionary dump.The retrieval corpus of the search toolkit is also constructed from around 3 million cleaned records of the same dump.We then leverage the dataset and toolkit to train a translation agent via reinforcement learning (RL) and to evaluate the accuracy of neologismaware machine translation.Furthermore, we propose an RL training framework featuring a novel reward design and an adaptive rollout generation strategy that exploits "translation difficulty" to further improve the translation quality of translation agents using our search toolkit 2 . Zhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa Tsuruoka |
ACL (1) | 2 |
| 2026 | FedPAD: Aggregation-free federated learning with prototype-based adaptive distillation
Kaiyan Zhao, He Zhu 0002, Xiaoguang Niu |
Knowl. Based Syst. | 2 |
| 2025 | HeMoRa: Unsupervised Heuristic Consensus Sampling for Robust Point Cloud RegistrationabstractHeuristic information for consensus set sampling is essential for correspondence-based point cloud registration, but existing approaches typically rely on supervised learning or expert-driven parameter tuning. In this work, we propose HeMoRa, a new unsupervised framework that trains a Heuristic information Generator (HeGen) to estimate sampling probabilities for correspondences using a Multi-order Reward Aggregator (MoRa) loss. The core of MoRa is to train HeGen through extensive trials and feedback, enabling unsupervised learning. While this process can be implemented using policy optimization, directly applying the policy gradient to optimize HeGen presents challenges such as sensitivity to noise and low reward efficiency. To address these issues, we propose a Maximal Reward Propagation (MRP) mechanism that enhances the training process by prioritizing noise-free signals and improving reward utilization. Experimental results show that equipped with HeMoRa, the consensus set sampler achieves improvements in both robustness and accuracy. For example, on the 3DMatch dataset with FCGF feature, the registration recall of our unsupervised methods (Ours+SM and Ours+SC2) even outperforms the state-of-the-art supervised method VBreg. Our code is available at HeMoRa. Shaocheng Yan, Kaiyan Zhao, Zhenjun Zhao, Yongjun Zhang 0002, Jiayuan Li 0001 |
CVPR | 3 |
| 2025 | BILE: An Effective Behavior-based Latent Exploration Scheme for Deep Reinforcement LearningabstractEfficient exploration of state spaces is critical for the success of deep reinforcement learning (RL). While many methods leverage exploration bonuses to encourage exploration instead of relying solely on extrinsic rewards, these bonus-based approaches often face challenges with learning efficiency and scalability, especially in environments with high-dimensional state spaces. To address these issues, we propose BehavIoral metric-based Latent Exploration (BILE). The core idea is to learn a compact representation within the behavioral metric space that preserves value differences between states. By introducing additional rewards to encourage exploration in this latent space, BILE drives the agent to visit states with higher value diversity and exhibit more behaviorally distinct actions, leading to more effective exploration of the state space. Additionally, we present a novel behavioral metric for efficient and robust training of the state encoder, backed by theoretical guarantees. Extensive experiments on high-dimensional environments, including realistic indoor scenarios in Habitat, robotic tasks in Robosuite, and challenging discrete Minigrid benchmarks, demonstrate the superiority and scalability of our method over other approaches. Kaiyan Zhao, Yan Li 0122, Leong Hou U |
IJCAI | 2 |
| 2025 | Efficient Diversity-based Experience Replay for Deep Reinforcement LearningabstractExperience replay is widely used to improve learning efficiency in reinforcement learning by leveraging past experiences. However, existing experience replay methods, whether based on uniform or prioritized sampling, often suffer from low efficiency, particularly in real-world scenarios with high-dimensional state spaces. To address this limitation, we propose a novel approach, Efficient Diversity-based Experience Replay (EDER). EDER employs a determinantal point process to model the diversity between samples and prioritizes replay based on the diversity between samples. To further enhance learning efficiency, we incorporate Cholesky decomposition for handling large state spaces in realistic environments. Additionally, rejection sampling is applied to select samples with higher diversity, thereby improving overall learning efficacy. Extensive experiments are conducted on robotic manipulation tasks in MuJoCo, Atari games, and realistic indoor environments in Habitat. The results demonstrate that our approach not only significantly improves learning efficiency but also achieves superior performance in high-dimensional, realistic environments. Kaiyan Zhao, Yan Li 0122, Leong Hou U, Xiaoguang Niu |
IJCAI | 1 |
| 2025 | BiCAM: A Bidirectional Contextualized Attentive Model for Analyzing the Correlation of Heterogeneous Security EventsabstractAs the Internet continues to evolve, modern information technology infrastructures are constantly under attack and need to be continuously monitored for timely responses. Different devices and detection platforms generate heterogeneous security events that are sent to security operations centers, where security operators investigate those events and identify potential threats. Unfortunately, it is impossible to manually analyze such a huge number of events, leading to “alert fatigue.” Despite a substantial amount of effort having been made to aggregate redundant related alerts, the effectiveness of previous works was essentially restrained by their limited relation learning and explaining abilities. In this work, we propose the bidirectional contextualized attentive model (BiCAM), a novel contextual analysis model that uses a self-supervised deep learning approach to automatically correlate security events in relation to their bidirectional context. It is developed by designing an encoder–decoder architecture that consists of bidirectional gated recurrent units and an attention mechanism to capture both sequential and nonsequential relations of previous and subsequent alerts and provide explainability information for the security operators. In addition, we introduce a bidirectional encoder representations from transformers (BERT)-based embedding method to deal with the heterogeneity of security events, enhancing our model's accommodation to the changes of detectors. We comprehensively evaluate our model on real-world datasets containing over 11M events generated by detectors from 8 different vendors. We found that our model enables accurate, unsupervised correlation extraction; and outperforms the state-of-the-art (SOTA) work when applying event relevance to semiautomatically classify security events (e.g., the$F1$-score of classification is improved by 4.3% and the false positive rate dropped to 1.39%). Lihua Yin, Kaiyan Zhao, Kexiang Qian, Daojuan Zhang |
IEEE Trans. Reliab. | 4 |
| 2024 | Leveraging Multi-lingual Positive Instances in Contrastive Learning to Improve Sentence EmbeddingabstractLearning multilingual sentence embeddings is a fundamental task in natural language processing.Recent trends in learning both monolingual and multilingual sentence embeddings are mainly based on contrastive learning (CL) among an anchor, one positive, and multiple negative instances.In this work, we argue that leveraging multiple positives should be considered for multilingual sentence embeddings because (1) positives in a diverse set of languages can benefit cross-lingual learning, and (2) transitive similarity across multiple positives can provide reliable structural information for learning.In order to investigate the impact of multiple positives in CL, we propose a novel approach, named MPCL, to effectively utilize multiple positive instances to improve the learning of multilingual sentence embeddings.Experimental results on various backbone models and downstream tasks demonstrate that MPCL leads to better retrieval, semantic similarity, and classification performance compared to conventional CL.We also observe that in unseen languages, sentence embedding models trained on multiple positives show better cross-lingual transfer performance than models trained on a single positive instance. Kaiyan Zhao, Qiyu Wu 0001, Xin-Qiang Cai, Yoshimasa Tsuruoka |
EACL (1) | 1 |
| 2024 | AARR-Net: An Attention Assistance Feature Fusion and Model Recursive Recovery Network for Category-Level 6D Object Pose Estimation
Kaiyan Zhao, Shaowu Wu, Xiaoguang Niu |
ICONIP (7) | 2 |
| 2024 | Rethinking Exploration in Reinforcement Learning with Effective Metric-Based Exploration BonusabstractEnhancing exploration in reinforcement learning (RL) through the incorporation of intrinsic rewards, specifically by leveraging *state discrepancy* measures within various metric spaces as exploration bonuses, has emerged as a prevalent strategy to encourage agents to visit novel states. The critical factor lies in how to quantify the difference between adjacent states as *novelty* for promoting effective exploration.
Nonetheless, existing methods that evaluate state discrepancy in the latent space under $L_1$ or $L_2$ norm often depend on count-based episodic terms as scaling factors for exploration bonuses, significantly limiting their scalability. Additionally, methods that utilize the bisimulation metric for evaluating state discrepancies face a theory-practice gap due to improper approximations in metric learning, particularly struggling with *hard exploration* tasks. To overcome these challenges, we introduce the **E**ffective **M**etric-based **E**xploration-bonus (EME). EME critically examines and addresses the inherent limitations and approximation inaccuracies of current metric-based state discrepancy methods for exploration, proposing a robust metric for state discrepancy evaluation backed by comprehensive theoretical analysis. Furthermore, we propose the diversity-enhanced scaling factor integrated into the exploration bonus to be dynamically adjusted by the variance of prediction from an ensemble of reward models, thereby enhancing exploration effectiveness in particularly challenging scenarios.
Extensive experiments are conducted on hard exploration tasks within Atari games, Minigrid, Robosuite, and Habitat, which illustrate our method's scalability to various scenarios. The project website can be found at https://sites.google.com/view/effective-metric-exploration. Kaiyan Zhao, Furui Liu, Leong Hou U |
NeurIPS | 2 |
| 2024 | Team-wise effective communication in multi-agent reinforcement learning
Kaiyan Zhao, Renzhi Dong, Yali Du 0001, Furui Liu, Mingliang Zhou 0001, Leong Hou U |
Auton. Agents Multi Agent Syst. | 2 |
| 2024 | CAG-Malconv: A Byte-Level Malware Detection Method With CBAM and Attention-GRUabstractWith the rise of generative artificial intelligence, malware creation has become more accessible, leading to a surge in malware and its variants. Traditional detection methods struggle to keep pace with this evolution. Dynamic analysis, though detailed, is resource intensive and susceptible to variations in computer hardware and simulation environments. Static analysis, on the other hand, faces the challenge of discerning valuable features from an extensive pool, especially for software across diverse architectures. To tackle these issues, we propose a binary sample classification approach based on raw bytes, named CAG-Malconv, which incorporates Convolutional Block Attention Module (CBAM) and Bidirectional Gated Recurrent Unit (BiGRU) to extract byte-level features. We evaluated it on two datasets with 48,000 samples of different file types and families. It outperforms state-of-the-art methods based on advanced features and raw bytes in terms of accuracy (ACC), Area Under the Curve (AUC), F1 score, and recall. Furthermore, it allows for the visualization of raw samples, facilitating the precise identification of malicious components like C&C URLs and encryption loops by analyzing activation patterns in hidden layers, thus streamlining malware investigative procedures. Honghui Fan, Lihua Yin, Shijie Jia 0001, Kaiyan Zhao |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2022 | LIBKDV: A Versatile Kernel Density Visualization Library for Geospatial AnalyticsabstractKernel density visualization (KDV) has been widely used in many geospatial analysis tasks, including traffic accident hotspot detection, crime hotspot detection, and disease outbreak detection. Although KDV can be supported by many scientific, geographical, and visualization software tools, none of these tools can support high-resolution KDV with large-scale datasets. Therefore, we develop the first versatile programming library, called LIBKDV, based on the set of our complexity-optimized algorithms. Given the high efficiency of these algorithms, LIBKDV not only accelerates the KDV computation but also enriches KDV-based geospatial analytics, including bandwidth-tuning analysis and spatiotemporal analysis, which cannot be natively and feasibly supported by existing software tools. In this demonstration, participants will be invited to use our programming library to explore interesting hotspot patterns on large-scale traffic accident, crime, and COVID-19 datasets. Tsz Nam Chan, Pak Lon Ip, Kaiyan Zhao, Leong Hou U, Byron Choi, Jianliang Xu |
Proc. VLDB Endow. | 3 |