VLDB 2026 Research / reviewers in the wild / expert
Yingnan Zhao 0002
dblp:92/4786-2
· DBLP profile ↗
18ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-1761-3669ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Adaptive Humanoid Control via Multi-Behavior Distillation and Reinforced Fine-TuningabstractHumanoid robots are promising to learn a diverse set of human-like locomotion behaviors, including standing up, walking, running, and jumping. However, existing methods predominantly require training independent policies for each skill, yielding behavior-specific controllers that exhibit limited generalization and brittle performance when deployed on irregular terrains and in diverse situations. To address this challenge, we propose Adaptive Humanoid Control (AHC) that adopts a two-stage framework to learn an adaptive humanoid locomotion controller across different skills and terrains. Specifically, we first train several primary locomotion policies and perform a multi-behavior distillation process to obtain a basic multi-behavior controller, facilitating adaptive behavior switching based on the environment. Then, we perform reinforced fine-tuning by collecting online feedback in performing adaptive behaviors on more diverse terrains, enhancing terrain adaptability for the adaptive behavior controller. We conduct experiments in both simulation and real-world experiments in Unitree G1 robots. The results show that our method exhibits strong adaptability across various situations and terrains. Yingnan Zhao 0002, Xinmiao Wang, Dan Lu 0004, Qilong Han, Peng Liu 0008, Chenjia Bai |
AAAI | 1 |
| 2026 | DIFT: Protecting Contrastive Learning Against Data Poisoning Backdoor AttacksabstractContrastive learning (CL) is a popular learning paradigm that excels in extracting meaningful representations from unlabeled data. Recent studies have shown that CL is highly vulnerable to backdoor attacks. Current defenses against backdoor attacks in CL are primarily reactive and post-training. That is, the detection and elimination of backdoors are executed in the deployment phase of a given well-trained model. However, these post-training defenses are usually prone to degrading model utility and resource-intensive, causing that the backdoor detection and elimination from a fully-trained model is quite challenging. To address this issue, we argue for a fundamental perspective, i.e., integrating the defense into the model's training phase, and propose a novel framework to mitigate the backdoor in CL, namely Density-Based Identification and Fine-Tuning (DIFT). Specifically, DIFT identifies potential poisoned samples during the early training phase via detecting embeddings with abnormal poisoning characteristic in the feature space. Then, to remove backdoors and preserve model utility, the detected poisoned samples are leveraged to fine-tune the model, and the remaining clean samples are further involved into training the model after the fine-tuning. DIFT, as a proactive training-time defense, avoids the problematic backdoor removal and the high computational cost associated with those reactive post-training methods. We empirically evaluate DIFT on various CL algorithms against backdoor attack. Experimental results demonstrate that our method exhibits promising defense effectiveness while maintaining model's clean data accuracy. Yulin Jin, Qingqing Ye 0001, Zhibiao Guo, Kun Fang 0004, Ruochen Du, Yingnan Zhao 0002, Haibo Hu 0001 |
AAAI | 7 |
| 2026 | Continuous alignment of multi-target preferences via instructed diffusion model
Yingnan Zhao 0002, Xinmiao Wang, Dan Lu 0004, Qilong Han, Chenjia Bai |
Pattern Recognit. | 1 |
| 2026 | An Imitative Reinforcement Learning Framework for Pursuit-Lock-Launch MissionsabstractUnmanned combat aerial vehicle (UCAV) within-visual-range (WVR) engagement, referring to a fight between two or more UCAVs at close quarters, plays a decisive role on the aerial battlefields. With the development of artificial intelligence, WVR engagement progressively advances toward intelligent and autonomous modes. However, autonomous WVR engagement policy learning is hindered by challenges such as weak exploration capabilities, low learning efficiency, and unrealistic simulated environments. To overcome these challenges, we propose a novel imitative reinforcement learning framework, which efficiently leverages expert data while enabling autonomous exploration. The proposed framework not only enhances learning efficiency through expert imitation but also ensures adaptability to dynamic environments via autonomous exploration with reinforcement learning. Therefore, the proposed framework can learn a successful policy of “pursuit-lock-launch” for UCAVs. To support data-driven learning, we establish an environment based on the Harfang3D sandbox. The extensive experimental results indicate that the proposed framework excels in this multistage task and significantly outperforms state-of-the-art reinforcement learning and imitation learning methods. Thanks to the ability of imitating experts and autonomous exploration, our framework can quickly learn the critical knowledge in complex aerial combat tasks, achieving up to a 100% success rate and demonstrating excellent robustness. Siyuan Li 0003, Rongchang Zuo, Bofei Liu, Yaoyu He, Peng Liu 0008, Yingnan Zhao 0002 |
ACM Trans. Auton. Adapt. Syst. | 6 |
| 2026 | RPD: Regional Prior Distillation for Breast Cancer Diagnosis in Ultrasound ImagesabstractBreast cancer is the leading cause of death among women worldwide. Ultrasound imaging is an important means for the early detection of breast cancer to improve the survival rate. Due to the shortage of experienced sonographers, computer-aided systems for breast cancer recognition become particularly important. Some recent studies analyze tumor types in lesion regions but rely on predefined ROIs. Some other studies recognize cancer in the whole ultrasound image, but always suffer from the extremely variable proportion, location and quantity of the tumor lesions. In this paper, we propose a Regional Prior Distillation (RPD) framework for breast cancer diagnosis in ultrasound images. To enhance the analysis of the tumor region, we propose an Image-Cross Attention (ICA) to fuse the predefined ROI prior information with ultrasound images by training a prior-fused model. To remove the constraint of predefined ROIs, we propose a Distribution Distillation Learning (DDL) to distill the prior-fused sample distribution from the prior-fused model into a diagnostic model, which analyzes the disease from only ultrasound images, based on the knowledge distillation paradigm of the teacher-student framework. Comprehensive experiments are conducted on multi-institutional datasets to validate the proposed RPD framework. The results demonstrate the following points. The ICA fuses regional prior information adequately, leading to a high-performance prior-fused model. The DDL distills the prior information effectively, enhancing the diagnostic model to focus on the tumor lesions. The performance of the diagnostic model surpasses that of current SOTA methods. In addition, the diagnostic model is robust to slight perturbations and achieves good generalization performance. Yingnan Zhao 0002, Dan Lu 0004, Yanchen Xu, Jiexiao Xue, Xi Chen 0110, Jingchi Jiang |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | Two Time-Scale DRL for Service Caching and Task Offloading in Cross-Domain Marine NetworksabstractWith increasing computational demands and limited network resources in marine environments, efficient service caching and task offloading have become critical. In such environments, Autonomous Underwater Vehicles (AUVs) rely on Unmanned Surface Vehicles (USVs) as relays, forming a cross-domain network comprising underwater acoustic and above-water RF links. However, the heterogeneity in bandwidth, latency, and bit error rates introduces challenges for reachability analysis and delay estimation. This paper addresses the joint optimization of caching, task offloading, and resource allocation in a cross-domain marine network composed of offshore base stations, USVs, and AUVs. To tackle the inherent heterogeneity in network links and decision timescales, we formulate the problem as a two-time-scale Hierarchical Markov Decision Process (H-MDP) and propose a Two Time-Scale Deep Reinforcement Learning (T2S-DRL) approach that integrates a hybrid policy network and a lightweight structure-aware action masking mechanism. The large time-scale agent optimizes caching decisions, while the short time-scale agent focuses on offloading and resource allocation. Extensive simulations show that our approach significantly reduces task execution delay and energy consumption, validating its effectiveness. Zhaoxiang Huang, Zhiwen Yu 0001, Liang Wang 0017, Yingnan Zhao 0002, Huan Zhou 0002, Bin Guo 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | MSC-TFN: Multi-Scale Convolutional and Transformer Fusion Network for Breast Tumor Classification in Ultrasound ImagesabstractUltrasound imaging has become an important modality for early breast cancer diagnosis. However, breast tumors in ultrasound images often exhibit location uncertainty, scale variations, and boundary ambiguity, which severely hinder accurate automatic classification. This paper proposes a Multiscale Convolutional and Transformer Fusion Network (MSCTFN). To address location uncertainty and boundary ambiguity, a Convolutional Block Attention Module (CBAM) is employed to process the extracted features, highlighting key tumor regions and enhancing boundary representations. To handle scale variations, a Multi-Scale Convolution (MSC) structure is designed to process features with varying receptive fields in parallel, which effectively models the multi-scale features. Additionally, a Dynamic Weighted FPN module is introduced to adaptively fuse multiscale features, significantly enhancing the model's discriminative ability and generalization performance. Experimental results on the BUSI and UDIAT breast ultrasound datasets validate the effectiveness of the proposed method, demonstrating strong generalization capability and promising clinical potential. Dan Lu 0004, Yanchen Xu, Yingnan Zhao 0002, Hongyang Zhao |
BIBM | 3 |
| 2025 | Auxiliary Reward Generation With Transition Distance Representation LearningabstractReinforcement learning (RL) has shown strengths in challenging sequential decision-making problems. The reward function in RL is crucial to the learning performance, as it quantifies the degree of task completion. In real-world problems, the rewards are predominantly human-designed, which requires laborious tuning, and is susceptible to human cognitive biases. To achieve automatic auxiliary reward generation, we propose a novel representation learning approach that can measure the “transition distance” between states. Building upon these representations, we introduce an auxiliary reward generation technique for both single-task and skill-chaining scenarios without the need for human knowledge. Furthermore, we theoretically show that the proposed auxiliary rewards maintain the policy invariance property, i.e., the generated rewards will not hurt the policy optimality under the original rewards. In the experiment section, we evaluate the proposed approach in both online and offline learning settings in a wide range of tasks, including robot manipulation and locomotion. The experiment results demonstrate the effectiveness of measuring the transition distance and the induced improvement by auxiliary rewards, which promotes better learning efficiency and increases convergent stability. Beyond that, we demonstrate that the learned manipulation policy with the auxiliary rewards in a simulator can be transferred to the real robot, as shown inhttps://sites.google.com/view/transition-distance-rp/tdrp. Note to Practitioners—The motivation for this paper arises from the need for a technique that enhances robot skill-learning efficiency and performance in both single-task and skill-chaining scenarios. Our research primarily focuses on robot arm manipulation tasks. To accelerate the policy learning process and improve policy performance for executing these tasks, we introduce an auxiliary reward generation technique for both single-task and skill-chaining scenarios without requiring human expertise. This technique leverages the proposed novel representation learning approach, which can measure the “transition distance” between states. During each policy training round, the robot receives a dense reshaped reward created by our approach. Using the policy trained by our method, we successfully control a real Franka Panda robot arm to complete various manipulation tasks. Siyuan Li 0003, Shijie Han, Yingnan Zhao 0002, Yiqin Yang, Qianchuan Zhao, Peng Liu 0008 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Joint Task Offloading and Migration Optimization in UAV-Enabled Dynamic MEC NetworksabstractUAV-enabled multi-access edge computing (MEC) is expanding possibilities for integrated space-air-ground networks, especially in the 5G era and beyond. In this scenario, tasks from mobile users (MUs) are offloaded to nearby UAVs for execution, with results returned upon completion. However, the unpredictable mobility of MUs, coupled with dynamic network conditions and fluctuating resource availability, can degrade the reliability of communication links, leading to increased delivery latency, particularly for tasks involving large computational results. To meet stringent QoS requirements, adaptive task migration across UAVs is essential to minimize latency. To address this issue, in this paper, we first investigateComputationTaskMiGration (CTMiG) problem in UAV-enabled dynamic MEC networks, focusing on joint optimization of task-serving (offloading and migration) decisions to reduce latency for all MUs. We propose the ILCTS algorithm, an imitation learning-based joint optimization method that adaptively adjusts scheduling strategies in response to environmental changes. An improved PPO algorithm is first proposed to train a policy and generate expert data, followed by generative adversarial imitation learning to imitate the data and continuously explore new ones through online learning to enhance the policy. Experimental results demonstrate that our algorithm achieves superior performance in training accuracy and average latency compared to other representative methods. Liang Wang 0017, Bingnan Shen, Lianbo Ma 0004, Yao Zhang 0005, Yingnan Zhao 0002, Hongzhi Guo 0005, Zhiwen Yu 0001, Bin Guo 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Review of Deep Learning-Based Entity Alignment Methods
Dan Lu 0004, Guoyu Han, Yingnan Zhao 0002, Qilong Han |
GPC (1) | 3 |
| 2023 | Anomaly Detection of Industrial Data Based on Multivariate Multi Scale Analysis
Dan Lu 0004, Siao Li, Yingnan Zhao 0002, Qilong Han |
GPC (1) | 3 |
| 2023 | UEBCS: Software Development Technology Based on Component Selection
Yingnan Zhao 0002, Xuezhao Qi, Dan Lu 0004 |
GPC (1) | 1 |
| 2023 | Research on Script-Based Software Component Development
Yingnan Zhao 0002, Dan Lu 0004 |
GPC (1) | 1 |
| 2023 | Variational Diversity Maximization for Hierarchical Skill Discovery
Yingnan Zhao 0002, Peng Liu 0008, Wei Zhao 0008, Xianglong Tang |
Neural Process. Lett. | 1 |
| 2023 | Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement LearningabstractEfficient exploration remains a challenging problem in reinforcement learning, especially for tasks where extrinsic rewards from environments are sparse or even totally disregarded. Significant advances based on intrinsic motivation show promising results in simple environments but often get stuck in environments with multimodal and stochastic dynamics. In this work, we propose a variational dynamic model based on the conditional variational inference to model the multimodality and stochasticity. We consider the environmental state-action transition as a conditional generative process by generating the next-state prediction under the condition of the current state, action, and latent variable, which provides a better understanding of the dynamics and leads to a better performance in exploration. We derive an upper bound of the negative log likelihood of the environmental transition and use such an upper bound as the intrinsic reward for exploration, which allows the agent to learn skills by self-supervised exploration without observing extrinsic rewards. We evaluate the proposed method on several image-based simulation tasks and a real robotic manipulating task. Our method outperforms several state-of-the-art environment model-based exploration approaches. Chenjia Bai, Peng Liu 0008, Kaiyu Liu, Lingxiao Wang 0003, Yingnan Zhao 0002, Lei Han 0001, Zhaoran Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Generating attentive goals for prioritized hindsight reinforcement learning
Peng Liu 0008, Chenjia Bai, Yingnan Zhao 0002, Chenyao Bai, Wei Zhao 0008, Xianglong Tang |
Knowl. Based Syst. | 3 |
| 2020 | Obtaining accurate estimated action values in categorical distributional reinforcement learning
Yingnan Zhao 0002, Peng Liu 0008, Chenjia Bai, Wei Zhao 0008, Xianglong Tang |
Knowl. Based Syst. | 1 |
| 2019 | An exploratory rollout policy for imagination-augmented agents
Peng Liu 0008, Yingnan Zhao 0002, Wei Zhao 0008, Xianglong Tang, Zichan Yang |
Appl. Intell. | 2 |