Ya Duan

dblp:375/1970 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0008-7274-390XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
YearPublicationVenuePosition
2026 Neural Rendering and Flow-Assisted Unsupervised Multi-View Stereo for Real-Time Monocular Tracking and Scene Perception
abstract
The existing camera tracking and perception methods mainly rely on sparse SLAM, which limits the dense perception ability of the scene and affects the reliability of auxiliary decision-making. Different from this, this work proposes a real-time tracking and unsupervised dense sensing framework. Firstly, the dense depth value of the scene is predicted by unsupervised multi-view stereo to remove the dependence on labeled data. Then, the quality of synthetic pseudo-reference image is quantified according to the predicted depth map and used as a weighted guidance to train the unsupervised model, thus reducing the ambiguity of feature matching in areas such as specular reflection. Moreover, the sparse optical flow of the keyframes is solved by real-time and robust ORB feature matching operator, which assists the high-precision training of unsupervised depth inference model. To increase the prediction accuracy of occluded area, a novel rendering consistency loss via neural radiance fields is designed to constrain the geometric characteristics of object surface. Finally, dense direct image alignment is performed from a global model to improve the tracking robustness, which is incrementally constructed from dense depth prediction. Extensive experiments on synthetic datasets and real datasets validate the effectiveness and practicability of the proposed work, which is an effective supplement to the existing SLAM work.
Kevin W. Tong, Yandong Cai, Yu-Wen Jie, Ya Duan, Yuhong Hou, Qi Wu 0003
IEEE Trans Autom. Sci. Eng.4
2025 Semantic Encoding Algorithm for Classification and Retrieval of Aviation Safety Reports
abstract
Automated analysis of aviation safety reports is helpful in effectively preventing future accidents and improving emergency response capabilities. To date, there are no publicly available large-scale aviation text similarity datasets, which hinders the successful application of NLP techniques in the aviation domain. We present an automatically created aviation text similarity dataset consisting of more than 500,000 pairs for fine-tuning pretrained language models. Since technical terms have specialized meanings that differ from everyday language, we propose an efficient semantic encoding algorithm to improve the ability of embeddings to adequately represent aviation terms. We provide new solutions and revised evaluation metrics for the classification and the retrieval of safety reports, confirming the reliability of our dataset and the superiority of our algorithm.Note to Practitioners—Text representation is an essential task in natural language processing(NLP). A crucial step towards the successful application of NLP in safety reports analysis is to ensure that aviation texts are adequately encoded. Aiming at the problem of poor ability of current embeddings to represent technical terms, we automatically create an aviation text similarity dataset and propose a semantic encoding algorithm for aviation terms. It is clear that the proposed method has great potential in representation of technical terms, thus providing assistance for downstream tasks such as text classification, information retrieval and question answering.
Yubing Gao, Guangyu Zhu 0001, Ya Duan, Jianfeng Mao
IEEE Trans Autom. Sci. Eng.3
2025 Semi-Supervised Image Domain Adaption for Aerial Refueling Drogue Detection on Embedded Chip Under Foggy Conditions
abstract
The application of aerial refueling technology to UAVs can reduce the dependence on the pilot’s operation, which has unique advantages in carrying out battlefield reconnaissance, monitoring suspicious targets and collecting intelligence through all-weather work. The existing vision-based drogue detection methods are assumed to be carried out under daily lighting conditions, but special weather, such as fog, makes it difficult to identify the characteristics of the drogue, which will greatly degrade the model performance or even fail. Moreover, the traditional computing architecture is difficult to be directly applied to real airborne equipment, so it is necessary to adopt AI processor module with faster and better computing power and supporting parallel computing to meet the requirements of low delay and high security in aerial refueling. Therefore, this work proposes a robust detection network based on image domain adaption. Firstly, an end-to-end image defogging module is designed to deal with foggy image enhancement under weak supervision. Then, knowledge distillation with the teacher-student network is applied to guide the student model to obtain the instance-level features of the unlabeled target domain. In addition, the detection model is compiled and transplanted on the system-on-chip chip of JFMQL100TAI. The comprehensive experimental results on public datasets and real refueling datasets validate the effectiveness and feasibility of the proposed work, which effectively complements the drogue detection of special autonomous aerial refueling tasks. Note to Practitioners—As a widely used refueling technology in the field of national defense, probe-and-drogue refueling has developed from manual control docking to monitoring auxiliary docking. However, it is difficult for pilots to accurately and quickly obtain the relative position of the refueling drogue through visual perception. In this work, a semi-supervised drogue detection network for special aerial refueling task is designed. The proposed work has good application potential in refueling scenes, which can provide fast and accurate drogue positioning under foggy conditions.
Kevin W. Tong, Ai Gu, Xiangyang Deng, Yandong Cai, Ya Duan, Yuhong Hou
IEEE Trans Autom. Sci. Eng.6
2025 Adaptive Guidance in Dynamic Environments: A Deep Reinforcement Learning Approach for Highly Maneuvering Targets
abstract
In future battlefields, missiles are expected to become highly precise and efficient strike weapons, with missile intelligence emerging as a critical development trend. To address the problem of optimizing 3-D missile interception guidance laws, this article introduces the deep Q-network (DQN) algorithm on the foundation of proportional navigation guidance (PNG) and proposes an adaptive proportional guidance algorithm based on deep reinforcement learning (DRL). The proposed algorithm uses air combat situational information as the state space and incorporates parameters such as the missile-target relative distance and line-of-sight (LOS) angle into the reward function design. The optimal proportional navigation coefficient$K^{*}$for low-overload maneuvering targets is determined through network search, and the longitudinal and lateral control commands of the missile are decoupled by designing the proportional coefficient increment$\Delta K$, constructing a discretized action space. Simulation results show that, compared to the PNG with a constant$K^{*}$, the proposed method significantly improves the hit probability of high-overload maneuvering targets while maintaining the hit rate for low-overload maneuvering targets. As an exploration of future intelligent combat scenarios, this guidance law design method holds both theoretical significance and practical application value.
Longjun Zhu, Yandong Cai, Kevin W. Tong, Shuai Wu 0004, Fengtao Xiang, Ya Duan, Yuhong Hou, Guangyu Zhu 0001, Qi Wu 0003
IEEE Trans. Comput. Soc. Syst.6
2024 mvEchoSeg: One-shot In-context Learning for Multi-view Echocardiography Segmentation
abstract
Echocardiography is the clinical standard for evaluation of cardiac morphology, and function, and providing hemodynamic parameters in patients with known or suspected heart disease. Due to the diversity and complexity of diagnostic tasks, comprehensive interpretation of echocardiography often requires multi-view imaging for integrated metric analysis. Deep learning methods have become the mainstream approach for echocardiography segmentation. However, achieving segmentation of multi-view echocardiography still requires a substantial amount of annotation. To remedy this, we propose a one-shot in-context learning network mvEchoSeg for multi-view echocardiography segmentation. This network requires only one annotated image for each view. Specifically, we propose a Task Prompt Identifier (TPI) module to identify the task type of the image and allocate the most precise task prompt for it, with minimal adaptation of CLIP and few-shot strategies. Additionally, we leverage a unified In-Context Model (ICM) capable of performing a diverse set of echocardiography segmentation tasks automatically. Using fine-tuning and low-rank adapters improved the performance of the pre-trained model, achieving significant results with minimal training cost. Furthermore, we collect a multi-view echocardiography dataset (MVECD) with 8 views to evaluate our method. The results show an improvement of more than 10% in the DICE score compared to the SOTA foundational medical image segmentation models. To our knowledge, this is the first exploration of a one-shot model for multi-view echocardiography segmentation. Our codes and models are available at https://github.com/stellating/mvEchoSeg.
Ya Duan, Wenfeng Song, Nannan Li 0002, Aili Li, Shuai Li 0001
BIBM2