Zun Liu

dblp:209/8829 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model
abstract
Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions—remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has shown promising prospects. However, the reasoning of such methods is limited to the linguistic modality, lacking visual reasoning capabilities. Moreover, existing reasoning modules are optimized separately from navigation policies, leading to incompatibility and potential conflicts in optimization objectives. To tackle these challenges, we introduce UNeMo, a novel framework designed for the collaborative optimization of visual state reasoning and navigational decision-making. It introduces a Multimodal World Model (MWM) that takes visual features, language instructions, and navigational actions as inputs to jointly predict subsequent visual states, enabling cross-modal reasoning. Via a Hierarchical Prediction-Feedback (HPN) mechanism, MWM collaborates with navigation policies: the first layer generates actions using current vision-and-language features; MWM then infers post-action visual states to guide the second layer’s fine-grained decisions. This forms a dynamic bidirectional promotion mechanism where MWM reasoning optimizes navigation policies, while policy decisions feedback to improve MWM’s reasoning accuracy. Experiments on R2R and REVERIE datasets show UNeMo outperforms state-of-the-art methods by 2.1% and 0.7% in navigation accuracy for unseen scenes, validating its effectiveness.
Changxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu, Runhao Zeng, Zun Liu
AAAI6
2026 CP-AGN: A causal prior-guided adaptive graph network for topology inference in non-cooperative networks
Ang Dong, Yangming Guo, Zun Liu
Neurocomputing6
2026 PLS-FUSION: Tightly-Coupled Stereo Visual-Inertial SLAM With Point and Line Features
abstract
In this paper, we propose the PLS-FUSION, where P stands for point features, L for line features, and S for stereo vision. Together, these components form a tightly-coupled stereo visual-inertial Simultaneous Localization and Mapping (SLAM) system that leverages both point and line features to enhance the robustness and accuracy of tracking. Compared with the classic SLAM methods which only use the point features, we extract both the point features and the line features of the stereo camera in the front end and use a modified Line Segment Detector (LSD) algorithm of PL-VINS to improve the speed of extracting the line features. We use Plücker coordinates and orthonormal representation to represent the line features and add re-projection errors of the line features between the stereo camera in the back end. Our proposal has been tested with several popular datasets (EuRoC and TUMVI) and in the real world with our flight platform based on PX4 and QGroundControl (QGC). The experiments validate that PLS-FUSION has a more robust performance than state-of-the-art methods such as PL-VINS and VINS-FUSION.
Zun Liu, Jinxiao Tan, Jianqiang Li 0001
IEEE Trans Autom. Sci. Eng.1
2025 Dense small target detection algorithm for UAV aerial imagery
Yangming Guo, Zun Liu, Zhuqing Wang
Image Vis. Comput.4
2025 Supervised Momentum Contrastive Learning-Based Coarse-to-Fine Fusion Path Planning
David Chieng, Boon-Giin Lee, Junkai Ji, Zun Liu, Jianqiang Li 0001
IEEE Trans Autom. Sci. Eng.5
2024 GGI-DDI: Identification for key molecular substructures by granule learning to interpret predicted drug-drug interactions
Hui Yu 0011, Omayo Silver, Zun Liu, JingTao Yao 0001, Jianyu Shi
Expert Syst. Appl.5
2024 Multi-view clustering with semantic fusion and contrastive learning
Hui Yu 0011, Hui-Xiang Bian, Zi-Ling Chong, Zun Liu, Jianyu Shi
Neurocomputing4
2023 The Devil is in the Crack Orientation: A New Perspective for Crack Detection
abstract
Cracks are usually curve-like structures that are the focus of many computer-vision applications (e.g., road safety inspection and surface inspection of the industrial facilities). The existing pixel-based crack segmentation methods rely on time-consuming and costly pixel-level annotations. And the object-based crack detection methods exploit the horizontal box to detect the crack without considering crack orientation, resulting in scale variation and intra-class variation. Considering this, we provide a new perspective for crack detection that models the cracks as a series of sub-cracks with the corresponding orientation. However, the vanilla adaptation of the existing oriented object detection methods to the crack detection tasks will result in limited performance, due to the boundary discontinuity issue and the ambiguities in sub-crack orientation. In this paper, we propose a first-of-its-kind oriented sub-crack detector, dubbed as CrackDet, which is derived from a novel piecewise angle definition, to ease the boundary discontinuity problem. And then, we propose a multi-branch angle regression loss for learning sub-crack orientation and variance together. Since there are no related benchmarks, we construct three fully annotated datasets, namely, ORC, ONPP, and OCCSD, which involve various cracks in road pavement and industrial facilities. Experiments show that our approach outperforms state-of-the-art crack detectors.
Zhuangzhuang Chen, Jin Zhang 0013, Zhuonan Lai, Guanming Zhu, Zun Liu, Jie Chen 0027, Jianqiang Li 0001
ICCV5
2023 A Hierarchical Reinforcement Learning Algorithm Based on Attention Mechanism for UAV Autonomous Navigation
abstract
Unmanned Aerial Vehicles (UAVs) are increasingly being used in many challenging and diversified applications. Meanwhile, UAV’s ability of autonomous navigation and obstacle avoidance becomes more and more critical. This paper focuses on filling up the gap between deep reinforcement learning (DRL) theory and practical application by involving attention mechanism and hierarchical mechanism to solve some severe problems encountered in the practical application of DRL. More specifically, in order to improve the robustness of DRL, we use averaged estimation function instead of the normal value estimation function. Then, we design a recurrent network and a temporal attention mechanism to improve the performance of the algorithm. Third, we propose a hierarchical framework to improve its performance on long-term tasks. Some realistic simulation environments, as well as the real-world, are used to evaluate the proposed UAV autonomous navigation method. The results demonstrate that our DRL-based navigation method performs well in different environments and outperforms the original DrQ algorithm.
Zun Liu, Yuanqiang Cao, Jianyong Chen, Jianqiang Li 0001
IEEE Trans. Intell. Transp. Syst.1
2022 Geometry-Aware Guided Loss for Deep Crack Recognition
abstract
Despite the substantial progress of deep models for crack recognition, due to the inconsistent cracks in varying sizes, shapes, and noisy background textures, there still lacks the discriminative power of the deeply learned features when supervised by the cross-entropy loss. In this paper, we propose the geometry-aware guided loss (GAGL) that enhances the discrimination ability and is only applied in the training stage without extra computation and memory during inference. The GAGL consists of the feature-based geometry-aware projected gradient descent method (FGA-PGD) that approximates the geometric distances of the features to the class boundaries, and the geometry-aware update rule that learns an anchor of each class as the approximation of the feature expected to have the largest geometric distance to the corresponding class boundary. Then the discriminative power can be enhanced by minimizing the distances between the features and their corresponding class anchors in the feature space. To address the limited availability of related benchmarks, we collect a fully annotated dataset, namely, NPP2021, which involves inconsistent cracks and noisy backgrounds in real-world nuclear power plants. Our proposed GAGL outperforms the state of the arts on various benchmark datasets including CRACK2019, SDNET2018, and our NPP2021.
Zhuangzhuang Chen, Jin Zhang 0013, Zhuonan Lai, Jie Chen 0027, Zun Liu, Jianqiang Li 0001
CVPR5
2022 Drug-drug interaction prediction with learnable size-adaptive molecular substructures
abstract
Drug-drug interactions (DDIs) are interactions with adverse effects on the body, manifested when two or more incompatible drugs are taken together. They can be caused by the chemical compositions of the drugs involved. We introduce gated message passing neural network (GMPNN), a message passing neural network which learns chemical substructures with different sizes and shapes from the molecular graph representations of drugs for DDI prediction between a pair of drugs. In GMPNN, edges are considered as gates which control the flow of message passing, and therefore delimiting the substructures in a learnable way. The final DDI prediction between a drug pair is based on the interactions between pairs of their (learned) substructures, each pair weighted by a relevance score to the final DDI prediction output. Our proposed method GMPNN-CS (i.e. GMPNN + prediction module) is evaluated on two real-world datasets, with competitive results on one, and improved performance on the other compared with previous methods. Source code is freely available at https://github.com/kanz76/GMPNN-CS.
Arnold K. Nyamabo, Hui Yu 0011, Zun Liu, Jianyu Shi
Briefings Bioinform.3
2022 Predict multi-type drug-drug interactions in cold start scenario
abstract
BACKGROUND: Prediction of drug-drug interactions (DDIs) can reveal potential adverse pharmacological reactions between drugs in co-medication. Various methods have been proposed to address this issue. Most of them focus on the traditional link prediction between drugs, however, they ignore the cold-start scenario, which requires the prediction between known drugs having approved DDIs and new drugs having no DDI. Moreover, they're restricted to infer whether DDIs occur, but are not able to deduce diverse DDI types, which are important in clinics. RESULTS: In this paper, we propose a cold start prediction model for both single-type and multiple-type drug-drug interactions, referred to as CSMDDI. CSMDDI predict not only whether two drugs trigger pharmacological reactions but also what reaction types they induce in the cold start scenario. We implement several embedding methods in CSMDDI, including SVD, GAE, TransE, RESCAL and compare it with the state-of-the-art multi-type DDI prediction method DeepDDI and DDIMDL to verify the performance. The comparison shows that CSMDDI achieves a good performance of DDI prediction in the case of both the occurrence prediction and the multi-type reaction prediction in cold start scenario. CONCLUSIONS: Our approach is able to predict not only conventional binary DDIs but also what reaction types they induce in the cold start scenario. More importantly, it learns a mapping function who can bridge the drugs attributes to their network embeddings to predict DDIs. The main contribution of CSMDDI contains the development of a generalized framework to predict the single-type and multi-type of DDIs in the cold start scenario, as well as the implementations of several embedding models for both single-type and multi-type of DDIs. The dataset and source code can be accessed at https://github.com/itsosy/csmddi .
Zun Liu, Xing-Nan Wang, Hui Yu 0011, Jianyu Shi, Wenmin Dong
BMC Bioinform.1
2022 Integrated Air-Ground Vehicles for UAV Emergency Landing Based on Graph Convolution Network
abstract
With unmanned aerial vehicle (UAV) technologies advanced rapidly, many applications have emerged in cities. However, those applications do not widely spread as the safety consideration hinders the UAV from integrating into the civilian environment. This work focuses on investigating the UAV emergency landing problem which is a critical safety functionality of UAV. This work proposed a graph convolution network (GCN)-based decision network to learn by imitating the human pilots’ landing strategy. To alleviate the needs of a large amount of real-world data for model training, the proposed model allows to be trained in a simulated environment and then transferred to the real-world scenario due to the separation of domain-specific terrain classes and domain-independent topological structures among down-looking camera images. The GCN-based decision network can be coupled with a topological heuristic to improve the performance of action prediction in an emergency situation. To evaluate the proposed method, this work implemented a simulation environment for collecting data and testing the UAV emergency landing. The empirical results in both simulated and real-world scenarios show that the proposed methods can outperform the state-of-the-art counterparts in terms of predictive accuracy and success landing rate.
Jie Chen 0027, Jianqiang Li 0001, Weiming Du, Zhuangzhuang Chen, Zun Liu, Huihui Wang 0001, Victor C. M. Leung
IEEE Internet Things J.6