Yongjun Zhang 0006

dblp:43/5828-6 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Pano3R: Training Free Panoramic 3D Reconstruction
abstract
Panoramic 3D reconstruction is essential for immersive scene understanding in robotics, AR, and autonomous driving. However, most existing methods are designed for pinhole images and generalize poorly to 360° inputs due to the scarcity of panoramic training data and the high cost of retraining. We present Pano3R, the first training-free framework for panoramic 3D reconstruction that adapts existing pinhole-based models without any retraining. Pano3R consists of two stages. Specifically, the pre-processing stage applies a position-aware pairing strategy to decompose each panorama into a minimal set of perspective views. These views are selected to ensure sufficient co-visible regions while minimizing the number of projections. The test-time optimization stage incorporates a pose-prior-guided global alignment strategy to improve global consistency and mitigate accumulated errors. Our method enables accurate 360° reconstruction under both single- and multi-view input conditions. Extensive experiments demonstrate that Pano3R consistently improves reconstruction accuracy and pose estimation quality, establishing a strong and practical benchmark for training-free panoramic 3D reconstruction.
Shiming Song 0003, Yongjun Zhang 0006, Yuanze Wang, Mengzhu Wang, Yuetian Wang, Zhuojing Tian, Jinming Song, Dian-xi Shi
ECAI2
2025 Memory Prompt for Multi-Modal Visual Object Tracking
abstract
Multi-modal trackers have drawn widespread attention for robust tracking in challenging scenarios. However, existing multi-modal trackers often rely solely on spatial matching between the initial target template and the search regions, or incorporate only single-frame historical information, failing to fully exploit temporal correlations in tracking sequences. Additionally, most trackers that introduce temporal modeling require either retraining the entire network or designing specialized modules for temporal feature extraction, which incurs additional computational costs. To alleviate these limitations, inspired by human visual memory, we propose MPTrack, a novel tracker that directly reuses pre-extracted historical target features as memory prompts, establishing temporal dependencies without redundant feature extraction or specially designed temporal extraction networks. Our proposed Memory Prompt Fusion module effectively combines initial target templates with multiple historical memory cues to generate enhanced templates, enabling the perception of long-term appearance dynamics while mitigating potential interference from individual memory. Simultaneously, to avoid the computational cost of full-model training, we design a lightweight memory adapter that allows the frozen backbone network to efficiently adapt to the memory-enhanced template. Extensive experiments demonstrate that our method effectively incorporates temporal information and achieves promising results across different multi-modal tracking scenarios, including RGB+Thermal, RGB+Event, and RGB+Depth tracking tasks.
Yongjun Zhang 0006, Jianqiang Xia, Yushe Cao, Junze Zhang, Dian-xi Shi
ECAI2
2024 Improved Communication and Collision-Avoidance in Dynamic Multi-Agent Path Finding
abstract
Multi-Agent Path Finding (MAPF) is a classic problem with a wide range of applications. To cope with more complex situations in reality, Dynamic MAPF (DMAPF) has received much attention. The existing DMAPF definition lacks completeness or considers too simple situations. In this paper, we comprehensively model DMAPF based on realistic scenarios. Consequently, dynamic scenarios bring many problems. The dynamics of agent tasks bring the problem of more difficult coordination and cooperation of the multi-agent system, and the dynamics of obstacles bring the problem of increased collisions. To address these problems, this paper proposes a fully decentralised multi-agent reinforcement learning method CO3, which uses COmmon knowledge in selective COmmunication and proposes obstacle COllision avoidance mechanism. Firstly, common knowledge for communication improves cooperation between agents, which improves system performance and reduces collisions between agents. Secondly, the obstacle collision avoidance mechanism consists of a collision avoidance helper module and a critical region. The collision avoidance helper module improves the agents’ alertness to nearby obstacles, and the critical region gives an early warning to the agents to beware of distant obstacles. The obstacle collision avoidance mechanism can effectively reduce collisions between agents and obstacles. Finally, experiments show that CO3 can solve the DMAPF problem quite well, and the number of collisions is significantly lower than other learning-based methods in a dynamic environment.
Jing Xie 0021, Yongjun Zhang 0006, Qianying Ouyang, Huanhuan Yang, Dian-xi Shi, Songchang Jin
IJCNN2
2023 Faster Target Encirclement with Utilization of Obstacles via Multi-Agent Reinforcement Learning
Yuxi Zheng, Yongjun Zhang 0006, Chenran Zhao, Huanhuan Yang, Tongyue Li, Qianying Ouyang
ACML2
2023 Chinese Medical Named Entity Recognition Based on Pre-training Model
Shaowu Yang, Yongjun Zhang 0006, Dian-xi Shi
GPC (1)4
2022 Deep Reinforcement Learning for Multi-UAV Exploration Under Energy Constraints
Yating Zhou, Dian-xi Shi, Huanhuan Yang, Haomeng Hu, Shaowu Yang, Yongjun Zhang 0006
CollaborateCom (2)6
2022 FusionSeg: Motion Segmentation by Jointly Exploiting Frames and Events
Zhe Liu 0029, Shaowu Yang, Dian-xi Shi, Yongjun Zhang 0006
PRICAI (3)6
2022 E-HANet: Event-based Hybrid Attention Network for Optical Flow Estimation
abstract
Optical flow estimation is an essential task in computer vision. Standard cameras are prone to blurred images or over-saturated regions under extreme conditions. The event camera is a novel vision sensor inspired by the biological retina. It has the advantages of high time resolution, low delay and high dynamic range. We propose a new event representation and a novel deep learning network E-HANet (Event-based Hybrid Attention Network) for event-based dense optical flow estimation. To take full advantage of the complementarity between positive and negative events, we introduce the stacked positive and negative event slices as input. The feature extractor based on the channel attention is able to model the features from different event slices and fuse them by weight. We then present the hybrid attention weighting module to globally aggregate motion features. Compared to the event-based state-of-the-art, our approach reduces the average end-point error by 3% on the MVSEC dataset and 20% on the DSEC-Flow dataset.
Qimin Wang, Yongjun Zhang 0006, Shaowu Yang, Zhe Liu 0029, Luoxi Jing
SMC2
2022 Multi actor hierarchical attention critic with RNN-based feature extraction
Dian-xi Shi, Chenran Zhao, Huanhuan Yang, Gongju Wang, Shaowu Yang, Yongjun Zhang 0006
Neurocomputing9
2022 PLC-VIO: Visual-Inertial Odometry Based on Point-Line Constraints
abstract
Visual–inertial odometry (VIO) is widely studied and used in autonomous robots. This article proposes a novel tightly coupled monocular VIO system based on point-line constraints (PLC-VIO). In the front end, PLC-VIO presents a line segment extraction and merging algorithm based on the EDLines method and achieves real-time feature tracking based on the geometric constraints between feature points and lines. In the back end, PLC-VIO reconstructs new 3-D landmarks of feature lines through points on the line and optimizes the states by minimizing a cost function that combines the preintegrated inertial measurement unit (IMU) error term together with the point and line reprojection error terms in a sliding window optimization framework. A loop closure module is also integrated, which enables relocalization and drift elimination. The corresponding experimental evaluations are conducted using public datasets to validate the effectiveness and robustness of the proposed system, and the results show that PLC-VIO can achieve good performance when compared with other state-of-the-art systems and, at the same time, with no compromise to real-time performance.Note to Practitioners—Visual–inertial odometry (VIO) can estimate the states of the rigid body (including position, attitude, and velocity) that is widely used in robotic navigation, autonomous driving, virtual reality (VR), and augmented reality (AR). Aiming at the problem of estimating the states of autonomous robots in the GPS-denied environment, this article proposes a novel VIO system based on the point-line constraints (PLC-VIO). PLC-VIO can not only achieve accurate pose estimation for robots due to the introduction of the line features but also make no concession to real-time performance. Furthermore, PLC-VIO can also enrich the texture features of the environment during the 3-D mapping construction. The corresponding experiments are implemented in public datasets to evaluate the effectiveness, efficiency, and robustness of the proposed system. We believe that PLC-VIO can be widely used in robotic navigation and AR/VR fields to provide accurate position and environment information in real time.
Zhe Liu 0029, Dian-xi Shi, Ruihao Li 0001, Yongjun Zhang 0006, Xiaoguang Ren
IEEE Trans Autom. Sci. Eng.5
2021 ContriQ: Ally-Focused Cooperation and Enemy-Concentrated Confrontation in Multi-Agent Reinforcement Learning
abstract
Centralized training with decentralized execution (CTDE) is an important setting for cooperative multi-agent reinforcement learning (MARL) due to communication constraints during execution and scalability constraints during training, which has shown superior performance but still suffers from challenges. One branch is to understand the mutual interplay between agents. Due to the communication constraints in practice, agents cannot exchange perceptual information, and thus, many approaches use a centralized attention network with scalability constraints. Contrary to these common approaches, we propose to learn to cooperate in a decentralized way by applying attention mechanism on the local observation so that each agent could focus on allied agents with a decentralized model, and therefore promote understanding. Another branch is to model how agents cooperate and simplify the learning process. Previous approaches that focus on value decomposition have achieved innovative results but still suffer from problems. These approaches either limit the representation expressiveness of their value function classes or relax the IGM consistency to achieve scalability, which may lead to poor performance. We combine value composition with game abstraction by modeling the relationships between agents as a bi-level graph. We propose a novel value decomposition network based on it through a bi-level attention network, which indicates the contribution of allied agents attacking enemies and the priority of attacking each enemy under the situation of each time step, respectively. We show that our method substantially outperforms existing state-of-the-art methods on battle games in StarCraft Ⅱ, and attention analysis is also comprehensively discussed with sights.
Chenran Zhao, Dian-xi Shi, Huanhuan Yang, Shaowu Yang, Yongjun Zhang 0006
ACML6
2021 Attention-Aware Actor for Cooperative Multi-agent Reinforcement Learning
Chenran Zhao, Dian-xi Shi, Yaqianwen Su, Yongjun Zhang 0006, Shaowu Yang
CollaborateCom (2)5
2021 Joint Communication-Motion Planning for UAV Relaying in Urban Areas
abstract
In this paper, we consider a challenging surveillance scenario where there could exist line-of-sight (LOS) propagations and non-line-of-sight (NLOS) propagations in air-to-ground (ATG) channel and air-to-air (ATA) channel due to obstacles in urban areas, and a ground mobile robot is deployed to survey this area and transmit collected data to a remote base station via an unmanned aerial vehicle (UAV) relay. In this scenario, we aim to plan the optimal transmit power and trajectory of the UAV relay to minimize energy consumption while maintaining the communication quality. Existing works typically rely on the free-space path loss model and the statistical channel model, thus neglect the positions and shapes of obstacles and may fail in practical NLOS scenarios. In this paper, we first exploit the end-to-end packet error rate (PER)-based communication model, which captures the LOS propagation and NLOS propagation. Then, taken the information of obstacles in environment into consideration, we propose an UAV relay-assisted joint communication-motion planning (UAV-JCMP) method for minimizing the total energy consumption in urban areas. By decomposing the concave problem into two subproblems and dividing its domain into several convex subdomains according to LOS conditions, we further get the optimal solution. At last, numerical results demonstrate that substantial energy-efficient improvements can be achieved over methods that only optimize communication energy consumption and methods using statistical channel model. We further discuss the robustness of UAV-JCMP method towards terrain measurement error.
Sining Yang, Dian-xi Shi, Yingxuan Peng, Yongjun Zhang 0006
SECON5
2021 Multi-agent deep reinforcement learning with type-based hierarchical group communication
Dian-xi Shi, Gongju Wang, Yongjun Zhang 0006
Appl. Intell.6
2020 Networked Multi-robot Collaboration in Cooperative-Competitive Scenarios Under Communication Interference
Dian-xi Shi, Yongjun Zhang 0006, Liujing Wang, Fujiang She
CollaborateCom (1)4
2020 IDNet: A Single-Shot Object Detector Based on Feature Fusion
abstract
This paper proposes a novel single shot network for object detection. The proposed network, termed IDNet, explores the strategies of the feature fusion to alleviate the scale variation problem in object detection. IDNet mainly consists of two feature fusion modules: an indirect feature fusion module (IF) and a direct feature fusion module (DF). The IF shares long-range dependencies within pyramidal layers and based on these information, IDNet learns to emphasize informative regions and suppress the less useful ones on each layer. The DF is a feature fusion strategy based on modified lateral connection inspired by feature pyramid networks (FPN). It utilizes the averaging operation to reduce the change of feature maps' order of magnitude during fusing features to further improve the performance for detecting small instances. Comprehensive experiments are performed and the results indicate the effectiveness of IDNet, which reaches 80.3 mAP on PASCAL VOC 2007 benchmark.
Yuning Cui 0001, Dian-xi Shi, Yongjun Zhang 0006, Qianchong Sun
ICTAI3
2020 Dual-template Siamese Network with Cross-Correlation Fusion for Tracking
abstract
Siamese network based trackers treat tracking as a maximum matching between the detection region and the template, in which the template is either fixed or updated. The template-fixed trackers degrade accuracy since the appearance variations of the object in the subsequent frames are lacked; while the template-updated trackers depress robustness once the tracking drift occurs. In this paper, we apply a dual-template tracking strategy in the Siamese region proposal network to compensate for the separate merits of fixed template and mutative template. As for the dual-template, the one template is fixed and the other template is constantly updated. Meanwhile, we integrate a learnable convolutional neural network to fuse two Cross-Correlation responses generated from two template tracking branches. Based on Siamese network, we try to utilize the complementary fusion of different template responses to promote tracking performances. Extensive experiments on five tracking benchmarks of OTB100, VOT2016, VOT2018, GOT-10k and UAV123 demonstrate that our approach achieves the state-of-the-art tracking performances.
Dian-xi Shi, Ying Kang, Zunlin Fan, Songchang Jin, Yongjun Zhang 0006
ICTAI6
2020 Multi-Agent Feature Learning and Integration for Mixed Cooperative and Competitive Environment
abstract
At present, most of the centralized training with decentralized execution (CTDE) multi-agent reinforcement learning (MARL) algorithms have good results in the research of homogeneous scenarios. Heterogeneous multi-agent scenarios with different roles, cooperation modeling and credit assignment problems lead difficulty to learn effective collective strategies. In this paper, we propose a method of feature learning and feature integration about cooperation. Specifically, in the aspect of feature learning, through graph attention network, the relationship between agents is simplified to graph adjacency matrix representation, so that their feature vectors have relationship attributes. At the same time, for feature integration, we use batch normalization (BN) method to concatenate trained feature. We expect that agent relations can be modeled by end-to-end design. Meanwhile, attention mechanism can enhance the communication between interrelated agents. Through the experiments, our method has a significant result on improving the cooperative-competitive scenario of heterogeneous multi-agent. Moreover, we can visualize the output to analyze the reasonable collaborative and emphases attack policy.
Dian-xi Shi, Yongjun Zhang 0006, Liujing Wang
ICTAI4
2020 Selective Feature Network for Object Detection
abstract
Scale variation is one of the important challenges in object detection. Many state-of-the-art objectors tackle this problem by utilizing the feature pyramids. However, the current methods of producing feature pyramids are still inefficient to integrate the semantic information from other layers. In this work, our motivation is to build a feature pyramid efficiently with the selected contextual feature by integrating the informative features and suppressing the useless ones. To achieve this goal, we propose a novel single-stage detection network termed Selective Feature Network(SFNet) which consists of a semantic-enhanced module and a selective feature module. The semantic-enhanced module improves the semantics of basic pyramids via a lightweight architecture. In conjunction with that, a selective feature module is employed to combine features across different channels and scales by attention mechanism. The resulting contextual feature is then injected into the pyramidal features. Comprehensive experiments are performed on PASCAL VOC and MS COCO datasets. Results demonstrate that, with a VGG16 based SFNet, our approach obtains significant improvements over the competitors without losing real-time processing speed.
Yuning Cui 0001, Dian-xi Shi, Yongjun Zhang 0006, Qianchong Sun
IJCNN3
2020 GHGC: Goal-based Hierarchical Group Communication in Multi-Agent Reinforcement Learning
abstract
In large-scale multi-agent systems, the existence of a large number of agents with different target tasks and connected by complex game relationships causes great difficulty for policy learning. Therefore, simplifying the learning process is an important issue. In multi-agent systems, agents with the same target tasks or attributes often interact more with each other and exhibit behaviors more similar. That means there are stronger collaborations between these agents. Most existing multi-agent reinforcement learning (MARL) algorithms expect to learn the collaborative strategies of all agents directly in order to maximize the common rewards. This causes the difficulty of policy learning to increase exponentially as the number and types of agents increase. To address this problem, we propose a goal-based hierarchical group communication (GHGC) algorithm. This algorithm divides the agents into different groups, and maintains the group's cognitive consistency through knowledge sharing. Subsequently, we introduce a group communication and value decomposition method to ensure cooperation between the various groups. Experiments demonstrate that our model outperforms state-of-the-art MARL methods on the widely adopted StarCraft II benchmarks across different scenarios, and also possesses potential value for large-scale real-world applications.
Dian-xi Shi, Gongju Wang, Yongjun Zhang 0006
SMC6
2020 Friend-or-Foe Deep Deterministic Policy Gradient
abstract
One of the toughest challenges in the multi-agent deep reinforcement learning (MADRL) is that when the opponents' policies change rapidly, the collaborative agents can't learn well to respond to the opponents' policies effectively. This may lead to a local optimum w.r.t. the learned policy of the collaborative agents may be only locally optimal to the opponents' current policies. To address this problem, we propose a novel algorithm termed Friend-or-Foe Deep Deterministic Policy Gradient (FD2PG), in which the cooperative agents can be trained more robust and have stronger cooperation ability in continuous action space. These collaborative agents can generalize easily and respond correctly, even if their opponents' policies alter. Inspired by the classic Friend-or-Foe Q-learning algorithm (FFQ), we introduce the idea of minimizing the foes and maximizing the friends into the centralized training distributed execution framework, multi-agent deep deterministic policy gradient algorithm (MADDPG), to enhance collaborative agents' robustness and cooperativity. Besides, we introduce a Minimax Multi-Agent Learning (MMAL) method to explore two special equilibriums (the adversarial equilibrium and the coordination equilibrium), which can guarantee the convergence of FD2PG and improve optimization. Extensive fine-grained experiments, including four representative scenario experiments and two scale-performance correlation experiments, were conducted to demonstrate the superior performance of FD2PG comparing with existing baselines.
Dian-xi Shi, Gongju Wang, Yongjun Zhang 0006
SMC6
2019 Non-Convex Transfer Subspace Learning for Unsupervised Domain Adaptation
abstract
Transfer subspace learning aims to learn robust subspace for the target domain by leveraging knowledge from the source domain. The traditional methods often adopt the convex norm to approximate the original sparse and low-rank constraints, which make the optimization problem be easily solved. However, such relax approximation leads to the performance deviation of the original non-convex model. In this paper, we propose a novel Non-convex Transfer Subspace Learning~(NTSL) method to provide a tighter approximation to the original sparse and low-rank constraints. Specifically, we design an objective function that leverages the Schatten p-norm and ℓ_2, p-norm to preserve the structure between the source and target domains. With Schatten p-norm, the objective function better approximates the rank minimization problem than the nuclear norm and preserves the structure of domains. Besides, the ℓ_2, p-norm can reduce the effect of noise and improve the robustness to outliers. Meanwhile, we develop an efficient algorithm to solve the non-convex minimization problem. Extensive experimental results on cross-domain tasks show the effectiveness of our proposed method.
Tingjin Luo, Wenjing Yang 0002, Yongjun Zhang 0006, Yuhua Tang
ICME5
2019 FA-Harris: A Fast and Asynchronous Corner Detector for Event Cameras
abstract
Recently, the emerging bio-inspired event cameras have demonstrated potentials for a wide range of robotic applications in dynamic environments. In this paper, we propose a novel fast and asynchronous event-based corner detection method which is called FA-Harris. FA-Harris consists of several components, including an event filter, a Global Surface of Active Events (G-SAE) maintaining unit, a corner candidate selecting unit, and a corner candidate refining unit. The proposed G-SAE maintenance algorithm and corner candidate selection algorithm greatly enhance the real-time performance for corner detection, while the corner candidate refinement algorithm maintains the accuracy of performance by using an improved event-based Harris detector. Additionally, FA-Harris does not require artificially synthesized event-frames and can operate on asynchronous events directly. We implement the proposed method in C++ and evaluate it on public Event Camera Datasets. The results show that our method achieves approximately 8× speed-up when compared with previously reported event-based Harris detector, and with no compromise on the accuracy of performance.
Ruoxiang Li, Dian-xi Shi, Yongjun Zhang 0006, Kaiyue Li, Ruihao Li 0001
IROS3
2018 MulAttenRec: A Multi-level Attention-Based Model for Recommendation
Wenjing Yang 0002, Yongjun Zhang 0006, Haotian Wang 0001, Yuhua Tang
ICONIP (2)3
2018 Multi-UAV Collaborative Monocular SLAM Focusing on Data Sharing
Zhuoyue Yang, Dian-xi Shi, Yongjun Zhang 0006, Shaowu Yang, Ruoxiang Li
ICONIP (7)3
2017 The Curve Boundary Design and Performance Analysis for DGM Based on OpenFOAM
Yongquan Feng, Xinhai Xu, Yuhua Tang, Yongjun Zhang 0006
ICA3PP5