Yunzhou Zhang

dblp:95/4904 · DBLP profile ↗
← Back
95ranked-venue papers
6as first author
77since 2021 · last 2026
0000-0003-0610-3732ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 44 since 2021Systems, architecture and hardware · 26 · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 15 since 2021Computer networks · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Local depth constraints and contextual information fusion to generate correspondences for visual localization
Man Qi, Shuying Zhao, Yunzhou Zhang, Fawei Ge, Li Wang 0160, Xichen Zhang
Neurocomputing3
2026 End-to-End Vectorized HD Map Construction Based on Graph Structure Modeling and Graph Transformer Optimization
abstract
High-definition (HD) maps play a crucial role in autonomous driving by providing a reliable foundation for behavior prediction and path planning. Recent methods model map elements as point sets and employ Transformer-based detection frameworks for end-to-end vectorized map construction. However, these approaches do not fully exploit geometric structural properties, thereby limiting the accuracy of predicting the shape and position of map elements. To address this limitation, we propose a novel method calledGraphMapTR, which models map elements as graph structures. Each map element is represented as a subgraph where key points serve as nodes and edges capture structural continuity. To optimize the node query embeddings within subgraphs, we design a Graph Transformer-based decoder comprising three core components:Multi-head Self-Attention Between Subgraphs (MHSA-BSG), which facilitates efficient feature interaction across different subgraphs while enhancing node attention to their respective subgraph regions;Local Geometric-aware Self-Attention Within Subgraphs (LGSA-WSG), which dynamically aggregates global node relationships and local edge structures to enrich node representations; andDeformable Cross-Attention (DCA-BEV), which refines node embeddings through interaction with BEV features. In LGSA-WSG, priors from the rasterized map segmentation branch are incorporated to generate local geometric consistency scores (LGC-scores), which constitute edge structural bias. Extensive experiments demonstrate thatGraphMapTRsignificantly outperforms state-of-the-art methods. In particular, it outperforms the state-of-the-art algorithm MapTRv2 by 5.0% mAP and 3.4% mAP on the nuScenes and Argoverse2 datasets, respectively.
Wenjing Bai, Yunzhou Zhang, Wei Liu 0022, Zuotao Ning, Shuai Cheng 0001
IEEE Trans. Intell. Transp. Syst.2
2026 Dynamic Query Management and Internal Consistency Representation Based Transformer for Online Vectorized HD Map Construction
abstract
The online vectorized map construction technique employs a neural network model to forecast the vectorized representation of a specific region around automobiles, using data obtained from sensors mounted on automobiles. Due to advances in end-to-end object detection with transformers framework, the research on query-based online mapping has attracted substantial attention. However, the fixed number of queries and the random initialization of query embeddings constrain the model's performance. Moreover, the transformer architecture for object detection is based on the assumption that queries are identically distributed and independent, a premise that is not entirely applicable to map point queries which possess established subordinate relationships with map element instances. To address these issues, we initially incorporate a simplified transformer layer that utilizes semantic priors in bird's-eye view features for query initialization. The queries are then sent to a transformer-based map decoder for optimization and combined with a dynamic query management mechanism to eliminate low-confidence queries, hence maintaining computational efficiency. Furthermore, to guarantee that point queries within each instance preserve a consistent representation and avoid feature confusion among map element instances, we proposed an instance internal consistency map decoder. We conduct extensive experiments on commonly used map construction datasets to evaluate the proposed method. The experimental results demonstrate that our proposed method achieves state-of-the-art performance on the nuScenes and Argoverse 2 datasets.
Wenjing Bai, Yunzhou Zhang, Wei Liu 0022, Shangwei Du, Jun Hu 0020, Shuai Cheng 0001, Zuotao Ning
IEEE Trans. Multim.2
2026 Anisotropic Optical Flow Guided Adaptive Multi-Stage Video Inpainting
abstract
Video inpainting is attracting more attention due to the potential applications of video object removal and video content restoration. Current approaches either use end-to-end methods to generate missing pixels directly or perform indirect transfer for known regions based on motion field guidance. However, such approaches cannot handle both high-resolution images and diverse degrees of scene variation between adjacent video frames, and they cannot achieve clear and accurate inpainting effects for large continuous missing areas. To this end, we propose an adaptive multi-stage interval video inpainting algorithm guided by anisotropic optical flow. First, we customize an optical flow inpainting method guided by single image inpainting, enabling optical flow to maintain a strong self-healing ability over a large range of missing areas. Then, the interval mechanism adaptively determines the required temporal neighbors for missing pixels by assessing video attributes and inpainted optical flow results. After the missing pixels complete the multi-candidate information fusion in their associated temporal neighbors, we obtain spatio-temporally consistent and accurate results. Finally, extensive experiments on the YouTubeVOS, DAVIS, A2D2, and custom datasets show that our proposed approach has achieved state-of-the-art performance with good environmental migration ability.
Lei Rong, Yunzhou Zhang, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Multim.2
2025 VSS-SLAM: Voxelized Surfel Splatting for Geometally Accurate SLAM
abstract
[1] Visual Simultaneous Localization and Mapping (SLAM) helps robots estimate their poses and perceive the environment in unknown settings. Recent work has demonstrated that implicit neural radiance fields and 3D Gaussian Splatting (3DGS) offer higher fidelity scene representation than traditional map representations. We propose VSS-SLAM, which utilizes voxelized surfels as the map representation for incremental mapping in unknown environments. This representation effectively addresses the issue of redundant and disordered primitives encountered in previous methods, thereby enhancing geometric accuracy during reconstruction. Specifically, our approach divides the scene using voxels and stores geometric and appearance information in feature vectors at the voxel vertices. Before rendering, these feature vectors are decoded to generate the corresponding surfels. Additionally, we align camera poses through image and depth rendering. Extensive experiments on the Replica and TUM-RGBD datasets demonstrate that VSS-SLAM delivers high-fidelity reconstruction and accurate pose estimation in both simulated and real-world environments. Source code will soon be available.
Xuanhua Chen, Yunzhou Zhang, Xingshuo Wang
ICRA2
2025 MDC-Seg: Multi-Directional Convolution-Based Semantic Segmentation for LiDAR Point Clouds
abstract
LiDAR point clouds 3D semantic segmentation enables efficient and accurate environmental sensing for intelligent vehicles and autonomous robots, greatly advancing these domains. Existing advanced methods that use 3D sparse convolutional often suffer from a small Effective Receptive Field (ERF), which limits context sensing and challenging highperformance segmentation. Building on this observation, we propose MDC-Seg for efficient ERF enlargement. We design Multi-directional Convolution (MDConv), which simultaneously performs sparse feature encoding on the Bird's Eye View (BEV) and Range View (RV) planes to enlarge the ERF of 3D sparse convolution. To enhance feature fusion in MDConv, we introduce an attention mechanism and design an efficient multifeature fusion (EMFF) module suitable for both 3D and 2D sparse features. To improve segmentation accuracy, we design a point-voxel constraint (PVC) module to handle edge voxels containing multiple point cloud categories, optimizing the final inference results. These modules add minimal memory and inference time but significantly improve performance compared to the baseline. Extensive experiments on the SemanticKITTI benchmark demonstrate MDC-Seg's excellent performance, with supplementary tests on nuScenes further confirming its superiority by yielding good results. The source code is available at https://github.com/OYgreat-river/MDC-Seg.
Xin Ouyang, Xiaolong Qian, Yunzhou Zhang, You Shen, Guiyuan Wang, Wei Liu 0022
ICRA3
2025 LE-Object: Language Embedded Object-Level Neural Radiance Fields for Open-Vocabulary Scene
abstract
Recent advancements in Visual Language Models (VLMs) have significantly driven research in open-vocabulary 3D scene reconstruction, showcasing strong potential in open-set retrieval and semantic understanding. However, existing approaches face challenges in open-world environments: they either suffer from insufficient precision in semantic segmentation, leading to inadequate fine-grained scene understanding, or they are limited to object-level reconstruction, failing to capture intricate object details and lack applicability in open-world settings. To address these issues, we introduce LE-Object, an object-centric Neural Implicit Radiance Field (NeRF) method for open-world scenarios to achieve fine-grained scene understanding and high-fidelity object reconstruction. LE-Object integrates spatial features (SF) from object point clouds with visual features (VF) from VLMs to perform object association, ensuring spatiotemporal consistency in object mask segmentation, and extends VLM features from 2D images into 3D space, enabling precise open-world semantic inference and detailed object reconstruction. Experimental results demonstrate that LE-Object excels in zero-shot semantic segmentation and open-world object reconstruction, offering innovative solutions for global navigation and local object manipulation in open-world applications.
Mengting Wang, Yunzhou Zhang, Xingshuo Wang, Zhiteng Li
ICRA2
2025 APA-BI: Adaptive Partition Aggregation and Bidirectional Integration for UAV-View Geo-Localization
abstract
The task of UAV-view geo-localization is to match a query image with database images to estimate the current geographic location of the query image. This is particularly useful in environments where GPS is not available or when the device fails. Although deep learning methods make sufficient progress in UAV-view geo-localization, they still face challenges in improving the distinguishability of features. For instance, some feature aggregation methods do not consider semantic integrity, and robust elements in the image are not given enough attention. This paper proposes a UAV-view geo-localization method (APA-BI) to tackle the above issues. Specifically, we propose an adaptive partition aggregation method to ensure feature integrity at the semantic level by increasing the receptive field of the classifier module. At the same time, we design a bidirectional integration module to further enhance feature distinguishability by extracting robust tubular topological structures from images. Experimental results on public datasets demonstrate that APA-BI achieves impressive retrieval accuracy and outperforms most state-of-the-art methods. Moreover, the test results of APA-BI in real-world scenarios also show excellent performance.
Xichen Zhang, Shuying Zhao, Yunzhou Zhang, Fawei Ge
ICRA3
2025 JRN-Geo: A Joint Perception Network Based on RGB and Normal Images for Cross-View Geo-Localization
abstract
Cross-view geo-localization plays a critical role in Unmanned Aerial Vehicle (UAV) localization and navigation. However, significant challenges arise from the drastic viewpoint differences and appearance variations between images. Existing methods predominantly rely on semantic features from RGB images, often neglecting the importance of spatial structural information in capturing viewpoint-invariant features. To address this issue, we incorporate geometric structural information from normal images and introduce a Joint perception network to integrate RGB and Normal images (JRN-Geo). Our approach utilizes a dual-branch feature extraction framework, leveraging a Difference-Aware Fusion Module (DAFM) and Joint-Constrained Interaction Aggregation (JCIA) strategy to enable deep fusion and joint-constrained semantic and structural information representation. Furthermore, we propose a 3D geographic augmentation technique to generate potential viewpoint variation samples, enhancing the network's ability to learn viewpoint-invariant features. Extensive experiments on the University-1652 and SUES-200 datasets validate the robustness of our method against complex viewpoint variations, achieving state-of-the-art performance.
Yunzhou Zhang, Tingsong Huang, Fawei Ge, Man Qi, Xichen Zhang
ICRA2
2025 Multiple Object Tracking with Dynamic Adaptive Object Motion Estimation
abstract
One indicator for evaluating autonomous vehicles’ capability is the accuracy of perceiving the surrounding environment. As an essential part of perception, MOT (multi-object tracking) algorithms provide vital guarantees for safe driving. However, many MOT algorithms based on the motion model only consider the information from previous frame when predicting the motion state of objects without taking into account the long-term motion state. Moreover, their motion model generally uses constant speed or acceleration models, which may cause tracking loss when the object suddenly changes its motion or is occluded by other objects. In this paper, we propose DA-MOT (Multiple Object Tracking with Dynamic Adaptive Object Motion Estimation), which utilizes information from lidar and camera sensors to calculate objects’ dynamic and static states under different sensor information. We modify the KF motion model parameters based on the object’s motion for better tracking performance. Furthermore, we design a re-association mechanism to re-assign IDs for inaccurate associations. We conducted experiments on the KITTI dataset, and the results show a significant improvement in accuracy. DA-MOT algorithm has about 1.5% improvement in MOTA metrics compared to other MOT algorithms in scenes with large changes in object state and can run about 1000 fps on the Intel Core Intel® Xeon(R) Gold 5217 CPU.
Borui Cheng, Yunzhou Zhang
IROS2
2025 MSPA-LIO: LiDAR-Inertial Odometry with Multi-Scale Plane Adjustment
abstract
Most current LiDAR-based odometry methods use point-to-local plane registration to constrain poses, ignoring the explicit plane structure in the environment. Due to noise interference and uneven distribution of point cloud, local planes are prone to tilt, resulting in registration errors. Therefore, we propose MSPA-LIO, a LiDAR-Inertial odometry with multi-scale plane adjustment, which uses geometric constraints and plane adjustment at both local voxel plane scale and large plane scale to improve odometry accuracy and enhance map consistency. In order to make full use of the planar structure in the environment, we propose an explicit large plane extraction method based on the voxel-based. We use large planes to correct the direction of the associated voxel planes, thereby overcoming the misregistration problem caused by local plane tilt. To further improve the odometry accuracy, we perform plane adjustments at the voxel plane scale and the large plane scale to make the pose and map more consistent. Experiments conducted on the VECtor Dataset and the Newer College Dataset demonstrate that our proposed algorithm outperforms four state-of-the-art algorithms.
Shuying Zhao, Yunzhou Zhang, Hengwang Ding, Sizhan Wang
IROS3
2025 Camera-invariance correlation learning and inter-domain-specific distinct representation for person re-identification
Shangdong Zhu, Yunzhou Zhang, Peng Duan 0002
Eng. Appl. Artif. Intell.2
2025 SUGrasping: a semantic grasping framework based on multi-head 3D U-Net
He Cao, Yunzhou Zhang, Zhexue Ge, Xiaozheng Liu
Multim. Tools Appl.2
2025 Learning Local Features by Jointly Semantic-Guided and Task Rewards
abstract
Learning local features is a fundamental task for many computer vision applications. Existing methods often struggle to maintain robustness and accuracy in extracting local features, especially in complex environments with numerous interfering objects. Although some studies have integrated semantic information into local feature extraction networks to enhance discrimination, their effectiveness remains limited. Therefore, this paper fully considers the importance of semantic information for feature extraction and proposes a semantically enhanced local feature extraction network framework. This framework includes a local feature network, a semantic segmentation network, and a reinforcement learning framework. Semantic information is incorporated into feature heatmaps and feature descriptors to improve the accuracy of feature points. Subsequently, the local feature network is continuously optimized by a reinforcement learning algorithm based on semantic information and matching ground truth to enhance robustness, ensuring that the final local features achieve optimal performance. Extensive experiments on three publicly available datasets validate the effectiveness of the proposed local feature network.
Li Wang 0160, Yunzhou Zhang, Fawei Ge, Wenjing Bai
IEEE Trans. Circuits Syst. Video Technol.2
2025 Joint Representation Learning Based on Feature Center Region Diffusion and Edge Radiation for Cross-View Geo-Localization
abstract
The essence of the cross-view geo-localization task is to accurately identify the same object across images captured from different viewpoints. Due to variations in image acquisition methods and viewing angles, the content information of the images can differ significantly, which may result in localization failure. Therefore, cross-view geo-localization remains a challenging task. To solve this issue, a joint representation learning network based on feature center region diffusion and edge radiation is proposed in this article. First, to extract the crucial information from the global features, we design the central diffusion module that identifies important regions within the features and enhances feature robustness. Additionally, we design an edge radiation mechanism that expands the receptive field and further highlights crucial information in the image to support the central diffusion module in achieving more stable performance. On this basis, we propose an adaptive triple InfoNCE loss function to assist network training, improving the discriminability of the extracted features. Finally, the proposed network is tested on two mainstream datasets, and experimental results demonstrate that the proposed model outperforms the state-of-the-art methods, which can prove its effectiveness.
Fawei Ge, Yunzhou Zhang, Li Wang 0160, Yixiu Liu, Pengju Si, You Shen
IEEE Trans. Geosci. Remote. Sens.2
2025 Semantic Boundary Constrained Network for Visual Place Recognition Under Adverse Conditions
abstract
Accurate localization of autonomous vehicles is crucial for autonomous driving and safety, especially in complex urban environments where high-precision GPS is not available. Visual place recognition (VPR) uses visual cues to identify the current location from a known database, serving as an auxiliary means for precise localization in autonomous driving. In real-world applications, variations in scene appearance due to changes in illumination and seasons present significant challenges for VPR. Current VPR methods often fail in adverse visual environments due to their inability to provide robust scene descriptions. Therefore, the extraction of stable and effective information from images is relatively important. In this paper, we propose a novel feature extraction network, termed SBCNet. This network is designed to capture semantic boundaries and texture details within images by training an auxiliary semantic boundary detection task. By focusing on these fundamental elements, the model’s perceptual capacity for structural features can be enhanced. Moreover, we introduce a semantic edge attention module that generates spatial attention maps based on semantic edges and texture details, allowing for the comprehensive utilization of pivotal structural cues. With this explicit guide, the network prioritizes local regions with appearance invariance during the feature extraction process. Experimental results demonstrate that our method maintains robust performance under various adverse visual conditions. Even in low-light environments, such as those encountered at night, our method exhibits commendable performance.
Yunzhou Zhang, Jian Ning, Kunmo Li, Dehao Zou
IEEE Trans. Intell. Transp. Syst.2
2025 Learning Local Features by Reinforcing Spatial Structure Information
abstract
Learning-based local feature extraction algorithms have advanced considerably in terms of robustness. While excelling at enhancing feature robustness, some outstanding algorithms tend to neglect discriminability—a crucial aspect in vision tasks. With the increase of deep learning convolutional layers, we observe an amplification of semantic information within images, accompanied by a diminishing presence of spatial structural information. This imbalance primarily contributes to the subpar feature discriminability. Therefore, this paper introduces a novel network framework aimed at imbuing feature descriptors with robustness and discriminative power by reinforcing spatial structural information. Our approach incorporates a spatial structure enhancement module into the network architecture, spanning from shallow to deep layers, ensuring the retention of rich structural information in deeper layers, thereby enhancing discriminability. Finally, we evaluate our method, demonstrating superior performance in visual localization and feature-matching tasks.
Li Wang 0160, Yunzhou Zhang, Fawei Ge, Wenjing Bai
IEEE Trans. Multim.2
2024 CTA-LO: Accurate and Robust LiDAR Odometry Using Continuous-Time Adaptive Estimation
abstract
Accurate and robust LiDAR odometry is a crucial technology for robot localization. However, motion distortion and ranging error make it a bottleneck. Most existing methods are limited in accuracy and robustness because they simply compensate for motion distortion by constant velocity motion assumption without accurate model of ranging error. In this paper, we propose a high-precision and robust LiDAR odometry (LO), which utilizes continuous-time estimation to remove LiDAR distortion and builds the spot uncertainty model to quantify the ranging error. Generally, the number of variables in continuous-time estimation is several times higher than that in discrete-time ones, leading to insufficient constraints on the LiDAR odometry. To solve this problem, we propose a marginalization method to retain prior scans’ constraints by exploiting the local support property of the B-spline. To further improve the odometry accuracy, we propose a residual adaptive weighting method and a probabilistic point cloud map based on the spot uncertainty model of LiDAR points. The experimental results show that our method outperforms state-of-the-art LiDAR odometry in accuracy and robustness.
Yuezhang Lv, Yunzhou Zhang, Jian Ning
ICRA2
2024 Enhancing Visual Place Recognition with Multi-modal Features and Time-constrained Graph Attention Aggregation
abstract
Visual place recognition(VPR) is a crucial technology for autonomous driving and robotic navigation. However, severe appearance and perspective changes often lead to degradation of algorithm performance. Current methods mainly utilize single-modality RGB images, which are sensitive to environmental changes. To address this challenge, we propose a novel multi-modal visual place recognition method by incorporating depth information as auxiliary data to enhance the robustness of the VPR algorithm. The pipeline involves dual-branch feature extraction and shared multi-modal feature fusion based on transformer(SFFM) to enable full interaction between semantic and structural information. Furthermore, we introduces a time-constrained graph attention aggregation(TC-GAT) that propagates node information across time and space to deal with perceptual aliasing. Extensive experiments on the Oxford Robotcar and MSLS datasets demonstrate that the proposed algorithm is not only effective in appearance changes but also competitive in opposing viewpoints.
Yunzhou Zhang, Jian Ning, Dehao Zou, Meiqi Pei
ICRA2
2024 VPE-SLAM: Neural Implicit Voxel-permutohedral Encoding for SLAM
abstract
NeRF can reconstruct incredibly realistic environmental maps in dense simultaneous localization and mapping, providing robots with more comprehensive scene map information. However, NeRF often struggles with geometric distortions in indoor reconstructions. To correct geometric distortions, we develop VPE-SLAM, based on the proposed voxel-permutohedral encoding, which can incrementally reconstruct maps of unknown scenes. Specifically, voxel-permutohedral encoding combines a sparse voxel feature grid created by an octree and multi-resolution permutohedral tetrahedral feature grids to represent the scene effectively. Especially when dealing with object edges, our method can effectively encode the geometry and texture of edges by the hybrid structural grid. We propose a novel local bundle adjustment module that utilizes a sliding window mechanism to manage adjacent keyframes requiring optimization. Furthermore, the proposed method establishes local map consistency by repeatedly optimizing keyframes that were initially under-optimized through a compensation strategy. The consistency of the local map can enhance the adaptability of our method to challenging scenes. Extensive experiments demonstrate that our method can achieve accurate camera tracking and produce high-quality reconstruction results on the Replica and ScanNet datasets. The source code will be available at https://github.com/NeuCV-IRMI/VPE-SLAM.
Yunzhou Zhang, You Shen, Lei Rong, Sizhan Wang, Xin Ouyang
ICRA2
2024 L-VIWO: Visual-Inertial-Wheel Odometry based on Lane Lines
abstract
To achieve precise localization for autonomous vehicles and mitigate the problem of accumulated drift error in odometry, this paper proposes L-VIWO, a Visual-Inertial-Wheel Odometry based on lane lines. This method effectively utilizes the lateral constraints provided by lane lines to eliminate and relieve the incrementally accumulated pose errors. Firstly, we introduce a lane line tracking method that enables multi-frame tracking of the same lane line, thereby obtaining multi-frame data of a lane line. Then, we utilize multi-frame data of the lane lines and the curvature characteristics of adjacent lane lines to optimize the positions of the lane line sample points, thus building a reliable lane line map. Finally, we use the built local lane line map to correct the position of the vehicle. Based on the corrected position and prior pose from the odometry, we build a graph optimization model to optimize the pose of the vehicle. Through localization experiments on the KAIST dataset, it has been demonstrated that the proposed method effectively enhances the localization accuracy of odometry, thus confirming the effectiveness of the method.
Yunzhou Zhang, Xichen Zhang, Zeyu Long
ICRA2
2024 LA-LIO: Robust Localizability-Aware LiDAR-Inertial Odometry for Challenging Scenes
abstract
Modern robotic systems are increasingly deployed in complex and diverse environments, and reliable localization under challenging conditions becomes crucial for the safe and efficient operation of these systems. The odometry based on LiDAR is prone to system collapse caused by computational divergence under conditions of aggressive motion and information deficiency in spatial geometry. To enhance the robustness of systems in challenging scenes, this work proposes LA-LIO, robust localizability-aware LiDAR inertial odometry. It mainly consists of three parts. Firstly, this paper presents a LiDAR degeneration detection method that enables stable degeneration assessment. Secondly, a method for segmenting LiDAR point clouds is proposed to alleviate the issue of excessive distortion in point clouds under aggressive motion scenes. The last is an Errors State Kalman Filter (ESKF) method with adaptive weights to utilize the existing spatial information as much as possible to improve the stability of the system in degenerated scenarios. The proposed method is evaluated and compared in multiple experiments, demonstrating the performance and reliability improvements of this approach in challenging environments.
Yunzhou Zhang, Qingdong Xu, Jun Liu 0087, Guiyuan Wang, Wei Liu 0022
IROS2
2024 ESO-SLAM: Tightly-Coupled and Simultaneous Estimation of Self and Multi-Object Pose via Sensor Fusion
abstract
Simultaneous Localization and Mapping (SLAM) is widely used in applications such as robotics and autonomous driving, with methods involving multi-sensor fusion demonstrating excellent performance. However, they simply reject dynamic features and ignore the mutual benefits of self and dynamic objects, which greatly limits their application in actual high-dynamic scenes. To address this issue, we propose ESO-SLAM, a tightly-coupled system for simultaneous self and multi-object pose estimation achieved through sensor fusion. This system employs a multi-probability fusion tracker based on filter to establish more robust object-level data association. Building upon this, we introduce a method that combines 3D Kalman filter velocity priors and camera optical flow decoupling for dynamic point cloud removal, aiming at improving the accuracy of self-pose estimation in odometry. Finally, we jointly refine the poses of the robot and objects using multiple constraint factors within our proposed framework. Experimental results on the KITTI raw dataset demonstrate that our approach achieves better pose accuracy for both self and tracked objects compared to baseline and state-of-the-art techniques. Furthermore, the proposed method exhibits feasibility in real-time performance to ensure its practical application value.
Yunzhou Zhang, Yuezhang Lv, Sizhan Wang, Guiyuan Wang
IROS2
2024 Neighborhood Consensus Guided Matching Based Place Recognition with Spatial-Channel Embedding
abstract
As a crucial part of mobile robotics and autonomous driving, Visual Place Recognition (VPR) is usually addressed by recognizing its similar reference images from a pre-obtained database. However, VPR always suffers from environmental changes, such as weather, illumination, perceptual-aliasing and so on. To address this, we firstly introduce a robust and discriminative global descriptor aggregation technique that normalizes the spatial and channel dimensions of features. A Spatial-Channel Embedding (SCE) module is proposed to learn the spatial and scale information of features which make global features more discriminative. Meanwhile, the traditional re-ranking methods (e.g. RANSAC) for geometric consistency verification are time-consuming. Here we propose a Neighborhood Consensus Guided Matching (NCGM) module, which uses Neighborhood Consensus to filter the features from patch-level matching to achieve more accurate matching while reduces the time consumption. Through extensive experiments on multiple benchmarks, we demonstrate that our method outperforms several state-of-the-art methods while maintaining lower time consumption and storage requirements.
Kunmo Li, Yunzhou Zhang, Jian Ning, Guiyuan Wang, Wei Liu 0022
IROS2
2024 HSS-SLAM: Human-in-the-Loop Semantic SLAM Represented by Superquadrics
abstract
The advancement of object detection algorithms has catalyzed the development of object-level semantic SLAM. However, due to missed and false detections, object-level semantic SLAM fails to represent the objects within the scene adequately. Therefore, this paper proposes a novel object-level semantic SLAM termed HSS-SLAM. We incorporate human-in-the-loop into our method, establishing an interaction module to facilitate human editing and rectifying semantic information. Additionally, to minimize the manual correction workload, a lightweight and intuitive method for semantic extension is proposed, augmenting the semantic richness of the global map with a few operations. Furthermore, our method adopts superquadrics for object representation, enabling detailed descriptions of various object shapes. This mitigates the limitation of conventional semantic mapping, where objects are difficult to distinguish due to the reliance on a single-shape representation. Subsequently, precise estimation of superquadric parameters and camera poses is achieved through joint optimization. Extensive experiments conducted on TUM RGB-D and Scenes V2 datasets demonstrate that the proposed approach exhibits competitive performance, surpassing current methods in both object representation and camera localization accuracy.
Yunzhou Zhang, You Shen, Tengda Zhang, Guolu Chen
IROS2
2024 FI-SLAM: Feature Fusion and Instance Reconstruction for Neural Implicit SLAM
abstract
Recent advancements in neural implicit fields for Simultaneous Localization and Mapping (SLAM) have provided breakthroughs. However, the benefits of reconstruction results to the perception ability of robot are minimal. Therefore, we propose FI-SLAM, a dense semantic instance SLAM system based on neural implicit representation, which significantly aids robots in better understanding the scene. FI-SLAM employs a coordinate and plane joint encoding method, which reduces the difficulty of feature storage by flattening the feature space. Furthermore, to improve representation efficiency, we use the method of adjacent feature level linear interpolation to describe features. We propose a feature fusion (FF) method to merge the object features with the scene features. The fused feature vector enhances the reconstruction accuracy of the local scene while ensuring the global reconstruction effect. It has improved the global reconstruction effect of the scene and the accuracy of camera tracking. Numerous experiments on synthetic and real-world datasets demonstrate that our method can assure accurate tracking precision, high-fidelity reconstruction results, and complete semantic instance maps. In summary, the algorithm we proposed heavily augments the scene perception capabilities of robot.
Xingshuo Wang, Yunzhou Zhang, Mengting Wang, Zhiteng Li, Xuanhua Chen
IROS2
2024 Pos2VPR: Fast Position Consistency Validation with Positive Sample Mining for Hierarchical Place Recognition
abstract
Visual place recognition (VPR) is a challenging issue for robotics and autonomous systems, focusing on utilizing visual information for robot localization. Currently, hierarchical architecture is being employed by growing works, which embraces RANSAC-based geometric verification for re-ranking. However, RANSAC is time-consuming and only employs geometric information while neglecting other potential information that could be useful for re-ranking. Here we propose a fast position consistency via local patch (PCLP) algorithm to take the position of task-relevant patch-descriptor into account. Without training, it only costs little time but performs better than other re-ranking methods that rely on geometric consistency verification. In this paper, we present a unified place recognition framework that incorporates an aggregation module to extract global features for retrieval and a PCLP validation module to filter local patch for reranking. Meanwhile, we propose a RANSAC-based tightly coupled learning (R-TCL) strategy to discover the best positive sample for training robust models. Unlike common sample mining methods, we introduce RANSAC into the sample mining process, achieving trade-off between efficiency and accuracy. Due to improved positive sample mining strategy and novel position validation module, our model is named as Pos2VPR. Remarkably, Pos2VPR outperforms state-of-the-art methods on four major datasets with extremely short running time.
Dehao Zou, Xiaolong Qian, Yunzhou Zhang
IROS3
2024 Bilateral guidance network for one-shot metal defect segmentation
Dexing Shan, Yunzhou Zhang, Xiaozheng Liu, Sonya A. Coleman, Dermot Kerr
Eng. Appl. Artif. Intell.2
2024 Learning robust representation and sequence constraint for retrieval-based long-term visual place recognition
Yanhai Tan, Yunzhou Zhang, Fawei Ge, Shangdong Zhu
Eng. Appl. Artif. Intell.3
2024 SADGFeat: Learning local features with layer spatial attention and domain generalization
Wenjing Bai, Yunzhou Zhang, Li Wang 0160, Wei Liu 0022, Jun Hu 0020
Image Vis. Comput.2
2024 Self-supervised assisted multi-task learning network for one-shot defect segmentation with fake defect generation
Ziqiang Hu, Hao Chu, Yunzhou Zhang, Dexing Shan, You Shen
Pattern Recognit. Lett.3
2024 Multibranch Joint Representation Learning Based on Information Fusion Strategy for Cross-View Geo-Localization
abstract
Cross-view geo-localization refers to recognizing images of the same geographic target obtained from different platforms (such as drone-view, satellite-view and ground-view). However, cross-view geo-localization is challenging as image capture using different platforms coupled with extreme viewpoint variations can cause significant changes to the visual image content. Existing methods mainly focus on mining the fine-grained features or the contextual information in neighboring areas, but ignore the complete information of the entire image and the association of contextual information of adjacent regions. Therefore, a multi-branch joint representation learning network model based on information fusion strategies is proposed to solve this cross-view geo-localization problem. Firstly, we obtain feature information from the image through global information fusion branch and local information fusion branch to help the network learn the discernable information in the different images. In addition, a local-guided-global information fusion branch is introduced to make local information assist global features to enhance the learning of potential information in the images. Secondly, we introduced different information fusion strategies in each branch to increase the extraction of contextual information through expanding the global receptive field, thus improving the performance of the model. Finally, a series of experiments is carried out on four prevailing benchmark datasets, namely University-1652, SUES-200, CVUAS and CVACT datasets. The quantitative comparisons from the experiments clearly indicate that the proposed network framework has great performance. For example, compared with some state-of-the-art methods, the quantitative improvements of the R@1 and AP on the University-1652 datasets are 1.91%, 2.18% and 1.55%, 2.99% in both tasks, respectively.
Fawei Ge, Yunzhou Zhang, Yixiu Liu, Guiyuan Wang, Sonya A. Coleman, Dermot Kerr, Li Wang 0160
IEEE Trans. Geosci. Remote. Sens.2
2024 Multilevel Feedback Joint Representation Learning Network Based on Adaptive Area Elimination for Cross-View Geo-Localization
abstract
Cross-view geo-localization refers to the task of matching the same geographic target using images obtained from different platforms, such as drone-view and satellite-view. However, the view angle of images obtained through different platforms will vary greatly, which can bring great challenges to the cross-view geo-localization task. Therefore, we propose a multi-level feedback joint representation learning network based on adaptive area elimination to solve the cross-view geo-localization problem. In our network model, we first process the extracted global features to obtain part-level and patch-level features. We then utilize these features as feedback to the global features to extract the contextual information in the global features and improve the robustness of the extracted features. In addition, as images obtained from different platforms differ, there will always be some interference when matching images. Therefore, we introduce an adaptive area elimination strategy to erase the interference information in the global features and assist the model in obtaining crucial information. On this basis, the feature correlation loss function is designed to constrain learning when using global feature information, thereby eliminating the possible interference, which can improve the network model performance. Finally, a series of experiments is carried out using two well-known benchmarks, namely University-1652 and SUES-200, and the experimental results show that the proposed network model achieves competitive results, thereby demonstrating the effectiveness of proposed model.
Fawei Ge, Yunzhou Zhang, Li Wang 0160, Wei Liu 0022, Yixiu Liu, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Geosci. Remote. Sens.2
2024 Fast and Robust LiDAR-Inertial Odometry by Tightly-Coupled Iterated Kalman Smoother and Robocentric Voxels
abstract
This paper presents a fast LiDAR-inertial odometry (LIO) that is robust to aggressive motion. To achieve robust tracking in aggressive motion scenes, we exploit the continuous scanning property of LiDAR to adaptively divide the full scan into multiple partial scans (named sub-frames) according to the motion intensity. And to avoid the degradation of sub-frames resulting from insufficient constraints, we propose a robust state estimation method based on a tightly-coupled iterated error state Kalman smoother (ESKS) framework. Furthermore, we propose a robocentric voxel map (RC-Vox) to improve the system’s efficiency. The RC-Vox allows efficient maintenance of map points and k nearest neighbor (k-NN) queries by mapping local map points into a fixed-size, two-layer 3D array structure. Extensive experiments are conducted on 27 sequences from 4 public datasets, our own dataset and real-world scenes. The results show that our system can achieve stable tracking in aggressive motion scenes (angular velocity up to 21.8 rad/s) that cannot be handled by other state-of-the-art methods, while our system can achieve competitive performance with these methods in general scenes. Furthermore, thanks to the RC-Vox, our system is much faster than the most efficient LIO system currently published.
Jun Liu 0087, Yunzhou Zhang, Zhengnan He, Wei Liu 0022, Xiangren Lv
IEEE Trans. Intell. Transp. Syst.2
2024 DynaQuadric: Dynamic Quadric SLAM for Quadric Initialization, Mapping, and Tracking
abstract
Dynamic SLAM is a key technology for autonomous driving and robotics, and accurate pose estimation of surrounding objects is important for semantic perception tasks. Current quadric SLAM methods are based on the assumption of a static environment and can only reconstruct static quadrics in the scene, which limits their applications in complex dynamic scenarios. In this paper, we propose a visual SLAM system that is capable of reconstructing dynamic objects as quadrics, with a unified framework for jointly optimizing pose estimation, multi-object tracking (MOT), and quadric parameters. We propose a robust object-centric quadric initialization algorithm for both static and moving objects, which decouples the prior estimation of the object pose from the quadric parameters. The object is initialized with a coarse sphere, and quadric parameters are further refined. We design a novel factor graph that tightly optimizes camera pose, object pose, map points and quadric parameters within the sliding window-based optimization. To the best of our knowledge, we are the first to propose a dynamic SLAM that combines quadric representations and MOT in a tightly coupled optimization. We perform qualitative and quantitative experiments on both simulated and real-world datasets, and demonstrate the robustness and accuracy in terms of camera localization, dynamic quadric initialization, mapping and tracking. Our system demonstrates the potential application of object perception with quadric representation in complex dynamic scenes.
Rui Tian 0002, Yunzhou Zhang, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.2
2024 Fast, Robust, Accurate, Multi-Body Motion Aware SLAM
abstract
Simultaneous ego localization and surrounding object motion awareness are significant issues for the navigation capability of unmanned systems and virtual-real interaction applications. Robust and accurate data association at object and feature levels is one of the key factors in solving this problem. However, currently available solutions ignore the complementarity among different cues in the front-end object association and the negative effects of poorly tracked features on the back-end optimization. It makes them not robust enough in practical applications. Motivated by these observations, we make up rigid environment as a unified whole to assist state decoupling by integrating high-level semantic information, ultimately enabling simultaneous multi-states estimation. A filter-based multi-cues fusion object tracker is proposed for establishing more stable object-level data association. Combined with the object’s motion priors, the motion-aided feature tracking algorithm is proposed to improve the feature-level data association performance. Furthermore, a novel state estimation factor graph is designed which integrates a specific feature observation uncertainty model and the intrinsic priors of tracked object, and solved through sliding-window optimization. Our system is evaluated using the KITTI dataset and achieves comparable performance to state-of-the-art object pose estimation systems both quantitatively and qualitatively. We have also validated our system on simulation environment and a real-world dataset to confirm the potential application value in different practical scenarios.
Linghao Yang, Yunzhou Zhang, Rui Tian 0002, Shiwen Liang, You Shen, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.2
2024 CapsLoc3D: Point Cloud Retrieval for Large-Scale Place Recognition Based on 3D Capsule Networks
abstract
Point cloud-based place recognition can be used for global localization in large-scale scenes and loop-closure detection in simultaneous localization and mapping (SLAM) systems in the absence of GPS. Current learning-based approaches aim to extract global and local features from 3D point clouds to encode them as descriptors for point cloud retrieval. The key problems are that the occlusion of point clouds by dynamic objects in the scene affects the point cloud structure, a single perceptual field of the network cannot adequately extract point cloud features, and the correlation between features is not fully utilized. To overcome this, we propose a novel network called CapsLoc3D. We first use the static point cloud generation module to remove the occlusion effects of dynamic objects, and then obtain the point cloud descriptors by processing with the CapsLoc3D network which contains the point spatial transformation module, multi-scale feature fusion module, Capsnet module and a GeM Pooling layer. After validation using the Oxford RobotCar, KITTI, and NEU datasets, experiments show that our method performs better and also has good generalization performance and computational efficiency compared with current state-of-the-art algorithms.
Yunzhou Zhang, Ming Liao, Rui Tian 0002, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.2
2024 MVSE-Net: A Multi-View Deep Network With Semantic Embedding for LiDAR Place Recognition
abstract
Place recognition is a critical technology in robot navigation and autonomous driving, remains challenging due to inefficient point cloud computation, limited feature representation capability, and poor robustness to long-term environmental changes. We propose MVSE-Net, a feature extraction network with embedded semantic information for multi-view feature fusion. MVSE-Net can convert point cloud data acquired by LiDAR in real time into global descriptors for retrieval. Processing a point cloud by projecting it onto a 2D image can greatly improve computational efficiency. We projected the point cloud into a range-view (RV) image and a bird’s-eye-view (BEV) image in forward and top view, respectively. The semantic segmentation network is then used to process the RV image, and the feature extraction part of the semantic model is connected to the transformer attention module to further refine the features for the place recognition task. The point cloud containing the semantic segmentation results is then converted into a semantic BEV image, and the multi-channel BEV image is processed using a group convolutional network. Finally, the features of the two branches are fused into a global feature representation by post-fusion. Our experiments on three publicly available datasets demonstrate that MVSE-Net exhibits high recall and strong generalization in LiDAR place recognition, outperforming previous state-of-the-art methods.
Yunzhou Zhang, Lei Rong, Rui Tian 0002, Sizhan Wang
IEEE Trans. Intell. Transp. Syst.2
2024 Double-Domain Adaptation Semantics for Retrieval-Based Long-Term Visual Localization
abstract
Due to seasonal and illumination variance, long-term visual localization tasks in dynamic environments is a crucial problem in the field of autonomous driving and robotics. At present, image-based retrieval is an effective method to solve this problem. However, it is difficult to completely distinguish changes in the same location over times by relying on content information alone. In order to solve these above problems, a double-domain network model combining semantic information and content information is proposed for visual localization task. In addition, this approach only needs to use the virtual KITTI 2 dataset for training. To reduce the domain difference between real scene and virtual image, the cross-predictive semantic segmentation mechanism is introduced to solve this problem. In addition, the obtained model achieves good domain adaptation and further has well generalization on other real datasets by introducing a domain loss function and a triplet semantic loss function. A series of experiments on the Extended CMU-Seasons dataset and the Oxford RobotCar-Seasons dataset demonstrates that the proposed network model outperformes the state-of-the-art baselines for retrieval-based visual localization in challenging environments.
Fawei Ge, Yunzhou Zhang, Li Wang 0160, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Multim.2
2024 LARNet: Towards Lightweight, Accurate and Real-Time Salient Object Detection
abstract
Salient object detection (SOD) has rapidly developed in recent years, and detection performance has greatly improved. However, the price of these improvements is increasingly complex networks that require more computing resources and sacrifice real-time performance. This makes it difficult to deploy these approaches on devices with limited computing resources (such as mobile phones, embedded platforms, etc.). Considering recently developed lightweight SOD models, their detection and real-time performance are always compromised in demanding practical application scenarios. To solve these problems, we propose a novel lightweight SOD method called LARNet and its corresponding extremely lightweight method LARNet$^{*}$according to application requirements. These methods balance the relationship between lightweight requirements, detection accuracy and real-time performance. First, we propose a saliency backbone network tailored for SOD, which removes the need for pre-training with ImageNet and effectively reduces feature redundancy. Subsequently, we propose a novel context gating module (CGM), which simulates the physiological mechanism of human brain neurons and visual information processing, and realizes the deep fusion of multi-level features at the global level. Finally, the saliency map is output after fusion of multi-level features. Extensive experiments on popular benchmark datasets demonstrate that the proposed LARNet (LARNet$^{*}$) achieves 98 (113) FPS on a GPU and 3 (6) FPS on a CPU. With approximately 680 K (90 K) parameters, the model has significant performance advantages over (extremely) lightweight methods, even surpassing some heavyweight models.
Zhenyu Wang 0010, Yunzhou Zhang, Yan Liu 0080, Cao Qin, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Multim.2
2023 Joint Segmentation and Grasp Pose Detection with Multi-Modal Feature Fusion Network
abstract
Efficient grasp pose detection is essential for robotic manipulation in cluttered scenes. However, most methods only utilize point clouds or images for prediction, ignoring the advantages of different features. In this paper, we present a multi-modal fusion network for joint segmentation and grasp pose detection. We design a point cloud and image co-guided feature fusion module that can be used to fuse features and adaptively estimate the importance of the point-pixel feature pairs. Moreover, we develop a seed point sampling algorithm that simultaneously considers the distance, semantics and attention scores. For selected seed points, we adopt a local feature aggregation module to fully utilize the local spatial features in the grasp region. Experimental results on the GraspNet-lBillion Dataset show that our network outperforms several state-of-the-art methods. We also conduct real robot grasping experiments to demonstrate the effectiveness of our approach.
Xiaozheng Liu, Yunzhou Zhang, He Cao, Dexing Shan
ICRA2
2023 SAMLoc: Structure-Aware Constraints With Multi-Task Distillation for Long-Term Visual Localization
abstract
Real-time and robust long-term visual localization is a crucial technology for autonomous driving. Season and illumination variance make this problem more challenging. At present, most of excellent visual localization algorithms cannot run in real-time on devices with limited computing resources. In this paper, we propose SAMLoc, a structure-aware and self-supervised visual localization system, for fast and robust 6-DoF localization. To obtain structural features in the scene, we propose local and global structure-aware constraints using edge information. Then, we integrate the structure-aware constraints into the hierarchical localization network of multi-task distillation, which significantly reduces the feature extraction time while ensuring localization accuracy. As a result, real-time and robust large-scale localization can be achieved on mobile devices. Experimental results on public datasets show that our system can achieve high localization accuracy and have satisfactory real-time performance. Compared with several state-of-the-art visual localization systems, our framework achieves a competitive localization performance.
Jian Ning, Yunzhou Zhang, Sonya A. Coleman, Kunmo Li, Dermot Kerr
ICRA2
2023 VIW-Fusion: Extrinsic Calibration and Pose Estimation for Visual-IMU-Wheel Encoder System
abstract
The data fusion of camera, IMU, and wheel encoder measurements has proved its effectiveness in localizing ground robots, and obtaining accurate sensor extrinsic parameters is its premise. We propose an extrinsic parameter calibration algorithm and a multi-sensor-based pose estimation algorithm for the camera-IMU-wheel encoder system. First, we propose a joint calibration algorithm for the extrinsic parameters of the camera-IMU-wheel encoder system, which improves the accuracy and robustness of the camera-wheel encoder calibration. We then extend the visual-inertial odometry (VIO) to incorporate the measurements from the wheel encoder and weight the wheel encoder measurements according to angular velocity in global optimization to improve the performance. We further propose a novel method for VIO initialization by integrating wheel encoder information, which significantly reduces the scale error in initialization. We conduct extrinsic parameter calibration experiments on a real self-driving car and validate the performance of our multi-sensor-based localization system on the KAIST dataset and a dataset collected by our self-driving vehicles by performing an exhaust comparison with the state-of-the-art algorithms. Our implementations are open source11https://github.com/chunxiaoqiao/VIW-Fusion.git.
Chunxiao Qiao, Shuying Zhao, Yunzhou Zhang
IROS3
2023 BSH-Det3D: Improving 3D Object Detection with BEV Shape Heatmap
abstract
The progress of LiDAR-based 3D object detection has significantly enhanced developments in autonomous driving and robotics. However, due to the limitations of LiDAR sensors, object shapes suffer from deterioration in occluded and distant areas, which creates a fundamental challenge to 3D perception. Existing methods estimate specific 3D shapes and achieve remarkable performance. However, these methods rely on extensive computation and memory, causing imbalances between accuracy and real-time performance. To tackle this challenge, we propose a novel LiDAR-based 3D object detection model named BSH-Det3D, which applies an effective way to enhance spatial features by estimating complete shapes from a bird's eye view (BEV). Specifically, we design the Pillar-based Shape Completion (PSC) module to predict the probability of occupancy whether a pillar contains object shapes. The PSC module generates a BEV shape heatmap for each scene. After integrating with heatmaps, BSH-Det3D can provide additional information in shape deterioration areas and generate high-quality 3D proposals. We also design an attention-based densification fusion module (ADF) to adaptively associate the sparse features with heatmaps and raw points. The ADF module integrates the advantages of points and shapes knowledge with negligible overheads. Extensive experiments on the KITTI benchmark achieve state-of-the-art (SOTA) performance in terms of accuracy and speed, demonstrating the efficiency and flexibility of BSH-Det3D. The source code is available on https://github.com/mystorm16/BSH-Det3D.
You Shen, Yunzhou Zhang, Yanmin Wu, Zhenyu Wang 0010, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IROS2
2023 A novel seminar learning framework for weakly supervised salient object detection
Yan Liu 0080, Yunzhou Zhang, Zhenyu Wang 0010, Fei Yang 0007, Sonya A. Coleman, Dermot Kerr
Eng. Appl. Artif. Intell.2
2023 GW-net: An efficient grad-CAM consistency neural network with weakening of random erasing features for semi-supervised person re-identification
Shangdong Zhu, Yunzhou Zhang
Image Vis. Comput.2
2023 Scale space tracker with multiple features
Jining Bao, Yunzhou Zhang, Shangdong Zhu
Multim. Tools Appl.2
2023 Correction to: Person search via class activation map transferring
Ruilong Li, Yunzhou Zhang, Shangdong Zhu, Shuangwei Liu
Multim. Tools Appl.2
2023 TRF-Net: a transformer-based RGB-D fusion network for desktop object instance segmentation
He Cao, Yunzhou Zhang, Dexing Shan, Xiaozheng Liu
Neural Comput. Appl.2
2023 WUSL-SOD: Joint weakly supervised, unsupervised and supervised learning for salient object detection
Yan Liu 0080, Yunzhou Zhang, Zhenyu Wang 0010, Sonya A. Coleman, Dermot Kerr
Neural Comput. Appl.2
2023 MMPL-Net: multi-modal prototype learning for one-shot RGB-D segmentation
Dexing Shan, Yunzhou Zhang, Xiaozheng Liu, Shitong Liu, Sonya A. Coleman, Dermot Kerr
Neural Comput. Appl.2
2023 Graph Wasserstein Autoencoder-Based Asymptotically Optimal Motion Planning With Kinematic Constraints for Robotic Manipulation
abstract
This paper presents a learning based motion planning method for robotic manipulation, aiming to solve the asymptotically-optimal motion planning problem with nonlinear kinematics in a complex environment. The core of the proposed method is based on a novel neural network model, i.e., graph wasserstein autoencoder (GraphWAE) network, which is used to represent the implicit sampling distributions of the configuration space (C-space) for sampling-based planning algorithms. Through learning the implicit distributions, we can guide the planning process to search or extend in the desired region to reduce the collision checks dramatically for fast and high-quality motion planning. The theoretical analysis and proofs are given to demonstrate the probabilistic completeness and asymptotic optimality of the proposed method. Numerical simulations and experiments are conducted to validate the effectiveness of the proposed method through a series of planning problems from 2D, 6D and 12D robot C-spaces in the challenging scenes. Results indicate that the proposed method can achieve better planning performance than the state-of-the-art planning algorithms. Note to Practitioners—The motivation of this work is to develop a fast and high-quality asymptotically optimal motion planning method for practical applications such as autonomous driving, robotic manipulation and others. Due to the time consumption caused by collision detection, current planning algorithms usually take much time to converge to the optimal motion path especially in the complicated environment. In this paper, we present a neural network model based on GraphWAE to learn the biasing sampling distributions as the sample generation source to further reduce or avoid collision checks of sampling-based planning algorithms. The proposed method is general and can be also deployed in other sampling-based planning algorithms for improving planning performance in different robot applications.
Chongkun Xia, Yunzhou Zhang, Sonya A. Coleman, Ching-Yen Weng, Houde Liu, Shichang Liu, I-Ming Chen 0001
IEEE Trans Autom. Sci. Eng.2
2023 ELWNet: An Extremely Lightweight Approach for Real-Time Salient Object Detection
abstract
Existing lightweight salient object detection (SOD) methods aim to solve the problem of high computational costs that is prevalent with heavyweight methods. However, compared with heavyweight methods, the detection accuracy of lightweight methods is greatly reduced while real-time performance is not significantly improved. Therefore, we aim to establish a trade off between computational cost and detection performance by improving the network efficiency. We propose a fast and extremely lightweight end-to-end wavelet neural network (ELWNet) for real-time salient object detection. ELWNet can achieve salient object detection and segmentation at approximately 70FPS (GPU), 19FPS (CPU) with 76K parameters and 0.38G FLOPs. We introduce wavelet transform theory into a neural network, proposing a wavelet transform module (WTM), a wavelet transform fusion module (WTFM), a novel feature residual mechanism, and construct an efficient architecture. The wavelet transform theory is integrated into the neural network to realize the interaction between the features in the frequency and the time domain. Meanwhile, ELWNet does not rely on a pre-trained model, which significantly reduces redundant features. We validate the performance of ELWNet using five well-known datasets, and demonstrate state-of-the-art performance compared with 24 other SOD models in terms of being lightweight, detection accuracy and real-time capabilities. Our method maintains high detection performance while reducing the number of model parameters by approximately 99% compared with heavyweight methods.
Zhenyu Wang 0010, Yunzhou Zhang, Yan Liu 0080, Delong Zhu 0001, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Circuits Syst. Video Technol.2
2023 Unseen-Material Few-Shot Defect Segmentation With Optimal Bilateral Feature Transport Network
abstract
Industrial defect segmentation is important to ensure product quality and production safety. The main challenges in industrial applications are insufficient defect samples, large intraclass variation, and the interference of background information. However, most current texture defect segmentation methods rely on large-scale datasets and can only deal with one specific type of texture defect, which reduces the application efficiency and application scope of defect segmentation algorithms. To this end, we propose an optimal bilateral feature transport network (OBFTNet) for few-shot texture defect segmentation, which can accurately segment texture defects in multiple unseen materials (domains), such as steel, wood, and leather. OBFTNet can perform bilateral prediction for background and defect regions of unseen material by dynamically predicting task-specific semantic correspondences conditioned on a small guidance set. Specifically, we introduce background images (defect-free images) as supplementary learning information for reverse prediction and model the semantic correspondence between the guidance (support and background images) and the query images in few-shot segmentation as an optimal bilateral feature transport problem and generate a set of optimal bilateral correlation tensors. Using 4-D and 2-D convolutions, the model gradually reduces optimal bilateral correlation tensors to precise segmentation masks. Experimental results show that our proposed method outperforms several state-of-the-art techniques with very few labeled samples and the method generalizes well to industrial defects on unseen materials.
Dexing Shan, Yunzhou Zhang, Sonya A. Coleman, Dermot Kerr, Shitong Liu, Ziqiang Hu
IEEE Trans. Ind. Informatics2
2023 Object SLAM With Robust Quadric Initialization and Mapping for Dynamic Outdoors
abstract
Object SLAM is a popular approach for autonomous driving and robotics, but accurate object perception in outdoor environments remains a challenge. State-of-the-art object SLAM algorithms rely on assumptions and are sensitive to observation noise, limiting their application in real-world scenarios. To address these challenges, we propose a novel object SLAM system that utilizes a quadric initialization algorithm based on constrained quadric optimization, which does not rely on planar assumptions and is robust to partial observations. Additionally, we introduce an automatic object data association algorithm capable of detecting motion states while associating objects across frames. To further enhance the accuracy of the quadric mapping, an extra thread is used to refine the ellipsoid parameters within a local sliding window composed of keyframes. Our system utilizes a joint optimization framework that optimizes camera poses, object landmarks, and point clouds in the local mapping thread for further global optimization while maintaining a consistent map. Experimental results on the real-world KITTI dataset show that the proposed system is more robust and significantly outperforms current state-of-the-art methods in quadric initialization and mapping in outdoor scenarios. Moreover, our system achieves real-time performance, making it suitable for practical applications.
Rui Tian 0002, Yunzhou Zhang, Zhenzhong Cao, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.2
2023 Structure-Aware Feature Disentanglement With Knowledge Transfer for Appearance-Changing Place Recognition
abstract
Long-term visual place recognition (VPR) is challenging as the environment is subject to drastic appearance changes across different temporal resolutions, such as time of the day, month, and season. A wide variety of existing methods address the problem by means of feature disentangling or image style transfer but ignore the structural information that often remains stable even under environmental condition changes. To overcome this limitation, this article presents a novel structure-aware feature disentanglement network (SFDNet) based on knowledge transfer and adversarial learning. Explicitly, probabilistic knowledge transfer (PKT) is employed to transfer knowledge obtained from the Canny edge detector to the structure encoder. An appearance teacher module is then designed to ensure that the learning of appearance encoder does not only rely on metric learning. The generated content features with structural information are used to measure the similarity of images. We finally evaluate the proposed approach and compare it to state-of-the-art place recognition methods using six datasets with extreme environmental changes. Experimental results demonstrate the effectiveness and improvements achieved using the proposed framework. Source code and some trained models will be available at http://www.tianshu.org.cn.
Cao Qin, Yunzhou Zhang, Yingda Liu, Delong Zhu 0001, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Neural Networks Learn. Syst.2
2023 An Object SLAM Framework for Association, Mapping, and High-Level Tasks
abstract
Object SLAM is considered increasingly significant for robot high-level perception and decision-making. Existing studies fall short in terms of data association, object representation, and semantic mapping and frequently rely on additional assumptions, limiting their performance. In this article, we present a comprehensive object SLAM framework that focuses on object-based perception and object-oriented robot tasks. First, we propose an ensemble data association approach for associating objects in complicated conditions by incorporating parametric and nonparametric statistic testing. In addition, we suggest an outlier-robust centroid and scale estimation algorithm for modeling objects based on the iForest and line alignment. Then a lightweight and object-oriented map is represented by estimated general object models. Taking into consideration the semantic invariance of objects, we convert the object map to a topological map to provide semantic descriptors to enable multimap matching. Finally, we suggest an object-driven active exploration strategy to achieve autonomous mapping in the grasping scenario. A range of public datasets and real-world results in mapping, augmented reality, scene matching, relocalization, and robotic manipulation have been used to evaluate the proposed object SLAM framework for its efficient performance.
Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Zhiqiang Deng, Wenkai Sun, Jian Zhang 0018
IEEE Trans. Robotics2
2022 Noise-Tolerant Learning with Silhouette Coefficient for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (re-ID) attracts growing attention due to its broad prospects in practical applications. State-of-the-art unsupervised re-ID approaches combine clustering-based pseudo-label prediction with feature fine-tuning. However, pseudo labels generated directly by clustering are not always reliable and inevitably contain noisy labels. To tackle these challenges, we propose a novel noise inhibition framework to estimate the confidence of each pseudo label and actively correct noisy labels. By introducing the silhouette coefficient, our method can estimate the pseudo-label confidence without any extra model or data, and calculate the correction matrix to correct clustering results directly. However, the silhouette coefficient is usually applied on the hyper-parameters selection of clustering algorithms. In order to make the silhouette coefficient more suitable for estimation and correction tasks, we calibrate the Jaccard distance matrix to alleviate the negative influence of the cluster size on the silhouette coefficient. Our proposed method brings significant improvement and achieves the state-of-the-art performance on benchmark datasets.
Shuying Zhao, Yunzhou Zhang, Yixiu Liu, Shangdong Zhu, Sonya A. Coleman
ICME3
2022 Salient Object Detection via Bilateral Feature Fusion and Score Sorting Attention Mechanism
abstract
Deep learning based salient object detection methods have recently received significant attention. However, current methods still suffer from shortcomings such as informative background information being ignored which a significant problem for image saliency understanding. Additionally, it is also a challenge to suppress the noisy features in the network. By analyzing the difference between high-level and low-level features from ResNet-50, we utilize a Bilateral Feature Fusion (BFF) module to deal with the problem caused by ignoring informative background information. Benefitting from the BFF module, our proposed network can capture more meaningful foreground and background cues, which helps to get a more accurate saliency map. Moreover, we adopt a Score Sorting Attention (SSA) module which suppresses noisy and irrelevant features. Experimental results on five benchmark datasets demonstrate that our proposed method performs better than other state-of-the-art methods. The ablation studies also prove our contributions.
Shuying Zhao, Yunzhou Zhang, Yan Liu 0080, Zhenyu Wang 0010, Sonya A. Coleman
ICME3
2022 SemLoc: Accurate and Robust Visual Localization with Semantic and Structural Constraints from Prior Maps
abstract
Semantic information and geometrical structures of a prior map can be leveraged in visual localization to bound drift errors and improve accuracy. In this paper, we propose SemLoc, a pure visual localization system, for accurate localization in a prior semantic map. To tightly couple semantic and structure information from prior maps, a hybrid constraint is presented by using the Dirichlet distribution. Then, with the local landmarks and their semantic states tracked in the frontend, the camera poses and data associations are jointly optimized through Expectation-Maximization (EM) algorithm. We validate the effectiveness of our approach in both monocular and stereo modes on the public KITTI dataset. Experimental results demonstrate that our system can greatly reduce drift errors with an satisfying real-time performance. Compared with several state-of-the-art visual localization systems, the proposed framework achieves a competitive localization performance.
Shiwen Liang, Yunzhou Zhang, Rui Tian 0002, Delong Zhu 0001, Linghao Yang, Zhenzhong Cao
ICRA2
2022 CFP-SLAM: A Real-time Visual SLAM Based on Coarse-to-Fine Probability in Dynamic Environments
abstract
The dynamic factors in the environment will lead to the decline of camera localization accuracy due to the violation of the static environment assumption of SLAM algorithm. Recently, some related works generally use the combination of semantic constraints and geometric constraints to deal with dynamic objects, but problems can still be raised, such as poor real-time performance, easy to treat people as rigid bodies, and poor performance in low dynamic scenes. In this paper, a dynamic scene-oriented visual SLAM algorithm based on object detection and coarse-to-fine static probability named CFP-SLAM is proposed. The algorithm combines semantic constraints and geometric constraints to calculate the static probability of objects, keypoints and map points, and takes them as weights to participate in camera pose estimation. Extensive evaluations show that our approach can achieve almost the best results in high dynamic and low dynamic scenarios compared to the state-of-the-art dynamic SLAM methods, and shows quite high real-time ability.
Xinggang Hu, Yunzhou Zhang, Zhenzhong Cao, Yanmin Wu, Zhiqiang Deng, Wenkai Sun
IROS2
2022 Semantic Topological Descriptor for Loop Closure Detection within 3D Point Clouds In Outdoor Environment
abstract
Loop closure detection has the potential to correct the drift of trajectories and build a global consistent map in LiDAR SLAM, however it remains a challenging problem in outdoor environment due to the sparsity of 3D point clouds data, large-scale scenes and moving objects. Inspired by the way humans perceive the environment through recognizing objects and identifying their relations, this paper presents a novel descriptor that contains semantic and topological information for loop closure detection. Unlike most existing methods that extract features from the raw point clouds or use all semantic objects, we directly discard point clouds representing pedestrians and vehicles after semantic segmentation. Then, we propose a semantic topological graph representation from the remaining point clouds and convert this graph into a descriptor. Additionally, we propose a two-stage algorithm for matching descriptors to efficiently determine the loop. Our method has been extensively evaluated using the KITTI dataset and outperforms state-of-the-art methods, especially in the challenging situations such as viewpoint changes and dynamic scenes.
Ming Liao, Yunzhou Zhang, Sonya A. Coleman, Dermot Kerr
IROS2
2022 VAC-Net: Visual Attention Consistency Network for Person Re-identification
abstract
Person re-identification (ReID) is a crucial aspect of recognising pedestrians across multiple surveillance cameras. Even though significant progress has been made in recent years, the viewpoint change and scale variations still affect model performance. In this paper, we observe that it is beneficial for the model to handle the above issues when boost the consistent feature extraction capability among different transforms (e.g., flipping and scaling) of the same image. To this end, we propose a visual attention consistency network (VAC-Net). Specifically, we propose Embedding Spatial Consistency (ESC) architecture with flipping, scaling and original forms of the same image as inputs to learn a consistent embedding space. Furthermore, we design an Input-Wise visual attention consistent loss (IW-loss) so that the class activation maps(CAMs) from the three transforms are aligned with each other to enforce their advanced semantic information remains consistent. Finally, we propose a Layer-Wise visual attention consistent loss (LW-loss) to further enforce the semantic information among different stages to be consistent with the CAMs within each branch. These two losses can effectively improve the model to address the viewpoint and scale variations. Experiments on the challenging Market-1501, DukeMTMC-reID, and MSMT17 datasets demonstrate the effectiveness of the proposed VAC-Net.
Yunzhou Zhang, Shangdong Zhu, Yixiu Liu, Sonya A. Coleman, Dermot Kerr
ICMR2
2022 LBCF: A Large-Scale Budget-Constrained Causal Forest Algorithm
abstract
Offering incentives (e.g., coupons at Amazon, discounts at Uber and video bonuses at Tiktok) to user is a common strategy used by online platforms to increase user engagement and platform revenue. Despite its proven effectiveness, these marketing incentives incur an inevitable cost and might result in a low ROI (Return on Investment) if not used properly. On the other hand, different users respond differently to these incentives, for instance, some users never buy certain products without coupons, while others do anyway. Thus, how to select the right amount of incentives (i.e. treatment) to each user under budget constraints is an important research problem with great practical implications. In this paper, we call such problem as a budget-constrained treatment selection (BTS) problem.
Meng Ai, Biao Li 0002, Heyang Gong, Qingwei Yu, Shengjie Xue, Yuan Zhang 0024, Yunzhou Zhang, Peng Jiang 0002
WWW7
2022 Complementary characteristics fusion network for weakly supervised salient object detection
Yan Liu 0080, Yunzhou Zhang, Zhenyu Wang 0010, Fei Yang 0007, Cao Qin, Sonya A. Coleman, Dermot Kerr
Image Vis. Comput.2
2022 TF-SOD: a novel transformer framework for salient object detection
Zhenyu Wang 0010, Yunzhou Zhang, Yan Liu 0080, Sonya A. Coleman, Dermot Kerr
Neural Comput. Appl.2
2022 Salient object detection by aggregating contextual information
Yan Liu 0080, Yunzhou Zhang, Shichang Liu, Sonya A. Coleman, Zhenyu Wang 0010
Pattern Recognit. Lett.2
2022 A CAM-Guided Parameter-Free Attention Network for Person Re-Identification
abstract
Most existing attention mechanisms have no supervised signal during the training phase, which limits the model feature learning capability. To solve this problem, we propose a novel parameter-free attention mechanism based on class activation mapping. Attention mechanisms usually consist of spatial attention and channel attention, which indicates that “where” and “what” is more meaningful, respectively. Our attention also contains both types of attention. For Spatial Attention, we use class activation mapping as a supervision signal to guide the generation of it directly in space. Thus our spatial attention can pay more attention to the informative pedestrian parts of the scene and reduce background interference. For Channel Attention, the importance of each channel is obtained by the similarity between the aforementioned spatial attention and the feature map of each channel. In this manner, our channel attention is indirectly guided by class activation mapping. In addition, our attention is parameter-free, which reduces the risk of over-fitting. Finally, we conduct extensive evaluations on three popular benchmark datasets including Market1501, DukeMTMC-reID, and MSMT17, demonstrating the effectiveness of our approach on discriminative person representations.
Yunzhou Zhang, Sonya A. Coleman
IEEE Signal Process. Lett.2
2022 Data Assimilation Network for Generalizable Person Re-Identification
abstract
In this paper, a data assimilation network is proposed to tackle the challenges of domain generalization for person re-identification (ReID). Most of the existing research efforts only focus on single-dataset issues, and the trained models are difficult to generalize to unseen scenarios. This paper presents a distinctive idea to improve the generality of the model by assimilating three types of images: style-variant images, misaligned images and unlabeled images. The latter two are often ignored in the previous domain generalization ReID studies. In this paper, a non-local convolutional block attention module is designed for assimilating the misaligned images, and an attention adversary network is introduced to correct it. A progressive augmented memory is designed for assimilating the unlabeled images by progressive learning. Moreover, we propose an attention adversary difference loss for attention correction, and a labeling-guide discriminative embedding loss for progressive learning. Rather than designing a specific feature extractor that is robust to style shift as in most previous domain generalization work, we propose a data assimilation meta-learning procedure to train the proposed network, so that it learns to assimilate style-variant images. It is worth mentioning that we add an unlabeled augmented dataset to the source domain to tackle the domain generalization ReID tasks. Extensive experiments demonstrate that our approach significantly outperforms the state-of-the-art domain generalization methods.
Yixiu Liu, Yunzhou Zhang, Bir Bhanu, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Circuits Syst. Video Technol.2
2021 Object SLAM-Based Active Mapping and Robotic Grasping
abstract
This paper presents the first active object mapping framework for complex robotic manipulation and autonomous perception tasks. The framework is built on an object SLAM system integrated with a simultaneous multi-object pose estimation process that is optimized for robotic grasping. Aiming to reduce the observation uncertainty on target objects and increase their pose estimation accuracy, we also design an object-driven exploration strategy to guide the object mapping process, enabling autonomous mapping and high-level perception. Combining the mapping module and the exploration strategy, an accurate object map that is compatible with robotic grasping can be generated. Additionally, quantitative evaluations also indicate that the proposed framework has a very high mapping accuracy. Experiments with manipulation (including object grasping and placement) and augmented reality significantly demonstrate the effectiveness and advantages of our proposed framework.
Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Sonya A. Coleman, Wenkai Sun, Xinggang Hu, Zhiqiang Deng
3DV2
2021 Accurate and Robust Scale Recovery for Monocular Visual Odometry Based on Plane Geometry
abstract
Scale ambiguity is a fundamental problem in monocular visual odometry. Typical solutions include loop closure detection and environment information mining. For applications like self-driving cars, loop closure is not always available, hence mining prior knowledge from the environment becomes a more promising approach. In this paper, with the assumption of a constant height of the camera above the ground, we develop a light-weight scale recovery framework leveraging an accurate and robust estimation of the ground plane. The framework includes a ground point extraction algorithm for selecting high-quality points on the ground plane, and a ground point aggregation algorithm for joining the extracted ground points in a local sliding window. Based on the aggregated data, the scale is finally recovered by solving a least-squares problem using a RANSAC-based optimizer. Sufficient data and robust optimizer enable a highly accurate scale recovery. Experiments on the KITTI dataset show that the proposed framework can achieve state-of-the-art accuracy in terms of translation errors, while maintaining competitive performance on the rotation error. Due to the light-weight design, our framework also demonstrates a high frequency of 20 Hz on the dataset.
Rui Tian 0002, Yunzhou Zhang, Delong Zhu 0001, Shiwen Liang, Sonya A. Coleman, Dermot Kerr
ICRA2
2021 Multi-level cross-view consistent feature learning for person re-identification
Yixiu Liu, Yunzhou Zhang, Bir Bhanu, Sonya A. Coleman, Dermot Kerr
Neurocomputing2
2021 A visual place recognition approach using learnable feature map filtering and graph attention networks
Cao Qin, Yunzhou Zhang, Yingda Liu, Sonya A. Coleman, Huijie Du, Dermot Kerr
Neurocomputing2
2021 MFC-Net : Multi-feature fusion cross neural network for salient object detection
Zhenyu Wang 0010, Yunzhou Zhang, Yan Liu 0080, Shichang Liu, Sonya A. Coleman, Dermot Kerr
Image Vis. Comput.2
2021 Semantic loop closure detection based on graph matching in multi-objects scenes
Cao Qin, Yunzhou Zhang, Yingda Liu, Guanghao Lv
J. Vis. Commun. Image Represent.2
2021 Person search via class activation map transferring
Ruilong Li, Yunzhou Zhang, Shangdong Zhu, Shuangwei Liu
Multim. Tools Appl.2
2021 Semi-supervised learning for person re-identification based on style-transfer-generated data by CycleGANs
Shangdong Zhu, Yunzhou Zhang, Sonya A. Coleman, Ruilong Li, Shuangwei Liu
Mach. Vis. Appl.2
2020 Scale-Invariant Siamese Network For Person Re-Identification
abstract
Most existing methods for person re-identification (ReID) almost match people at a single scale and ignore that people are often distinguishable at the right spatial locations and scales. Unlike previous works designing complex convolutional neural network (CNN) architecture or concatenating multi-branch scale-specific features, we aim to employ a simple network to learn scale-invariant features. Concretely, we first propose a shared two-branch framework with two-scale images from the same identity as inputs, which is beneficial for ReID network to focus on common features in different-scale images. Furthermore, we introduce a novel attention loss to enforce discriminative regions between two branches more consistent in the visual level. Finally, we conduct extensive evaluations on three largescale datasets and report competitive performance.
Yunzhou Zhang, Shuangwei Liu, Jining Bao, Ying Wei 0007
ICIP1
2020 EAO-SLAM: Monocular Semi-Dense Object SLAM Based on Ensemble Data Association
abstract
Object-level data association and pose estimation play a fundamental role in semantic SLAM, which remain unsolved due to the lack of robust and accurate algorithms. In this work, we propose an ensemble data associate strategy for integrating the parametric and nonparametric statistic tests. By exploiting the nature of different statistics, our method can effectively aggregate the information of different measurements, and thus significantly improve the robustness and accuracy of data association. We then present an accurate object pose estimation framework, in which an outliers-robust centroid and scale estimation algorithm and an object pose initialization algorithm are developed to help improve the optimality of pose estimation results. Furthermore, we build a SLAM system that can generate semi-dense or lightweight object-oriented maps with a monocular camera. Extensive experiments are conducted on three publicly available datasets and a real scenario. The results show that our approach significantly outperforms state-of-the-art techniques in accuracy and robustness. The source code is available on https://github.com/yanmin-wu/EAO-SLAM.
Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Yonghui Feng, Sonya A. Coleman, Dermot Kerr
IROS2
2020 A new patch selection method based on parsing and saliency detection for person re-identification
Yixiu Liu, Yunzhou Zhang, Sonya A. Coleman, Bir Bhanu, Shuangwei Liu
Neurocomputing2
2020 Multi-level and multi-scale horizontal pooling network for person re-identification
Yunzhou Zhang, Shuangwei Liu, Sonya A. Coleman, Dermot Kerr
Multim. Tools Appl.1
2019 Dual Reverse Attention Networks for Person Re-Identification
abstract
In this paper, we enhance feature representation ability of person re-identification (Re-ID) by learning invariances to hard examples. Unlike previous works of hard examples mining and generating in image level, we propose a dual reverse attention networks (DRANet) to generate hard examples in the convolutional feature space. Specifically, we use a classification branch of attention mechanism to model that `what' in channel and `where' in spatial dimensions are informative in the feature maps. Meanwhile, we introduce two branches of reverse attention modules in parallel way, which convert informative feature maps into hard examples of uninformative ones. In the proposed framework, both classification and dual reverse attention branches are learned in a joint way. Experimental results on three mainstream datasets demonstrate the efficacy of the proposed method.
Shuangwei Liu, Yunzhou Zhang
ICIP3
2019 Adversarially Erased Learning for Person Re-identification by Fully Convolutional Networks
abstract
The generalization ability of deep person re-identification networks is subject to inadequate person data and occlusions. To relieve this dilemma, we propose a feature-level augmentation strategy, Adversarially Erased Learning Module (AELM), using two adversarial classifiers. Specifically, we utilize a classifier to identify discriminative regions and erase them to increase the variant of features. Meanwhile, we input the erased feature maps to another classifier to discover new body regions, which effectively resist occlusion of key parts. To easily perform end-to-end training for AELM, we propose a novel Identity model based on Fully Convolutional Networks (IFCN) to directly obtain body response heatmap during the forward pass by selecting corresponding class-specific feature map. Thus, the discriminative regions can be identified and erased in a convenient way. Moreover, to capture discriminative region for AELM, we present a Complementary Attention Module (CoAM) combined with channel and spatial attention to automatically focus on which feature types and positions are meaningful in the feature maps. In this paper, CoAM and AELM are cascaded into one module which is applied to the outputs of different convolutional layers to integrate mid- and high-level semantic features. Experimental results on three challenging benchmarks demonstrate the effectiveness of the proposed method.
Shuangwei Liu, Yunzhou Zhang, Sonya A. Coleman, Dermot Kerr, Shangdong Zhu
IJCNN2
2019 Fully convolutional multi-scale dense networks for monocular depth estimation
abstract
Monocular depth estimation is of vital importance in understanding the 3D geometry of a scene. However, inferring the underlying depth is ill‐posed and inherently ambiguous. In this study, two improvements to existing approaches are proposed. One is about a clean improved network architecture, for which the authors extend Densely Connected Convolutional Network (DenseNet) to work as end‐to‐end fully convolutional multi‐scale dense networks. The dense upsampling blocks are integrated to improve the output resolution and selected skip connection is incorporated to connect the downsampling and the upsampling paths efficiently. The other is about edge‐preserving loss functions, encompassing the reverse Huber loss, depth gradient loss and feature edge loss, which is particularly suited for estimation of fine details and clear boundaries of objects. Experiments on the NYU‐Depth‐v2 dataset and KITTI dataset show that the proposed model is competitive to the state‐of‐the‐art methods, achieving 0.506 and 4.977 performance in terms of root mean squared error respectively.
Yunzhou Zhang, Jiahua Cui, Yonghui Feng, Linzhuo Pang
IET Comput. Vis.2
2019 Three-dimensional robot localization using cameras in wireless multimedia sensor networks
Sheng Feng, Shigen Shen, Longjun Huang, Adam C. Champion, Shui Yu 0001, Chengdong Wu 0001, Yunzhou Zhang
J. Netw. Comput. Appl.7
2019 Learning sampling distribution for motion planning with local reconstruction-based self-organizing incremental neural network
Chongkun Xia, Yunzhou Zhang, I-Ming Chen 0001
Neural Comput. Appl.2
2019 Innovative architecture of single chip edge device based on virtualization technology
Yunzhou Zhang, Mo Zhang, Haoqi Liu, Gang Zhang 0001
Pervasive Mob. Comput.1
2019 Probabilistic coverage in directional sensor networks
Pengju Si, Chengdong Wu 0001, Yunzhou Zhang, Hao Chu, He Teng
Wirel. Networks3
2018 Multiple Pairwise Ranking with Implicit Feedback
abstract
As users implicitly express their preferences to items on many real-world applications, the implicit feedback based collaborative filtering has attracted much attention in recent years. Pairwise methods have shown state-of-the-art solutions for dealing with the implicit feedback, with the assumption that users prefer the observed items to the unobserved items. However, for each user, the huge unobserved items are not equal to represent her preference. In this paper, we propose a Multiple Pairwise Ranking (MPR) approach, which relaxes the simple pairwise preference assumption in previous works by further tapping the connections among items with multiple pairwise ranking criteria. Specifically, we exploit the preference difference among multiple pairs of items by dividing the unobserved items into different parts. Empirical studies show that our algorithms outperform the state-of-the-art methods on real-world datasets.
Runlong Yu, Yunzhou Zhang, Yuyang Ye 0002, Le Wu 0001, Chao Wang 0086, Qi Liu 0003, Enhong Chen
CIKM2
2018 Fractional Order Flight Control of Quadrotor UAS: an OS4 Benchmark Environment and a Case Study
abstract
The OS4 quadrotor is a classic quadrotor simulation platform. So far, many different kinds of controllers have been designed based on its plant model. Most of the research only provided a numerical simulation to verify their designed controllers. Only a few researches have put the proposed controllers back to OS4 quadrotor to verify, but they didn't share the project folder to let others continue their work. More open-source and well-documented codes are needed to accelerate the application of fractional order controllers in industry. This paper updated the OS4 folder for the latest MATLAB version. A case of study demonstrated the workflow to design a fractional order proportional derivative controller for the simulated drone. Comparisons showed that fractional order controllers perform better in a nonlinear system like OS4 than integer order PID controllers. An impulse disturbance scenario is also used as a testbed. Project folder can be accessed from: https://ww2.mathworks.cn/matlabcentral/fileexchange/67882-os4-foc. Related videos can be found from this link: https://youtu.be/heuz4tFqf64.
Bo Shang, Yunzhou Zhang, Chengdong Wu 0001, YangQuan Chen
ICARCV2
2018 Reasonable Grasping Based on Hierarchical Decomposition Models of Unknown Objects
abstract
Reasonable grasping for unknown objects is an interesting and important problem for autonomous robots in unstructured environment. Current grasping methods for unknown objects mostly focus on precision and stability. Until now there are few specific studies or reports describing reasonable grasp of unknown objects for service robots such as home service robots and nursing robots. In the paper we proposed a reasonable grasping method for unknown objects based on hierarchical decomposition point cloud models using a vision sensor. As a complex task, vision-based grasp is composed of a series solution of a mixture of subproblems. Therefore, we adopt an improved superquadrics fitting algorithm with the improved cuckoo search strategy (S-ICS) to achieve restoration and segmentation of incomplete point cloud data of unknown objects in a single visual angle. Then a reasonable region decision method based on hierarchical decomposition models is proposed to evaluate the reasonableness of grasped positions of unknown objects. Finally, we use a simulation to verify the effectiveness of the proposed method. Moreover, we also perform an extensive real-world grasping experiment on a set of unknown objects in daily use. The results also verify the effectiveness of our approach.
Chongkun Xia, Yunzhou Zhang, Yanli Shang, Tongbo Liu
ICARCV2
2018 Image Segmentation Based on Semantic Knowledge and Hierarchical Conditional Random Fields
Cao Qin, Yunzhou Zhang, Meiyu Hu, Hao Chu, Lei Wang 0207
PRCV (1)2
2016 ICRS: inter-layer compression method combined with generation of a spatial image pyramid
Yunzhou Zhang, Mo Zhang, Jinnian Wang, Gang Zhang 0001
Multim. Tools Appl.1
2013 Navigation for Indoor Mobile Robot Based on Wireless Sensor Network
Yunzhou Zhang, Guanting Fan, Jixian Zhou
WASA1
2013 Virtual edge based coverage hole detection algorithm in wireless sensor networks
abstract
With the knowledge of locations of each node in randomly deployed wireless sensor networks, the detection of coverage holes is researched in this paper. An improved hole detection algorithm is proposed based on the Boolean sensing model. The algorithm screens out hole-boundary nodes by Voronoi Diagram. In order to achieve location information of coverage holes, we introduce a new method called Virtual Edge to calculate boundary nodes. The simulation shows that compared to such popular hole detection methods as Voronoi Diagram algorithm and Simplicial Complex algorithm, the algorithm proposed can get more accurate location, shape and area information of coverage holes.
Yunzhou Zhang
WCNC1