Yu Zhang 0018

dblp:50/671-18 · DBLP profile ↗
← Back
42ranked-venue papers
5as first author
29since 2021 · last 2026
0000-0002-0043-4904ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 IO-LIO: Information-Oriented Voxel Mapping for Efficient and Precise LiDAR-Inertial Odometry
abstract
In dynamic urban scenarios involving autonomous vehicles and vehicle-infrastructure interactions, LiDAR-inertial odometry (LIO) methods are widely adopted to provide real-time vehicle pose estimation with robust and accurate positioning, particularly in urban environments where GPS signals are frequently disrupted. However, existing methods typically struggle with balancing real-time computational efficiency and localization accuracy, limiting their practical applicability. To address this challenge, we propose IO-LIO, an information-oriented voxel-based LIO system specifically designed for diverse real-world transportation scenarios. IO-LIO enhances real-time performance through concurrent processing and targeted computational optimizations, while efficiently extracting and utilizing high-value voxel information. Specifically, a lookup table-based method is introduced for rapid point cloud undistortion, coupled with an incremental distribution computation approach for information-preserving voxel downsampling. The mapping module adopts an adaptive voxel merging strategy with a multilevel Least Recently Used (M-LRU) mechanism, effectively reducing memory usage. Moreover, we propose a weighted GICP residual construction method, where residuals are weighted by voxel information values, quantitatively improving trajectory accuracy compared to state-of-the-art methods. Additionally, IO-LIO systematically addresses the critical but often overlooked challenge of memory management in LIO systems through a multi-threaded, template-based object pool. Extensive experiments were conducted in realistic ITS-relevant scenarios, including autonomous vehicles operating in dense urban traffic, vehicles navigating campus and park environments, and pedestrian-based backpack mapping. These experiments were complemented by a dedicated ablation study that quantified the incremental benefit of each module. Collectively, the results demonstrate that IO-LIO significantly surpasses state-of-the-art methods in localization accuracy, real-time performance, and reliability, highlighting its strong potential for practical deployment in intelligent transportation applications.
Junyuan Lu, Shichun Yi, Weiquan Liu, Yue Wang 0020, Rong Xiong, Yu Zhang 0018
IEEE Trans. Intell. Transp. Syst.7
2025 HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography Estimation
abstract
Feature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved substantial performance enhancements and established a widely-accepted coarse-to-fine paradigm. However, the majority of existing methods focus on improving coarse feature representation rather than the fine-matching module. Prior fine-matching techniques, which rely on point-to-patch matching probability expectation or direct regression, often lack precision and do not guarantee the continuity of feature points across sequential images. To address this limitation, this paper concentrates on enhancing the fine-matching module in the semi-dense matching framework. We employ a lightweight and efficient homography estimation network to generate the perspective mapping between patches obtained from coarse matching. This patch-to-patch approach achieves the overall alignment of two patches, resulting in a higher sub-pixel accuracy by incorporating additional constraints. By leveraging the homography estimation between patches, we can achieve a dense matching result with low computational cost. Extensive experiments demonstrate that our method achieves higher accuracy compared to previous semi-dense matchers. Meanwhile, our dense matching results exhibit similar end-point-error accuracy compared to previous dense matchers while maintaining semi-dense efficiency.
Xiaolong Wang 0013, Lei Yu 0005, Jiangwei Lao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Yu Zhang 0018, Ming Yang 0007
AAAI8
2025 Adaptive Wavelet-Positional Encoding for High-Frequency Information Learning in Implicit Neural Representation
abstract
Implicit Neural Representation (INR) has shown great potential in constructing the complex nature signal as a continuous implicit function. However, the representation results are incomplete since different components of the signal correspond to different frequencies and neural network inherently tends to low-frequency convergence. In this paper, we propose the adaptive Wavelet-Positional Encoding (WPE) to precisely represent content under different frequency distributions for coordinate-based implicit representations. The High-Frequency Perception (HFP) method is first proposed to query locations of high-frequency components from input signals, which can be indicated as local centers of WPE. Then, motivated by wavelet series regression, we present to embed these queried low-dimensional coordinate inputs into wavelet-frequency space by WPE to represent fine details of target signals. Experiments demonstrate that the proposed method can be integrated into various INR methods without modifying training frameworks while significantly improving their performance in 1D signal fitting, 2D image regression, and even 3D scene representation.
Hongxu Zhao, Zelin Gao, Yue Wang 0020, Rong Xiong, Yu Zhang 0018
AAAI5
2025 RISED: Accurate and Efficient RGB-Colorized Mapping Using Image Selection and Point Cloud Densification
abstract
Recent advances in robotics have underscored the critical role of colorized point clouds in enhancing environmental perception accuracy. However, conventional multisensor fusion Simultaneous Localization and Mapping (SLAM) systems typically employ all available images indiscriminately for point cloud colorization, resulting in suboptimal outcomes with blurred textures. Notably, achieving precise texture-togeometry alignment remains a challenge despite the availability of accurate pose estimation. This study introduces RISED, an advanced colorized mapping system that tackles this challenge from two perspectives: projection accuracy and distribution uniformity. For projection accuracy, we analyze the influence of camera poses on colorization and carefully select the optimal viewpoint to minimize errors. Regarding distribution uniformity, point cloud densification is applied to eliminate LiDAR scanning traces. Furthermore, a novel evaluation method is introduced to provide comprehensive assessment of colorized point clouds, filling a gap in this field. Experimental results show that our method outperforms traditional approaches in RGB-colorized mapping. Specifically, our method achieves notable improvements in projection accuracy (55.2 %), geometric accuracy (63.1 %), and surface coverage (30.8 %).
Changjian Jiang, Zeyu Wan, Ruilan Gao, Yue Wang 0020, Rong Xiong, Yu Zhang 0018
ICRA7
2025 LHMM: A Tightly-Coupled LiDAR-Inertial Hybrid-Map Matching Approach for Robust and Efficient Global Localization
abstract
LiDAR map matching (LMM) faces two key challenges: the enormous number of point clouds imposes constraints on storage and computation, and traditional two-stage frameworks suffer from initial guess errors during degeneration. This paper presents LHMM, a hybrid-map framework that first compresses the prior map and then performs tightly coupled pose estimation within a Maximum A Posteriori (MAP) estimation formulation. First, a skeletonization-based prior map compression method is proposed, which retains only stable structural features, reducing the map storage while enabling fast runtime association through a dual-mode map representation. Second, constraints from IMU, skeleton-feature prior map, and local voxel map are jointly optimized within a unified MAP formulation, recovering the full system state in a single step and preventing error cascades. The local map benefits from a hole-aware keyframe mechanism, focusing on regions with environmental changes or areas with partial map coverage, thereby reducing computation compared to full mapping. Extensive evaluations across multiple datasets demonstrate that LHMM not only reduces storage and computational overhead but also outperforms state-of-the-art methods in terms of localization accuracy and robustness. We will open-source the code1.
Junyuan Lu, Qishu Wu, Yu Zhang 0018
IROS3
2025 EliGen: Entity-Level Controlled Image Generation with Regional Attention
abstract
Recent advancements in diffusion models have significantly advanced text-to-image generation, yet global text prompts alone remain insufficient for achieving fine-grained control over individual entities within an image. To address this limitation, we present EliGen, a novel framework for Entity-level controlled image Generation. Firstly, we put forward regional attention, a mechanism for diffusion transformers that requires no additional structures, seamlessly integrating entity prompts and arbitrary-shaped spatial masks. By contributing a high-quality dataset with fine-grained spatial and semantic entity-level annotations, we train EliGen to achieve robust and accurate entity-level manipulation, surpassing existing methods in both spatial precision and image quality. Additionally, we propose an inpainting fusion pipeline, extending EliGen’s capabilities to multi-entity image inpainting tasks. We further demonstrate EliGen’s flexibility by integrating it with other open-source models such as IP-Adapter, In-Context LoRA and MLLM, unlocking new creative possibilities. The source code, model, and dataset will be published.
Zhongjie Duan, Yingda Chen, Yu Zhang 0018
MMAsia5
2025 LD-Seg: Training-Free Novel Instance Segmentation Based on LVLM-Driven Vision Foundation Models
Yingnan Guo, Yongliang Lin, Hanqing Yang 0002, Yu Zhang 0018
PRCV (17)5
2025 Recurrent spiking neural networks as models of the entorhinal-hippocampal system for path integration: Grid cells and beyond
Ruilan Gao, Changjian Jiang, Yu Zhang 0018
Neurocomputing3
2025 Resolving Symmetry Ambiguity in Correspondence-Based Methods for Instance-Level Object Pose Estimation
abstract
Estimating the 6D pose of an object from a single RGB image is a critical task that becomes additionally challenging when dealing with symmetric objects. Recent approaches typically establish one-to-one correspondences between image pixels and 3D object surface vertices. However, the utilization of one-to-one correspondences introduces ambiguity for symmetric objects. To address this, we propose SymCode, a symmetry-aware surface encoding that encodes the object surface vertices based on one-to-many correspondences, eliminating the problem of one-to-one correspondence ambiguity. We also introduce SymNet, a fast end-to-end network that directly regresses the 6D pose parameters without solving a PnP problem. We demonstrate faster runtime and comparable accuracy achieved by our method on the T-LESS and IC-BIN benchmarks of mostly symmetric objects. The code is available at https://github.com/lyltc1/SymNet.
Yongliang Lin, Yongzhi Su, Sandeep Inuganti, Yan Di, Naeem Ajilforoushan, Hanqing Yang 0002, Yu Zhang 0018, Jason R. Rambach
IEEE Trans. Image Process.7
2024 Knowledge-Aware Self-supervised Educational Resources Recommendation
Jing Chen 0037, Yu Zhang 0018, Zhenghao Liu 0001, Minghe Yu 0001, Bin Xu 0003, Ge Yu 0001
WISA2
2024 Memory-Efficient Reversible Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are potential competitors to artificial neural networks (ANNs) due to their high energy-efficiency on neuromorphic hardware. However, SNNs are unfolded over simulation time steps during the training process. Thus, SNNs require much more memory than ANNs, which impedes the training of deeper SNN models. In this paper, we propose the reversible spiking neural network to reduce the memory cost of intermediate activations and membrane potentials during training. Firstly, we extend the reversible architecture along temporal dimension and propose the reversible spiking block, which can reconstruct the computational graph and recompute all intermediate variables in forward pass with a reverse process. On this basis, we adopt the state-of-the-art SNN models to the reversible variants, namely reversible spiking ResNet (RevSResNet) and reversible spiking transformer (RevSFormer). Through experiments on static and neuromorphic datasets, we demonstrate that the memory cost per image of our reversible SNNs does not increase with the network depth. On CIFAR10 and CIFAR100 datasets, our RevSResNet37 and RevSFormer-4-384 achieve comparable accuracies and consume 3.79x and 3.00x lower GPU memory per image than their counterparts with roughly identical model complexity and parameters. We believe that this work can unleash the memory constraints in SNN training and pave the way for training extremely large and deep SNNs.
Yu Zhang 0018
AAAI2
2024 HiPose: Hierarchical Binary Surface Encoding and Correspondence Pruning for RGB-D 6DoF Object Pose Estimation
abstract
In this work, we present a novel dense-correspondence method for 6DoF object pose estimation from a single RGB-D image. While many existing data-driven methods achieve impressive performance, they tend to be time-consuming due to their reliance on rendering-based refinement approaches. To circumvent this limitation, we present HiPose, which establishes 3D-3D correspondences in a coarse-to-fine manner with a hierarchical binary surface encoding. Unlike previous dense-correspondence methods, we estimate the correspondence surface by employing point-to-surface matching and iteratively constricting the surface until it becomes a correspondence point while gradually removing outliers. Extensive experiments on public benchmarks LM-O, YCB-V, and T-Less demonstrate that our method surpasses all refinement-free methods and is even on par with expensive refinement-based approaches. Crucially, our approach is computationally efficient and enables real-time critical applications with high accuracy requirements.
Yongliang Lin, Yongzhi Su, Praveen Nathan, Sandeep Inuganti, Yan Di, Martin Sundermeyer, Fabian Manhardt, Didier Stricker, Jason R. Rambach, Yu Zhang 0018
CVPR10
2024 RGBD-based Image Goal Navigation with Pose Drift: A Topo-metric Graph based Approach
abstract
Image-goal navigation in unknown environments with sensor error is of considerable difficulty for autonomous robots. In this paper, we propose a drift-resisting topo-metric graph to map the environment and localize the robot using only relative poses. The error-sharing mechanism under this representation effectively reduces the impact of accumulated drifts commonly encountered in navigation tasks. A Reinforcement Learning based policy was proposed for sub-goal selection on this topo-metric graph, which improves navigation efficiency by handling task-driven features taking both image correlation and topological layout into account. We adopt a modular system design with this map representation and graph policy, leaving the low-level motion planning problems to classical controllers for better stability and generalizability. Experimental results demonstrate that our method can achieve robust navigation performance in a variety of unknown environments and even 50% higher success rate over existing methods in complex environments with odometry drift.
Shuhao Ye, Yuxiang Cui, Hao Sha 0002, Yu Zhang 0018, Rong Xiong, Yue Wang 0020
ICRA5
2024 ERASOR++: Height Coding Plus Egocentric Ratio Based Dynamic Object Removal for Static Point Cloud Mapping
abstract
Mapping plays a crucial role in location and navigation within automatic systems. However, the presence of dynamic objects in 3D point cloud maps generated from scan sensors can introduce map distortion and long traces, thereby posing challenges for accurate mapping and navigation. To address this issue, we propose ERASOR++, an enhanced approach based on the Egocentric Ratio of Pseudo Occupancy for effective dynamic object removal. To begin, we introduce the Height Coding Descriptor, which combines height difference and height layer information to encode the point cloud. Subsequently, we propose the Height Stack Test, Ground Layer Test, and Surrounding Point Test methods to precisely and efficiently identify the dynamic bins within point cloud bins, thus overcoming the limitations of prior approaches. Through extensive evaluation on open-source datasets, our approach demonstrates superior performance in terms of precision and efficiency compared to existing methods. Furthermore, the techniques described in our work hold promise for addressing various challenging tasks or aspects through subsequent migration.
Yu Zhang 0018
ICRA2
2024 Versatile correlation learning for size-robust generalized counting: A new perspective
Hanqing Yang 0002, Sijia Cai, Bing Deng, Mohan Wei, Yu Zhang 0018
Knowl. Based Syst.5
2024 Context-Aware and Semantic-Consistent Spatial Interactions for One-Shot Object Detection Without Fine-Tuning
abstract
One-shot object detection (OSOD) without fine-tuning has recently garnered considerable attention and research focus. It aims to directly detect novel-class objects in the target image by providing merely one support image patch without undergoing the fine-tuning stage. However, most existing methods adopt image pair matching regardless of the scale inconsistency and spatial semantic mismatch of image pairs, which limits their ability to acquire high-quality target-support related features. This paper addresses these limitations by incorporating cross-scale contexts and semantic-consistent cues that are robust against the challenges of scarce and ambiguous matching. Specifically, we first introduce a simple yet effective Aggregation-Transformer-based Pyramid (ATP) module to explore the long-range cross-scale spatial interactions by employing the customized size-aware aggregation approach and the vanilla transformer encoder, thus the coarse-to-fine local image patterns are optimally utilized. Furthermore, we formulate the 4D contrastive cross-correlation tensor for instance-level features matching and suggest a Geometric Consistent Correlation (GCC) module that utilizes the bidirectional spatial-aware convolutions to extract the long-range semantic correspondences for target-support pairs. Additionally, a Channel Contrastive Learning (CCL) branch is adopted to complement the inter-channel interactions between target-support pairs for the GCC module. Extensive experiments demonstrate that our approach significantly outperforms the previous state-of-the-art methods by 6.5% and 2.1% on PASCAL VOC and COCO datasets for unseen classes, respectively.
Hanqing Yang 0002, Sijia Cai, Bing Deng, Jieping Ye, Guosheng Lin, Yu Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.6
2024 Learning Active Force-Torque Based Policy for Sub-mm Localization of Unseen Holes
abstract
Hole localization is crucial in the peg-in-hole process. Our goal is to enable robots to operate effectively in contact-rich environments with tight tolerances, and adapt to new tasks involving unseen peg-hole pairs. Most existing “black-box” methods train a policy that performs the task directly from perceptual inputs, which requires extensive real-world interactions for task adaptation. Departing from this direct mapping paradigm, our work propose to formulate the task as a force matching and localization problem, where the objective is to establish correspondences between current and template force–torque observation maps for localization purpose. The formulation enables the design of a decoupled map-locator-policy framework, offering improved success rates, efficiency, and augmented generalization capabilities, surpassing current state-of-the-art methods. Experiments demonstrate the effectiveness of the proposed method, achieving a 90% success rate across 12 unseen 3-D models and a variety of unseen tight workpieces. Within a mere 5-min adaption process, the performance can be further improved by more than 95%.
Yu Zhang 0018, Rong Xiong, Yue Wang 0020
IEEE Trans. Ind. Informatics4
2023 Adaptive Positional Encoding for Bundle-Adjusting Neural Radiance Fields
abstract
Neural Radiance Fields have shown great potential to synthesize novel views with only a few discrete image observations of the world. However, the requirement of accurate camera parameters to learn scene representations limits its further application. In this paper, we present adaptive positional encoding (APE) for bundle-adjusting neural radiance fields to reconstruct the neural radiance fields from unknown camera poses (or even intrinsics). Inspired by Fourier series regression, we investigate its relationship with the positional encoding method and therefore propose APE where all frequency bands are trainable. Furthermore, we introduce period-activated multilayer perceptrons (PMLPs) to construct the implicit network for the high-order scene representations and fine-grained gradients during backpropagation. Experimental results on public datasets demonstrate that the proposed method with APE and PMLPs can outperform the state-of-the-art methods in accurate camera poses and high-fidelity view synthesis.
Zelin Gao, Weichen Dai 0001, Yu Zhang 0018
ICCV3
2023 Continuous-Time LiDAR-Inertial-Vehicle Odometry Method with Lateral Acceleration Constraint
abstract
In this paper, we propose a continuous-time-based LiDAR-inertial-vehicle odometry method, which can tightly fuse the data from Light Detection And Ranging (LiDAR), inertial measurement units (IMU), and vehicle measurements. The lateral acceleration constraint is further added to trajectory estimation to make the estimated trajectory follow the motion characteristics of vehicles. In addition, since vehicle model parameters vary with different motion conditions and tyre pressure, we estimate vehicle correction factors that rectify changes in vehicle model parameters online, and also analyze the observability of these vehicle correction factors. In experiments, the proposed method is evaluated and compared with state-of-the-art methods in the public dataset. The experimental results show that the proposed method achieves more accurate results in all sequences since we add additional sensor measurements and utilize the characteristic of vehicle motion to restrict the trajectory estimation. The ablation study also proved the effectiveness of continuous-time representation, online correction factor estimation, and incorporation of lateral acceleration constraint.
Weichen Dai 0001, Zeyu Wan, Yu Zhang 0018
ICRA5
2023 Fine-Grained Cross-View Geo-Localization Using a Correlation-Aware Homography Estimator
abstract
In this paper, we introduce a novel approach to fine-grained cross-view geo-localization. Our method aligns a warped ground image with a corresponding GPS-tagged satellite image covering the same area using homography estimation. We first employ a differentiable spherical transform, adhering to geometric principles, to accurately align the perspective of the ground image with the satellite map. This transformation effectively places ground and aerial images in the same view and on the same plane, reducing the task to an image alignment problem. To address challenges such as occlusion, small overlapping range, and seasonal variations, we propose a robust correlation-aware homography estimator to align similar parts of the transformed ground image with the satellite image. Our method achieves sub-pixel resolution and meter-level GPS accuracy by mapping the center point of the transformed ground image to the satellite image using a homography matrix and determining the orientation of the ground camera using a point above the central axis. Operating at a speed of 30 FPS, our method outperforms state-of-the-art techniques, reducing the mean metric localization error by 21.3\% and 32.4\% in same-area and cross-area generalization tasks on the VIGOR benchmark, respectively, and by 34.4\% on the KITTI benchmark in same-area evaluation.
Xiaolong Wang 0013, Runsen Xu, Zhuofan Cui, Zeyu Wan, Yu Zhang 0018
NeurIPS5
2023 A Deep Reinforcement Learning Based Real-Time Solution Policy for the Traveling Salesman Problem
abstract
The rapid development of logistics and navigation has led to increasing demand for solving route optimization problems in real-time. The traveling salesman problem (TSP) tends to require fast and reliable online solutions, which may not be met by traditional iterative optimization algorithms. In this work, a real-time solution policy is proposed for TSP. The idea is to build a mapping between city information and optimal solutions using deep neural networks. Therefore, when given a new set of city coordinates, the optimal route can be directly and quickly calculated without iteration. Considering the recent advancement in computer vision with deep convolutional neural networks (DCNNs), an image representation is proposed to convert TSP to a computer vision problem. A problem decomposition method is introduced to reduce the mapping complexity. Taking advantage of the powerful fitting capabilities of DCNN, a deep reinforcement learning method is designed without any labeling requirement. The proposed method is superior for real-time applications compared with other algorithms.
Zhengxuan Ling, Yu Zhang 0018, Xi Chen 0075
IEEE Trans. Intell. Transp. Syst.2
2022 Balanced and Hierarchical Relation Learning for One-shot Object Detection
abstract
Instance-level feature matching is significantly important to the success of modern one-shot object detectors. Re-cently, the methods based on the metric-learning paradigm have achieved an impressive process. Most of these works only measure the relations between query and target objects on a single level, resulting in suboptimal performance overall. In this paper, we introduce the balanced and hierarchical learning for our detector. The contributions are two-fold: firstly, a novel Instance-level Hierarchical Relation (IHR) module is proposed to encode the contrastive-level, salient-level, and attention-level relations simultane-ously to enhance the query-relevant similarity representation. Secondly, we notice that the batch training of the IHR module is substantially hindered by the positive-negative sample imbalance in the one-shot scenario. We then in-troduce a simple but effective Ratio-Preserving Loss (RPL) to protect the learning of rare positive samples and sup-press the effects of negative samples. Our loss can adjust the weight for each sample adaptively, ensuring the desired positive-negative ratio consistency and boosting query-related IHR learning. Extensive experiments show that our method outperforms the state-of-the-art method by 1.6% and 1.3% on PASCAL VOC and MS COCO datasets for unseen classes, respectively. The code will be available at https://github.com/hero-y/BHRL.
Hanqing Yang 0002, Sijia Cai, Hualian Sheng, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Yu Zhang 0018
CVPR8
2022 BorderPointsMask: One-stage instance segmentation with boundary points representation
Hanqing Yang 0002, Liyang Zheng, Saba Ghorbani Barzegar, Yu Zhang 0018, Bin Xu 0003
Neurocomputing4
2022 Isomorphic model-based initialization for convolutional neural networks
Hanqing Yang 0002, Yu Zhang 0018
J. Vis. Commun. Image Represent.5
2022 RGB-D SLAM in Dynamic Environments Using Point Correlations
abstract
In this paper, a simultaneous localization and mapping (SLAM) method that eliminates the influence of moving objects in dynamic environments is proposed. This method utilizes the correlation between map points to separate points that are part of the static scene and points that are part of different moving objects into different groups. A sparse graph is first created using Delaunay triangulation from all map points. In this graph, the vertices represent map points, and each edge represents the correlation between adjacent points. If the relative position between two points remains consistent over time, there is correlation between them, and they are considered to be moving together rigidly. If not, they are considered to have no correlation and to be in separate groups. After the edges between the uncorrelated points are removed during point-correlation optimization, the remaining graph separates the map points of the moving objects from the map points of the static scene. The largest group is assumed to be the group of reliable static map points. Finally, motion estimation is performed using only these points. The proposed method was implemented for RGB-D sensors, evaluated with a public RGB-D benchmark, and tested in several additional challenging environments. The experimental results demonstrate that robust and accurate performance can be achieved by the proposed SLAM method in both slightly and highly dynamic environments. Compared with other state-of-the-art methods, the proposed method can provide competitive accuracy with good real-time performance.
Weichen Dai 0001, Yu Zhang 0018, Ping Li 0017, Zheng Fang 0001, Sebastian A. Scherer
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Can Deep Learning Solve Parametric Mathematical Programming? An Application to 0-1 Linear Programming Through Image Representation
abstract
Deep learning has been widely applied in many fields. Efficient optimization algorithms contribute a lot to the enhancement of deep learning. However, reverse studies on how deep learning can solve optimization problems, especially mathematical programming, are relatively scarce. In this work, we aim to initiate a discussion on using deep learning to solve parametric mathematical programming, which can be converted to find a mapping from parameter space to solution space. Given that deep convolutional neural networks (DCNNs) have been successfully applied in the field of computer vision, converting a mathematical programming problem into an image representation may provide the solution. This work takes the 0–1 knapsack problem (KP) as an example to build an original image representation method. After modifying the image details, the proposed method is extended to the general 0–1 linear programming (LP) problem, and a deep learning architecture is designed to solve the transferred computer vision problem. The efficiency of the proposed DCNNs method is validated on a large number of problems, and results show that this method can solve the 0–1 LP accurately and efficiently without iteration. The solution speed is 20 times faster than that of traditional optimization solvers.
Zhengxuan Ling, Yu Zhang 0018, Xi Chen 0075
IEEE Trans. Syst. Man Cybern. Syst.3
2021 A Multi-spectral Dataset for Evaluating Motion Estimation Systems
abstract
Visible images have been widely used for motion estimation. Thermal images, in contrast, are more challenging to be used in motion estimation since they typically have lower resolution, less texture, and more noise. In this paper, a novel dataset for evaluating the performance of multi-spectral motion estimation systems is presented. All the sequences are recorded from a handheld multi-spectral device. It consists of a standard visible-light camera, a long-wave infrared camera, an RGB-D camera, and an inertial measurement unit (IMU). The multi-spectral images, including both color and thermal images in full sensor resolution (640 × 480), are obtained from a standard and a long-wave infrared camera at 32Hz with hardware-synchronization. The depth images are captured by a Microsoft Kinect2 and can have benefits for learning cross-modalities stereo matching. For trajectory evaluation, accurate ground-truth camera poses obtained from a motion capture system are provided. In addition to the sequences with bright illumination, the dataset also contains dim, varying, and complex illumination scenes. The full dataset, including raw data and calibration data with detailed data format specifications, is publicly available.
Weichen Dai 0001, Yu Zhang 0018, Shenzhou Chen, Donglei Sun, Da Kong
ICRA2
2021 Towards improving classification power for one-shot object detection
Hanqing Yang 0002, Yongliang Lin, Yu Zhang 0018, Bin Xu 0003
Neurocomputing4
2021 Solving Optimization Problems Through Fully Convolutional Networks: An Application to the Traveling Salesman Problem
abstract
In the new wave of artificial intelligence, deep learning is impacting various industries. As a closely related area, optimization algorithms greatly contribute to the development of deep learning. But the reverse applications are still insufficient. Is there any efficient way to solve certain optimization problem through deep learning? The key is to convert the optimization to a representation suitable for deep learning. In this article, a traveling salesman problem (TSP) is studied. Considering that deep learning is good at image processing, an image representation method is proposed to transfer a TSP to an image. Based on samples of a ten city TSP, a fully convolutional network (FCN) is used to learn the mapping from a feasible region to an optimal solution. The training process is analyzed and interpreted through stages. A visualization method is presented to show how an FCN can understand the training task of a TSP. Once the training is completed, no significant effort is required to solve a new TSP and the prediction is obtained on the scale of milliseconds. The results show good performance in finding the global optimal solution. Moreover, the developed FCN model has been demonstrated on TSP’s with different city numbers, proving excellent generalization performance.
Zhengxuan Ling, Xinyu Tao, Yu Zhang 0018, Xi Chen 0075
IEEE Trans. Syst. Man Cybern. Syst.3
2020 Influence of Periodic Role Switching Intervals on Pair Programming Effectiveness
Bin Xu 0003, Kening Gao, Yu Zhang 0018, Ge Yu 0001
WISA4
2020 Robust adaptive control of hypersonic flight vehicle with asymmetric AOA constraint
Yuyan Guo, Bin Xu 0003, Weixin Han, Shuai Li 0002, Yueping Wang, Yu Zhang 0018
Sci. China Inf. Sci.6
2020 Novel 3D point set registration method based on regionalized Gaussian process map reconstruction
abstract
Point set registration has been a topic of significant research interest in the field of mobile intelligent unmanned systems. In this paper, we present a novel approach for a three-dimensional scan-to-map point set registration. Using Gaussian process (GP) regression, we propose a new type of map representation, based on a regionalized GP map reconstruction algorithm. We combine the predictions and the test locations derived from the GP as the predictive points. In our approach, the correspondence relationships between predictive point pairs are set up naturally, and a rigid transformation is calculated iteratively. The proposed method is implemented and tested on three standard point set datasets. Experimental results show that our method achieves stable performance with regard to accuracy and efficiency, on a par with two standard methods, the iterative closest point algorithm and the normal distribution transform. Our mapping method also provides a compact point-cloud-like map and exhibits low memory consumption.
Bo Li 0071, Yu Zhang 0018, Ping Li 0017
Frontiers Inf. Technol. Electron. Eng.2
2019 Multi-Spectral Visual Odometry without Explicit Stereo Matching
abstract
Multi-spectral sensors consisting of a standard (visible-light) camera and a long-wave infrared camera can simultaneously provide both visible and thermal images. Since thermal images are independent from environmental illumination, they can help to overcome certain limitations of standard cameras under complicated illumination conditions. However, due to the difference in the information source of the two types of cameras, their images usually share very low texture similarity. Hence, traditional texture-based feature matching methods cannot be directly applied to obtain stereo correspondences. To tackle this problem, a multi-spectral visual odometry method without explicit stereo matching is proposed in this paper. Bundle adjustment of multi-view stereo is performed on the visible and the thermal images using direct image alignment. Scale drift can be avoided by additional temporal observations of map points with the fixed-baseline stereo. Experimental results indicate that the proposed method can provide accurate visual odometry results with recovered metric scale. Moreover, the proposed method can also provide a metric 3D reconstruction in semi-dense density with multi-spectral information, which is not available from existing multi-spectral methods.
Weichen Dai 0001, Yu Zhang 0018, Donglei Sun, Naira Hovakimyan, Ping Li 0017
3DV2
2019 A Method of Link Prediction Using Meta Path and Attribute Information
Yu Zhang 0018, Kening Gao, Ge Yu 0001
WISA1
2019 Vision Information and Laser Module Based UAV Target Tracking
abstract
This paper investigates the target tracking mission of an Unmanned Aerial Vehicle (UAV) equipped with a camera and a laser module. Firstly, utilizing Deep Neural Network (DNN) and Kernelized Correlation Filters (KCF), target recognition and location in the pixel coordinate system is achieved based on vision. Furthermore, by combining the laser ranging information and the distance estimation algorithm based on image, the distance between the UAV and the target is well estimated. To ensure the target tracking, a PID controller based on the distance error is applied to the UAV. The effectiveness of the system is verified on an actual UAV target tracking scenario.
Chang Liu 0049, Yansui Song, Yuyan Guo, Bin Xu 0003, Yu Zhang 0018, Zhen Li 0011
IECON5
2019 Uncalibrated downward-looking UAV visual compass based on clustered point features
Yu Zhang 0018, Ping Li 0017, Bin Xu 0003
Sci. China Inf. Sci.2
2018 Feature Regions Segmentation Based RGB-D Visual Odometry in Dynamic Environment
abstract
A novel RGB-D visual odometry method for dynamic environment is proposed. Majority of visual odometry systems can only work in static environments, which limits their applications in real world. In order to improve the accuracy and robustness of visual odometry in dynamic environment, a Feature Regions Segmentation algorithm is proposed to resist the disturbance caused by the moving objects. The matched features are divided into different regions to separate the moving objects from the static background. The features in the largest region which belong to the static background are used to estimate the camera pose finally. The effectiveness of our visual odometry method is verified in a dynamic environment of our lab. Furthermore, an exhaustive experimental evaluation is conducted on benchmark datasets including static environments and dynamic environments compared with the state-of-art visual odometry systems. The accuracy comparison results show that the proposed algorithm outperforms those systems in large scale dynamic environments. Our method tracks the camera movement correctly while others failed. In addition, our method can give the same good performances in static environment. Experiments demonstrate that the proposed RGB-D visual odometry can obtain accurate and robust estimation results in dynamic environments.
Yu Zhang 0018, Weichen Dai 0001, Ping Li 0017, Zheng Fang 0001
IECON1
2016 A novel biologically inspired ELM-based network for image recognition
Yu Zhang 0018, Ping Li 0017
Neurocomputing1
2015 Neural control of hypersonic flight dynamics with actuator fault and constraint
Shixing Wang 0002, Yu Zhang 0018, Yuqiang Jin
Sci. China Inf. Sci.2
2015 Neural discrete back-stepping control of hypersonic flight vehicle with equivalent prediction model
Bin Xu 0003, Yu Zhang 0018
Neurocomputing2
2015 MLP technique based reinforcement learning control of discrete pure-feedback systems
Yu Zhang 0018, Shixing Wang 0002
Neurocomputing1
2012 Using Non-topological Node Attributes to Improve Results of Link Prediction in Social Networks
abstract
This paper examines the importance of non-topological node attributes for link prediction in social networks. Rank method and supervised learning method were introduced to show the role of the node attributes in link prediction respectively. A rule for choosing the appropriate node attributes was discussed and a method for aggregating two node attributes was proposed. The result of the experiments on a blog dataset showed that using non-topological node attributes make a better performance in link prediction.
Yu Zhang 0018, Bin Xu 0003, Kening Gao, Ge Yu 0001
WISA1