Zhiyuan Zhang 0002

dblp:72/1760-2 · DBLP profile ↗
← Back
16ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0002-2605-1549ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 L2MLP: A novel MLP-based locality learning method for point cloud analysis
Zhiyuan Zhang 0002, Zhihui Li 0002, Panhe Hu, Junpeng Shi
Pattern Recognit.1
2026 Corrigendum to "L2MLP: A novel MLP-based locality learning method for point cloud analysis" [Pattern Recognition 171 (2026) 112115]
Zhiyuan Zhang 0002, Zhihui Li 0002, Panhe Hu, Junpeng Shi
Pattern Recognit.1
2026 Generative-contrastive learning for open set radar emitter identification
Dongming Wu 0003, Junpeng Shi, Zhiyuan Zhang 0002, Zhihui Li 0002, Fangling Zeng
Signal Process.3
2026 Detection Drives an End-to-End Fusion of Infrared and Visible Images Based on Diffusion Models
abstract
Infrared and visible image fusion methods have shown promising results, yet existing approaches either compromise downstream detection performance through independent fusion processes or sacrifice computational efficiency and flexibility by requiring joint training of fusion and detection models. To address these challenges, we propose a detection-driven image fusion network based on diffusion models (termed as DDIF), which optimizes the fused images specifically for object detection tasks. Our method features the following three aspects: 1) we reformulate the image fusion process as an inverse problem solved by a non-differentiable optimization process wherein the fused result preserves the source modality information while conforming to the image prior provided by the diffusion model; 2) we design a Response Guide Learning Module (RGLM) to learn response maps, which determine the contribution of each modality in the fusion process according to the downstream detection task; 3) we establish explicit gradient relationships to ensure compatibility between RGLM training and the non-differentiable optimization process, enabling end-to-end training. Notably, a moderate coupling mechanism is formed in our framework as the subsequent detection model is pre-trained and frozen, enabling flexible integration with various advanced detection networks while maintaining computational efficiency. Extensive experiments indicate that our method achieves superior detection performance compared to SOTA approaches and produces high-quality image fusion results.
Zhiyuan Zhang 0002, Chaohua Shi, Junpeng Shi, Yongxiang Liu
IEEE Trans. Image Process.2
2025 MRM-RETrack: Hybrid Multi-scale Residual and Mamba for RGB-Event Tracking
Bin Fan 0002, Zhexiong Wan, Zhiyuan Zhang 0002, Yuchao Dai
PRCV (18)4
2024 Unsupervised Pose Decoder: Learn to Disentangle the Pose Attribute for Point Cloud Shape Analysis
abstract
Pose is a fundamental attribute of 3D point cloud shape, which profoundly impacts point cloud analysis tasks. However, it is very tricky to directly solve the pose attribute since it is deeply coupled with geometry shape. To this end, the representation separation strategy has been proposed, where the global representation is modeled as a combination of the pose-related part representation and the geometry shape part representation. However, these methods still can not model the representation of the pose attribute well. As a reply, we design a new pose decoder in this paper, learning to disentangle the pose attribute by exploiting its complement,i.e. the geometry shape part representation. Specifically, a Siamese structure is introduced constituting of two shared branches, where two consistent point clouds with different pose attributes are input. The geometry shape part representation and the global representation are learned in each branch network to solving the pose-related part representation for disentangling the pose distribution. Then, we emphasize the completeness and no-redundancy of geometry shape part representation by designing two constraints. 1) We recover the learned geometry shape part representation to a point cloud and enforce it to maintain the same geometry shape as the original input point cloud to guarantee all geometry shape information is retained. 2) We develop two geometry shape part representations embedded from two branches to be the same so as to filter the pose information out. These two constraints are incorporated into the unsupervised loss function to train our pose decoder. Our pose decoder can be integrated into different point cloud shape analysis methods. We evaluate our pose decoder in point cloud classification and part segmentation tasks to handle the pose diversity problem of the input point cloud, which significantly improves the robustness. Besides, the obtained respective poses of input point clouds can be used to register them naturally, making the unsupervised method achieving superior performance.
Zhiyuan Zhang 0002, Zhihui Li 0002, Mingyang Du, Junpeng Shi
IEEE Trans. Geosci. Remote. Sens.1
2022 End-to-End Learning the Partial Permutation Matrix for Robust 3D Point Cloud Registration
abstract
Even though considerable progress has been made in deep learning-based 3D point cloud processing, how to obtain accurate correspondences for robust registration remains a major challenge because existing hard assignment methods cannot deal with outliers naturally. Alternatively, the soft matching-based methods have been proposed to learn the matching probability rather than hard assignment. However, in this paper, we prove that these methods have an inherent ambiguity causing many deceptive correspondences. To address the above challenges, we propose to learn a partial permutation matching matrix, which does not assign corresponding points to outliers, and implements hard assignment to prevent ambiguity. However, this proposal poses two new problems, i.e. existing hard assignment algorithms can only solve a full rank permutation matrix rather than a partial permutation matrix, and this desired matrix is defined in the discrete space, which is non-differentiable. In response, we design a dedicated soft-to-hard (S2H) matching procedure within the registration pipeline consisting of two steps: solving the soft matching matrix (S-step) and projecting this soft matrix to the partial permutation matrix (H-step). Specifically, we augment the profit matrix before the hard assignment to solve an augmented permutation matrix, which is cropped to achieve the final partial permutation matrix. Moreover, to guarantee end-to-end learning, we supervise the learned partial permutation matrix but propagate the gradient to the soft matrix instead. Our S2H matching procedure can be easily integrated with existing registration frameworks, which has been verified in representative frameworks including DCP, RPMNet, and DGR. Extensive experiments have validated our method, which creates a new state-of-the-art performance.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He
AAAI1
2022 Context-Aware Video Reconstruction for Rolling Shutter Cameras
abstract
With the ubiquity of rolling shutter (RS) cameras, it is becoming increasingly attractive to recover the latent global shutter (GS) video from two consecutive RS frames, which also places a higher demand on realism. Existing solutions, using deep neural networks or optimization, achieve promising performance. However, these methods generate intermediate GS frames through image warping based on the RS model, which inevitably result in black holes and noticeable motion artifacts. In this paper, we alleviate these issues by proposing a context-aware GS video reconstruction architecture. It facilitates the advantages such as occlusion reasoning, motion compensation, and temporal abstraction. Specifically, we first estimate the bilateral motion field so that the pixels of the two RS frames are warped to a common GS frame accordingly. Then, a refinement scheme is proposed to guide the GS frame synthesis along with bilateral occlusion masks to produce high-fidelity GS video frames at arbitrary times. Furthermore, we derive an approximated bilateral motion field model, which can serve as an alternative to provide a simple but effective GS frame initialization for related tasks. Experiments on synthetic and real data show that our approach achieves superior performance over state-of-the-art methods in terms of objective metrics and subjective visual quality. Code is available at https://github.com/GitCVfb/CVR.
Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002, Qi Liu 0054, Mingyi He
CVPR3
2022 Differential SfM and image correction for a rolling shutter stereo rig
Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002
Image Vis. Comput.3
2022 A Representation Separation Perspective to Correspondence-Free Unsupervised 3-D Point Cloud Registration
abstract
3-D point cloud registration in remote sensing field has been greatly advanced by deep learning-based methods, where the rigid transformation is either directly regressed from the two point clouds (correspondences-free approaches) or computed from the learned correspondences (correspondences-based approaches). Existing correspondence-free methods generally learn the holistic representation of the entire point cloud, which is fragile for partial and noisy point clouds. In this letter, we propose a correspondence-free unsupervised point cloud registration (UPCR) method from the representation separation perspective. First, we model the input point cloud as a combination of pose-invariant representation and pose-related representation. Second, the pose-related representation is used to learn the relative pose w.r.t. a “latent canonical shape” for thesourceandtargetpoint clouds, respectively. Third, the rigid transformation is obtained from the above two learned relative poses. Our method not only filters out the disturbance in pose-invariant representation but also is robust to partial-to-partial point clouds or noise. Experiments on benchmark datasets demonstrate that our unsupervised method achieves comparable if not better performance than state-of-the-art supervised registration methods.The source code will be made public.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He
IEEE Geosci. Remote. Sens. Lett.1
2022 Self-supervised rigid transformation equivariance for accurate 3D point cloud registration
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He
Pattern Recognit.1
2022 Fast and Robust Differential Relative Pose Estimation With Radial Distortion
abstract
In this letter, we address the differential two-view geometry problem of estimating the relative pose between two consecutive frames in the presence of radial distortion. This problem is of both theoretical and practical interests and has not been solved. We derive its parameterization and present an effective and robust generalized eigenvalue solver based on the hidden variable technique. Furthermore, we propose a nonlinear refinement scheme within the maximum likelihood criterion to produce more accurate estimates of the relative pose and radial distortion. Compared with the standard differential solutions without modeling the radial distortion, our approach can recover more geometrically correct point correspondences for a pair of radially distorted images. Moreover, our differential solution runs an order of magnitude faster than the discrete solution in terms of recovering the full camera motion. Experiment results on both synthetic and real data demonstrate the effectiveness of our model and method in dealing with the radial distortion.
Bin Fan 0002, Yuchao Dai, Zhiyuan Zhang 0002, Mingyi He
IEEE Signal Process. Lett.3
2022 Searching Dense Point Correspondences via Permutation Matrix Learning
abstract
Although 3D point cloud data has received widespread attentions as a general form of 3D signal expression, applying point clouds to the task of dense correspondence estimation between 3D shapes has not been investigated widely. Furthermore, even in the few existing 3D point cloud-based methods, an important and widely acknowledged principle,i.e. one-to-one matching, is usually ignored. In response, this paper presents a novel end-to-end learning-based method to estimate the dense correspondence of 3D point clouds, in which the problem of point matching is formulated as a zero-one assignment problem to achieve a permutation matching matrix to implement the one-to-one principle fundamentally. Note that the classical solutions of this assignment problem are always non-differentiable, which is fatal for deep learning frameworks. Thus we design a special matching module, which solves a doubly stochastic matrix at first and then projects this obtained approximate solution to the desired permutation matrix. Moreover, to guarantee end-to-end learning and the accuracy of the calculated loss, we calculate the loss from the learned permutation matrix but propagate the gradient to the doubly stochastic matrix directly which bypasses the permutation matrix during the backward propagation. Our method can be applied to both non-rigid and rigid 3D point cloud data and extensive experiments show that our method achieves state-of-the-art performance for dense correspondence learning.The code will be released.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Qi Liu 0054
IEEE Signal Process. Lett.1
2022 Learning a Task-Specific Descriptor for Robust Matching of 3D Point Clouds
abstract
Existing learning-based point feature descriptors are usually task-agnostic, which pursue describing the individual 3D point clouds as accurate as possible. However, the matching task aims at describing the corresponding points consistently across different 3D point clouds. Therefore these too accurate features may play a counterproductive role due to the inconsistent point feature representations of correspondences caused by the unpredictable noise, partiality, deformation, etc., in the local geometry. In this paper, we propose to learn a robust task-specific feature descriptor to consistently describe the correct point correspondence under interference. Born with anEncoder and aDynamicFusion module, our method EDFNet develops from two aspects. First, we augment the matchability of correspondences by utilizing their repetitive local structure. To this end, a special encoder is designed to exploit two input point clouds jointly for each point descriptor. It not only captures the local geometry of each point in the current point cloud by convolution, but also exploits the repetitive structure from paired point cloud by Transformer. Second, we propose a dynamical fusion module to jointly use different scale features. There is an inevitable struggle between robustness and discriminativeness of the single scale feature. Specifically, the small scale feature is robust since little interference exists in this small receptive field. But it is not sufficiently discriminative as there are many repetitive local structures within a point cloud. Thus the resultant descriptors will lead to many incorrect matches. In contrast, the large scale feature is more discriminative by integrating more neighborhood information. But it is easier to be disturbed since there is much more interference in the large receptive field. Compared with the conventional fusion strategy that handles multiple scale features equally, we analyze the consistency of them to judge the clean ones and perform larger aggregation weights on them during fusion. Then, a robust and discriminative feature descriptor is achieved by focusing on multiple clean scale features. Extensive evaluations validate that EDFNet learns a task-specific descriptor, which achieves state-of-the-art or comparable performance for robust matching of 3D point clouds.
Zhiyuan Zhang 0002, Yuchao Dai, Bin Fan 0002, Jiadai Sun, Mingyi He
IEEE Trans. Circuits Syst. Video Technol.1
2022 VRNet: Learning the Rectified Virtual Corresponding Points for 3D Point Cloud Registration
abstract
3D point cloud registration is fragile to outliers, which are labeled as the points without corresponding points. To handle this problem, a widely adopted strategy is to estimate the relative pose based only on some accurate correspondences, which is achieved by building correspondences on the identified inliers or by selecting reliable ones. However, these approaches are usually complicated and time-consuming. By contrast, the virtual point-based methods learn the virtual corresponding points (VCPs) for allsourcepoints uniformly without distinguishing the outliers and the inliers. Although this strategy is time-efficient, the learned VCPs usually exhibit serious collapse degeneration due to insufficient supervision and the inherent distribution limitation. In this paper, we propose to exploit the best of both worlds and present a novel robust 3D point cloud registration framework. We follow the idea of the virtual point-based methods but learn a new type of virtual points called rectified virtual corresponding points (RCPs), which are defined as the point set with the same shape as thesourceand with the same pose as thetarget. Hence, a pair of consistent point clouds,i.e.sourceand RCPs, is formed by rectifying VCPs to RCPs (VRNet), through which reliable correspondences betweensourceand RCPs can be accurately obtained. Since the relative pose betweensourceand RCPs is the same as the relative pose betweensourceandtarget, the input point clouds can be registered naturally. Specifically, we first construct the initial VCPs by using an estimated soft matching matrix to perform a weighted average on thetargetpoints. Then, we design a correction-walk module to learn an offset to rectify VCPs to RCPs, which effectively breaks the distribution limitation of VCPs. Finally, we develop a hybrid loss function to enforce the shape and geometry structure consistency of the learned RCPs and thesourceto provide sufficient supervision. Extensive experiments on several benchmark datasets demonstrate that our method achieves advanced registration performance and time-efficiency simultaneously.The code will be made public.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Bin Fan 0002, Mingyi He
IEEE Trans. Circuits Syst. Video Technol.1
2020 Deep learning based point cloud registration: an overview
abstract
Point cloud registration aims at finding a rigid transformation to align one point cloud to another one. It is a fundamental problem in computer vision and robotics, which has been widely used in various applications, such as 3D reconstruction, SLAM (simultaneous localization and mapping), and autonomous driving. Over the last decades, many researchers have devoted themselves to tackle this challenging problem. Recently, the success of deep learning in high-level vision tasks has been extended to different geometric vision tasks. Various kinds of deep learning based point cloud registration methods have been proposed to exploit different aspects of the problem. However, a comprehensive overview of these approaches is still missing. To this end, in this paper, we summarize recent progress and present a comprehensive overview for deep learning based point cloud registration. We classify the popular approaches into different categories such as, correspondences-based or correspondences-free, effective modules: feature extractor, matching, outlier rejection, and motion estimation. Furthermore, we discuss the merits and demerits in detail. We provide a systematic and compact framework towards currently proposed methods and discuss future research directions.
Zhiyuan Zhang 0002, Yuchao Dai, Jiadai Sun
Virtual Real. Intell. Hardw.1