Haijiang Zhu

dblp:23/6663 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 A multi-mode structured light 3D imaging system with multi-source information fusion for underwater pipeline detection
Qinghan Hu, Haijiang Zhu, Zhengqiang Fan
Knowl. Based Syst.2
2025 F2Unet: F-Shaped U-Net Architecture for Medical Image Segmentation Combining Fourier Transforms
Feiyue Qi, Yiwen Dai, Kaiye Xu, Zhuohang Wang, Haijiang Zhu
ICANN (2)6
2025 Floorplan-SLAM: A Real-Time, High-Accuracy, and Long-Term Multi-Session Point-Plane SLAM for Efficient Floorplan Reconstruction
abstract
Floorplan reconstruction provides structural priors essential for reliable indoor robot navigation and high-level scene understanding. However, existing approaches either require time-consuming offline processing with a complete map, or rely on expensive sensors and substantial computational resources. To address the problems, we propose FloorplanSLAM, which incorporates floorplan reconstruction tightly into a multi-session SLAM system by seamlessly interacting with plane extraction, pose estimation, back-end optimization, and loop & map merging, achieving real-time, high-accuracy, and long-term floorplan reconstruction using only a stereo camera. Specifically, we present a robust plane extraction algorithm that operates in a compact plane parameter space and leverages spatially complementary features to accurately detect planar structures, even in weakly textured scenes. Furthermore, we propose a floorplan reconstruction module tightly coupled with the SLAM system, which uses continuously optimized plane landmarks and poses to formulate and solve a novel optimization problem, thereby enabling real-time and high-accuracy floorplan reconstruction. Note that by leveraging the map merging capability of multi-session SLAM, our method supports long-term floorplan reconstruction across multiple sessions without redundant data collection. Experiments on the VECtor and the self-collected datasets indicate that Floorplan-SLAM significantly outperforms state-of-the-art methods in terms of plane extraction robustness, pose estimation accuracy, and floorplan reconstruction fidelity and speed, achieving real-time performance at 25–45 FPS without GPU acceleration, which reduces the floorplan reconstruction time for a 1000 m2scene from 16 hours and 44 minutes to just 9.4 minutes.
Haolin Wang 0005, Zeren Lv, Hao Wei 0008, Haijiang Zhu, Yihong Wu 0002
IROS4
2025 Maximum Clique-Based Floorplan Association for Robust Multi-Session Stereo SLAM in Challenging Indoor Environments
abstract
Existing multi-session visual simultaneous localization and mapping (SLAM) systems struggle severely to achieve robust localization and map merging under extreme viewpoint and illumination variations, particularly when handling completely opposite viewpoints and drastic day-night lighting changes. These challenges stem largely from the limited viewpoint/illumination invariance of conventional low-level visual features and their inability to capture a global structural context. In this paper, we make the critical observation that a life-long floorplan not only encodes rich geometric and semantic information—serving as a robust high-level structural representation—but is also inherently more robust to severe viewpoint and illumination variations than purely visual data. Building on this insight, we propose a novel hierarchical framework for multi-session SLAM that integrates a floorplan-based map as a global feature to achieve robust indoor localization and map merging under drastic viewpoint and illumination shifts. In particular, we innovatively formulate floorplan association as a maximum clique problem augmented with trajectory data to achieve robust floorplan-level global localization. We further introduce a novel coarse-to-fine localization and map merging strategy that seamlessly integrates floorplan alignment, multistage point cloud registration, and feature matching, fully leveraging the macro-level stability of global features and the micro-level precision of local features to achieve keyframe-level fine localization. Extensive experiments on both public and self-collected datasets demonstrate that our method consistently outperforms state-of-the-art (SOTA) approaches reliant solely on low-level visual or geometric features. Crucially, it delivers superior accuracy and robustness even in the face of completely opposite viewpoints and extreme day–night illumination changes. This work underscores the promise of fusing macro-level floorplan representations with conventional SLAM frameworks to advance long-term, robust indoor localization and map merging under the most challenging conditions.
Haolin Wang 0005, Hao Wei 0008, Zeren Lv, Haijiang Zhu, Yihong Wu 0002
IROS4
2025 Touching the limit of Rolling Multilayer Perceptron for efficient two-dimensional medical image segmentation
Haijiang Zhu
Eng. Appl. Artif. Intell.2
2025 NPMFF-Net: A training-free unified framework for point cloud classification and segmentation
Hualong Zeng, Haijiang Zhu, Huaiyuan Yu
Knowl. Based Syst.2
2024 Rolling-Unet: Revitalizing MLP's Ability to Efficiently Extract Long-Distance Dependencies for Medical Image Segmentation
abstract
Medical image segmentation methods based on deep learning network are mainly divided into CNN and Transformer. However, CNN struggles to capture long-distance dependencies, while Transformer suffers from high computational complexity and poor local feature learning. To efficiently extract and fuse local features and long-range dependencies, this paper proposes Rolling-Unet, which is a CNN model combined with MLP. Specifically, we propose the core R-MLP module, which is responsible for learning the long-distance dependency in a single direction of the whole image. By controlling and combining R-MLP modules in different directions, OR-MLP and DOR-MLP modules are formed to capture long-distance dependencies in multiple directions. Further, Lo2 block is proposed to encode both local context information and long-distance dependencies without excessive computational burden. Lo2 block has the same parameter size and computational complexity as a 3×3 convolution. The experimental results on four public datasets show that Rolling-Unet achieves superior performance compared to the state-of-the-art methods.
Haijiang Zhu, Huaiyuan Yu
AAAI2
2024 Mask-TS Net: Mask Temperature Scaling Uncertainty Calibration for Polyp Segmentation
Yudian Zhang, Kaiye Xu, Haijiang Zhu
ICPR (4)4
2024 ODC-SA Net: Orthogonal Direction Enhancement and Scale Aware Network for Polyp Segmentation
Yudian Zhang, Kaiye Xu, Haijiang Zhu
PRCV (14)4
2024 Novel camera self-calibration method with clustering prior and nonlinear optimization from an image sequence
Xiaohui Jiang, Haijiang Zhu, Ning An 0002, Binjian Xie, Hao Wei 0008, Fulin Tang, Yihong Wu 0002
Multim. Tools Appl.2
2024 MSAMS-Net: accurate lung lesion segmentation from COVID-19 CT images
Haijiang Zhu, Xiaoyu Gao
Multim. Tools Appl.2
2023 PoinLin-Net: Point Cloud Completion Network Based on Geometric Feature Extraction and Linformer Structure
Dejie Li, Kejin Huang, Yinchu Wang, Haijiang Zhu
ICANN (1)4
2023 PST-Net: Point Cloud Completion Network Based on Local Geometric Feature Reuse and Neighboring Recovery with Taylor Approximation
abstract
Transformer has recently been introduced into point cloud completion and achieved inspirational performance on 3D point cloud generation. However, the low-level local geometric features are ignored in the existing feature extraction network, and this leads to the loss of geometry in the recovery results. Meanwhile, FoldingNet models that estimate neighborhood information from predicted center points cannot effectively recover geometric information. In this paper, we propose a point cloud completion network based on local geometric feature reuse and neighboring recovery with Taylor approximation (PST-Net). Specifically, a feature extraction network named Skip-DGCNN is constructed to integrate local and global geometric features to reduce the geometry loss during the feature extraction. In addition, we propose a computational model through Taylor approximation to recover the geometry information in the neighborhood of the prediction center. Moreover, we design the TSMB module corresponding to Taylor's approximation to maintain the end-to-end training mode. The proposed method is extensively evaluated and compared with previous methods on three datasets including PCN, ShapeNet-55 and ShapeNet-34. The proposed model outperforms the state-of-the-art (SOTA) and PoinTr on ShapeNet-55 and ShapeNet-34. The complexity analysis on the PCN dataset shows that the number of FLOPs of our approach is 60.79% lower than that of the SOTA. Visual comparisons demonstrate that the proposed method can effectively and accurately complete the geometry of missing parts.
Yinchu Wang, Haijiang Zhu, Guanghui Wang 0001
IJCNN2
2022 ARB U-Net: An Improved Neural Network for Suprapatellar Bursa Effusion Ultrasound Image Segmentation
Le Mao, Haijiang Zhu, Xiaoyu Gao
ICANN (3)5
2021 CNNapsule: A Lightweight Network with Fusion Features for Monocular Depth Estimation
Yinchu Wang, Haijiang Zhu, Mengze Liu
ICANN (1)2
2019 Homography Estimation Based on Error Elliptical Distribution
abstract
How to estimate accurately the homography is always a challenging problem in computer vision. In the reported literature, the measurement error of the image points is usually assumed to obey isotropic Gaussian distribution. However, real data very seldom follows this assumption. This paper proposes an estimation of homography under the assumption of image point errors following elliptical distribution, which is more coincident with real data. In the proposed method, the adaptive-scale elliptical residual kernel consensus (ASERKC) robust estimator is used to filter out inliers which are utilized to compute homography. Then, the elliptical weighted L-M (EW L-M) algorithm is optimized the homography. The experimental results show that the proposed method may present a more accurate homography. Especially when we applied it to incremental structure-from-motion (SFM), we find that the exact homography matrix is useful to select a better initial image pairs which can help obtain a more complete 3D points cloud.
Lulu Mao, Haijiang Zhu, Fuqing Duan
ICASSP2
2019 3D reconstruction for ultrasonic C-scan images of tissue-mimicking phantom based on an improved K-nearest neighbor filtering
Haijiang Zhu, Longbiao He, Guanghui Wang 0001
Multim. Tools Appl.1
2018 Improved graph-cut segmentation for ultrasound liver cyst image
Haijiang Zhu, Zhanhong Zhuang, Xuejing Wang, Wenhua Xu
Multim. Tools Appl.1
2017 Segmentation of liver cyst in ultrasound image based on adaptive threshold algorithm and particle swarm optimization
Haijiang Zhu, Zhanhong Zhuang, Fan Zhang 0007, Xuejing Wang
Multim. Tools Appl.1
2016 Estimation of fisheye camera external parameter based on second-order cone programming
abstract
Although second‐order cone programming (SOCP) has been applied to optimise camera parameters in computer vision, it is occasionally been used to refine fisheye camera external parameters as well. This study presents a fisheye camera external parameter estimation based on SOCP in convex optimisation. The homography constraint between two spherical images are first exploited to derive an equation with respect to a given error threshold. Then, the fisheye camera external estimation is transformed into an SOCP optimisation problem through reformulating the parameter estimation equation. The SOCP method has been implemented in Matlab and the optimisation toolbox has been made publicly available. The fisheye camera external parameter optimisation method has been validated by some experiments with synthetic and real data. Comparison experiments between the proposed method and other methods in the literature are also carried out, and the results show that the SOCP method is better for the corrected images.
Haijiang Zhu, Fan Zhang 0007, Jing Wang 0016, Xuejing Wang
IET Comput. Vis.1
2016 Improved maximally stable extremal regions based method for the segmentation of ultrasonic liver images
Haijiang Zhu, Junhui Sheng, Fan Zhang 0007, Jing Wang 0016
Multim. Tools Appl.1
2014 Approximate model of fisheye camera based on the optical refraction
Haijiang Zhu, Xuejing Wang
Multim. Tools Appl.1
2013 Using vanishing points to estimate parameters of fisheye camera
abstract
This study presents an approach for estimating the fisheye camera parameters using three vanishing points corresponding to three sets of mutually orthogonal parallel lines in one single image. The authors first derive three constraint equations on the elements of the rotation matrix in proportion to the coordinates of the vanishing points. From these constraints, the rotation matrix is calculated under the assumption of the image centre known. The experimental results with synthetic images and real fisheye images validate this method. In contrast to the existing methods, the authors method needs less image information and does not know the three‐dimensional reference point coordinates.
Haijiang Zhu, Xuejing Wang
IET Comput. Vis.1
2012 Estimating fisheye camera parameters from homography
Haijiang Zhu, Shigang Li 0001
Sci. China Inf. Sci.1
2005 Camera calibration with moving one-dimensional objects
Fuchao Wu, Zhanyi Hu, Haijiang Zhu
Pattern Recognit.3
2004 Camera Calibration from the Quasi-affine Invariance of Two Parallel Circles
Yihong Wu 0002, Haijiang Zhu, Zhanyi Hu, Fuchao Wu
ECCV (1)2