Huaibo Song

dblp:55/343 · DBLP profile ↗
← Back
23ranked-venue papers
0as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive fine-grained feature enhancement network for non-contact calf diarrhea detection
Liuru Pu, Haoyu Kang, Xiangfeng Kong, Xiaopeng Du, Huaibo Song
Eng. Appl. Artif. Intell.6
2026 HDFENet: High-frequency and dual-directional feature enhancement network for maize tassels counting
Liuru Pu, Haowen Pan, Huaibo Song, Bo Jiang 0017
Expert Syst. Appl.5
2026 CSAFNet: Cross-modal spatial alignment and fusion network for RGB-T crowd counting
Liuru Pu, Huaibo Song, Bo Jiang 0017
Pattern Recognit.3
2025 Dual-decoupling inter-correction multitemporal framework for high-, medium-, and low-resolution optical remote sensing image reconstruction
Changqing Huang, Yonghua Jiang 0001, Jingyin Wang, Guo Zhang 0001, Huaibo Song, Xinghua Li 0002
Appl. Intell.6
2025 Joint depth-segmentation learning with segment priors for non-contact seedling height and stem thickness estimation
abstract
To achieve precise and rapid computation of seedling height and stem diameter — key phenotypic traits for monitoring seedling growth and selecting superior varieties — this study proposes a SAM-Integrated Adaptive Fusion Depth Network (SAFD-Net). SAFD-Net integrates segmentation masks generated by Segment Anything Model (SAM) with an Adaptive Prior Extraction (APE) module to produce priors focused on individual seedling characteristics, and it fuses these priors with deep features through an Adaptive Attention Fusion (AAF) module. A Local Depth Generation (LDG) module refines depth details to improve estimation accuracy, and an Adaptive Multi-scale Fusion (AMF) module merges LDG outputs at different scales to produce high-precision depth maps. From these maps, seedling region depth, pixel height, and pixel stem diameter are extracted to compute actual seedling height and stem diameter. Comparisons with various depth estimation networks demonstrate that SAFD-Net outperforms existing models in both depth estimation and seedling measurement. Experimental evaluations on seedlings from three crops with distinct phenotypic characteristics further show that the method maintains high accuracy under varying shooting distances, lighting conditions, multiple targets, and tilt angles, offering a novel approach for phenotypic monitoring during seedling cultivation. Code is released at https://github.com/Songlei7664/SAFD-Net .
Lei Song 0010, Bo Jiang 0017, Huaibo Song
Eng. Appl. Artif. Intell.3
2025 One-stage keypoint detection network for end-to-end cow body measurement
Guangyuan Yang, Yongliang Qiao, Hongxing Deng, Qinfeng Shi, Huaibo Song
Eng. Appl. Artif. Intell.5
2025 Multi-Target spraying behavior detection based on an improved YOLOv8n and ST-GCN model with Interactive of video scenes
Liuru Pu, Zhixin Hua, Mengxuan Han, Huaibo Song
Expert Syst. Appl.5
2025 Local plane estimation network with multi-scale fusion for efficient monocular depth estimation
Lei Song 0010, Bo Jiang 0017, Huaibo Song
Expert Syst. Appl.3
2025 Multi-camera multi-cow tracking under non-overhead views
Xingshi Xu, Hongxing Deng, Yuying Shang, Shujin Zhang, Diyi Chen, Huaibo Song
Expert Syst. Appl.7
2025 Adaptive Clustering and Frequency Division Network for Efficient Monocular Depth Estimation
abstract
Monocular depth estimation infers the relative depth of objects by analyzing visual cues in images, ultimately enhancing the comprehension of complex scenes in computer vision systems. Although existing Transformer architectures effectively capture long-range visual dependencies, two significant challenges persist: (a) insufficient integration of spatial context leads to inconsistent depth estimation, particularly under varying perspectives or lighting conditions; (b) the model’s difficulty in capturing global features hinders the parsing of subtle object differences, causing confusion in depth information and reducing sensitivity to variations in object distance and scene layout. To address the aforementioned issues, an Adaptive Clustering Mechanism (ACM) module coupled with a Deformable Frequency Division Fusion (DFDF) module was introduced. Specifically, the ACM module refines and adjusts features via cosine similarity, thereby enhancing cluster center similarity and stabilizing depth estimation. The DFDF module leverages frequency decomposition to extract differential features between objects, enhancing high-frequency information to improve the discrimination of subtle features. Integrating these components, the Frequency Division Adaptive Clustering Enhancement (FDACE) module emerges as the decoder’s core within the Adaptive Clustering and Frequency Division Network (ACFD-Net), facilitating both the precise generation of depth information and the efficient recovery of spatial resolution. Furthermore, we present a progressive depth estimation strategy that seamlessly integrates non-gradient output features from FDACE modules across various scales, and conducts independent optimizations, merging multi-scale information with localized details, and progressively calibrating depth estimates to enhance congruence with actual scenes. The ACM and DFDF modules concentrate on pivotal features, selectively enhancing high-frequency information, thereby minimizing redundant computations and optimizing resource allocation, which significantly boosts computational efficiency. Experimental results demonstrate that ACFD-Net significantly improves both the accuracy and efficiency of depth estimation.Code is released athttps://github.com/Songlei7664/ACFD-Net.
Lei Song 0010, Huaibo Song, Bo Jiang 0017
IEEE Trans. Circuits Syst. Video Technol.2
2024 Collaborative dual-harmonization reconstruction network for large-ratio cloud occlusion missing information in high-resolution remote sensing images
Yonghua Jiang 0001, Guo Zhang 0001, Huaibo Song, Xinghua Li 0002
Eng. Appl. Artif. Intell.5
2024 Plant leaf disease identification by parameter-efficient transformer with adapter
Xingshi Xu, Guangyuan Yang, Yuying Shang, Zhixin Hua, Huaibo Song
Eng. Appl. Artif. Intell.7
2024 Automated measurement of beef cattle body size via key point detection and monocular depth estimation
Yuchen Wen, Shujin Zhang, Xingshi Xu, Baoling Ma, Huaibo Song
Expert Syst. Appl.6
2024 E-YOLO: Recognition of estrus cow based on improved YOLOv8n model
Zhixin Hua, Yuchen Wen, Shujin Zhang, Xingshi Xu, Huaibo Song
Expert Syst. Appl.6
2024 IIMT-net: Poly-1 weights balanced multi-task network for semantic segmentation and depth estimation using interactive information
Mengfei He, Zhiyou Yang, Guangben Zhang, Yan Long 0005, Huaibo Song
Image Vis. Comput.5
2024 Global and Local Dual Fusion Network for Large-Ratio Cloud Occlusion Missing Information Reconstruction of a High-Resolution Remote Sensing Image
abstract
Large-ratio cloud occlusion significantly hampers the utilization of high-resolution remote sensing imagery. The existing reconstruction methods (1) overlook the problem of reconstructed and composite images sharing high-and low-level semantic and visual attributes in non-reconstructed regions, exacerbating the pronounced boundary effects; (2) neglect appearance discrepancies between reconstructed and non-reconstructed regions, leading to spectral degradation, and texture loss; and (3) overlook the problem of reconstructing large-ratio missing information. To address these issues, a global and local dual fusion network is proposed in this study for large-ratio cloud occlusion removal in high-resolution remote sensing images. The global foreground–background aware attention module tackles shared high-level semantic features, whereas the local visual feature enhancement module addresses appearance differences. The global and local dual fusion network combines the Sobel and reconstruction loss functions for effective reconstruction by employing a two-stage fusion strategy. Compared to the classical recurrent feature reasoning network, spatiotemporal generator network, spatial-temporal-spectral convolutional neural network, and bishift network, the proposed model demonstrates superior quantitative and visual reconstruction outcomes for the 40%, 50%, and 70% missing ratios of Gaofen-1 (2 m).
Yonghua Jiang 0001, Jingyin Wang, Guo Zhang 0001, Huaibo Song, Jun Yang 0012, Xinghua Li 0002
IEEE Geosci. Remote. Sens. Lett.6
2024 CenterNet-LW-SE net: integrating lightweight CenterNet and channel attention mechanism for the detection of Camellia oleifera fruits
Hongxing Deng, Baoling Ma, Huaibo Song
Multim. Tools Appl.6
2023 Extracting cow point clouds from multi-view RGB images with an improved YOLACT++ instance segmentation
Guangyuan Yang, Shujin Zhang, Yuchen Wen, Xingshi Xu, Huaibo Song
Expert Syst. Appl.6
2020 Combining an information-maximization-based attention mechanism and illumination invariance theory for the recognition of green apples in natural scenes
Sashuang Sun, Mei Jiang, Dongjian He, Yan Long 0005, Huaibo Song, Zhenjiang Zhou
Multim. Tools Appl.6
2019 Combining SUN-based visual attention model and saliency contour detection algorithm for apple image segmentation
Dongjian He, Huaibo Song, Hongting Xiong
Multim. Tools Appl.3
2017 Extracting the symmetry axes of partially occluded single apples in natural scene using convex hull theory and shape context algorithm
Leilei Niu, Weicong Zhou, Dongjian He, Haihui Zhang, Huaibo Song
Multim. Tools Appl.6
2016 Recognition and localization of occluded apples using K-means clustering algorithm and convex hull theory: a comparison
Huaibo Song, Zhihui Tie, Weiyuan Zhang, Dongjian He
Multim. Tools Appl.2
2015 Bottom-up saliency estimation using sparse representation and structural redundancy reduction
Yan Long 0005, Dongjian He, Huaibo Song
Multim. Tools Appl.3