Baoxiang Huang

dblp:214/1746 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-0380-419XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Whale identification and size estimation in satellite imagery via intelligent subtle perception
Baoxiang Huang, Milena Radenkovic 0001, Ge Chen 0002
Expert Syst. Appl.2
2026 Automatic detection of dolphin click signals based on acoustic spectrogram decoupling and fusion
Zongwei Liu, Xiaoke Liu, Yaqian Shi, Baoxiang Huang, Huan Yang 0001
Multim. Syst.6
2025 Contour and texture preservation underwater image restoration via low-rank regularizations
Guojia Hou, Weidong Zhang 0007, Baoxiang Huang, Zhenkuan Pan 0001
Expert Syst. Appl.4
2025 DCAFusion: A novel general image fusion framework based on reference image reconstruction and dual-cross attention mechanism
abstract
In this study, a novel end-to-end image fusion method , DCAFusion, is proposed. The method is based on Swin Transformer and introduces a dual cross-attention mechanism for the task of fusing infrared and visible, multi-focus and medical images. Although infrared and visible datasets contain source image pairs, they lack corresponding labels and cannot be trained by a unified supervised learning framework. To address this problem, DCAFusion designs image reconstruction blocks that generate reconstructed images as labels to guide model feature learning and provide dynamic information retention. The reconstructed images enable the full reference loss function to intervene in a supervised learning manner and participate in the computation of cross-attention scores through information mapping to more efficiently integrate complementary information between source images. In the comparison experiments of infrared and visible image fusion, DCAFusion's fusion result reaches 13.9420 and 1.0895 in the two metrics, ahead of the second-ranked 13.2547 and 0.9983, respectively. These metrics also maintain the lead in comparison experiments of other fusion tasks, which proves DCAFusion's unique advantages in fusion results.
Lixing Fang, Meng Hou, Baoxiang Huang, Ge Chen 0002, Jie Yang 0056
Inf. Sci.3
2025 CHGAFF-YOLO: A Cascade Hybrid Global Adaptive Feature Fusion Framework for Real-Time Ocean Internal Wave Detection
abstract
Internal waves (IWs) are widely present in the global ocean and play a crucial role in ocean dynamics, material transport, and climate change research. However, due to the large-scale variations and complex morphological features of internal waves objects, existing detection methods have limitations in cross scale feature extraction and information fusion. To address the aforementioned issues, this letter proposes a Cascade Hybrid Global Adaptive Feature Fusion YOLO (CHGAFF-YOLO) model for ocean internal waves detection. First, we employ the HGNetV2 structure as the backbone network of YOLO to better capture global information and extract complex features in a lightweight manner. Second, we introduce a Cascade Hybrid Cross-head Attention Module (CHCAM), which integrates in-group cascade attention with cross-head parallel self-attention mechanisms to achieve optimized multi-scale feature extraction and enhance global feature representation capability. Finally, we design a Four-Head Adaptive Feature Fusion Module (FHAFF), which dynamically fuses feature information from different scales by constructing four detection heads, further enhancing cross-scale information interaction. Extensive experimental results show that our method significantly outperforms existing approaches on both the Sentinel-1 SAR and MODIS satellite internal waves remote sensing datasets.
Xianwei Huo, He Gao, Baoxiang Huang, Ge Chen 0002
IEEE Geosci. Remote. Sens. Lett.3
2024 Intelligent Sparse2Dense Profile Reconstruction for Predicting Global Subsurface Chlorophyll Maxima
abstract
Subsurface chlorophyll maxima (SCM) is a crucial ecological indicator for marine ecosystems. Previous studies have indicated that this phenomenon is globally widespread. Although the biogeochemical Argo assimilation results have yielded positive results, the sparse data prevents them from being effectively used in oceanographic operations. Considering the dependence of ocean parameter, a deep learning model termed AT-GRU based on gated recurrent units is proposed. By incorporating an attention mechanism, the model can effectively address missing data in the biogeochemical-Argo (BGC-Argo) profiles, achieving the transition of data from sparse to dense (Sparse2Dense) and improving the accuracy of estimating subsurface chlorophyll-a (Chla) concentration. Specifically, the dataset of satellite remote sensing data and the associated BGC-Argo profiles is first established. AT-GRU is employed to reconstruct Chla concentration profiles from 1 to 300 m, utilizing several sources of ocean surface data. Next, an in-depth investigation is conducted to determine the characteristics of SCM. The objective is to enable wider research on SCM by analyzing vertical Chla profiles in four geographical locations. Finally, the general improvement of skill performance metrics, with R-squared reaching 0.84, demonstrates the feasibility of the proposed methodology through extensive experiments. In addition, we apply AT-GRU to global surface satellite data from January 2023 and compare the results with numerical modeling data to further validate the performance. This study presents promising opportunities for leveraging artificial intelligence in subsurface oceanic phenomena with the idea of Sparse2Dense and holds significant implications for the field of marine ecology.
Yongjun Yu, Baoxiang Huang, Milena Radenkovic 0001, Ge Chen 0002
IEEE Trans. Geosci. Remote. Sens.2
2024 Global Oceanic Mesoscale Eddies Trajectories Prediction With Knowledge-Fused Neural Network
abstract
Efficient eddy trajectory prediction driven by multi-information fusion can facilitate the scientific research of oceanography, while the complicated dynamics mechanism makes this issue challenging. Benefiting from ocean observing technology, the eddy trajectory dataset can be qualified for data-intensive research paradigms. In this paper, the dynamics mechanism is used to inspire the design idea of the eddy trajectory prediction neural network (termed EddyTPNet) and is also transformed into prior knowledge to guide the learning process. This study is among the first to implement eddy trajectory prediction with physics informed neural network. First, an in-depth analysis of the kinematic characteristics indicates that the longitude and latitude of the trajectory should be decoupled; Second, the directional dispersion prior knowledge of global eddy propagation is embedded into the decoder of the EddyTPNet to improve the performance; Finally, EddyTPNet predicts global eddy trajectories through pre-training and adapts to complex local regions via model transfer. Extensive experimental results demonstrate that EddyTPNet can reliably forecast the motion of eddies for the next 7 days, ensuring a low daily mean geodetic error. This exploratory study provides valuable insights into solving the prediction problem of ocean phenomena by using knowledge-based time series neural networks.
Baoxiang Huang, Ge Chen 0002, Linyao Ge, Milena Radenkovic 0001, Guojia Hou
IEEE Trans. Geosci. Remote. Sens.2
2023 Global Oceanic Eddy-Front Associations From Synergetic Remote Sensing Data by Deep Learning
abstract
Recently, fronts (eddies) at the margins of eddies (fronts) have been discovered by observing sea surface temperature (SST) and sea level anomaly (SLA) data. They can both induce strong vertical motions and submesoscale processes, and are important for the vertical exchange of ocean mass and energy as well as ocean ecological processes. However, it raises an important challenge about the global spatiotemporal distribution of eddy-induced fronts and frontal eddies. This letter proposes a deep learning (DL) approach, dubbed eddy-front association detection network (EFADN), that is appropriate for mining eddy-front associations (EFAs) to extract the features of eddy-induced fronts (anticyclonic and cyclonic eddy-induced fronts) and frontal eddies [frontal anticyclonic eddies (AEs) and frontal cyclonic eddies (CEs)] from SLA and SST satellite data during 2006–2015 in the global ocean. The EFADN model integrates encoder-decoder and attention structures. The introduced spatial attention (SA) module in attention structure utilizes large-scale convolutional kernels to extract spatial information, which enlarges the receptive field to enhance the recognition of topological structures between eddies and fronts, improving the ability of EFA detection. The results of comparative experiments demonstrate that EFADN surpasses the state-of-the-art (SOTA) eddy detection model. Ablation studies underscore the crucial importance of all modules within EFADN for achieving accurate detection of EFAs. Moreover, the spatiotemporal distribution characteristics of eddy-induced fronts and frontal eddies are displayed. They are widely dispersed in the western boundary current (WBC) and Antarctic Circumpolar Current (ACC) regions, and they are active in the boreal summer while weak in the austral summer.
Fenglin Tian, Shuang Long, Baoxiang Huang, Ge Chen 0002
IEEE Geosci. Remote. Sens. Lett.4
2023 Oceanic Eddy Identification Using Pyramid Split Attention U-Net With Remote Sensing Imagery
abstract
Oceanic eddy is the ubiquitous ocean flow phenomenon, which has been the key factor in the transportation of ocean energy and materials. Consequently, oceanographic understanding can be enhanced by the intelligent identification of eddy. State-of-the-art deep learning technologies are gradually improving identification methods. This letter proposes the pyramid split attention (PSA) eddy detection U-Net architecture (PSA-EDUNet) that targets oceanic eddy identification from ocean remote sensing imagery. As for the PSA-EDUNet, its inspiration comes from U-Net, which contains encoder and decoder parts, making the integration of inferior and senior features efficient and ensuring the feature information will not be lost in large quantities through nonlinear connection mode. Meanwhile, the PAS module is introduced to enhance feature extraction. In terms of the fusion data, the sea surface feature is the main criterion of eddy identification, including sea surface temperature (SST) and sea level anomaly (SLA). The experiments are implemented on the Kuroshio Extension (KE) and the South Atlantic regions, the results demonstrate that the proposed method can outperform other methods, especially for eddy edges and small-scale eddies.
Baoxiang Huang, Jie Yang 0056, Milena Radenkovic 0001, Ge Chen 0002
IEEE Geosci. Remote. Sens. Lett.2
2023 Medium-Range Trajectory Prediction Network Compliant to Physical Constraint for Oceanic Eddy
abstract
Predicting the trajectory of ocean eddies can promote the understanding of the transport of matter and energy in the ocean. However, accurately and rapidly predicting the trajectory of eddies poses a significant challenge due to their intricate nonlinear motion within a physical environment. Regrettably, existing data-driven methods primarily focus on the migration and combination of models, as well as the fusion processing of diverse observational data on oceanic eddies. These ways often overlook the crucial aspect of modeling the underlying motion mechanism of the eddies. We believe that the expeditious and precise prediction of eddies is closely intertwined with the physical mechanism and historical time series. Consequently, a medium-range eddy trajectory prediction neural network (ETPNet) compliant with the physical constraint is proposed, which embeds the physical regulation, intrinsic relations, and mutual interactions into the network via constraints. Then, a novel variant of the long short-term memory (LSTM) cell is designed to enhance the dynamic interaction and representation ability of the features, constraints, and knowledge. Finally, a geographically informed comprehensive loss function for marine tasks is formulated, namely mean absolute geodetic error (MAGE), which optimizes the network in Euclidean and sphere space. The proposed network is evaluated by predicting the future seven days trajectory of anticyclone eddies in the$15^{\circ }\text{N}$to$40^{\circ }\text{N}$. The extensive experiments and evaluations demonstrate that the proposed network guided by the comprehensive loss function can implement a state-of-the-art performance. The code is available athttps://github.com/AI4Ocean/ETPNet
Linyao Ge, Baoxiang Huang, Ge Chen 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 Semantic-aware multi-task learning for image aesthetic quality assessment
abstract
In recent years, image aesthetic quality assessment has attracted considerable attention due to the massive growth of digital images in social platforms and the Internet. However, automatically assessing aesthetic quality of an image is a challenging task, because image aesthetic is affected by various factors, and the criteria for judging the aesthetic of images with diverse semantic information are different. To this end, a Semantic-Aware Multi-task convolution neural network (SAM-CNN) for evaluating image aesthetic quality is proposed in this paper. The network can fuse intermediate features of different layers at different scales in CNN to obtain a more comprehensive and accurate aesthetic expression, under the joint supervision of image aesthetic quality assessment task and semantic classification task in a multi-task learning manner. Besides, by applying the attention mechanism, semantic information with a large receptive field extracted from deep layers is utilised to guide the network to focus on the key parts of features to be fused, to improve the effectiveness of feature fusion. Experimental results on the AVA dataset and Photo.net dataset demonstrate the effectiveness and superiority of the proposed SAM-CNN.
Weiliang Yan, Huan Yang 0001, Baoxiang Huang, Zhenkuan Pan 0001
Connect. Sci.4
2022 Vertical Structure-Based Classification of Oceanic Eddy Using 3-D Convolutional Neural Network
abstract
The eddy identification is an important part of human cognition of the ocean. Significant achievements have been made by using sea level anomaly (SLA) data observed by the altimeter. However, the abundant eddies, which do not cause sea surface characteristic anomalies, cannot be identified. In this study, the eddy subsurface vertical structure-oriented 3-D neural network is developed to classify the oceanic eddies. This study is among the first that explores the ability of deep learning in eddy identification with vertical structure. First, the purified eddy profiles dataset is constructed based on the fact that the structure derived from vertical profiles is highly correlated with the sea surface topography detected by altimetry. Then, the eddy vertical structure-oriented 3-D neural network based on the residual network (ResNet) is constructed, which can classify the eddies as anticyclonic eddies (AEs), cyclonic eddies (CEs), and noneddies (NEs) effectively. Furthermore, the spatial and temporal features can be combined in the proposed network as external factors. Meanwhile, through 3-D convolutions and 3-D pooling, the proposed network is capable of modeling 3-D eddy data and can be extended to the deeper network structure. Finally, the classification experiments are implemented to validate the performance of the proposed methodology. The most striking result emerging from experiments is that the proposed method can expand the capacity of eddy identification by using vertical profiles as calibrated by altimetry with competitive classification performance. Together these results provide important insights into the application of artificial intelligence in oceanic eddy research.
Baoxiang Huang, Linyao Ge, Ge Chen 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 A no-Reference Stereoscopic Image Quality Assessment Network Based on Binocular Interaction and Fusion Mechanisms
abstract
In contemporary society full of stereoscopic images, how to assess visual quality of 3D images has attracted an increasing attention in field of Stereoscopic Image Quality Assessment (SIQA). Compared with 2D-IQA, SIQA is more challenging because some complicated features of Human Visual System (HVS), such as binocular interaction and binocular fusion, must be considered. In this paper, considering both binocular interaction and fusion mechanisms of the HVS, a hierarchical no-reference stereoscopic image quality assessment network (StereoIF-Net) is proposed to simulate the whole quality perception of 3D visual signals in human cortex, including two key modules: BIM and BFM. In particular, Binocular Interaction Modules (BIMs) are constructed to simulate binocular interaction in V2-V5 visual cortex regions, in which a novel cross convolution is designed to explore the interaction details in each region. In the BIMs, different output channel numbers are designed to imitate various receptive fields in V2-V5. Furthermore, a Binocular Fusion Module (BFM) with automatic learned weights is proposed to model binocular fusion of the HVS in higher cortex layers. The verification experiments are conducted on the LIVE 3D, IVC and Waterloo-IVC SIQA databases and three indices including PLCC, SROCC and RMSE are employed to evaluate the assessment consistency between StereoIF-Net and the HVS. The proposed StereoIF-Net achieves almost the best results compared with advanced SIQA methods. Specifically, the metric values on LIVE 3D, IVC and WIVC-I are the best, and are the second-best on the WIVC-II.
Jianwei Si, Baoxiang Huang, Huan Yang 0001, Weisi Lin, Zhenkuan Pan 0001
IEEE Trans. Image Process.2
2021 A full-reference stereoscopic image quality assessment index based on stable aggregation of monocular and binocular visual features
abstract
Abstract In stereoscopic image quality assessment, human visual system has been universally taken into account to detect perceptual characteristics. A novel full‐reference stereoscopic image assessment metric by considering both monocular and binocular visual features of human visual system is proposed. In particular, a new region segmentation algorithm is firstly proposed to divide 3D images into occluded and non‐occluded regions. The just noticeable difference model is employed on the occluded regions to formulate the monocular vision, while the binocular just noticeable difference model is applied to the non‐occluded regions to reveal the binocular vision of the human visual system. In the proposed region segmentation, disparity information and Euclidean distance between stereo pairs are both adopted to solve the unstable segmentation problem of traditional methods. A new pooling strategy based on global edge features is then presented to aggregate the just noticeable difference and binocular just noticeable difference evaluation maps. In addition, some local image features as supplementary of just noticeable difference to describe visual characteristics of the human visual system are also extracted. Finally, an overall quality score is calculated based on the above‐mentioned features to measure the visual quality of distorted stereo pairs. Experimental results show that the proposed metric achieves high consistency with the human visual system, and outperforms state‐of‐the‐art algorithms on stereoscopic image quality assessment.
Jianwei Si, Huan Yang 0001, Baoxiang Huang, Zhenkuan Pan 0001, Honglei Su
IET Image Process.3
2021 Nonlocal graph theory based transductive learning for hyperspectral image classification
Baoxiang Huang, Linyao Ge, Ge Chen 0002, Milena Radenkovic 0001, Jinming Duan 0001, Zhenkuan Pan 0001
Pattern Recognit.1
2020 Efficient image structural similarity quality assessment method using image regularised feature
abstract
Image regularised features play a critical role in image processing domain, by integrating regularised feature and structural similarity, a new full‐reference image assessment method (IRF_SSIM) is proposed in this study. As well known, the gradient operator always be used to capture the edge information of the image, while the total variational regularised features can be adopted to calculate the detailed change information of image contrast and texture, as well as noise removal and edge retention. Therefore, the IRF_SSIM method extends the gradient features into the image regularised features to measure the structural changes in the image. In addition, image quality is also affected by variations of luminance and contrast. For a more comprehensive image quality assessment, the IRF_SSIM method considers the changes in structure, luminance and contrast simultaneously. In other words, the total image quality is estimated by structural similarity calculated by integrating the effects of image structure, luminance and contrast changes. Comparing with the representative methods, the experimental results illustrate that the IRF_SSIM method is highly consistent with the subjective assessment results.
Baoxiang Huang, Huan Yang 0001, Guojia Hou, Jinming Duan 0001
IET Image Process.2
2020 A novel dark channel prior guided variational framework for underwater image restoration
Guojia Hou, Jingming Li, Guodong Wang 0001, Huan Yang 0001, Baoxiang Huang, Zhenkuan Pan 0001
J. Vis. Commun. Image Represent.5
2020 Variational level set method for image segmentation with simplex constraint of landmarks
Baoxiang Huang, Zhenkuan Pan 0001, Huan Yang 0001, Li Bai 0001
Signal Process. Image Commun.1
2018 Hue preserving-based approach for underwater colour image enhancement
abstract
In this study, a novel underwater colour image enhancement approach based on hue preserving is presented by combining hue–saturation–intensity (HSI) and HS–value (HSV) colour models. In this study, the proposed wavelet‐domain filtering (WDF) and constrained histogram stretching (CHS) algorithms are operated on HSI and HSV colour models, respectively. The degraded image is first converted from red–green–blue colour model into the HSI colour model, wherein the hue component H is preserved and WDF algorithm is executed on the S and I components. Similarly, the image is further converted into the HSV colour model, wherein H component is kept invariant as well and CHS algorithm is applied on the S and V components. The authors' key contribution is that the H preserving method can improve image quality in terms of contrast, colour rendition, non‐uniform illumination, and denoising. In addition, experimental results show that the proposed approach outperforms several other state‐of‐the‐art algorithms.
Guojia Hou, Zhenkuan Pan 0001, Baoxiang Huang, Guodong Wang 0001, Xin Luan
IET Image Process.3
2016 Unsupervised color texture segmentation using active contour model and oscillating information
abstract
It is common that textures occur in real-word color image, moreover, textures could cause difficulties in image segmentation. For the purpose of solving those difficulties, we put forward a new model. In this model we only need the structural and oscillating components’ information of the real color image. This model is based on the VO model, MTV and active contour models. We will use the fast Split Bregman algorithm to solve this model. The results of our model is mentioned in numerical experiments.
Guodong Wang 0001, Zhenkuan Pan 0001, Baoxiang Huang
ICMV4