Hua Li 0012

dblp:80/6898-12 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-0740-0691ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DiffLLFace: Learning Alternate Illumination-Diffusion Adaptation for Low-Light Face Super-Resolution and Beyond
abstract
Facial image acquisition under constrained illumination and with limited-resolution imaging devices often results in coupled photometric and geometric degradations, manifesting as low-light and low-resolution (LLR) conditions. Prevailing research predominantly follows fragmented optimization paradigms that address low-light image enhancement (LLIE) and face super-resolution (FSR) as isolated tasks. This approach overlooks the compound nature of the degradations, thereby significantly limiting their applicability in practical scenarios. To bridge this gap, we present DiffLLFace, a unified framework that harnesses diffusive generative capabilities with illumination-aware trajectories to achieve robust FSR from LLR observations. The core of our method lies in its alternate illumination-diffusion adaptation, which operates throughout the generation process. This mechanism not only captures degradation patterns in both brightness and structure to harmonize latent representations but also dynamically calibrates the illumination prior with the generative knowledge inherent to diffusion models. As such, DiffLLFace attains precise control over conditional adaptation and illumination rectification. We further devise a simple yet effective non-parametric Fourier enhancement strategy, which provides structural appearance clues that work in concert with the alternate adaptation to ensure texture and color consistency. Extensive experiments demonstrate the superiority of DiffLLFace over existing methods and remarkable generalizability on complex natural scenes. Code is available at https://github.com/KaishengPang/DiffLLFace.
Runmin Cong, Kaisheng Pang, Feng Li 0037, Hua Li 0012, Huihui Bai 0001, Sam Kwong, Wei Zhang 0021
IEEE Trans. Image Process.4
2026 ViT-UWA: Vision Transformer Underwater-Adapter for Dense Predictions Beneath the Water Surface
abstract
Vision Transformer (ViT) and its variants have witnessed a significant success in computer vision. However, their performance may degrade in underwater dense prediction tasks due to challenges like complex underwater environments, quality degradation, and light scattering in underwater images. To solve this problem, we propose the Vision Transformer Underwater-Adapter (ViT-UWA), the first detail-focused and adapted ViT backbone for underwater dense prediction tasks, without requiring task-specific pretraining. In ViT-UWA, we first introduce High-frequency Components Prior (HFCP) to add high-frequency information of underwater images to the plain ViT, which can help recover and capture lost high-frequency information of underwater images. Then, we propose a Detail Aware Module (DAM) to obtain a detail-focused multi-scale convolutional feature pyramid, which can be used in kinds of dense prediction tasks. Through the ViT-DAM Cross Fusion (VDCF), we achieve bidirectional feature cross fusion between ViT and DAM. We evaluate ViT-UWA on multiple underwater dense prediction tasks, including semantic segmentation, instance segmentation, and object detection. With only ImageNet-22K pretraining, our ViT-UWA-B yields state-of-the-art 46.4 box AP and 44.2 mask AP on USIS10K dataset, which demonstrates the superiority of our method. Our code is available at https://github.com/Linqirui/ViT-UWA.
Yuheng Jia, Qirui Lin, Hua Li 0012, Sam Kwong, Runmin Cong
IEEE Trans. Image Process.3
2026 Expose Camouflage in the Water: Underwater Camouflaged Instance Segmentation and Dataset
abstract
With the development of underwater exploration and marine protection, underwater vision tasks are widespread. Due to the degraded underwater environment, characterized by color distortion, low contrast, and blurring, camouflaged instance segmentation (CIS) faces greater challenges in accurately segmenting objects that blend closely with their surroundings. Traditional camouflaged instance segmentation methods, trained on terrestrial-dominated datasets with limited underwater samples, may exhibit inadequate performance in underwater scenes. To address these issues, we introduce the first underwater camouflaged instance segmentation (UCIS) dataset, abbreviated as UCIS4K, which comprises 3,953 images of camouflaged marine organisms with instance-level annotations. In addition, we propose an Underwater Camouflaged Instance Segmentation network based on Segment Anything Model (UCIS-SAM). Our UCIS-SAM includes three key modules. First, the Channel Balance Optimization Module (CBOM) enhances channel characteristics to improve underwater feature learning, effectively addressing the model's limited understanding of underwater environments. Second, the Frequency Domain True Integration Module (FDTIM) is proposed to emphasize intrinsic object features and reduce interference from camouflage patterns, enhancing the segmentation performance of camouflaged objects blending with their surroundings. Finally, the Multi-scale Feature Frequency Aggregation Module (MFFAM) is designed to strengthen the boundaries of low-contrast camouflaged instances across multiple frequency bands, improving the model's ability to achieve more precise segmentation of camouflaged objects. Extensive experiments on the proposed UCIS4K and public benchmarks show that our UCIS-SAM outperforms state-of-the-art approaches. The code and dataset are released at https://github.com/wchchw/UCIS4K.
Chuhong Wang, Hua Li 0012, Chongyi Li, Huazhong Liu, Xiongxin Tang, Sam Kwong
IEEE Trans. Image Process.2
2025 TMANet: Triple Multi-Scale Attention based Network with Boundary Association Loss for Superpixel Segmentation
abstract
Superpixel segmentation with deep learning has been proposed in recent years and is widely employed to reduce the input image primitives for subsequent computer vision tasks. In this paper, we propose a Triple Multi-Scale Attention based Network (TMANet) for superpixel segmentation. First, aiming to extract more detailed context information, we design a Triple Multi-Scale Attention (TMA) to adapt the varying object scale and reduce inevitable redundant information in the encoder. Moreover, we observe that the dark parts with low probability in the association map generated by the TMANet are closer to the superpixel boundary. Therefore, we devise Boundary Association (BA) loss based on the association map to obtain fine boundaries and contours. Extensive experiments on public datasets show that TMANet outperforms the state-of-the-art methods to a certain extent. The application in saliency object detection of remote sensing also demonstrates the superiority of the proposed method.
Shijie Lian, Hua Li 0012
ICASSP3
2025 OptiDiff: Unsupervised Deep-Sea Image Enhancement via Optical Priors Guided Stable Diffusion
abstract
Deep-sea images suffer from extreme light attenuation and non-uniform illumination caused by artificial light sources. To get rid of limitation aroused by low-quality training data, we propose an unsupervised approach for deep-sea image enhancement based on the prior-to-image framework, termed as OptiDiff. Instead of learning mapping from paired underwater dataset, the framework for OptiDiff is trained on air image dataset. Specifically, four optical-invariant priors (OIPs) are used for guiding the stable diffusion model to recover degraded underwater image. One of the utilized OIPs is particularly designed for recover blurred details in background. Besides, to simulate the domain shift between air and underwater images, a channel attenuation strategy derived from characteristics of real-world underwater image is equipped with the framework. Extensive experimental results demonstrate the superiority of the proposed OptiDiff on restoring image from low contrast, low visibility, and severe blur, also showing robustness on shallow-water datasets. Code is available at https://github.com/Miaaaaaa1024/OptiDiff.
Wenhui Wu 0001, Yuemiao Wang, Hua Li 0012, Yuanhao Gong
ICME3
2025 FSCDiff: Frequency-Spatial Entangled Conditional Diffusion model for Underwater Salient Object Detection
abstract
Salient object detection (SOD) plays a crucial role in image understanding and visual guidance. However, due to the complexity of underwater environments, the accuracy of underwater salient object detection is often low. To improve the accuracy and robustness of underwater salient object detection, different from the existing spatial domain aware RGB-D methods that rely on pixel-level probabilities, we propose a novel Fourier-Spatial Entangled Conditional Diffusion model (FSCDiff) for underwater salient object detection. The FSCDiff aims to address the insufficient representation and boundary shift issues in underwater salient object detection by leveraging Fourier-domain information and the powerful multi-step iterative generation capability of diffusion models. The FSCDiff framework consists of two key components: the Dual-Domain Entanglement Enhancement Block (DTEB) and the Stable Time-step Mask Prediction Module (STMP). DTEB utilizes Fourier-spatial entanglement learning to fully exploit the Fourier and spatial domain information of RGB images and depth maps, thereby optimizing feature representation. STMP takes advantage of the excellent multi-step iterative mechanism of diffusion models to enhance the accuracy and robustness of the segmentation results. Comprehensive experimental results indicate that our FSCDiff method outperforms the state-of-the-art approaches on the USOD10K and USOD datasets. The source code is available at: https://github.com/lgwplay/FSCDiff.
Hua Li 0012, Gaowei Lin, Sam Kwong, Runmin Cong
ACM Multimedia1
2025 Tucker-Based High-Accuracy Multi-Modal Clustering for Social Information Network
abstract
With the explosion of social media platforms, a substantial amount of data is generated from social information network. Tensor-based multi-modal clustering methods have been widely applied in various scenarios of social information network by mining potential correlative relationships from large-scale heterogeneous data. Nevertheless, the accuracy and efficiency of tensor-based multi-modal clustering methods are seriously restricted by noise data and the curse of dimensionality. Therefore, this paper presents a Tucker-based multi-modal clustering (TuMC) and an improved TuMC (ITuMC) to enhance the accuracy and efficiency of multi-modal clustering. First, we propose two Tucker-based attribute weight ranking learning approaches to calculate weight tensor efficiently. Then, we present a calculation approach for Tucker-based selective weighted tensor distance (SWTD) and a TuMC method. Meanwhile, an ITuMC method is explored by optimizing the calculation efficiency of the SWTD to further improve clustering speed. Finally, we present a Tucker-based multi-modal clustering and service framework for social information network. Extensive experimental results based on social Geolife GPS trajectory and electricity consumption datasets demonstrate that the TuMC and ITuMC methods can cluster multi-source heterogeneous data with both higher accuracy and efficiency under complex social information network by DVI, AR and execution time measurement.
Huazhong Liu, Xiaotong Zhou, Jihong Ding, Laurence T. Yang, Hua Li 0012
IEEE Trans. Big Data7
2024 ESNet: Evolution and Succession Network for High-Resolution Salient Object Detection
abstract
Preserving details and avoiding high computational costs are the two main challenges for the High-Resolution Salient Object Detection (HRSOD) task. In this paper, we propose a two-stage HRSOD model from the perspective of evolution and succession, including an evolution stage with Low-resolution Location Model (LrLM) and a succession stage with High-resolution Refinement Model (HrRM). The evolution stage achieves detail-preserving salient objects localization on the low-resolution image through the evolution mechanisms on supervision and feature; the succession stage utilizes the shallow high-resolution features to complement and enhance the features inherited from the first stage in a lightweight manner and generate the final high-resolution saliency prediction. Besides, a new metric named Boundary-Detail-aware Mean Absolute Error (${MAE}_{{BD}}$) is designed to evaluate the ability to detect details in high-resolution scenes. Extensive experiments on five datasets demonstrate that our network achieves superior performance at real-time speed (49 FPS) compared to state-of-the-art methods.
Hongyu Liu 0003, Runmin Cong, Hua Li 0012, Qianqian Xu 0001, Qingming Huang, Wei Zhang 0021
ICML3
2024 Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale Dataset
abstract
With the breakthrough of large models, Segment Anything Model (SAM) and its extensions have been attempted to apply in diverse tasks of computer vision. Underwater salient instance segmentation is a foundational and vital step for various underwater vision tasks, which often suffer from low segmentation accuracy due to the complex underwater circumstances and the adaptive ability of models. Moreover, the lack of large-scale datasets with pixel-level salient instance annotations has impeded the development of machine learning techniques in this field. To address these issues, we construct the first large-scale underwater salient instance segmentation dataset (USIS10K), which contains 10,632 underwater images with pixel-level annotations in 7 categories from various underwater scenes. Then, we propose an Underwater Salient Instance Segmentation architecture based on Segment Anything Model (USIS-SAM) specifically for the underwater domain. We devise an Underwater Adaptive Visual Transformer (UA-ViT) encoder to incorporate underwater domain visual prompts into the segmentation network. We further design an out-of-the-box underwater Salient Feature Prompter Generator (SFPG) to automatically generate salient prompters instead of explicitly providing foreground points or boxes as prompts in SAM. Comprehensive experimental results show that our USIS-SAM method can achieve superior performance on USIS10K datasets compared to the state-of-the-art methods. Datasets and codes are released on https://github.com/LiamLian0727/USIS10K.
Shijie Lian, Hua Li 0012, Laurence T. Yang, Sam Kwong, Runmin Cong
ICML3
2024 Stereo Superpixel Segmentation via Decoupled Dynamic Spatial-Embedding Fusion Network
abstract
Stereo superpixel segmentation aims at grouping the discretizing pixels into perceptual regions through left and right views more collaboratively and efficiently. Existing superpixel segmentation algorithms mostly utilize color and spatial features as input, which may impose strong constraints on spatial information while utilizing the disparity information in terms of stereo image pairs. To alleviate this issue, we propose a stereo superpixel segmentation method with a decoupling mechanism of spatial information in this work. To decouple stereo disparity information and spatial information, the spatial information is temporarily removed before fusing the features of stereo image pairs, and a decoupled stereo fusion module (DSFM) is designed to handle the stereo features alignment as well as occlusion problems. Moreover, since the spatial information is vital to superpixel segmentation, we further design a dynamic spatiality embedding module (DSEM) to re-add spatial information, and the weights of spatial information will be adaptively adjusted through the dynamic fusion (DF) mechanism in DSEM for achieving a finer segmentation. Comprehensive experimental results demonstrate that our method can achieve the state-of-the-art performance on the KITTI2015 and Cityscapes datasets, and also verify the efficiency when applied in salient object detection on NJU2K dataset. The source code will be available publicly after paper is accepted.
Hua Li 0012, Junyan Liang, Runmin Cong, Wenhui Wu 0001, Sam Kwong
IEEE Trans. Multim.1
2023 WaterMask: Instance Segmentation for Underwater Imagery
abstract
Underwater image instance segmentation is a fundamental and critical step in underwater image analysis and understanding. However, the paucity of general multiclass instance segmentation datasets has impeded the development of instance segmentation studies for underwater images. In this paper, we propose the first underwater image instance segmentation dataset (UIIS), which provides 4628 images for 7 categories with pixel-level annotations. Meanwhile, we also design WaterMask for underwater image instance segmentation for the first time. In Water-Mask, we first devise Difference Similarity Graph Attention Module (DSGAT) to recover lost detailed information due to image quality degradation and downsampling to help the network prediction. Then, we propose Multi-level Feature Refinement Module (MFRM) to predict foreground masks and boundary masks separately by features at different scales, and guide the network through Boundary Mask Strategy (BMS) with boundary learning loss to provide finer prediction results. Extensive experimental results demonstrates that WaterMask can achieve significant gains of 2.9, 3.8 mAP over Mask R-CNN when using ResNet-50 and ResNet-101. Code and Dataset are available at https://github.com/LiamLian0727/WaterMask.
Shijie Lian, Hua Li 0012, Runmin Cong, Suqi Li, Wei Zhang 0021, Sam Kwong
ICCV2
2023 Multi-View Super Resolution for Underwater Images Utilizing Atmospheric Light Scattering Model
abstract
The underwater environment is complex and the underwater light propagation undergoes absorption, scattering and reflection. This leads to the fact that the underwater light imaging cannot be generalized from land-based. How to use these imaging features to work better with super-resolution tasks for underwater imagery applications is still rarely studied. In this paper, we introduce the medium transmission (MT) maps to advance super-resolution tasks for underwater images. A multi-view network is designed to fuse information from the original underwater images and the MT maps, which provides information on the underlying physical properties of the water, such as the attenuation coefficients in different parts of water. By integrating information from multiple views, the proposed network can capture more of the underlying structure and features of the scene, leading to higher-quality super-resolved images. Besides, a new loss function, namely MT Loss, is developed according to the lack of details in special region of the underwater images. This loss function emphasizes the regions with less influence from the underwater environment during the underwater imaging process and therefore the network outputs a more detailed image. Finally, we compare our algorithm with state-of-the-art methods, and extensive results show that our network achieves better qualitative and quantitative performance.
Jin Hao, Wenli Duan, Guangfei Li, Shiyan Chen, Wenhui Wu 0001, Hua Li 0012
ICPADS6
2023 FSNet: Frequency Domain Guided Superpixel Segmentation Network for Complex Scenes
abstract
Existing superpixel segmentation algorithms mainly focus on natural image with high-quality, while neglecting the inevitable environment constraint in complex scenes. In this paper, we propose an end-to-end frequency domain guided superpixel segmentation network (FSNet) to generate superpixels with sharp boundary adherence for complex scenes by fusing the deep features in spatial and frequency domains. To utilize the frequency domain information of the image, an improved frequency information extractor (IFIE) is proposed to extract the frequency domain information with sharp boundary features. Moreover, considering the over-sharp feature may damage the semantic information of superpixel, we further design a dense hybrid atrous convolution (DHAC) block to preserve semantic information via capturing wider and deeper semantic information in spatial domain. Finally, the extracted deep features in spatial and frequency domains will be fused to generate semantic perceptual superpixels with sharp boundary adherence. Extensive experiments on multiple challenging datasets with complex boundaries demonstrate that our method achieves the state-of-the-art performance both quantitatively and qualitatively, and we further verify the superiority of the proposed method when applied in salient object detection.
Hua Li 0012, Junyan Liang, Wenhui Wu 0001
ACM Multimedia1
2023 Atmospheric Scattering Model Induced Statistical Characteristics Estimation for Underwater Image Restoration
abstract
Underwater images often suffer from color deviation and low contrast due to selective absorption and light scattering, whose degradation is generally described by an Atmospheric Scattering Model (ASM). However, it is challenging to design hand-craft priors to estimate the transmission map and global light within ASM. To avoid the estimation on these two variables, in this paper, we establish a statistical characteristics relationship between underwater and recovered images based on ASM. With this relationship, a novel lightweight model is proposed for efficient Underwater Image Restoration (UIR). Within our proposed model, the UIR problem is disentangled into global restoration and local compensation, for which two modules are developed. Extensive experimental results demonstrate that our proposed method can effectively improve color deviation and low contrast while preserving details, and outperform state-of-the-art methods.
Shuaibo Gao, Wenhui Wu 0001, Hua Li 0012, Linwei Zhu, Xu Wang 0006
IEEE Signal Process. Lett.3
2021 Stereo Superpixel Segmentation Via Dual-Attention Fusion Networks
abstract
Stereo image pairs can improve performance of many tasks benefiting from the additional information obtained from a second viewpoint when compared with single images. Existing superpixel segmentation algorithms for stereo images mostly adopt single images as input, and neglect the correspondence between the left and right views. In this work, we consider to exploit the depth information between stereo image pairs, and propose an end-to-end dual-attention fusion network for stereo images to generate parallax-consistency superpixels. We first utilize a deep convolution network to extract the deep features of stereo images. Then, to effectively utilize the additional information from the other view, features of the left and right views is integrated by a parallax attention and channel attention mechanism. Finally, the stereo superpixels are generated by a differentiable clustering algorithm, which is end-to-end trainable with deep learning networks. Comprehensive experimental results demonstrate that our method can outperform the state-of-the-art performance on the KITTI2015 and Cityscapes dataset.
Yajuan Du, Hua Li 0012, Yucong Dai
ICME3
2021 Stereo superpixel: An iterative framework based on parallax consistency and collaborative optimization
Hua Li 0012, Runmin Cong, Sam Kwong, Chuanbo Chen, Qianqian Xu 0001, Chongyi Li
Inf. Sci.1
2021 Superpixel Segmentation Based on Spatially Constrained Subspace Clustering
abstract
Superpixel segmentation aims at dividing the input image into some representative regions containing pixels with similar and consistent intrinsic properties, without any prior knowledge about the shape and size of each superpixel. In this article, to alleviate the limitation of superpixel segmentation applied in practical industrial tasks that detailed boundaries are difficult to be kept, we regard each representative region with independent semantic information as a subspace, and correspondingly formulate superpixel segmentation as a subspace clustering problem to preserve more detailed content boundaries. We show that a simple integration of superpixel segmentation with the conventional subspace clustering does not effectively work due to the spatial correlation of the pixels within a superpixel, which may lead to boundary confusion and segmentation error when the correlation is ignored. Consequently, we devise a spatial regularization and propose a novel convex locality-constrained subspace clustering model that is able to constrain the spatial adjacent pixels with similar attributes to be clustered into a superpixel and generate the content-aware superpixels with more detailed boundaries. Finally, the proposed model is solved by an efficient alternating direction method of multipliers solver. Experiments on different standard datasets demonstrate that the proposed method achieves superior performance both quantitatively and qualitatively compared with some state-of-the-art methods.
Hua Li 0012, Yuheng Jia, Runmin Cong, Wenhui Wu 0001, Sam Kwong, Chuanbo Chen
IEEE Trans. Ind. Informatics1
2020 A parallel down-up fusion network for salient object detection in optical remote sensing images
Chongyi Li, Runmin Cong, Chunle Guo, Hua Li 0012, Chunjie Zhang 0001, Feng Zheng 0001, Yao Zhao 0001
Neurocomputing4
2019 Superpixel Segmentation Based on Square-Wise Asymmetric Partition and Structural Approximation
abstract
Superpixel segmentation aims at grouping discretizing pixels into high-level correlative units and reducing the complexity of subsequent tasks, e.g., saliency detection and object tracking. Existing superpixel segmentation algorithms mainly focus on maintaining the geometrical information, while neglecting the irregular structure of superpixels. In this paper, a superpixel segmentation method is proposed to generate approximately structural superpixels with sharp boundary adherence and comprehensive semantic information. The superpixel segmentation is formulated as a square-wise asymmetric partition problem, where the semantic perceptual superpixels are recorded in a square level to preserve abundant semantic information and save storage simultaneously. Moreover, in order to achieve regular-shape superpixel units to better adhere to image boundaries and contours, a combinatorial optimization strategy is devised to achieve an optimal combination of squares and isolated pixels. Experimental comparisons with some state-of-the-art superpixel segmentation methods on the public benchmarks demonstrate the effectiveness of the proposed method quantitatively and qualitatively. In addition, we have applied the method to brain tissue segmentation to illustrate superior performance.
Hua Li 0012, Sam Kwong, Chuanbo Chen, Yuheng Jia, Runmin Cong
IEEE Trans. Multim.1