Feng Huang 0007

dblp:19/6809-7 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-4652-4312ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Explainable Artificial Intelligence Enhance Image Semantic Communication System in 6G-IoT
abstract
The emerging 6G-IoT paradigm is driving communication toward intelligent services, semantic communication enables efficient semantic sharing via artificial intelligence (AI), significantly boosting communication efficiency. However, current semantic systems suffer from black-box decision-making, while existing explainable artificial intelligence (XAI) methods face two key challenges: explicability granularity mismatch and closed-loop optimization gap. To address these, we propose a semantic communication framework integrated with XAI (XAI-SCS). Specifically, we first design an explainable semantic codec architecture enhanced by Kolmogorov–Arnold Networks (KAN), where traditional fixed activation functions are replaced with learnable and parameterized ones, enabling function-level visualization to improve model explainability. Second, we develop an explainable semantic transmission module driven by contrastive learning that enhances the robustness of semantic transmission, and incorporating a semantic separability metric to quantify channel impacts on semantic integrity. Third, we introduced a KAN-enhanced causal semantic decoder, which integrates counterfactual interventions to generate pixel-level difference maps. We also propose a contrastive explanation consistency metric to evaluate the sensitivity of key features, enhancing the quality of the reconstruction. The experimental results show that our approaches enhance explainability across the entire decision process, achieve a significant accuracy improvement of up to 55% on the CIFAR-10 dataset with a bandwidth compression ratio of 1/25, and also obtain competitive image reconstruction quality in increasing compression levels. The source code is publicly available at: https://github.com/guyuangui/XAI-SCS.git.
Mingkai Chen 0001, Yuangui Gu, Xiaoming He 0004, Feng Huang 0007, Lei Wang 0009
IEEE Internet Things J.5
2026 DFINet: Dynamic feedback iterative network for infrared small target detection
Jing Wu 0023, Changhai Luo, Zhaobing Qiu, Liqiong Chen, Rixiang Ni, Feng Huang 0007
Pattern Recognit.7
2026 Corrections to "Exploring Fuzzy Priors From Multimapping GAN for Robust Image Dehazing"
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, Li Zhao 0005, En Fan, Feng Huang 0007
IEEE Trans. Fuzzy Syst.6
2025 RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-Resolution
abstract
Benefiting from their powerful generative capabilities, pretrained diffusion models have garnered significant attention for real-world image super-resolution (Real-SR). Existing diffusion-based SR approaches typically utilize semantic information from degraded images and restoration prompts to activate prior for producing realistic high-resolution images. However, general-purpose pretrained diffusion models, not designed for restoration tasks, often have suboptimal prior, and manually defined prompts may fail to fully exploit the generated potential. To address these limitations, we introduce RAP-SR, a novel restoration prior enhancement approach in pretrained diffusion models for Real-SR. First, we develop the High-Fidelity Aesthetic Image Dataset (HFAID), curated through a Quality-Driven Aesthetic Image Selection Pipeline (QDAISP). Our dataset not only surpasses existing ones in fidelity but also excels in aesthetic quality. Second, we propose the Restoration Priors Enhancement Framework, which includes Restoration Priors Refinement (RPR) and Restoration-Oriented Prompt Optimization (ROPO) modules. RPR refines the restoration prior using the HFAID, while ROPO optimizes the unique restoration identifier, improving the quality of the resulting images. RAP-SR effectively bridges the gap between general-purpose models and the demands of Real-SR by enhancing restoration prior. Leveraging the plug-and-play nature of RAP-SR, our approach can be seamlessly integrated into existing diffusion-based SR methods, boosting their performance. Extensive experiments demonstrate its broad applicability and state-of-the-art results.
Qingnan Fan, Feng Huang 0007, Wenqi Ren
AAAI5
2025 LVPTrack: High Performance Domain Adaptive UAV Tracking with Label Aligned Visual Prompt Tuning
abstract
Visual object tracking is essentially crucial for unmanned aerial vehicles (UAVs). Despite the substantial progress, most of the existing UAV trackers are designed for well-conditioned daytime data, while for the scenarios in challenging weather condition, e.g. foggy or nighttime environment, the tremendous domain gap leads to significant performance degradation. To address this issue, in this paper, we propose a novel robust UAV tracker termed LVPTrack, which conducts high quality label-aligned visual prompt tuning to adapt to various challenging weather conditions. Specifically, we first synthesize the sequential foggy and nighttime video frames to assist the model training. A domain adaptive teacher-student network is utilized to distill the hierarchical visual semantic of the target objects in cross-domain scenarios. Then we propose a target-aware pseudo-label voting (PLV) strategy to alleviate the target-level misalignment in the dual domains. Furthermore, we propose a dynamic aggregated prompt (DAP) module to facilitate the appearance variation adaptation of the target object in challenging scenarios. Extensive experiments demonstrate that our tracker achieves superior performance over existing state-of-the-art UAV trackers.
Hongjing Wu, Siyuan Yao, Feng Huang 0007, Linchao Zhang, Zhuoran Zheng, Wenqi Ren
AAAI3
2025 Text-IRSTD: Leveraging Semantic Text to Promote Infrared Small Target Detection in Complex Scenes
Feng Huang 0007, Shuyuan Zheng, Zhaobing Qiu, Huanxian Liu, Huanxin Bai, Liqiong Chen
ICCV1
2025 Uncertainty and diversity-based active learning for UAV tracking
Yingqin Liang, Feng Huang 0007, Zhaobing Qiu, Xiu Shu, Qiao Liu 0001, Di Yuan 0002
Neurocomputing2
2025 Exploring Fuzzy Priors From Multimapping GAN for Robust Image Dehazing
abstract
Single image dehazing has been extensively studied. While convolutional neural networks (CNNs) have driven notable progress in single image dehazing, their performance remains fundamentally constrained by the limited local receptive fields of convolutional operations, which impede the capture of global structural dependencies. In contrast, generative adversarial networks (GANs) have demonstrated exceptional capabilities in image synthesis, offering global insights into structure, texture, and color. The fuzzy prior, a probabilistic knowledge acquired through adversarial training in GANs, plays a pivotal role in robust dehazing. Motivated by this, we propose the fuzzy prior guided dehazing network (FPGDN). Our framework begins with a novel module that distills the fuzzy prior by translating an edge map into a color image, simultaneously capturing global structural, local textural, and color information. Subsequently, a dehazing network is constructed, leveraging this fuzzy prior. While the fuzzy prior captures rich color and texture features, the generated images may exhibit color shifts relative to the original scene. To remedy this, a CNN network is employed to capture local nuances and refine the dehazing outcome. Extensive experiments substantiate that the proposed FPGDN achieves superior dehazing performance on a variety of real and synthetic hazy images.
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, Li Zhao 0005, En Fan, Feng Huang 0007
IEEE Trans. Fuzzy Syst.6
2025 Aperture Array-Based Self-Supervised Hyperspectral Super-Resolution Imaging
abstract
Current hyperspectral imaging (HSI) technologies struggle with prolonged acquisition time, bulky system design, and high computational complexity. We propose Aperture Array-Based Self-Supervised Hyperspectral Super-Resolution Imaging, a novel computational HSI framework that integrates an aperture array optical structure with a physics-driven, self-supervised deep learning model. The aperture array enables spectral division while leveraging aperture parallax for spatial super-resolution (SR), achieving high spectral fidelity and spatial SR in a compact form factor. To fully exploit the imaging characteristics of the proposed system, we design a Divide-and-Conquer Self-Supervised Hyperspectral Super-Resolution Imaging (DSHSI) algorithm, which is data-efficient, non-iterative, and free of large-scale ground-truth datasets. Unlike existing hyperspectral restoration methods that suffer from long inference times and poor generalization, DSHSI integrates physics-driven priors into self-supervised learning framework, enabling robust performance on unseen scenes and out-of-distribution degradations. DSHSI jointly performs denoising, SR, and deblurring in a unified manner, reducing runtime by 25×, FLOPs by 9× and model parameters by 3.6× than DualSR on public remote sensing datasets, while achieving a 1.9 dB PSNR gain, enabling fast and efficient hyperspectral reconstruction. In the optical experiments, it achieves the best NIQE and the spectral distortion metric that is closely aligned with the spectral response of observed HSI data. Our code is publicly available at https://github.com/THUHoloLab/DSHSR.
Yating Chen, Feng Huang 0007, Liangcai Cao
IEEE Trans. Geosci. Remote. Sens.2
2025 Point-to-Point Regression: Accurate Infrared Small Target Detection With Single-Point Annotation
abstract
Infrared small target detection (IRSTD) plays a vital role in various fields, especially in military early warning and maritime rescue. Its main goal is to accurately locate targets at long distances. Current deep learning (DL)-based methods mainly rely on mask-to-mask or box-to-box regression training approaches, making considerable progress in detection accuracy. However, these methods rely on large amounts of training data with expensive manual annotation. Although some researchers attempt to reduce the cost using single-point weak supervision (SPWS), the limited labeling accuracy significantly degrades the detection performance. To address these issues, we propose a novel point-to-point regression high-resolution dynamic network (P2P-HDNet), which can accurately locate the target center using only single-point annotation. Specifically, we first devise the high-resolution cross-feature extraction module (HCEM) to provide richer target detail information for the deep feature maps. Notably, HCEM maintains high resolution throughout the feature extraction process to minimize information loss. Then, the dynamic coordinate fusion module (DCFM) is devised to fully fuse the multidimensional features and enhance the positional sensitivity. Finally, we devise an adaptive target localization detection head (ATLDH) to further suppress clutter and improve the localization accuracy by regressing the Gaussian heatmap and adaptive nonmaximal suppression strategy. Extensive experimental results show that P2P-HDNet can achieve better detection accuracy than the state-of-the-art (SOTA) methods with only single-point annotation. In addition, our code and datasets will be available at:https://github.com/Anton-Nrx/P2P-HDNet.
Rixiang Ni, Jing Wu 0023, Zhaobing Qiu, Liqiong Chen, Changhai Luo, Feng Huang 0007, Qiujiang Liu, Binxing Wang, Youli Li
IEEE Trans. Geosci. Remote. Sens.6
2025 PFAN: progressive feature aggregation network for lightweight image super-resolution
Liqiong Chen, Xiangkun Yang, Ying Shen 0004, Jing Wu 0023, Feng Huang 0007, Zhaobing Qiu
Vis. Comput.6
2024 INformer: Inertial-Based Fusion Transformer for Camera Shake Deblurring
abstract
Inertial measurement units (IMU) in the capturing device can record the motion information of the device, with gyroscopes measuring angular velocity and accelerometers measuring acceleration. However, conventional deblurring methods seldom incorporate IMU data, and existing approaches that utilize IMU information often face challenges in fully leveraging this valuable data, resulting in noise issues from the sensors. To address these issues, in this paper, we propose a multi-stage deblurring network named INformer, which combines inertial information with the Transformer architecture. Specifically, we design an IMU-image Attention Fusion (IAF) block to merge motion information derived from inertial measurements with blurry image features at the attention level. Furthermore, we introduce an Inertial-Guided Deformable Attention (IGDA) block for utilizing the motion information features as guidance to adaptively adjust the receptive field, which can further refine the corresponding blur kernel for pixels. Extensive experiments on comprehensive benchmarks demonstrate that our proposed method performs favorably against state-of-the-art deblurring approaches.
Wenqi Ren, Linrui Wu, Yanyang Yan, Shengyao Xu, Feng Huang 0007, Xiaochun Cao
IEEE Trans. Image Process.5
2023 RSHAN: Image super-resolution network based on residual separation hybrid attention module
Ying Shen 0004, Weihuang Zheng, Liqiong Chen, Feng Huang 0007
Eng. Appl. Artif. Intell.4
2023 Spectral Clustering Super-Resolution Imaging Based on Multispectral Camera Array
abstract
Although multispectral and hyperspectral imaging acquisitions are applied in numerous fields, the existing spectral imaging systems suffer from either low temporal or spatial resolution. In this study, a new multispectral imaging system-camera array based multispectral super resolution imaging system (CAMSRIS) is proposed that can simultaneously achieve multispectral imaging with high temporal and spatial resolutions. The proposed registration algorithm is used to align pairs of different peripheral and central view images. A novel, super-resolution, spectral-clustering-based image reconstruction algorithm was developed for the proposed CAMSRIS to improve the spatial resolution of the acquired images and preserve the exact spectral information without introducing false information. The reconstructed results showed that the spatial and spectral quality and operational efficiency of the proposed system are better than those of a multispectral filter array (MSFA) based on different multispectral datasets. The PSNR of the multispectral super-resolution images obtained by the proposed method were respectively higher by 2.03 and 1.93 dB than those of GAP-TV and DeSCI, and the execution time was significantly shortened by approximately 54.55 s and 9820.19 s when the CAMSI dataset was used. The feasibility of the proposed system was verified in practical applications based on different scenes captured by the self-built system.
Feng Huang 0007, Yating Chen, Xianyu Wu
IEEE Trans. Image Process.1
2023 Dynamic Representation Learning via Recurrent Graph Neural Networks
abstract
A large number of real-world systems generate graphs that are structured data aligned with nodes and edges. Graphs are usually dynamic in many scenarios, where nodes or edges keep evolving over time. Recently, graph representation learning (GRL) has received great success in network analysis, which aims to produce informative and representative features or low-dimensional embeddings by exploring node attributes and network topology. Most state-of-the-art models for dynamic GRL are composed of a static representation learning model and a recurrent neural network (RNN). The former generates the representations of a graph or nodes from one static graph at a discrete time step, while the latter captures the temporal correlation between adjacent graphs. However, the two-stage design ignores the temporal dynamics between contiguous graphs during the learning processing of graph representations. To alleviate this problem, this article proposes a representation learning model for dynamic graphs, called DynGNN. Differently, it is a single-stage model that embeds an RNN into a graph neural network to produce better representations in a compact form. This takes the fusion of temporal and topology correlations into account from low-level to high-level feature learning, enabling the model to capture more fine-grained evolving patterns. From the experimental results on both synthetic and real-world networks, the proposed DynGNN yields significant improvements in multiple tasks compared to the state-of-the-art counterparts.
Chun-Yang Zhang, Zhiliang Yao, Feng Huang 0007, C. L. Philip Chen
IEEE Trans. Syst. Man Cybern. Syst.4