Hongxia Gao

dblp:73/2494 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dynamic expandable framework for incremental anomaly detection
Yuxuan Tan, Hongxia Gao, Tongtong Liu 0003, Xiaoqin Wen
Neural Networks2
2026 Entropy-increasing linear attention for multi-class unsupervised anomaly detection
Tongtong Liu 0003, Hongxia Gao, Yuxuan Tan, Jinpeng Li 0002, Jinhui Zhao
Pattern Recognit.2
2025 GDFDNet: A Novel Graph-Based Dynamically Fused Dual-Stream Network for Accuracy Prohibited Items Detection
abstract
In various security inspection scenarios, prohibited items detection in X-ray images is of great significance for safeguarding public safety and effectively reducing the potential risks of crimes and terrorist activities. However, existing detection methods still face the challenge of overlapping image features when distinguishing prohibited items from complex backgrounds. To tackle this issue, this paper proposes a novel Graph-based Dynamically Fused Dual-Stream Network (GDFDNet). This network introduces highly decoupled HSV color space features, which are dynamically fused with RGB color space features to fully explore the edge and material characteristics of overlapping prohibited items. Specifically, we adopt a dynamically fused RGB-HSV dual-stream framework and design two key modules: a Graph-based Edge-Aware module (GEA) and a Graph-based Material-Aware Fusion module (GMAF). The former is utilized to enhance multi-scale edge features of prohibited items, while the latter is responsible for performing the dynamic fusion of features in dual color spaces guided by edge features, and then extracting valuable material information of prohibited items from the overlapping foreground and background features. Extensive experiments on the SIXray and PIDray datasets demonstrate that our proposed network significantly outperforms the existing state-of-the-art methods.
Hongxia Gao, Yaobin Huang, Zhenming Guan, Litao Li
ICASSP2
2024 VQCNIR: Clearer Night Image Restoration with Vector-Quantized Codebook
abstract
Night photography often struggles with challenges like low light and blurring, stemming from dark environments and prolonged exposures. Current methods either disregard priors and directly fitting end-to-end networks, leading to inconsistent illumination, or rely on unreliable handcrafted priors to constrain the network, thereby bringing the greater error to the final result. We believe in the strength of data-driven high-quality priors and strive to offer a reliable and consistent prior, circumventing the restrictions of manual priors. In this paper, we propose Clearer Night Image Restoration with Vector-Quantized Codebook (VQCNIR) to achieve remarkable and consistent restoration outcomes on real-world and synthetic benchmarks. To ensure the faithful restoration of details and illumination, we propose the incorporation of two essential modules: the Adaptive Illumination Enhancement Module (AIEM) and the Deformable Bi-directional Cross-Attention (DBCA) module. The AIEM leverages the inter-channel correlation of features to dynamically maintain illumination consistency between degraded features and high-quality codebook features. Meanwhile, the DBCA module effectively integrates texture and structural information through bi-directional cross-attention and deformable convolution, resulting in enhanced fine-grained detail and structural fidelity across parallel decoders. Extensive experiments validate the remarkable benefits of VQCNIR in enhancing image quality under low-light conditions, showcasing its state-of-the-art performance on both synthetic and real-world datasets. The code is available at https://github.com/AlexZou14/VQCNIR.
Wenbin Zou, Hongxia Gao, Tian Ye 0001, Liang Chen 0026, Weipeng Yang 0002, Shasha Huang, Sixiang Chen
AAAI2
2024 Adaptxray: Vision Transformer And Adapter In X-Ray Images For Prohibited Items Detection
abstract
Prohibited items detection refers to the non-contact inspection of passenger baggage for potential threats through X-ray image. Since the uncertainty of artificial security screening, previous research has mainly concentrated on direct transfer by universal detection frameworks based on natural image and design in enhance or aware module with salient features like edge and color. With the increasing complexity in both categories and quantities of X-ray security inspection, the unreliability of direct transfer and complication of task-specific design make the existing algorithms difficult to reliably and efficiently adapt the complex security inspection. To address this challenge, we propose the Adapter in X-ray (AdaptXray), which firstly explores pre-trained Vision Transformer with powerful representation and Parameter Efficient Transfer Learning method applying for prohibited items detection. Specifically, we design Color Prior Extractor to perceive local prior features from different color spaces. Subsequently, we develop Global-aware Self-Adapter to adaptively perceive and optimize the global universal features in the backbone. Additionally, we propose Local-aware Interactive Adapter to incorporate prior knowledge into the pretrained backbone. Thorough experimentation on two public baggage datasets, namely OPIXray and PIDray, demonstrates that the effectiveness of our proposed method, outperforming the existing renown CNN -based detection approaches.
Yaobin Huang, Hongxia Gao
ICIP2
2024 Apnet: Generating Precise Anomaly Prior Information for Mixed-Supervised Defect Detection
abstract
Mixed-supervised defect detection is an emerging paradigm, referring to using defect location information provided by unsupervised framework to enhance supervised defect detection. However, current research on unsupervised defect detection methods is limited to the MVTEC AD dataset and performs poorly on industrial images with simple structures. To address this challenge, we propose a general anomaly detection model called Anomaly Prior Network (APNet). APNet is composed of Repair Module and Discriminate Module, the Repair Module repairs defect texture into defect-free texture by vector quantization algorithm, and the Discriminate Module achieves precise anomaly localization by discovering the differences between original and repaired image on multi-scales features. In addition, we have introduced a large-scale mixed-supervised industrial dataset named the Inductance Core Defect (ICD) dataset, which consists of 20913 low-resolution (400×320) inductance core samples collected from a real pipeline. Extensive experiments conducted on ICD and MVTEC AD datasets verify the effectiveness of the proposed method compared to other advanced methods.
Guanji Li, Hongxia Gao
ICIP2
2024 Surface Anomaly Detection With Anomalous Feature Restriction And Difference-Aware Enhancement
abstract
In industrial automatic product quality inspection, visual anomaly detection is paramount. While unsupervised anomaly detection methods based on reconstruction have shown promising results, particularly in anomaly localization, these methods still suffer from challenges such as overfitting of pseudo-anomalous distribution by reconstruction networks and difficulty in distinguishing near-distribution anomalies by discriminative networks. In this paper, we propose a novel Anomalous Feature Restriction and Difference-Aware Enhancement Network (RE-Net), which aims to constrain abnormal features while enhancing the minute discrepancies between normal and abnormal features. This network comprises two key modules: the Abnormal Feature Restriction Module (AFRM) and the Difference-Aware Enhancement Module (DAEM). AFRM first explicitly constrains abnormal features by utilizing normal features in the reconstructed subnetwork to prevent the network from overfitting pseudo-abnormal distributions while ensuring the consistency of normal regions. Upon achieving a normal reconstruction from an anomalous input, DAEM is then used to enhance the perception of the difference between normal and abnormal in the discriminant subnetwork, thereby effectively improving the detection ability of highly camouflaged near-distribution anomalies. A series of comparative experiments on textured objects in the MVTec AD dataset show that our method achieves better anomaly detection results, reaching 99.9% image-level AUROC and 98.76% pixel-level AUROC.
Jinhui Zhao, Hongxia Gao, Tongtong Liu 0003
ICIP2
2024 Low-Light Image Enhancement via Weighted Low-Rank Tensor Regularized Retinex Model
abstract
Images captured under low light conditions are often affected by intense noise, which may become more pronounced during image enhancement, resulting in poor visual quality. The aim of this paper is to establish an effective low-light image enhancement model that can suppress noise and artifacts while preserving image details. To deal with intense noise, we propose a Weighted Low-Rank Tensor regularization Retinex (WLRT-Retinex) model, which introduces weighted low-rank tensor priors in the Retinex decomposition process to suppress noise and artifacts in the reflectance. Furthermore, since noise in dark areas is typically more severe, we introduce an illumination-aware weighting scheme in the total variation regularization term of the reflectance, which helps achieve adaptive denoising and preserve details in bright areas. Experiments on seven challenging datasets demonstrate the effectiveness of the proposed method, achieving better or comparable performance compared with state-of-the-art methods. Our code is available at https://github.com/YangWeipengscut/WLRT-Retinex.
Weipeng Yang 0002, Hongxia Gao, Wenbin Zou, Tongtong Liu 0003, Shasha Huang, Jianliang Ma
ICMR2
2024 Wave-Mamba: Wavelet State Space Model for Ultra-High-Definition Low-Light Image Enhancement
abstract
Ultra-high-definition (UHD) technology has attracted widespread attention due to its exceptional visual quality, but it also poses new challenges for low-light image enhancement (LLIE) techniques. UHD images inherently possess high computational complexity, leading existing UHD LLIE methods to employ high-magnification downsampling to reduce computational costs, which in turn results in information loss. The wavelet transform not only allows downsampling without loss of information, but also separates the image content from the noise. It enables state space models (SSMs) to avoid being affected by noise when modeling long sequences, thus making full use of the long-sequence modeling capability of SSMs. On this basis, we propose Wave-Mamba, a novel approach based on two pivotal insights derived from the wavelet domain: 1) most of the content information of an image exists in the low-frequency component, less in the high-frequency component. 2) The high-frequency component exerts a minimal influence on the outcomes of low-light enhancement. Specifically, to efficiently model global content information on UHD images, we proposed a low-frequency state space block (LFSSBlock) by improving SSMs to focus on restoring the information of low-frequency sub-bands. Moreover, we propose a high-frequency enhance block (HFEBlock) for high-frequency sub-band information, which uses the enhanced low-frequency information to correct the high-frequency information and effectively restore the correct high-frequency details. Through comprehensive evaluation, our method has demonstrated superior performance, significantly outshining current leading techniques while maintaining a more streamlined architecture. The code is available at https://github.com/AlexZou14/Wave-Mamba.
Wenbin Zou, Hongxia Gao, Weipeng Yang 0002, Tongtong Liu 0003
ACM Multimedia2
2024 114Xray: A Large-Scale X-Ray Security Detection Benchmark and Aware Enhance Network for Real-World Prohibited Item Inspection in Baggage
Hongxia Gao, Zhenming Guan, Yaobin Huang, Hongyu Liao, Hongzhen Zheng, Runze Lin, Litao Li, Haolin Tang, Guoyuan Lin, Zhanhong Chen
PRCV (11)1
2024 Employing Multiple Priors in Retinex-Based Low-Light Image Enhancement
Weipeng Yang 0002, Hongxia Gao, Tongtong Liu 0003, Jianliang Ma, Wenbin Zou, Shasha Huang
EGSR (ST)2
2024 Stereo matching from monocular images using feature consistency
abstract
Abstract Synthetic images facilitate stereo matching. However, synthetic images may suffer from image distortion, domain bias, and stereo mismatch, which would significantly restrict the widespread use of stereo matching models in the real world. The first goal in this paper is to synthesize real‐looking images for minimizing the domain bias between the synthesized and real images. For this purpose, sharpened disparity maps are produced from a mono real image. Then, stereo image pairs are synthesized using these imperfect disparity maps and the single real image in the proposed pipeline. Although the synthesized images are as realistic as possible, the domain styles of the synthesized images are always very different from the real images. Thus, the second goal is to enhance the domain generalization ability of the stereo matching network. For that, the feature extraction layer is replaced with a teacher–student model. Then, a constraint of binocular contrast features is imposed on the output of the model. When tested on the KITTI, ETH3D, and Middlebury datasets, the accuracy of the method outperforms traditional methods by at least 30%. Experiments demonstrate that the approaches are general and can be conveniently embedded into existing stereo networks.
Zhongjian Lu, Hongxia Gao, Langwen Zhang, Congyu Zhang
IET Image Process.3
2023 Joint Edge-Guided and Spectral Transformation Network for Self-supervised X-Ray Image Restoration
Shasha Huang, Wenbin Zou, Hongxia Gao, Weipeng Yang 0002, Shicheng Niu, Tian Qi, Jianliang Ma
ICANN (2)3
2023 Feature-Aware Prohibited Items Detection for X-Ray Images
abstract
Prohibited items detection is a challenging problem in security inspection. Since the commonly used manual security screening is a subjective, inefficient and costly process, some researchers have used existing models designed for natural images and their modifications to replace it. However, these methods do not take full advantage of the unique properties of X-ray images to solve its clutter and occlusion problem in detection. In this paper, we propose a feature-aware prohibited items detection (FAPID) method for X-ray images to detect normal and overlapped heavily prohibited items. The overall framework consists of shape-guided feature enhancement module (SGFE) and prohibited item aware module (PIA). The SGFE module enhances the items’ structural integrity in the learned features by utilizing the strong shapes in X-ray images without losing texture features. And the PIA module learns discriminative fine contour and texture of prohibited items without the influence of overlapping noise by using frequency information and the feedback-like mechanism. Extensive experiments conducted on HiXray and OPIXray datasets verify the effectiveness of the proposed method compared to other SOTA methods.
Hongyu Liao, Hongxia Gao
ICIP3
2023 Joint Priors-Based Restoration Method for Degraded Images Under Medium Propagation
Wenbin Zou, Hongxia Gao, Weipeng Yang 0002, Shasha Huang, Jianliang Ma
PRCV (11)3
2023 Enhancing Low-Light Images: A Variation-based Retinex with Modified Bilateral Total Variation and Tensor Sparse Coding
abstract
Abstract Low‐light conditions often result in the presence of significant noise and artifacts in captured images, which can be further exacerbated during the image enhancement process, leading to a decrease in visual quality. This paper aims to present an effective low‐light image enhancement model based on the variation Retinex model that successfully suppresses noise and artifacts while preserving image details. To achieve this, we propose a modified Bilateral Total Variation to better smooth out fine textures in the illuminance component while maintaining weak structures. Additionally, tensor sparse coding is employed as a regularization term to remove noise and artifacts from the reflectance component. Experimental results on extensive and challenging datasets demonstrate the effectiveness of the proposed method, exhibiting superior or comparable performance compared to state‐of‐the‐art approaches. Code, dataset and experimental results are available at https://github.com/YangWeipengscut/BTRetinex .
Weipeng Yang 0002, Hongxia Gao, Wenbin Zou, Shasha Huang, Jianliang Ma
Comput. Graph. Forum2
2023 A Survey on Cyber-Physical Systems Security
abstract
Cyber–physical systems (CPSs) are new types of intelligent systems that integrate computing, control, and communication technologies, bridging the cyberspace and physical world. These systems enhance the capabilities of our critical infrastructure and are widely used in a variety of safety-critical systems. CPSs are susceptible to cyber attacks due to their vulnerabilities such that their security has become a critical issue. Therefore, it is important to classify and comprehensively investigate this issue. Most of the existing surveys on it are conducted from a single perspective. In this article, we present a comprehensive view of the security of CPSs from three perspectives: 1) the physical domain; 2) the cyber domain; and 3) the cyber–physical domain. In the physical domain, we review some attacks that directly damage the physical components of CPSs such as sensors and discuss corresponding defenses. We also review the attacks that CPSs in the cyber domain may face and study methods to detect and defend against them. In addition, we survey the intelligent attacks faced by CPSs and the corresponding defensive means. In the cyber–physical domain, we provide an overview of attacks that come from the cyber domain and eventually damage the physical parts, and discuss the corresponding detection and defense methods. Finally, we present the challenges and future research directions. Through this in-depth review, we attempt to summarize the current security threats to CPSs and the state-of-the-art security means to provide researchers with a comprehensive overview.
Zhenhua Yu 0001, Hongxia Gao, Xuya Cong, Houbing Song
IEEE Internet Things J.2
2023 Trustworthiness analysis and evaluation for command and control cyber-physical systems using generalized stochastic Petri nets
Zhenhua Yu 0001, Hongxia Gao
Inf. Sci.3
2023 Learning Relative Feature Displacement for Few-Shot Open-Set Recognition
abstract
Few-shot learning (FSL) usually assumes that the query is drawn from the same label space as the support set, while queries from unknown classes may emerge unexpectedly in many open-world application scenarios. Such an open-set issue will limit the practical deployment of FSL systems, which remains largely unexplored. In this paper, we investigate the problem of few-shot open-set recognition (FSOR) and propose a novel solution, called Relative Feature Displacement Network (RFDNet), which empowers FSL systems to reject queries from unknown classes while accurately classifying those from known classes. First, we suggest a different relative feature displacement learning (RFDL) paradigm for FSOR, i.e., meta-learning a feature displacement relative to a pretrained reference feature embedding, based on our insightful observations on the randomness drift issue of previous meta-learning based for FSOR methods, as well as the generalization ability of the feature embedding pretrained for general classification. Second, we design the RFDNet framework to implement the RFDL paradigm, which is mainly featured by a task-aware RFD generator and a marginal open-set loss. Comprehensive experiments on three public datasets, i.e., miniImageNet, CIFAR-FS and tieredImageNet, demonstrate that RFDNet can consistently outperform the state-of-the-art methods, achieving improvement of 5.2%, 2.0% and 1.7% respectively, in terms of AUROC for unknown-class rejection under the 5-way 5-shot setting.
Shule Deng, Jin-Gang Yu, Zihao Wu 0004, Hongxia Gao, Yansheng Li 0001, Yang Yang 0066
IEEE Trans. Multim.4
2022 A multi-stage restoration method for degraded images with light scattering and absorption
abstract
Degraded images with light scattering and absorption such as haze, underwater and sandstorm images always suffer from low contrast, detail loss, and color distortion. To solve these problems, herein, we propose a multi-stage image restoration method which includes three steps: dehazing, detail preserving and color correction. Dehazing includes estimations of ambient light and transmission map. Ambient light is estimated by fusing scene depth map and high illumination area, and transmission map is modified by adaptive linear transformation. Then, an objective function with Relative Total Variation (RTV) and edge preservation component is proposed to preserve details and suppress noise. Finally, an effective color correction method is introduced to correct color distortion and avoid over-correction. Experiments on haze, sandstorm and underwater images demonstrate that our method can obtain high quality results with high contrast, clear visibility, and natural color. Additional experiments suggest that our method can also enhance low-light images.
Ye Cai 0005, Hongxia Gao, Shicheng Niu, Tian Qi, Weipeng Yang 0002
ICPR2
2020 Feature Selection and Classification of Texture Images Based on Local Structure and Low-Rank Constraints
Rihong Li, Hongxia Gao, Jiaxiang Luo, Haiming Liu 0003, Weipeng Yang 0002
PRCV (2)2
2020 Image super-resolution based on conditional generative adversarial network
abstract
Generative adversarial network (GAN) is one of the most prevalent generative models that can synthesise realistic high‐frequency details. However, a mismatch between the input and the output may arise when GAN is directly applied to image super‐resolution. To alleviate this issue, the authors adopted a conditional GAN (cGAN) in this study. The cGAN discriminator attempted to guess whether the unknown high‐resolution (HR) image was produced by the generator with the aid of the original low‐resolution (LR) image. They propose a novel discriminator that only penalises at the scale of the patch and, thus, has relatively few parameters to train. The generator of cGAN is an encoder–decoder with skip connections to shuttle the shared low‐level information directly across the network. To better maintain the low‐frequency information and recover the high‐frequency information, they designed a generator loss function combining adversarial loss term and L1 loss term. The former term is beneficial to the synthesis of fine‐grained textures, while the latter is responsible for learning the overall structure of the LR input. The experiments revealed that the proposed method could generate HR images with richer details and less over‐smoothness.
Hongxia Gao, Zhanhong Chen, Binyang Huang, Zhifu Li
IET Image Process.1
2020 Exemplar-Based Recursive Instance Segmentation With Application to Plant Image Analysis
abstract
Instance segmentation is a challenging computer vision problem which lies at the intersection of object detection and semantic segmentation. Motivated by plant image analysis in the context of plant phenotyping, a recently emerging application field of computer vision, this paper presents the Exemplar-Based Recursive Instance Segmentation (ERIS) framework. A three-layer probabilistic model is firstly introduced to jointly represent hypotheses, voting elements, instance labels and their connections. Afterwards, a recursive optimization algorithm is developed to infer the maximum a posteriori (MAP) solution, which handles one instance at a time by alternating among the three steps of detection, segmentation and update. The proposed ERIS framework departs from previous works mainly in two respects. First, it is exemplar-based and model-free, which can achieve instance-level segmentation of a specific object class given only a handful of (typically less than 10) annotated exemplars. Such a merit enables its use in case that no massive manually-labeled data is available for training strong classification models, as required by most existing methods. Second, instead of attempting to infer the solution in a single shot, which suffers from extremely high computational complexity, our recursive optimization strategy allows for reasonably efficient MAP-inference in full hypothesis space. The ERIS framework is substantialized for the specific application of plant leaf segmentation in this work. Experiments are conducted on public benchmarks to demonstrate the superiority of our method in both effectiveness and efficiency in comparison with the state-of-the-art.
Jin-Gang Yu, Yansheng Li 0001, Changxin Gao, Hongxia Gao, Gui-Song Xia, Zhu Liang Yu, Yuanqing Li 0001
IEEE Trans. Image Process.4
2018 Deep Pixel Probabilistic Model for Super Resolution Based on Human Visual Saliency Mechanism
abstract
This work explores super resolution (SR) with a deep network based on a pixel probabilistic model, where particular small inputs and large magnification factors make the problem highly underspecified since fairly large amounts of high-frequency details are missing in low resolution (LR) source. In this paper, we develop a deep architecture comprising of a PixelCNN and a residual network (ResNet), in which PixelCNN predicts the serial dependencies of the pixel sequence and ResNet for capturing the global structure of LR input. A human visual saliency mechanism (HVSM) by employing accurate SR in salient regions and fast interpolation in nonsalient regions is integrated within the pixel probabilistic model to efficiently reduce the computational complexity while maintaining the desired visual quality. Additionally, we present a Bayesian optimization technique to automatically determine the optimal weight of loss function. Furthermore, a modified image quality assessment taking into account HVSM is introduced, trying to align with the human visual perception. Experiments demonstrate that the proposed algorithm could generate more plausible facial features than previous deep learning methods, offering finer details and significant improvement in visual quality.
Hongxia Gao, Zhanhong Chen, Ge Ma, Wang Xie, Zhifu Li
ICPR1
2018 Image Registration Based on Patch Matching Using a Novel Convolutional Descriptor
Wang Xie, Hongxia Gao, Zhanhong Chen
PRCV (2)2
2017 A stochastic iterative evolution CT reconstruction algorithm for limited-angle sparse projection data
abstract
Subject to data acquisition time, X-ray dose and geometric position of medical tomography system, only limited-angle or sparse projection data can be collected in many applications of medical computed tomography (CT). Lacking in completeness and symmetry of limited-angle sparse projection data leads to a huge search space for reconstruction algorithms. The conventional convex optimization methods typically suffer from poor convergence and unwanted local optima caused by the data insufficiency and the fixed descent path of iteration. To improve the reconstruction quality, a stochastic iterative evolution CT reconstruction algorithm is proposed, in which a population of solutions based on stochastic strategy is adopted to enhance the global search capability and the search direction is generated by combining gradient descent with fitness evaluation of stochastic population. Meanwhile, Markov Chain is introduced to predict the iterative evolution model and accelerate the proposed algorithm's convergence. Experiments results demonstrate the proposed algorithm's effectiveness and robustness in image reconstruction from limited-angle sparse projection data.
Hongxia Gao, Yinghao Luo, Yongfei Chen
BIBM2
2015 An image reconstruction model and hybrid algorithm for limited-angle projection data
abstract
More and more applications in industrial or medical CT have to adopt limited-angle projection data, such as restricted scanning, decreasing radiation dose and so on. Traditional reconstruction methods for incomplete projection data, including analytic reconstruction and iterative reconstruction, can't meet the demand of limited-angle reconstruction. To deal with TV reconstruction model's imbalance in reconstructing image information of different direction, this paper proposes a reconstruction model and its corresponding hybrid algorithm, which utilize horizontal direction gradient to assist TV reconstruction in vertical direction. Experiments results demonstrate the proposed method's effectiveness and performance in image reconstruction from limited-angle projection data.
Hongxia Gao, Yinghao Luo, Ge Ma, Lixuan Wu
BIBM1