Minglong Xue

dblp:207/8876 · DBLP profile ↗
← Back
29ranked-venue papers
14as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Better utilization of illumination prior via KANs for nighttime flare removal
Aoxiang Ning, Minglong Xue, Senming Zhong, Palaiahnakote Shivakumara, Mingliang Zhou 0001
Neural Networks2
2026 UR2P-Dehaze: Learning a Simple Image Dehaze Enhancer via Unpaired Rich Physical Prior
Minglong Xue, Shuaibin Fan, Palaiahnakote Shivakumara, Mingliang Zhou 0001
Pattern Recognit.1
2026 Unified image restoration and enhancement: Degradation calibrated cycle reconstruction diffusion model
Minglong Xue, Jinhong He, Palaiahnakote Shivakumara, Mingliang Zhou 0001
Pattern Recognit.1
2026 Nighttime flare removal via frequency decoupling
Minglong Xue, Aoxiang Ning, Jinhong He, Shuaibin Fan, Senming Zhong
Pattern Recognit. Lett.1
2026 DFCCNet: Unified Dual-domain Fusion and Color-aware Residual Correction for Robust Single Image Dehazing
Wenchao Yan, Minglong Xue, Palaiahnakote Shivakumara, Mingliang Zhou 0001
Signal Process.2
2026 Optimizing a 4D Lookup Table for Low-Light Video Enhancement via Wavelet Priori
abstract
Low-light video enhancement is highly demanding in maintaining spatiotemporal color consistency. Therefore, improving the accuracy of color mapping and keeping the latency low are challenging. On this basis, we propose incorporating wavelet-priori for the 4D lookup table (WaveLUT), which effectively enhances the color coherence between video frames and the accuracy of color mapping while maintaining low latency. Specifically, we use the wavelet low-frequency domain to construct an optimized lookup prior and achieve an adaptive enhancement effect through a designed wavelet-prior 4D lookup table. To effectively compensate for the a priori loss in the low light region, we further explore a dynamic fusion strategy that adaptively determines the spatial weights on the basis of the correlation between the wavelet lighting prior and the target intensity structure. In addition, during the training phase, we devise a Fourier-text driven appearance reconstruction method that dynamically balances brightness and content through multimodal semantics-driven Fourier spectra. Extensive experiments on a wide range of benchmark datasets show that this method effectively enhances the previous method's ability to perceive the color space and achieves metric-favourable and perceptually oriented real-time enhancement while maintaining high efficiency. The code is available athttps://github.com/hejh8/WaveLUT.
Jinhong He, Minglong Xue, Wenhai Wang, Mingliang Zhou 0001
IEEE Trans. Multim.2
2026 Unsupervised Dehazing of Real-World Images with Frequency-Aware Learning
abstract
Image dehazing remains a challenging task in real-world scenarios. Unlike synthetic datasets, real-world hazy images often exhibit nonuniform atmospheric degradation, severe haze accumulation, and significant texture loss, making it difficult for existing methods to restore realistic appearances and fine details consistently. To address these challenges, we propose a frequency-aware learning-based unpaired image dehazing network for real-world hazy scenes. First, we propose a content-aware state space modeling paradigm, focus on details through high-frequency enhancement, and incorporate global contextual understanding at the encoding stage, enabling adaptive representation of complex textures. Second, to effectively handle spatially nonuniform degradation, a hazy density estimation module is designed to guide multiple expert-gated feedback units, which dynamically select feature fusion paths. Finally, we propose a contour-guided differentiable frequency domain enhancement mechanism to explicitly recover edge and texture details in degraded regions. Extensive experiments on real-world hazy datasets under unsupervised settings demonstrate that our method achieves competitive performance, validating its effectiveness and strong practical potential under complex atmospheric conditions. The code is available at https://github.com/Fan-pixel/FAL-Net .
Minglong Xue, Shuaibin Fan, Wenchao Yan, Mingliang Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2026 Dynamic nonlinear networks for adaptive low-light image enhancement
Minglong Xue, Senming Zhong
Vis. Comput.1
2026 Wavelet transform-guided transformer light transfer network for zero-shot low-light image enhancement
Minglong Xue, Xukun Shang, Jinhong He, Senming Zhong
Vis. Comput.1
2025 Degradation-Consistent Learning via Bidirectional Diffusion for Low-Light Image Enhancement
abstract
Low-light image enhancement aims to improve the visibility of degraded images to better align with human visual perception. While diffusion-based methods have shown promising performance due to their strong generative capabilities. However, their unidirectional modelling of degradation often struggles to capture the complexity of real-world degradation patterns, leading to structural inconsistencies and pixel misalignments. To address these challenges, we propose a bidirectional diffusion optimization mechanism that jointly models the degradation processes of both low-light and normal-light images, enabling more precise degradation parameter matching and enhancing generation quality. Specifically, we perform bidirectional diffusion-from low-to-normal light and from normal-to-low light during training and introduce an adaptive feature interaction block (AFI) to refine feature representation. By leveraging the complementarity between these two paths, our approach imposes an implicit symmetry constraint on illumination attenuation and noise distribution, facilitating consistent degradation learning and improving the model's ability to perceive illumination and detail degradation. Additionally, we design a reflection-aware correction module (RACM) to guide color restoration post-denoising and suppress overexposed regions, ensuring content consistency and generating high-quality images that align with human visual perception. Extensive experiments on multiple benchmark datasets demonstrate that our method outperforms state-of-the-art methods in both quantitative and qualitative evaluations while generalizing effectively to diverse degradation scenarios.Code
Jinhong He, Minglong Xue, Zhipu Liu, Mingliang Zhou 0001, Aoxiang Ning, Palaiahnakote Shivakumara
ACM Multimedia2
2025 Artistic-style text detector and a new Movie-Poster dataset
Aoxiang Ning, Minglong Xue, Yiting Wei, Mingliang Zhou 0001, Senming Zhong
Expert Syst. Appl.2
2025 SC-BSN: Shifted Convolutions Based Blind-Spot Network for self-supervised image denoising
Guo Yang, Chengyun Song, Minglong Xue, Jian Yu 0002
Neurocomputing3
2025 Hybrid-Domain Attention Dense Network for Efficient Image Super-Resolution
abstract
Efficient Image Super Resolution (EISR) techniques are critical to meeting the real-time demands of resource-constrained devices. Current approaches focus too much on parameter reduction, thus slightly sacrificing image quality. To alleviate this issue, we propose a novel Hybrid Attention Dense Network (HADN) in this paper. HADN fully explores how to balance model performance and image recovery quality better. It incorporates two key designs: (1) We apply Hybrid-domain Attention Dense Blocks (HADB) with different feature layers cumulatively connected to enhance the representation of shallow features. (2) We designed a novel Hybrid-domain Attention Block (HAB) to reinforce the correlation between image edges and receptive fields for effective information fusion to improve the perception of detailed image features. Extensive experimental analyses affirm the competitiveness of the proposed method, showcasing its ability to balance performance and recovery effectively, even with smaller training datasets. The code is available at https://github.com/Yuii666/HADN .
Yanyi He, Jinhong He, Minglong Xue, Senming Zhong, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.3
2025 Addressing domain discrepancy: A dual-branch collaborative model to unsupervised dehazing
Shuaibin Fan, Minglong Xue, Aoxiang Ning, Senming Zhong
Pattern Recognit. Lett.2
2025 Zero-Shot Low-Light Image Enhancement via Joint Frequency Domain Priors Guided Diffusion
abstract
Due to the singularity of real-world paired datasets and the complexity of low-light environments, this leads to supervised methods lacking a degree of scene generalisation. Meanwhile, limited by poor lighting and content guidance, existing zero-shot methods cannot handle unknown severe degradation well. To address this problem, we will propose a new zero-shot low-light enhancement method to compensate for the lack of light and structural information in the diffusion sampling process by effectively combining the wavelet and Fourier frequency domains to construct rich a priori information. The key to the inspiration comes from the similarity between the wavelet and Fourier frequency domains: both light and structure information are closely related to specific frequency domain regions, respectively. Therefore, by transferring the diffusion process to the wavelet low-frequency domain and combining the wavelet and Fourier frequency domains by continuously decomposing them in the inverse process, the constructed rich illumination prior is utilised to guide the image generation enhancement process. Sufficient experiments show that the framework is robust and effective in various scenarios.
Jinhong He, Palaiahnakote Shivakumara, Aoxiang Ning, Minglong Xue
IEEE Signal Process. Lett.4
2025 KAN See in the Dark
abstract
Low-lightimage enhancement methods are difficult to fit the complex nonlinear relationship between normal and low-light images due to uneven illumination and noise effects. The recently proposed Kolmogorov-Arnold networks (KANs) feature spline-based convolutional layers and learnable activation functions, which can effectively capture nonlinear dependencies. In this paper, we design a KAN-Block based on KANs and innovatively apply it to low-light image enhancement. This method effectively alleviates the limitations of current methods constrained by linear network structures and lack of interpretability, further demonstrating the potential of KANs in low-level vision tasks. Given the poor perception of current low-light image enhancement methods and the stochastic nature of the inverse diffusion process, we further introduce frequency-domain perception for visually oriented enhancement. Extensive experiments demonstrate the competitive performance of our method on benchmark datasets.
Aoxiang Ning, Minglong Xue, Jinhong He, Chengyun Song
IEEE Signal Process. Lett.2
2025 Low-Light Image Enhancement via CLIP-Fourier Guided Wavelet Diffusion
abstract
Low-light image enhancement techniques have significantly progressed, but unstable image quality recovery and unsatisfactory visual perception are still significant challenges. To solve these problems, we propose a novel and robust low-light image enhancement method via CLIP-Fourier guided wavelet diffusion, abbreviated as CFWD. Specifically, the CFWD leverages multimodal visual-language information in the frequency domain space created by multiple wavelet transforms to guide the enhancement process. Multiscale supervision across different modalities facilitates the alignment of image features with semantic features during the wavelet diffusion process, effectively bridging the gap between the degraded and normal domains. Moreover, to further promote the effective recovery of the image details, we combine the Fourier transform based on the wavelet transform and construct a hybrid high-frequency perception module (HFPM) with a significant perception of the detailed features. This module avoids the diversity confusion of the wavelet diffusion process by guiding the fine-grained structure recovery of the enhancement results to achieve favourable metrics and perceptually oriented enhancement. Extensive quantitative and qualitative experiments on publicly available real-world benchmarks show that our approach outperforms existing state-of-the-art methods, achieving significant progress in image quality and noise suppression. The project code is available at https://github.com/hejh8/CFWD .
Minglong Xue, Jinhong He, Wenhai Wang, Mingliang Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2025 A novel Half-To-All MOTR approach for robust video text tracking with incomplete annotations
Peiqi Xie, Minglong Xue, Chengyun Song
Vis. Comput.2
2024 Zero-Reference Lighting Estimation Diffusion Model for Low-Light Image Enhancement
Jinhong He, Minglong Xue, Aoxiang Ning, Chengyun Song
ACML2
2024 A Novel Encoder-Decoder Network with Multi-domain Information Fusion for Video Deblurring
Peiqi Xie, Jinhong He, Chengyun Song, Minglong Xue
ICPR (32)4
2024 DLDiff: Image Detail-Guided Latent Diffusion Model for Low-Light Image Enhancement
abstract
Low-light image enhancement is an essential task in image restoration. Inspired by the diffusion model, the related methods have achieved remarkable results in low-level visual tasks. However, such methods are susceptible to large-scale images, generating problems such as overconsumption of resources and low recovery efficiency. To address this, we propose a detail-guided latent space low-light image enhancement diffusion model called DLDiff. Leveraging the generative power of the latent diffusion model, we explore ways to speed up inference better while producing excellent perceptual fidelity. Specifically, we initially employ the latent diffusion model to transform low-light image features into a latent space representation, thereby reducing computational resource consumption. Next, we design a lightweight detail prompt module that combines cross-convolution and vast-receptive-field convolution blocks. This module enhances the fine-grained details of the image, effectively supplements multiscale feature information, and minimizes feature loss in the latent space. Furthermore, we devise the content-aware loss group to facilitate learning noise and image information, enhancing the model's recovery capability, guiding stable sampling, and constraining diverse content generation. Through extensive experiments, we demonstrate the model's significant efficiency and quality advantages in low-light image enhancement tasks.
Minglong Xue, Yanyi He, Jinhong He, Senming Zhong
IEEE Signal Process. Lett.1
2023 UT-GAN: A Novel Unpaired Textual-Attention Generative Adversarial Network for Low-Light Text Image Enhancement
abstract
How to balance lighting and texture details to achieve the desired visual effect remains the bottleneck of existing low-light image enhancement methods. In this paper, we propose a novel Unpaired Textual-attention Generative Adversarial N network (UT-GAN) for low-light text image enhancement task. UT-GAN first uses the Zero-DCE net for initial illumination recovery and our TAM module is proposed to translate text information into a textual attention mechanism for the overall network, emphasizing attention to the details of text regions. Moreover, the method constructs an AGM-Net module to mitigate noise effects and fine-tune the illumination. Experiments show that UT-GAN outperforms existing methods in qualitative and quantitative evaluation on the widely used the low-light datasets LOL and SID.
Minglong Xue, Zhengyang He, Yanyi He, Peiqi Xie
ICIP1
2023 A Geometry-Aware Consistent Constraint for Height Estimation From a Single SAR Imagery in Mountain Areas
abstract
Height estimation from a single synthetic aperture radar (SAR) imagery has shown a great potential in real-time scene understanding and environment detection. It is a mathematical ill-posed problem for that a single 2-D image may be projected from multiple 3-D scenes. Then, the problem is that the accuracy of height estimation from a single SAR image is not high enough without prior knowledge, especially in mountain areas. Thus, we propose a geometry-aware consistent constraint (GACC) to improve the performance of single SAR height estimation. The simulated SAR images generated from height maps are used to establish high-precision transformation between map and radar coordinates in mountain areas. The main idea of GACC is that the simulated SAR image generated from estimated height map should be consistent with the ground-truth simulated SAR image. A sparse height (SH) map is included as the supplementary inputs to further improve the accuracy of estimated height maps. Comparison experiments are performed in Geermu, Huangshan, and Guiyang datasets, and the results show that the root-mean-square error (RMSE) of the estimated height map by U-shaped convolution neural network (Unet) with GACC and 0.0434% SH gets reduced by about 98%, 78%, and 94% compared with that by Unet in the three datasets.
Minglong Xue, Jian Li 0042, Qingli Luo
IEEE Geosci. Remote. Sens. Lett.1
2022 PSND: A Robust Parking Space Number Detector
abstract
It is necessary to detect the license plate and the parking space number in autonomous driving. Due to the complex environment of the parking space, it is challenging and interesting in parking space number detection. This paper proposes a robust Parking Space Number Detector (PSND), which enhances the Differentiable Binarization (DB) model by introducing a cascaded feature enhancement module and context attention block. Our method has better feature extraction ability and a better detection effect for long text than the DB model. Meanwhile, we collected and annotated a dataset containing 9000 parking space number images to train and test our model. Extensive experiments demonstrate that our method achieves better or competitive performance in terms of accuracy on various standard benchmarks, including MSRA-TD500, ICDAR2015, CTW1500, and Total-Text datasets while maintaining the real-time detection speed. The high recognition accuracy makes it possible for the model to be practically applied.
Chengyun Song, Minglong Xue
ICPR3
2022 Toward Optimal Learning Rate Schedule in Scene Classification Network
abstract
Stochastic gradient descent (SGD) and adaptive methods, including ADAM, RMSProp, AdaDelta, and AdaGrad, are two dominant optimization algorithms for training convolution neural network (CNN) in scene classification tasks. Recent work reveals that these adaptive methods lead to degraded performance in image classification tasks compared with SGD. In this letter, a learning rate schedule named switching from constant to step decay (SCTSD) is proposed to further improve the classification accuracy of SGD. SCTSD begins with a constant learning rate and switches to step decay learning rates when appropriate. Theoretical evidence is provided on the superiority of SCTSD compared with other manually tuned schedules. Comparison experiments have been conducted among adaptive methods, SCTSD, and other manual schedules with three CNN architectures. The experiment results on various scene classification data sets show that SCTSD has the highest accuracy on the test set and it is state of the art. In the end, some suggestions on hyperparameters selection of SCTSD are given for scene classification.
Minglong Xue, Jian Li 0042, Qingli Luo
IEEE Geosci. Remote. Sens. Lett.1
2021 A Novel Attention Enhanced Residual-In-Residual Dense Network for Text Image Super-Resolution
abstract
Natural scene text images captured by handheld devices usually cause low resolution (LR) problems, thus making sub-sequent detection and recognition tasks more challenging. To address this problem, LR text images are generally super-resolution (SR) processed first. In this paper, we propose a novel low-resolution text image super-resolution method. This method adopts the residual-in-residual dense network (RRDN) to extract deeper high-frequency features than the residual dense network (RDN). Then, enhances the spatial and channel features with an attention mechanism. According to the characteristics of the text, we added gradient loss to adversarial learning. Experiments show that our method performs well in both qualitative and quantitative aspects of the latest public text image super-resolution dataset. Similarly, the proposed super-resolution method for text images of natural scenes also achieves the latest results.
Minglong Xue, Zhiheng Huang, Ruo-Ze Liu
ICME1
2021 Arbitrarily-Oriented Text Detection in Low Light Natural Scene Images
abstract
Text detection in low light natural scene images is challenging due to poor image quality and low contrast. Unlike most existing methods that focus on well-lit (normally daylight) images, the proposed method considers much darker natural scene images. For this task, our method first integrates spatial and frequency domain features through fusion to enhance fine details in the image. Next, we use Maximally Stable Extremal Regions (MSER) for detecting text candidates from the enhanced images. We then introduce Cloud of Line Distribution (COLD) features, which capture the distribution of pixels of text candidates in the polar domain. The extracted features are sent to a Convolution Neural Network (CNN) to correct the bounding boxes for arbitrarily oriented text lines by removing false positives. Experiments are conducted on a dataset of low light images to evaluate the proposed enhancement step. The results show our approach is more effective compared to existing methods in terms of standard quality measures, namely, BRISQE, NIQE and PIQE. In addition, experimental results on a variety of standard benchmark datasets, namely, ICDAR 2013, ICDAR 2015, SVT, Total-Text, ICDAR 2017-MLT and CTW1500, show that the proposed approach not only produces better results for low light images, at the same time it is also competitive for daylight images.
Minglong Xue, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti, Zhibo Yang 0003
IEEE Trans. Multim.1
2019 A Text-Context-Aware CNN Network for Multi-oriented and Multi-language Scene Text Detection
abstract
The existing deep learning based state-of-theart scene text detection methods treat scene texts a type of general objects, or segment text regions directly. The latter category achieves remarkable detection results on arbitrary orientation and large aspect ratios of scene texts based on instance segmentation algorithms. However, due to the lack of context information with consideration of scene text unique characteristics, directly applying instance segmentation to text detection task is prone to result in low accuracy, especially producing false positive detection results. To ease this problem, we propose a novel text-context-aware scene text detection CNN structure, which appropriately encodes channel and spatial attention information to construct context-aware and discriminative feature map for multi-oriented and multi-language text detection tasks. With high representation ability of text context-aware feature map, the proposed instance segmentation based method can not only robustly detect multi-oriented and multi-language text from natural scene images, but also produce better text detection results by greatly reducing false positives. Experiments on ICDAR2015 and ICDAR2017-MLT datasets show that the proposed method has achieved superior performances in precision, recall and F-measure than most of the existing studies.
Minglong Xue, Tong Lu 0002, Yirui Wu, Palaiahnakote Shivakumara
ICDAR2
2019 Curved text detection in blurred/non-blurred video/scene images
Minglong Xue, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001
Multim. Tools Appl.1