Min Gan

dblp:71/1066 · DBLP profile ↗
← Back
75ranked-venue papers
16as first author
59since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 11 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 2 first-author · 23 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Partial label zero-shot learning with semantic mining and instance-label alignment
Jinfu Fan, Linqing Huang, Min Gan, C. L. Philip Chen
Eng. Appl. Artif. Intell.4
2026 Adaptive load forecasting under regional distribution shifts: A meta-learning framework
Hongxia Zhou, Liwen Tang, Feiyan Chen, Guang-Yong Chen, Min Gan
Eng. Appl. Artif. Intell.6
2026 ICH-ASNet: Automatic Prompt-based segmentation for intracranial hemorrhage in CT images
Tianzong Nie, Guang-Yong Chen, Min Gan
Expert Syst. Appl.3
2026 Learning view-adaptive implicit regularization for robust multi-view subspace clustering
Guang-Yong Chen, Hui-Lang Xu, Min Gan
Neurocomputing4
2026 WaveGateNet: A wavelet-guided gated network for frequency-spatial collaborative underwater image enhancement
Shaotuo Zhang, Zhaolong Gao, Jingchun Zhou, Min Gan, Jinjiang Li 0001
Neurocomputing5
2026 Hierarchical Dynamic Self-Supervised Learning for Robust Deep Non-negative Matrix Factorization in clustering tasks
Dengxiu Yu, Guang-Yong Chen, Shuqiang Wang, Min Gan
Pattern Recognit.5
2026 SAFAformer: Scale-Aware Frequency-Adaptive Guidance for Nighttime Flare Removal
abstract
Nighttime flare removal is challenging due to the difficulty of acquiring real-world paired data. Existing methods, trained on synthetic pipelines, often struggle to generalize to real-world scenarios. A key limitation of these pipelines is their focus on single-flare scenes, whereas real-world conditions frequently involve more complex cases, such as multi-flare and composite flare scenarios, which are difficult to simulate effectively. This discrepancy significantly hampers model performance in practical applications. Through detailed analysis, we uncover a fundamental characteristic of flare degradation: regardless of whether the scene is synthetic single-flare, real-world single-flare, or multi-flare, the degradation information exhibits a similar distribution across frequency subbands—predominantly concentrated in the low-frequency region, with a minor presence in the high-frequency region. Notably, the severity of the glare effect correlates with an even stronger concentration in the low-frequency domain. This finding suggests that targeted frequency modeling can bridge the gap between synthetic and real-world domains, forming a principled approach to improving generalization. Building on this insight, we propose the Scale-Aware Frequency-Adaptive Guidance Network for Nighttime Flare Removal (SAFAformer), which integrates a Frequency-Adaptive Guidance Module (FAGM) and a Scale-Aware Transformer Block (SATB) to leverage frequency-domain properties during training. Extensive experiments demonstrate that SAFAformer achieves state-of-the-art performance in flare removal compared to existing methods. Our code and pre-trained models are available on GitHub for validation.
Fan Zhang 0045, Min Gan, Guang-Yong Chen, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.4
2026 Partial Label Learning via Mutual Information Representation Learning
abstract
Partial label learning (PLL) is a paradigm in weakly supervised learning. The goal is to identify the ground-truth label from a set of candidate labels associated with a given sample. However, due to the ambiguity of labels, improving the accuracy of ground-truth label recognition is a challenge. In this paper, we propose an innovative training framework PLMR, for solving the PLL problem. Given the specificity of the PLL problem, its data is rich in valid information but significantly noisy. To overcome the problem, PLMR incorporates mutual information (MI) theory to mine the potential information of candidate labels and data features to distinguish between positive and negative sample pairs. This operation is to effectively utilize the raw data information in PLL and reduce the class conflict problem. In this way, the discriminative power of representation is improved according to the positive and negative pair selection strategy. At the same time, PLMR introduces cluster centers to optimize subsequent tasks, combining data augmentation samples with the K-means cross-attention mechanism to refine the optimized cluster centers. This is to improve the ability of the clustering centre to accurately represent the class information, thus improving the overall quality of the clusters. Further, through the ambiguity marking correction mechanism, weights are calculated based on the association between cluster centers and sample representations to guide model training. Experimental results show that PLMR demonstrates excellent classification performance on multiple datasets, verifying its effectiveness and sophistication.
Linqing Huang, JiangNan Li, Kangrui Ren, Jinfu Fan, QingKai Bu, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.7
2026 Dual Guidance of Visual and Semantic Information for Real-World Scene Text Image Super-Resolution: A Novel Approach and Benchmark Dataset
abstract
Scene Text Image Super-Resolution (STISR) methods improve recognition accuracy by refining text regions locally. However, most existing approaches are limited by short-range dependencies, hindering the modeling of long-range semantic relationships between characters, which reduces their effectiveness in complex text scenes. Moreover, current methods are predominantly evaluated on synthetic datasets, which do not adequately capture real-world challenges such as diverse text styles, complex backgrounds, and spatial distortions. To address these limitations, we propose a Dual-Guided Visual and Semantic (DGVS) framework. This innovative framework utilizes a recognizer-driven attention mechanism to decouple character sequences, effectively distinguishing text from background noise. Additionally, we incorporate a state-space model to establish global semantic reasoning links, enhancing the comprehension of contextual relationships within text images. Furthermore, we construct Real-World Text (RealWT), a novel real-world benchmark dataset that integrates diverse data sources, including online images and multi-device captures. This dataset incorporates factors like device variations, resolution differences, and degradation, offering a more realistic simulation of real-world conditions. It enables models to learn degradation patterns that closely mirror practical applications, offering a standardized benchmark for evaluating STISR performance. Extensive experiments demonstrate that our method outperforms existing approaches, as validated by evaluations on TextZoom and RealWT. Our dataset and code are available on https://github.com/yoursmith/sde-DGVS.
Rui-Lin Shi, Zishu Yao, Feiyan Chen, Guang-Yong Chen, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.5
2026 ASCFormer: An Adaptive Structure-Aware Cascaded Transformer for 3D Object Detection
abstract
3D object detection has achieved significant progress in outdoor LiDAR point clouds, however, the inherent irregularity and varying sparsity distribution of point occupancy present a key challenge. Existing transformer-based 3D detectors often treat all tokens within the attention window as equally important, regardless of varying sparsity, which not only fails to address the disparities between the varying beam densities but also results in increased memory and computational costs. In this work, we propose an adaptive structure-aware cascaded transformer (ASCFormer) that dynamically captures density-insensitive multiscale structure features to model long-range dependencies via cascaded learning. Our ASCFormer detector includes an adaptive structure-aware token learning module that embeds voxel-level foreground probability and grid-level local density into the grid tokens to enhance structural perception capability. Moreover, we integrate these factors to compute significance scores, which are then utilized in inverse transform sampling to select a subset of multiscale tokens with varying receptive field sizes. To improve the training convergence of the window-based transformer in 3D voxel space, we employ cascaded learning via cross-stage attention to enhance the feature representation capability and refine the localization precision of 3D bounding boxes. This design of structure-aware reweighting effectively enhances the cascade paradigm, making to more adaptable to the varying sparsity distribution of point clouds. Extensive experiments on the KITTI and Waymo Open datasets demonstrate that the proposed ASCFormer detector achieves exceptional performance compared with state-of-the-art 3D object detection methods. The source code is publicly available at https://github.com/Xinglong-Li1/ASCFormer.
Xiaowei Zhang 0003, Xinglong Li, Mingliang Zhou 0001, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.4
2026 Riemannian Acceleration for Sparse PCA With Separable Structure and Second-Order Information Exploration
abstract
Sparse Principal Component Analysis (SPCA) is a powerful technique for dimensionality reduction and feature extraction in high-dimensional data, with applications spanning various fields such as computer vision, pattern recognition, and data mining. However, the computational intensity of SPCA presents a significant challenge, necessitating the development of efficient and robust algorithms. In this paper, we shed light on the SPCA problem and uncover intriguing structures that enable us to design an efficient algorithm, which we have named SPCA_ACC. Firstly, we identify a separable structure in this problem, which prompts us to draw on the Variable Projection (VP) strategy and generalize it to separable nonlinear problem in Stiefel manifold. This strategy projects out part of the parameters to obtain a reduced problems, allowing the SPCA_ACC algorithm to optimize in a lower-dimensional parameter space. Secondly, we resolve the coupling between different parameters of the SPCA problem in the optimization process on a fixed coordinate-sparsity manifold, which opens the way to the use of second-order Riemannian accelerated VP strategy. Moreover, we systematically analyze the advantages of using VP to solve the SPCA problem from a theoretical perspective, and confirm the local quadratic convergence of our algorithm. Numerical experiments on datasets of different sizes and types demonstrate that our method achieves rapid convergence and significantly reduces computational costs.
Guang-Yong Chen, Hui-Lang Xu, Xiang-Xiang Su, Min Gan, Xing Chen 0002, C. L. Philip Chen
IEEE Trans. Image Process.4
2025 IniRetinex: Rethinking Retinex-type Low-Light Image Enhancer via Initialization Perspective
abstract
Retinex-based methods have become a general approach for solving low-light image enhancement (LLIE). However, traditional methods require post-processing of illumination (e.g., gamma correction), which lacks adaptability and disrupts the illumination structure. Retinex-based deep networks typically follow a ‘decomposition-adjustment-exposure control’ process, which is redundant and lacks robustness. One major issue is the inaccuracy in estimating and decomposing the initial illumination. Accurate initial illumination can prevent further post-processing instability. We propose IniRetinex, rethinking the Retinex-based LLIE method from the perspective of initialization. By using neural networks to provide reasonable initial illumination and solving for smooth illumination through optimization, higher performance LLIE is achieved. We construct a two-layer convolutional neural network to capture the low-frequency structure of the image, adaptively compensating for classical initial illumination and avoiding additional post-processing. The network requires no pre-training and can be implemented in an unsupervised manner with just a few iterations, making it highly efficient. Additionally, we propose a new illumination optimization strategy by introducing an additional proximal penalty term, improving illumination in areas with varying levels and enhancing image details. Extensive experiments on various low-light image datasets demonstrate that our method achieves state-of-the-art (SOTA) results on multiple benchmarks, offering higher stability and inference efficiency compared to current advanced methods.
Zishu Yao, Guang-Yong Chen, Jian-Nan Su, Min Gan
AAAI5
2025 Clinically Robust Polyp Segmentation: Enhanced Generalization and Perturbation Resistance
abstract
Colonoscopy is vital for detecting colorectal polyps, which are closely linked to colorectal cancer. Accurate segmentation of polyps in colonoscopic images is essential for diagnosis and surgical planning but is challenging due to variability in polyp size, shape, and unclear boundaries. The Segment Anything Model (SAM) has shown promise in polyp segmentation but relies heavily on user-provided prompts and involves a large number of parameters, limiting its practicality in clinical settings. To address the limitations of SAM in clinical practice, we introduced Low Rank and Perturbation Segment Anything Model (LP-SAM) to improve segmentation accuracy and generalization ability while reducing the parameter count and complexity of user input. LP-SAM showed enhanced generalization and a lightweight design, making it more suitable for clinical applications where precise user input may not always be feasible. Comparative evaluations demonstrate that LP-SAM outperforms state-of-the-art methods on datasets such as CVC-ColonDB, CVC-300, and ETIS.
Shanchuan Wang, Jian-Nan Su, Min Gan
ICASSP4
2025 Orthogonal NMF on a Higher-Order Manifold: A Unified Framework for Clustering
Guangyong Chen, Min Gan
PRCV (1)4
2025 IGSENet: A Clinically Robust Polyp Segmentation Method via Interactive Fusion and Guided-Selective Enhancement
Shanchuan Wang, Tianzong Nie, Jian-Nan Su, Min Gan
PRCV (13)4
2025 TSSA-Net: A Temporal-Spiking-Spatial-Attention Network for Frequency-Aware and Robust Time Series Forecasting
Guang-Yong Chen, Min Gan
PRICAI (5)4
2025 Illumination-aware and structure-guided transformer for low-light image enhancement
Zishu Yao, Min Gan
Comput. Vis. Image Underst.3
2025 Knowledge-prompted intracranial hemorrhage segmentation on brain computed tomography
Tianzong Nie, Feiyan Chen, Jian-Nan Su, Guang-Yong Chen, Min Gan
Expert Syst. Appl.5
2025 Adaptive decoupled strategy for robust and efficient low-rank matrix decomposition
Min Gan, Fan Zhang 0045, Xiang-Xiang Su, Guang-Yong Chen
Neurocomputing1
2025 Online Learning Under a Separable Stochastic Approximation Framework
abstract
We propose an online learning algorithm tailored for a class of machine learning models within a separable stochastic approximation framework. The central idea of our approach is to exploit the inherent separability in many models, recognizing that certain parameters are easier to optimize than others. This paper focuses on models where some parameters exhibit linear characteristics, which are common in machine learning applications. In our proposed algorithm, the linear parameters are updated using the recursive least squares (RLS) algorithm, akin to a stochastic Newton method. Subsequently, based on these updated linear parameters, the nonlinear parameters are adjusted using the stochastic gradient method (SGD). This dual-update mechanism can be viewed as a stochastic approximation variant of block coordinate gradient descent, where one subset of parameters is optimized using a second-order method while the other is handled with a first-order approach. We establish the global convergence of our online algorithm for non-convex cases in terms of the expected violation of first-order optimality conditions. Numerical experiments demonstrate that our method achieves significantly faster initial convergence and produces more robust performance compared to other popular learning algorithms. Additionally, our algorithm exhibits reduced sensitivity to learning rates and outperforms the recently proposedslimTrainalgorithm (Newman et al. 2022). For validation, the code has been made available on GitHub.
Min Gan, Xiang-Xiang Su, Guang-Yong Chen, Jing Chen 0007, C. L. Philip Chen
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 An Efficient Decoupled Optimization Algorithm for a Class of Regression Models
Guang-Yong Chen, Xiang-Xiang Su, Min Gan, C. L. Philip Chen
IEEE Signal Process. Lett.4
2025 Moving Average-Based Variable Projection for Separable Nonlinear Problems
abstract
The identification of separable nonlinear models, prevalent in tasks such as signal analysis, image processing, time series analysis, and machine learning, presents a non-convex optimization challenge that necessitates the development of efficient identification algorithms. The Variable Projection (VP) algorithm has been proven to be quite effective for addressing these problems; however, traditional VP relying on the Hessian matrix and its inverse are highly time-consuming and unsuitable for complex, large-scale applications. This letter introduces a novel approach that employs the exponential moving average of gradient and gradient estimation bias to indirectly estimate the curvature of the objective landscape, proposing a Moving Average-based Variable Projection method (MAVP). The proposed algorithm utilizes only gradient information and can properly tackle the coupling relationships between different parameters during the optimization process, thereby achieving faster convergence. Numerical results on nonlinear time series analysis and image reconstruction demonstrate that the MAVP algorithm exhibits significant efficiency and effectiveness.
Min Gan, Guang-Yong Chen, C. L. Philip Chen
IEEE Signal Process. Lett.2
2025 LPFSformer: Location Prior Guided Frequency and Spatial Interactive Learning for Nighttime Flare Removal
abstract
When capturing images under strong light sources at night, intense lens flare artifacts often appear, significantly degrading visual quality and impacting downstream computer vision tasks. Although transformer-based methods have achieved remarkable results in nighttime flare removal, they fail to adequately distinguish between flare and non-flare regions. This unified processing overlooks the unique characteristics of these regions, leading to suboptimal performance and unsatisfactory results in real-world scenarios. To address this critical issue, we propose a novel approach incorporating Location Prior Guidance (LPG) and a specialized flare removal model, LPFSformer. LPG is designed to accurately learn the location of flares within an image and effectively capture the associated glow effects. By employing Location Prior Injection (LPI), our method directs the model’s focus towards flare regions through the interaction of frequency and spatial domains. Additionally, to enhance the recovery of high-frequency textures and capture finer local details, we designed a Global Hybrid Feature Compensator (GHFC). GHFC aggregates different expert structures, leveraging the diverse receptive fields and CNN operations of each expert to effectively utilize a broader range of features during the flare removal process. Extensive experiments demonstrate that our LPFSformer achieves state-of-the-art flare removal performance compared to existing methods. Our code and a pre-trained LPFSformer have been uploaded to GitHub for validation.
Guang-Yong Chen, Jian-Nan Su, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.5
2025 Real-World Image Reflection Removal: An Ultra-High-Definition Dataset and an Efficient Baseline
abstract
Reflection removal is a crucial issue in image reconstruction, especially for high-definition images. Removing undesirable reflections can greatly enhance the performance of various visual systems, such as medical imaging, autonomous driving, and security surveillance. However, the resolution of existing reflection removal datasets is not high and the training data heavily relies on synthetic data, which hampers the performance of reflection removal methods and restricts the development of effective techniques tailored for high-definition images. Therefore, this paper introduces a new dataset, Real-world Reflection Removal in 4K (RR4K). This novel dataset, with its large capacity and high resolution of$6000\times 4000$pixels, represents a significant advancement in the field, ensuring a realistic and high quality benchmark. Furthermore, building upon the dataset, we propose an efficient method for single-image reflection removal, optimized for high-definition processing. This method employs the U-Net architecture, enhanced with large kernel distillation and scale-aware features, enabling it to effectively handle complex reflection scenarios while reducing computational demands. Comprehensive testing on the RR4K dataset and existing low-resolution datasets has demonstrated the method’s superior efficiency and effectiveness. We believe that our constructed RR4K dataset can better evaluate and design algorithms for removing undesirable reflection from real-world high-definition images. Our dataset and code are available athttps://github.com/jengchauwei/RR4K.
Guang-Yong Chen, Chao-Wei Zheng, Jian-Nan Su, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.5
2025 FDCE-Net: Underwater Image Enhancement With Embedding Frequency and Dual Color Encoder
abstract
Underwater images often suffer from various issues such as low brightness, color shift, blurred details, and noise due to light absorption and scattering caused by water and suspended particles. Previous underwater image enhancement (UIE) methods have primarily focused on spatial domain enhancement, neglecting the frequency domain information inherent in the images. However, the degradation factors of underwater images are closely intertwined in the spatial domain. Although certain methods focus on enhancing images in the frequency domain, they overlook the inherent relationship between the image degradation factors and the information present in the frequency domain. As a result, these methods frequently enhance certain attributes of the improved image while inadequately addressing or even exacerbating other attributes. Moreover, many existing methods heavily rely on prior knowledge to address color shift problems in underwater images, limiting their flexibility and robustness. In order to overcome these limitations, we propose the Embedding Frequency and Dual Color Encoder Network (FDCE-Net) in our paper. The FDCE-Net consists of two main structures: 1) Frequency Spatial Network (FS-Net) aims to achieve initial enhancement by utilizing our designed Frequency Spatial Residual Block (FSRB) to decouple image degradation factors in the frequency domain and enhance different attributes separately; 2) To tackle the color shift issue, we introduce the Dual-Color Encoder (DCE). The DCE establishes correlations between color and semantic representations through cross-attention and leverages multi-scale image features to guide the optimization of adaptive color query. The final enhanced images are generated by combining the outputs of FS-Net and DCE through a fusion network. These images exhibit rich details, clear textures, low noise and natural colors. Extensive experiments demonstrate that our FDCE-Net outperforms state-of-the-art (SOTA) methods in terms of both visual quality and quantitative metrics. The code of our model is publicly available at:https://github.com/Alexande-rChan/FDCE-Net.
Jingchun Zhou, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.4
2025 Fine-Grained Image Captioning by Ranking Diffusion Transformer
abstract
The CLIP visual feature-based image captioning models have developed rapidly and achieved remarkable results. However, existing models still struggle to produce descriptive and discriminative captions because they insufficiently exploit fine-grained visual cues and fail to model complex vision-language alignment. To address these limitations, we propose a Ranking Diffusion Transformer (RDT), which integrates a Ranking Visual Encoder (RVE) and a Ranking Loss (RL) for fine-grained image captioning. The RVE introduces a novel ranking attention mechanism that effectively mines diverse and discriminative visual information from CLIP features. Meanwhile, the RL leverages the ranking of generated caption quality as a global semantic supervisory signal, thereby enhancing the diffusion process and strengthening vision-language semantic alignment. We show that by collaborating RVE and RL via the novel RDT-and by gradually adding and removing noise in the diffusion process-more discriminative visual features are learned and precisely aligned with the language features. Experimental results on popular benchmark datasets demonstrate that our proposed RDT surpasses existing state-of-the-art image captioning models in the literature. The code is publicly available at: https://github.com/junwan2014/RDT.
Jun Wan 0005, Min Gan, Lefei Zhang, Jie Zhou 0009, Jun Liu 0036, Bo Du 0001, C. L. Philip Chen
IEEE Trans. Image Process.2
2025 KMT-PLL: K-Means Cross-Attention Transformer for Partial Label Learning
abstract
Partial label learning (PLL) studies the problem of learning instance classification with a set of candidate labels and only one is correct. While recent works have demonstrated that the Vision Transformer (ViT) has achieved good results when training from clean data, its applications to PLL remain limited and challenging. To address this issue, we rethink the relationship between instances and object queries to propose K-means cross-attention transformer for PLL (KMT-PLL), which can continuously learn cluster centers and be used for downstream disambiguation tasks. More specifically, K-means cross-attention as a clustering process can effectively learn the cluster centers to represent label classes. The purpose of this operation is to make the similarity between instances and labels measurable, which can effectively detect noise labels. Furthermore, we propose a new corrected cross entropy formulation, which can assign weights to candidate labels according to the instance-to-label relevance to guide the training of the instance classifier. As the training goes on, the ground-truth label is progressively identified, and the refined labels and cluster centers in turn help to improve the classifier. Simulation results demonstrate the advantage of the KMT-PLL and its suitability for PLL.
Jinfu Fan, Linqing Huang, Chaoyu Gong, Yang You 0001, Min Gan, Zhongjie Wang 0004
IEEE Trans. Neural Networks Learn. Syst.5
2025 Obstacle-Avoiding X-Architecture Bounded-Skew Tree Algorithm Under Timing Slack Constraints
abstract
As interconnect delay increasingly becomes the primary source of chip delay, timing analysis in the very large-scale integration (VLSI) routing process is becoming more crucial. Concurrently, to maintain computational synchronization in the chip, the bounded-skew constraint must be introduced. Additionally, the issue of obstacle-avoiding has gained attention due to the presence of routing obstacles on the chip. Furthermore, the introduction of X-architecture enables more efficient utilization of routing resources. In this article, we propose an obstacle-avoiding X-architecture bounded-skew tree (BST) algorithm under timing slack constraints, which, for the first time, simultaneously considers timing slack, bounded-skew, obstacle-avoidance, and X-architecture in a unified framework. First, an effective preprocessing strategy is presented to support fast information retrieval for the subsequent strategies. Second, a BST construction strategy is developed to ensure compliance with skew constraints by consulting and updating a dedicated skew table. Third, a local worst negative slack (WNS) optimization strategy is designed to improve the WNS of critical paths by balancing wirelength (WL) and radius. Fourth, an obstacle-avoiding strategy is implemented to navigate around routing obstacles while minimizing unnecessary WL overhead. Finally, a path refinement strategy is designed to select routing structures with maximal edge sharing to replace the initial structure, thereby further optimizing WL. Experimental results demonstrate that the proposed algorithm significantly improves both WL and the key timing metric WNS, while satisfying obstacle-avoidance and bounded-skew constraints.
Genggeng Liu, Ren Lu, Zhifeng Lin, Chuandong Chen, Min Gan, Jianli Chen, Wenzhong Guo
IEEE Trans. Syst. Man Cybern. Syst.6
2024 Hybrid network via key feature fusion for image restoration
Shuteng Hu, Jingchun Zhou, Jinfu Fan, Min Gan, C. L. Philip Chen
Eng. Appl. Artif. Intell.5
2024 FISTA acceleration inspired network design for underwater image enhancement
Bing-Yuan Chen, Jian-Nan Su, Guang-Yong Chen, Min Gan
J. Vis. Commun. Image Represent.4
2024 Texture-aware and color-consistent learning for underwater image enhancement
Shuteng Hu, Min Gan, C. L. Philip Chen
J. Vis. Commun. Image Represent.4
2024 Progressive encoding-decoding image dehazing network
Min Gan
Multim. Tools Appl.3
2024 Revealing the Dark Side of Non-Local Attention in Single Image Super-Resolution
abstract
Single Image Super-Resolution (SISR) aims to reconstruct a high-resolution image from its corresponding low-resolution input. A common technique to enhance the reconstruction quality is Non-Local Attention (NLA), which leverages self-similar texture patterns in images. However, we have made a novel finding that challenges the prevailing wisdom. Our research reveals that NLA can be detrimental to SISR and even produce severely distorted textures. For example, when dealing with severely degrade textures, NLA may generate unrealistic results due to the inconsistency of non-local texture patterns. This problem is overlooked by existing works, which only measure the average reconstruction quality of the whole image, without considering the potential risks of using NLA. To address this issue, we propose a new perspective for evaluating the reconstruction quality of NLA, by focusing on the sub-pixel level that matches the pixel-wise fusion manner of NLA. From this perspective, we provide the approximate reconstruction performance upper bound of NLA, which guides us to design a concise yet effective Texture-Fidelity Strategy (TFS) to mitigate the degradation caused by NLA. Moreover, the proposed TFS can be conveniently integrated into existing NLA-based SISR models as a general building block. Based on the TFS, we develop a Deep Texture-Fidelity Network (DTFN), which achieves state-of-the-art performance for SISR. Our code and a pre-trained DTFN are available on GitHub†for verification.
Jian-Nan Su, Min Gan, Guang-Yong Chen, Wenzhong Guo, C. L. Philip Chen
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Nonmonotone variable projection algorithms for matrix decomposition with missing data
Xiang-Xiang Su, Min Gan, Guang-Yong Chen
Pattern Recognit.2
2024 Dynamic Degradation Intensity Estimation for Adaptive Blind Super-Resolution: A Novel Approach and Benchmark Dataset
abstract
Blind Super-Resolution (BlindSR) aims to reconstruct high-resolution (HR) images from low-resolution (LR) images without prior knowledge of the image degradation process. This is a challenging problem in real-world applications, where the degradation can be complex and unknown. Recent unsupervised learning-based BlindSR methods can estimate the image degradation in an unsupervised manner, but they suffer from limited adaptability to different types and intensities of degradation. They tend to capture the average level of degradation across all training samples, resulting in over-smoothing or over-sharpening effects for some images. As a result, the final reconstruction may exhibit the mean effect. Moreover, existing synthetic datasets do not reflect the real-world degradation scenarios, making it difficult to evaluate the performance of BlindSR methods. To address these issues, we propose a novel Degradation Intensity Estimation Module (DIEM) method, which can estimate the pixel-level degradation information of the input image more specifically and use it to guide image reconstruction. Furthermore, we construct a benchmark dataset under real scenarios, which is closer to the real-world BlindSR problem than existing synthetic datasets, and can provide a more reasonable evaluation of BlindSR methods. Extensive experimental results demonstrate that our DIEM-guided BlindSR method can achieve state-of-the-art image reconstruction results. Our code and pre-trained models have been uploaded to GitHub† for validation.
Guang-Yong Chen, Wu-Ding Weng, Jian-Nan Su, Min Gan, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.4
2024 Unsupervised Degradation Aware and Representation for Real-World Remote Sensing Image Super-Resolution
abstract
Blind super-resolution (BlindSR) has recently attracted attention in the field of remote sensing. Due to the lack of paired data, most works assume that the acquired remote sensing images are high-resolution (HR) and use predefined degradation models to synthesize low-resolution (LR) images for training and evaluation. However, these acquired remote sensing images are often degraded by various factors, which still require super-resolution reconstruction to meet practical needs. Using them as ground truth images will limit the model’s ability to restore fine details, resulting in blurry and noisy reconstructions. To overcome these limitations, we propose an unsupervised degradation-aware network which transforms natural images into the degraded domain as real-world remote sensing images. It uses natural images containing rich texture information as a reference for fine-grained restoration of the network, enabling the network to produce clearer reconstructions. Furthermore, we discovered the remarkable capability of patch-wise discriminator to perceive the degradation type of different regions within the acquired remote sensing image. Inspired by this finding, we design a novel degradation representation module (DRM) that can estimate the degradation information from LR images and guide the network to perform adaptive restoration. Comprehensive experimental results demonstrate that our proposed unsupervised blind super-resolution framework (UDASR) achieves state-of-the-art restoration performance. Our code and pre-trained models have been uploaded to GitHub† for validation.
Wenzhong Guo, Wu-Ding Weng, Guang-Yong Chen, Jian-Nan Su, Min Gan, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.5
2024 Spatial-Frequency Dual-Domain Feature Fusion Network for Low-Light Remote Sensing Image Enhancement
abstract
Low-light remote sensing (RS) images generally feature high resolution and high spatial complexity, with continuously distributed surface features in space. This continuity in scenes leads to extensive long-range correlations in spatial domains within RS images. convolutional neural networks (CNNs), which rely on local correlations for long-distance modeling, struggle to establish long-range correlations in such images. On the other hand, transformer-based methods that focus on global information face high computational complexities when processing high-resolution RS images. From another perspective, the Fourier transform can compute global information without introducing a large number of parameters, enabling the network to more efficiently capture the overall image structure and establish long-range correlations. Therefore, we propose a dual-domain feature fusion network (DFFN) for low-light RS image enhancement. Specifically, this challenging task of low-light enhancement is divided into two more manageable subtasks: the first phase learns amplitude information to restore image brightness, and the second phase learns phase information to refine details. To facilitate information exchange between the two phases, we designed an information fusion affine block that combines data from different phases and scales. In addition, we have constructed two dark light RS datasets to address the current lack of datasets in dark light RS image enhancement. Extensive evaluations show that our method outperforms existing state-of-the-art methods. The code is available athttps://github.com/iijjlk/DFFN.
Zishu Yao, Jinfu Fan, Min Gan, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.4
2024 High-Similarity-Pass Attention for Single Image Super-Resolution
abstract
Recent developments in the field of non-local attention (NLA) have led to a renewed interest in self-similarity-based single image super-resolution (SISR). Researchers usually use the NLA to explore non-local self-similarity (NSS) in SISR and achieve satisfactory reconstruction results. However, a surprising phenomenon that the reconstruction performance of the standard NLA is similar to that of the NLA with randomly selected regions prompted us to revisit NLA. In this paper, we first analyzed the attention map of the standard NLA from different perspectives and discovered that the resulting probability distribution always has full support for every local feature, which implies a statistical waste of assigning values to irrelevant non-local features, especially for SISR which needs to model long-range dependence with a large number of redundant non-local features. Based on these findings, we introduced a concise yet effective soft thresholding operation to obtain high-similarity-pass attention (HSPA), which is beneficial for generating a more compact and interpretable distribution. Furthermore, we derived some key properties of the soft thresholding operation that enable training our HSPA in an end-to-end manner. The HSPA can be integrated into existing deep SISR models as an efficient general building block. In addition, to demonstrate the effectiveness of the HSPA, we constructed a deep high-similarity-pass attention network (HSPAN) by integrating a few HSPAs in a simple backbone. Extensive experimental results demonstrate that HSPAN outperforms state-of-the-art approaches on both quantitative and qualitative evaluations. Our code and a pre-trained model were uploaded to GitHub (https://github.com/laoyangui/HSPAN) for validation.
Jian-Nan Su, Min Gan, Guang-Yong Chen, Wenzhong Guo, C. L. Philip Chen
IEEE Trans. Image Process.2
2024 Online Identification of Nonlinear Systems With Separable Structure
abstract
Separable nonlinear models (SNLMs) are of great importance in system modeling, signal processing, and machine learning because of their flexible structure and excellent description of nonlinear behaviors. The online identification of such models is quite challenging, and previous related work usually ignores the special structure where the estimated parameters can be partitioned into a linear and a nonlinear part. In this brief, we propose an efficient first-order recursive algorithm for SNLMs by introducing the variable projection (VP) step. The proposed algorithm utilizes the recursive least-squares method to eliminate the linear parameters, resulting in a reduced function. Then, the stochastic gradient descent (SGD) algorithm is employed to update the parameters of the reduced function. By considering the tight coupling relationship between linear parameters and nonlinear parameters, the proposed first-order VP algorithm is more efficient and robust than the traditional SGD algorithm and alternating optimization algorithm. More importantly, since the proposed algorithm just uses the first-order information, it is easier to apply it to large-scale models. Numerical results on examples of different sizes confirm the effectiveness and efficiency of the proposed algorithm.
Guang-Yong Chen, Min Gan, Long Chen 0001, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.2
2024 Multiscale Cross-Connected Dehazing Network With Scene Depth Fusion
abstract
In this article, we propose a multiscale cross-connected dehazing network with scene depth fusion. We focus on the correlation between a hazy image and the corresponding depth image. The model encodes and decodes the hazy image and the depth image separately and includes cross connections at the decoding end to directly generate a clean image in an end-to-end manner. Specifically, we first construct an input pyramid to obtain the receptive fields of the depth image and the hazy image at multiple levels. Then, we add the features of the corresponding dimensions in the input pyramid to the encoder. Finally, the two paths of the decoder are cross-connected. In addition, the proposed model uses wavelet pooling and residual channel attention modules (RCAMs) as components. A series of ablation experiments shows that the wavelet pooling and RCAMs effectively improve the performance of the model. We conducted extensive experiments on multiple dehazing datasets, and the results show that the model is superior to other advanced methods in terms of peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and subjective visual effects. The source code and supplementary are available at https://github.com/CCECfgd/MSCDN-master.
Min Gan, Bi Fan, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.2
2024 Gradient Matching Federated Domain Adaptation for Brain Image Classification
abstract
Federated learning has shown its unique advantages in many different tasks, including brain image analysis. It provides a new way to train deep learning models while protecting the privacy of medical image data from multiple sites. However, previous studies suggest that domain shift across different sites may influence the performance of federated models. As a solution, we propose a gradient matching federated domain adaptation (GM-FedDA) method for brain image classification, aiming to reduce domain discrepancy with the assistance of a public image dataset and train robust local federated models for target sites. It mainly includes two stages: 1) pretraining stage; we propose a one-common-source adversarial domain adaptation (OCS-ADA) strategy, i.e., adopting ADA with gradient matching loss to pretrain encoders for reducing domain shift at each target site (private data) with the assistance of a common source domain (public data) and 2) fine-tuning stage; we develop a gradient matching federated (GM-Fed) fine-tuning method for updating local federated models pretrained with the OCS-ADA strategy, i.e., pushing the optimization direction of a local federated model toward its specific local minimum by minimizing gradient matching loss between sites. Using fully connected networks as local models, we validate our method with the diagnostic classification tasks of schizophrenia and major depressive disorder based on multisite resting-state functional MRI (fMRI), respectively. Results show that the proposed GM-FedDA method outperforms other commonly used methods, suggesting the potential of our method in brain imaging analysis and other fields, which need to utilize multisite data while preserving data privacy.
Jianpo Su, Min Gan, Hui Shen 0004, Dewen Hu
IEEE Trans. Neural Networks Learn. Syst.4
2024 A Robust Multilayer X-Architecture Global Routing System Based on Particle Swarm Optimization
abstract
Global routing is an extremely important stage of very large scale integration (VLSI) physical design. With the rise of nano-scale integrated circuit design, the multilayer global routing problem has attracted considerable research interest during the past few years. In this article, a multilayer X-architecture global routing (ML-XGR) system based on particle swarm optimization (PSO), called FZU-Router, is proposed to solve the ML-XGR problem for the first time. FZU-Router contains a multilayer X-architecture integer linear programming (MX-ILP) model and a multilayer X-architecture PSO (MX-PSO) algorithm, which are presented to formulate and solve the ML-XGR problem, respectively. Moreover, four effective strategies are designed to enhance the efficiency of FZU-Router: 1) a strategy for generating new routing modes is proposed to strengthen the robustness of encoding strategy of MX-PSO; 2) a strategy for combining MX-PSO with maze routing is proposed to improve the routability; 3) a strategy for reducing the channel capacity is proposed to make better use of optimization ability of MX-PSO; and 4) a strategy for dynamic resource assignment is proposed to make better use of routing resources and shorten the running time. Experimental results on multiple benchmarks confirm that the proposed FZU-Router leads to fewer total overflow and shorter total wirelength compared with the state-of-the-art routers.
Genggeng Liu, Zhen Zhuang, Zhenyu Pei, Min Gan, Xing Huang 0001, Wenzhong Guo
IEEE Trans. Syst. Man Cybern. Syst.5
2024 Timing-Driven Obstacle-Avoiding X-Architecture Steiner Minimum Tree Algorithm With Slack Constraints
abstract
SMT is an optimized model for solving the routing problem of a multipin net in very large-scale integrated circuits. As the appearance of various obstacles on chips, the obstacle-avoiding problem has attracted much attention in recent years. Meanwhile, since interconnect delay plays a major role in chip delay, timing analysis is another critical problem worthy of consideration when constructing an Steiner minimum tree (SMT). Furthermore, the introduction of theX-architecture allows for better utilization of routing resources. In this article, a timing-driven obstacle-avoiding X-architecture Steiner minimum tree algorithm with slack constraints (TD-OAXSMT-SC) is proposed to consider obstacle-avoiding, timing slack constraints, andX-architecture simultaneously for the first time. The TD-OAXSMT-SC algorithm consists of four major stages: 1) in the routing tree initialization stage, this article constructs anX-architecture Prim–Dijkstra spanning tree as the initial routing tree with minimum total delay; 2) in the particle swarm optimization (PSO)-based routing tree iteration stage, a novel discrete PSO algorithm based on genetic operators is proposed to obtain a high-quality routing tree; 3) in the routing tree standardization stage, two effective standardization strategies are proposed to obtain a routing tree that satisfies both obstacle-avoiding and timing slack constraints; and 4) in the routing tree optimization stage, the connection of interconnected wires is optimized in a global manner, thus obtaining an optimized routing tree. Experimental results show that the proposed TD-OAXSMT-SC algorithm outperforms the state-of-the-art methods in routing quality with slack constraints.
Genggeng Liu, Ren Lu, Xing Huang 0001, Min Gan, Wenzhong Guo
IEEE Trans. Syst. Man Cybern. Syst.5
2023 Global Learnable Attention for Single Image Super-Resolution
abstract
Self-similarity is valuable to the exploration of non-local textures in single image super-resolution (SISR). Researchers usually assume that the importance of non-local textures is positively related to their similarity scores. In this paper, we surprisingly found that when repairing severely damaged query textures, some non-local textures with low-similarity which are closer to the target can provide more accurate and richer details than the high-similarity ones. In these cases, low-similarity does not mean inferior but is usually caused by different scales or orientations. Utilizing this finding, we proposed a Global Learnable Attention (GLA) to adaptively modify similarity scores of non-local textures during training instead of only using a fixed similarity scoring function such as the dot product. The proposed GLA can explore non-local textures with low-similarity but more accurate details to repair severely damaged textures. Furthermore, we propose to adopt Super-Bit Locality-Sensitive Hashing (SB-LSH) as a preprocessing method for our GLA. With the SB-LSH, the computational complexity of our GLA is reduced from quadratic to asymptotic linear with respect to the image size. In addition, the proposed GLA can be integrated into existing deep SISR models as an efficient general building block. Based on the GLA, we constructed a Deep Learnable Similarity Network (DLSN), which achieves state-of-the-art performance for SISR tasks of different degradation types (e.g., blur and noise). Our code and a pre-trained DLSN have been uploaded to GitHub†for validation.
Jian-Nan Su, Min Gan, Guang-Yong Chen, Jia-Li Yin, C. L. Philip Chen
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Fine Perceptive GANs for Brain MR Image Super-Resolution in Wavelet Domain
abstract
Magnetic resonance (MR) imaging plays an important role in clinical and brain exploration. However, limited by factors such as imaging hardware, scanning time, and cost, it is challenging to acquire high-resolution MR images clinically. In this article, fine perceptive generative adversarial networks (FP-GANs) are proposed to produce super-resolution (SR) MR images from the low-resolution counterparts. By adopting the divide-and-conquer scheme, FP-GANs are designed to deal with the low-frequency (LF) and high-frequency (HF) components of MR images separately and parallelly. Specifically, FP-GANs first decompose an MR image into LF global approximation and HF anatomical texture subbands in the wavelet domain. Then, each subband generative adversarial network (GAN) simultaneously concentrates on super-resolving the corresponding subband image. In generator, multiple residual-in-residual dense blocks are introduced for better feature extraction. In addition, the texture-enhancing module is designed to trade off the weight between global topology and detailed textures. Finally, the reconstruction of the whole image is considered by integrating inverse discrete wavelet transformation in FP-GANs. Comprehensive experiments on the MultiRes_7T and ADNI datasets demonstrate that the proposed model achieves finer structure recovery and outperforms the competing methods quantitatively and qualitatively. Moreover, FP-GANs further show the value by applying the SR results in classification tasks.
Senrong You, Bai Ying Lei, Shuqiang Wang, Charles K. Chui, Albert C. Cheung, Yong Liu 0018, Min Gan, Guo-Cheng Wu 0001, Yanyan Shen
IEEE Trans. Neural Networks Learn. Syst.7
2022 Multiscale Low-Light Image Enhancement Network With Illumination Constraint
abstract
Images captured under low-light environments typically have poor visibility, affecting many advanced computer vision tasks. In recent years, there have been some low-light image enhancement models based on deep learning, but they have not been able to effectively mine the deep multiscale features in the image, resulting in poor generalization performance and instability of the model. The disadvantages are mainly reflected in the color distortion, color unsaturation and artifacts. Current methods unable to adjust the exposure effectively, resulting in uneven exposure or partial overexposure. To address these issues, we propose an end-to-end low-light image enhancement model, which is called multiscale low-light image enhancement network with illumination constraint (MLLEN-IC), to achieve preferable generalization ability and stable performance. On the one hand, we use the squeeze-and-excitation-Res2Net block (SE-Res2block) as a base unit to enhance the model’s ability by extracting deep multiscale features. On the other hand, to make the model more adaptable in low-light image enhancement tasks, we calculate the illumination constraint by the low-light itself to prevent overexposure, uneven exposure, and unsaturated colors. Extensive experiments are conducted to demonstrate MLLEN-IC not only adjusts light levels, but also has a more natural visual effect, and avoids problems such as color distortion, artifacts, and uneven exposure. In particular, MLLEN-IC has pretty generalization and stability performance. The source code and supplementary are available athttps://github.com/CCECfgd/MLLEN-IC.
Bi Fan, Min Gan, Guang-Yong Chen, C. L. Philip Chen
IEEE Trans. Circuits Syst. Video Technol.3
2022 Robust Standard Gradient Descent Algorithm for ARX Models Using Aitken Acceleration Technique
abstract
A robust standard gradient descent (SGD) algorithm for ARX models using the Aitken acceleration method is developed. Considering that the SGD algorithm has slow convergence rates and is sensitive to the step size, a robust and accelerative SGD (RA-SGD) algorithm is derived. This algorithm is based on the Aitken acceleration method, and its convergence rate is improved from linear convergence to at least quadratic convergence in general. Furthermore, the RA-SGD algorithm is always convergent with no limitation of the step size. Both the convergence analysis and the simulation examples demonstrate that the presented algorithm is effective.
Jing Chen 0007, Min Gan, Pritesh Narayan, Yanjun Liu 0001
IEEE Trans. Cybern.2
2022 Weighted Generalized Cross-Validation-Based Regularization for Broad Learning System
abstract
The broad learning system (BLS) is an emerging flat network, which has demonstrated its outstanding performance in classification and regression problems. The regularization plays an important role in the performance of the BLS. In real applications, since the BLS network is usually expanded dynamically, a predetermined regularization parameter may reduce the performance of the network. Using a fixed regularization in some cases, the classification accuracy of the BLS decreases dramatically when we expand the network. To alleviate this problem, we propose a method that automatically finds appropriate regularization parameters for different datasets, which is based on the weighted generalized cross-validation (WGCV). The experimental results indicate that the WGCV method improves the performance of the BLS, and alleviates the accuracy decrease of the incremental learning algorithm.
Min Gan, Hong-Tao Zhu, Guang-Yong Chen, C. L. Philip Chen
IEEE Trans. Cybern.1
2022 Frequency Principle in Broad Learning System
abstract
Deep neural networks have achieved breakthrough improvement in various application fields. Nevertheless, they usually suffer from a time-consuming training process because of the complicated structures of neural networks with a huge number of parameters. As an alternative, a fast and efficient discriminative broad learning system (BLS) is proposed, which takes the advantages of flat structure and incremental learning. The BLS has achieved outstanding performance in classification and regression problems. However, the previous studies ignored the reason why the BLS can generalize well. In this article, we focus on the interpretation from the viewpoint of the frequency domain. We discover the existence of the frequency principle in BLS, i.e., the BLS preferentially captures low-frequency components quickly and then fits the high frequencies during the incremental process of adding feature nodes and enhancement nodes. The frequency principle may be of great inspiration for expanding the application of BLS.
Guang-Yong Chen, Min Gan, C. L. Philip Chen, Hong-Tao Zhu, Long Chen 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 Nuisance Parameter Estimation Algorithms for Separable Nonlinear Models
abstract
Many inverse problems in machine learning, system identification, and image processing include nuisance parameters, which are important for the recovering of other parameters. Separable nonlinear optimization problems fall into this category. The special separable structure in these problems has inspired several efficient optimization strategies. A well-known method is the variable projection (VP) that projects out a subset of the estimated parameters, resulting in a reduced problem that includes fewer parameters. The expectation maximization (EM) is another separated method that provides a powerful framework for the estimation of nuisance parameters. The relationships between EM and VP were ignored in previous studies, though they deal with a part of parameters in a similar way. In this article, we explore the internal relationships and differences between VP and EM. Unlike the algorithms that separate the parameters directly, the hierarchical identification algorithm decomposes a complex model into several linked submodels and identifies the corresponding parameters. Therefore, this article also studies the difference and connection between the hierarchical algorithm and the parameter-separated algorithms like VP and EM. In the numerical simulation part, Monte Carlo experiments are performed to further compare the performance of different algorithms. The results show that the VP algorithm usually converges faster than the other two algorithms and is more robust to the initial point of the parameters.
Long Chen 0001, Jia-Bing Chen, Guang-Yong Chen, Min Gan, C. L. Philip Chen
IEEE Trans. Syst. Man Cybern. Syst.4
2022 Constrained Variable Projection Optimization for Stationary RBF-AR Models
abstract
Stationarity is fundamental for time-series modeling and prediction. In this article, we focus on the radial basis function network-based autoregressive (RBF-AR) models which have been widely used in practical applications. Compared to previous work, we give a less-restrictive sufficient condition for the asymptotic stationarity of the RBF-AR model. The parameter estimation of the RBF-AR model is converted to the optimization of a variable projection functional with constraints of stationarity to always derive a stationary model. The constrained evolutionary algorithm is used to solve the optimization problem. Numerical results demonstrate the effectiveness of the proposed method.
Min Gan, Guang-Yong Chen, C. L. Philip Chen
IEEE Trans. Syst. Man Cybern. Syst.2
2022 An Iterative Implementation of Variable Projection for Separable Nonlinear Optimization Problems
abstract
The separable nonlinear least-squares (SNLLS) problems considered in this article frequently appear in a wide range of research fields, such as machine learning, computer vision, system identification, and signal processing. The variable projection algorithm proposed by Golub and Pereyra, which reduces the dimension of the parameters by projecting the linear parameters out of the problem, is quite valuable in solving SNLLS problems. Previous implementations of the variable projection algorithm are based on matrix factorization. In this article, we propose an iterative implementation of the variable projection algorithm. Compared with previous implementations based on matrix decomposition, the proposed method can effectively avoid suffering from large condition number of the matrix or even matrix decomposition failure when dealing with ill-posed SNLLS problems. Numerical experiments on real-world data and synthetic data show the efficiency and robustness of the proposed iterative variable projection algorithm.
Guang-Yong Chen, Min Gan, Hong-Tao Zhu, Long Chen 0001, C. L. Philip Chen
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Q-Learning with Fisher Score for Feature Selection of Large-Scale Data Sets
Min Gan, Li Zhang 0004
KSEM1
2021 Iteratively local fisher score for feature selection
Min Gan, Li Zhang 0004
Appl. Intell.1
2021 BFGS method based variable projection approach for image restoration
abstract
Abstract In this paper, a variable projection approach based on the BFGS (Broyden–Fletcher–Goldfarb–Shanno) method for image reconstruction problems is proposed, which is an alternative to the common alternating minimisation scheme. The image restoration is expressed as a nonlinear least‐squares problem with reduced parameter space. To improve the efficiency of the algorithm, the BFGS method is proposed to be used to optimise the reduced objective function. The large‐scale problem considered in this paper is projected on to a small Krylov subspace using Lanczos bidiagonalisation. The regularisation parameter is selected by a weighted generalised cross validation criterion. Numerical examples demonstrate the efficiency and effectiveness of the proposed algorithm.
Qiong-Ying Chen, Yun-Zhi Huang, Min Gan, C. L. Philip Chen, Guang-Yong Chen
IET Image Process.3
2021 Diabetic Retinopathy Diagnosis Using Multichannel Generative Adversarial Network With Semisupervision
abstract
Diabetic retinopathy (DR) is one of the major causes of blindness. It is of great significance to apply deep-learning techniques for DR recognition. However, deep-learning algorithms often depend on large amounts of labeled data, which is expensive and time-consuming to obtain in the medical imaging area. In addition, the DR features are inconspicuous and spread out over high-resolution fundus images. Therefore, it is a big challenge to learn the distribution of such DR features. This article proposes a multichannel-based generative adversarial network (MGAN) with semisupervision to grade DR. The multichannel generative model is developed to generate a series of subfundus images corresponding to the scattering DR features. By minimizing the dependence on labeled data, the proposed semisupervised MGAN can identify the inconspicuous lesion features by using high-resolution fundus images without compression. Experimental results on the public Messidor data set show that the proposed model can grade DR effectively. Note to Practitioners-This article is motivated by the challenging problem due to the inadequacy of labeled data in medical image analysis and the dispersion of efficient features in high-resolution medical images. As for the inadequacy of labeled data in medical image analysis, the reasons mainly include the followings: 1) the high-quality annotation of medical imaging sample depends heavily on scarce medical expertise which is very expensive and 2) comparing with natural issues, it is more difficult to collect medical images because of privacy issues. It is of great significance to apply deep-learning techniques for diabetic retinopathy (DR) recognition. In this article, the multichannel generative adversarial network (GAN) with semisupervision is developed for DR-aided diagnosis. The proposed model can deal with DR classification problem with inadequacy of labeled data in the following ways: 1) the multichannel generative scheme is proposed to generate a series of subfundus images corresponding to the scattering DR features and 2) the proposed multichannel-based GAN (MGAN) model with semisupervision can make full use of both labeled data and unlabeled data. The experimental results demonstrate that the proposed model outperforms the other representative models in terms of accuracy, area under ROC curve (AUC), sensitivity, and specificity.
Shuqiang Wang, Yong Hu 0003, Yanyan Shen, Zhile Yang, Min Gan, Bai Ying Lei
IEEE Trans Autom. Sci. Eng.6
2021 Basis Function Matrix-Based Flexible Coefficient Autoregressive Models: A Framework for Time Series and Nonlinear System Modeling
abstract
We propose, in this paper, a framework for time series and nonlinear system modeling, called the basis function matrix-based flexible coefficient autoregressive (BFM-FCAR) model. It has very flexible nonlinear structure. We show that many famous nonlinear time series models can be derived under this framework by choosing the proper basis function matrices. Some probabilistic properties (the conditions of geometrical ergodicity) of the BFM-FCAR model are investigated. Taking advantage of the model structure, we present an efficient parameter estimation algorithm for the proposed framework by using the variable projection method. Finally, we show how new models are generated from the proposed framework.
Guang-Yong Chen, Min Gan, C. L. Philip Chen, Han-Xiong Li
IEEE Trans. Cybern.2
2021 Insights Into Algorithms for Separable Nonlinear Least Squares Problems
abstract
Separable nonlinear least squares (SNLLS) problems have attracted interest in a wide range of research fields such as machine learning, computer vision, and signal processing. During the past few decades, several algorithms, including the joint optimization algorithm, alternated least squares (ALS) algorithm, embedded point iterations (EPI) algorithm, and variable projection (VP) algorithms, have been employed for solving SNLLS problems in the literature. The VP approach has been proven to be quite valuable for SNLLS problems and the EPI method has been successful in solving many computer vision tasks. However, no clear explanations about the intrinsic relationships of these algorithms have been provided in the literature. In this paper, we give some insights into these algorithms for SNLLS problems. We derive the relationships among different forms of the VP algorithms, EPI algorithm and ALS algorithm. In addition, the convergence and robustness of some algorithms are investigated. Moreover, the analysis of the VP algorithm generates a negative answer to Kaufman's conjecture. Numerical experiments on the image restoration task, fitting the time series data using the radial basis function network based autoregressive (RBF-AR) model, and bundle adjustment are given to compare the performance of different algorithms.
Guang-Yong Chen, Min Gan, Shuqiang Wang, C. L. Philip Chen
IEEE Trans. Image Process.2
2021 Recursive Variable Projection Algorithm for a Class of Separable Nonlinear Models
abstract
In this article, we study the recursive algorithms for a class of separable nonlinear models (SNLMs) in which the parameters can be partitioned into a linear part and a nonlinear part. Such models are very common in machine learning, system identification, and signal processing. Utilizing the special structure of the SNLMs, we propose a recursive variable projection (RVP) algorithm, in which at each recursion, the linear parameters of the model are eliminated, and the nonlinear parameters are updated by the recursive Levenberg-Marquart algorithm. Then, based on the updated nonlinear parameters, the linear parameters are updated by the recursive least-squares algorithm. According to a convergence analysis of the RVP algorithm, the parameter estimation error is mean-square bounded. Numerical examples confirm the satisfactory performance of the proposed algorithm.
Min Gan, Guang-Yong Chen, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.1
2020 A Novel EM Identification Method for Hammerstein Systems With Missing Output Data
abstract
This article concerns a novel auxiliary-model-based expectation maximization (EM) estimation method for Hammerstein systems with data loss by extending the EM method to estimate models with multiple parameter vectors. The novel EM method relaxes the requirements on an autoregression model with one parameter vector, interactively maximizes the expectation over multiple parameter vectors in a more general model, and uses the output of an auxiliary model to substitute the missing outputs in the information vector in iteration processes. A numerical simulation is employed to demonstrate the effectiveness of the proposed novel EM method.
Dongqing Wang, Shuo Zhang 0007, Min Gan, Jianlong Qiu
IEEE Trans. Ind. Informatics3
2020 Term Selection for a Class of Separable Nonlinear Models
abstract
In this paper, we consider the term selection problem for a class of separable nonlinear models. The strategy is a two-step process in which the nonlinear parameters of the model are first optimized by a variable projection method, and then the least absolute shrinkage and selection operator are adopted to obtain a sparse solution by picking out the critical terms automatically. This process may be repeated several times. The proposed algorithm is tested on parameter estimation problems for an exponential model and a neural network-based model. The numerical results show that the proposed algorithm can pick out the appropriate terms from the overparameterized model and the obtained parsimonious model performs better than other methods.
Min Gan, Guang-Yong Chen, Long Chen 0001, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.1
2019 Quality control of imbalanced mass spectra from isotopic labeling experiments
abstract
BACKGROUND: Mass spectra are usually acquired from the Liquid Chromatography-Mass Spectrometry (LC-MS) analysis for isotope labeled proteomics experiments. In such experiments, the mass profiles of labeled (heavy) and unlabeled (light) peptide pairs are represented by isotope clusters (2D or 3D) that provide valuable information about the studied biological samples in different conditions. The core task of quality control in quantitative LC-MS experiment is to filter out low-quality peptides with questionable profiles. The commonly used methods for this problem are the classification approaches. However, the data imbalance problems in previous control methods are often ignored or mishandled. In this study, we introduced a quality control framework based on the extreme gradient boosting machine (XGBoost), and carefully addressed the imbalanced data problem in this framework. RESULTS: In the XGBoost based framework, we suggest the application of the Synthetic minority over-sampling technique (SMOTE) to re-balance data and use the balanced data to train the boosted trees as the classifier. Then the classifier is applied to other data for the peptide quality assessment. Experimental results show that our proposed framework increases the reliability of peptide heavy-light ratio estimation significantly. CONCLUSIONS: Our results indicate that this framework is a powerful method for the peptide quality assessment. For the feature extraction part, the extracted ion chromatogram (XIC) based features contribute to the peptide quality assessment. To solve the imbalanced data problem, SMOTE brings a much better classification performance. Finally, the XGBoost is capable for the peptide quality control. Overall, our proposed framework provides reliable results for the further proteomics studies.
Tianjun Li, Long Chen 0001, Min Gan
BMC Bioinform.3
2019 Adaptive RBF-AR Models Based on Multi-Innovation Least Squares Method
abstract
In the previous work, the parameters of radial basis function network based autoregressive (RBF-AR) models are estimated offline and no longer updated afterward. In this letter, an adaptive learning algorithm is proposed for the RBF-AR models. The proposed strategy is that the nonlinear parameters are previously determined by an off-line variable projection method; and once new samples are available, the linear parameters are updated. The linear adaptive algorithm adopted in this letter is the multi-innovation least squares method, due to its high performance. The simulation results show that with the adaption of the linear parameters, the prediction performance of the RBF-AR models may be significantly improved, which demonstrates the effectiveness of the proposed algorithm.
Min Gan, Xiao-Xian Chen, Feng Ding 0001, Guang-Yong Chen, C. L. Philip Chen
IEEE Signal Process. Lett.1
2019 Modified Gram-Schmidt Method-Based Variable Projection Algorithm for Separable Nonlinear Models
abstract
Separable nonlinear models are very common in various research fields, such as machine learning and system identification. The variable projection (VP) approach is efficient for the optimization of such models. In this paper, we study various VP algorithms based on different matrix decompositions. Compared with the previous method, we use the analytical expression of the Jacobian matrix instead of finite differences. This improves the efficiency of the VP algorithms. In particular, based on the modified Gram-Schmidt (MGS) method, a more robust implementation of the VP algorithm is introduced for separable nonlinear least-squares problems. In numerical experiments, we compare the performance of five different implementations of the VP algorithm. Numerical results show the efficiency and robustness of the proposed MGS method-based VP algorithm.
Guang-Yong Chen, Min Gan, Feng Ding 0001, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.2
2018 Generalized exponential autoregressive models for nonlinear time series: Stationarity, estimation and applications
Guang-Yong Chen, Min Gan
Inf. Sci.2
2018 On Some Separated Algorithms for Separable Nonlinear Least Squares Problems
abstract
For a class of nonlinear least squares problems, it is usually very beneficial to separate the variables into a linear and a nonlinear part and take full advantage of reliable linear least squares techniques. Consequently, the original problem is turned into a reduced problem which involves only nonlinear parameters. We consider in this paper four separated algorithms for such problems. The first one is the variable projection (VP) algorithm with full Jacobian matrix of Golub and Pereyra. The second and third ones are VP algorithms with simplified Jacobian matrices proposed by Kaufman and Ruano et al. respectively. The fourth one only uses the gradient of the reduced problem. Monte Carlo experiments are conducted to compare the performance of these four algorithms. From the results of the experiments, we find that: 1) the simplified Jacobian proposed by Ruano et al. is not a good choice for the VP algorithm; moreover, it may render the algorithm hard to converge; 2) the fourth algorithm perform moderately among these four algorithms; 3) the VP algorithm with the full Jacobian matrix perform more stable than that of the VP algorithm with Kuafman's simplified one; and 4) the combination of VP algorithm and Levenberg-Marquardt method is more effective than the combination of VP algorithm and Gauss-Newton method.
Min Gan, C. L. Philip Chen, Guang-Yong Chen, Long Chen 0001
IEEE Trans. Cybern.1
2015 Gradient Radial Basis Function Based Varying-Coefficient Autoregressive Model for Nonlinear and Nonstationary Time Series
abstract
We propose a gradient radial basis function based varying-coefficient autoregressive (GRBF-AR) model for modeling and predicting time series that exhibit nonlinearity and homogeneous nonstationarity. This GRBF-AR model is a synthesis of the gradient RBF and the functional-coefficient autoregressive (FAR) model. The gradient RBFs, which react to the gradient of the series, are used to construct varying coefficients of the FAR model. The Mackey-Glass chaotic time series are used to evaluate the performance of the proposed method. It is shown that the GRBF-AR model not only achieves much more parsimonious structure but also much better prediction performance than that of GRBF network.
Min Gan, C. L. Philip Chen, Han-Xiong Li, Long Chen 0001
IEEE Signal Process. Lett.1
2015 A Variable Projection Approach for Efficient Estimation of RBF-ARX Model
abstract
The radial basis function network-based autoregressive with exogenous inputs (RBF-ARX) models have much more linear parameters than nonlinear parameters. Taking advantage of this special structure, a variable projection algorithm is proposed to estimate the model parameters more efficiently by eliminating the linear parameters through the orthogonal projection. The proposed method not only substantially reduces the dimension of parameter space of RBF-ARX model but also results in a better-conditioned problem. In this paper, both the full Jacobian matrix of Golub and Pereyra and the Kaufman's simplification are used to test the performance of the algorithm. An example of chaotic time series modeling is presented for the numerical comparison. It clearly demonstrates that the proposed approach is computationally more efficient than the previous structured nonlinear parameter optimization method and the conventional Levenberg-Marquardt algorithm without the parameters separated. Finally, the proposed method is also applied to a simulated nonlinear single-input single-output process, a time-varying nonlinear process and a real multiinput multioutput nonlinear industrial process to illustrate its usefulness.
Min Gan, Han-Xiong Li, Hui Peng 0001
IEEE Trans. Cybern.1
2015 Fuzzy Restricted Boltzmann Machine for the Enhancement of Deep Learning
abstract
In recent years, deep learning caves out a research wave in machine learning. With outstanding performance, more and more applications of deep learning in pattern recognition, image recognition, speech recognition, and video processing have been developed. Restricted Boltzmann machine (RBM) plays an important role in current deep learning techniques, as most of existing deep networks are based on or related to it. For regular RBM, the relationships between visible units and hidden units are restricted to be constants. This restriction will certainly downgrade the representation capability of the RBM. To avoid this flaw and enhance deep learning capability, the fuzzy restricted Boltzmann machine (FRBM) and its learning algorithm are proposed in this paper, in which the parameters governing the model are replaced by fuzzy numbers. This way, the original RBM becomes a special case in the FRBM, when there is no fuzziness in the FRBM model. In the process of learning FRBM, the fuzzy free energy function is defuzzified before the probability is defined. The experimental results based on bar-and-stripe benchmark inpainting and MNIST handwritten digits classification problems show that the representation capability of FRBM model is significantly better than the traditional RBM. Additionally, the FRBM also reveals better robustness property compared with RBM when the training data are contaminated by noises.
C. L. Philip Chen, Chun-Yang Zhang, Long Chen 0001, Min Gan
IEEE Trans. Fuzzy Syst.4
2014 Detecting and monitoring abrupt emergences and submergences of episodes over data streams
Min Gan, Honghua Dai 0001
Inf. Syst.1
2014 An Efficient Variable Projection Formulation for Separable Nonlinear Least Squares Problems
abstract
We consider in this paper a class of nonlinear least squares problems in which the model can be represented as a linear combination of nonlinear functions. The variable projection algorithm projects the linear parameters out of the problem, leaving the nonlinear least squares problems involving only the nonlinear parameters. To implement the variable projection algorithm more efficiently, we propose a new variable projection functional based on matrix decomposition. The advantage of the proposed formulation is that the size of the decomposed matrix may be much smaller than those of previous ones. The Levenberg-Marquardt algorithm using finite difference method is then applied to minimize the new criterion. Numerical results show that the proposed approach achieves significant reduction in computing time.
Min Gan, Han-Xiong Li
IEEE Trans. Cybern.1
2012 A global-local optimization approach to parameter estimation of RBF-type models
Min Gan, Hui Peng 0001
Inf. Sci.1
2012 Design a Wind Speed Prediction Model Using Probabilistic Fuzzy System
abstract
Generation of wind is a very complicated process and influenced by large numbers of unknown factors. A probabilistic fuzzy system based prediction model is designed for the short-term wind speed prediction. By introducing the third probability dimension, the proposed prediction model can capture both stochastic and the deterministic uncertainties, and guarantee a better prediction in complex stochastic environment. The effectiveness of this intelligent wind speed prediction model is demonstrated by the simulations on a group of wind speed data. The robust modeling performance further discloses its potential in the practical prediction of wind speed under complex circumstance.
Han-Xiong Li, Min Gan
IEEE Trans. Ind. Informatics3
2011 Fast Mining of Non-derivable Episode Rules in Complex Sequences
Min Gan, Honghua Dai 0001
MDAI1
2010 A locally linear RBF network-based state-dependent AR model for nonlinear time series modeling
Min Gan, Hui Peng 0001, Xiaohong Chen 0001, Garba Inoussa
Inf. Sci.1