EDBT 2026 Demo / reviewers in the wild / expert
Hongying Liu 0001
dblp:43/8776-1
· DBLP profile ↗
67ranked-venue papers
13as first author
41since 2021 · last 2026
0000-0001-5961-5569ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 5 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 15 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large ModelsabstractAdamW has become one of the most effective optimizers for training large-scale models. We have also observed its effectiveness in the context of federated learning (FL). However, directly applying AdamW in federated learning settings poses significant challenges: (1) due to data heterogeneity, AdamW often yields high variance in the second-moment estimate v; (2) the local overfitting of AdamW may cause client drift; and (3) Reinitializing moment estimates (v, m) at each round slows down convergence. To address these challenges, we propose the first Federated AdamW algorithm, called FedAdamW, for training and fine-tuning various large models. FedAdamW aligns local updates with the global update using both a local correction mechanism and decoupled weight decay to mitigate local overfitting. FedAdamW efficiently aggregates the mean of the second-moment estimates to reduce their variance and reinitialize them. Theoretically, we prove that FedAdamW achieves a linear speedup convergence rate of O(p(L∆σ2l )/(SKRε2) + (L∆)/R) without heterogeneity assumption, where S is the number of participating clients per round, K is the number of local iterations, and R is the total number of communication rounds. We also employ PAC-Bayesian generalization analysis to explain the effectiveness of decoupled weight decay in local training. Empirically, we validate the effectiveness of FedAdamW on language and vision Transformer models. Compared to several baselines, FedAdamW significantly reduces communication rounds and improves test accuracy. Junkang Liu, Fanhua Shang, Hongying Liu 0001, Yuanyuan Liu 0001, Kewen Zhu, Zhouchen Lin |
AAAI | 3 |
| 2026 | SAVSR++: An All-Stage Scale-Aware and Temporal Omniscient Framework for Arbitrary-Scale Video Super-Resolution
Hongying Liu 0001, Zekun Li 0014, Wenbin Zuo, Fanhua Shang, Liang Wang 0001, Wei Feng 0005 |
Int. J. Comput. Vis. | 1 |
| 2026 | Boosting adversarial transferability via diversified sampling and flatness-aware momentum
Siyuan Deng, Fanhua Shang, Chengchao Zhang, Hongying Liu 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Causal Inference via Style Bias Deconfounding for Domain GeneralizationabstractDeep neural networks (DNNs) often struggle with out-of-distribution data, limiting their reliability in real-world visual applications. To address this issue, domain generalization methods have been developed to learn domain-invariant features from single or multiple training domains, enabling generalization to unseen testing domains. However, existing approaches usually overlook the impact of style frequency within the training set. This oversight predisposes models to capture spurious visual correlations caused by style confounding factors, rather than learning truly causal representations, thereby undermining inference reliability. In this work, we introduce Style Deconfounding Causal Learning (SDCL), a novel causal inference-based framework that explicitly addresses style as a confounding factor to enhance domain generalization in image modalities. Our approaches begins with constructing a structural causal model (SCM) tailored to the domain generalization problem and applies a backdoor adjustment strategy to account for style influence. Building on this foundation, we design a style-guided expert module (SGEM) to adaptively clusters style distributions during training, capturing the global confounding style. Additionally, a backdoor causal learning module (BDCL) performs causal interventions during feature extraction, ensuring fair integration of global confounding styles into sample predictions, effectively reducing style bias. The SDCL framework is highly versatile and can be seamlessly integrated with state-of-the-art data augmentation techniques. Extensive experiments across diverse natural and medical image recognition tasks validate its efficacy, demonstrating superior performance in both multi-domain and the more challenging single-domain generalization scenarios. Di Lin 0002, Hao Chen 0011, Hongying Liu 0001, Wei Feng 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Enhancing the impact of model performance gains for semi-supervised medical image segmentation
Wenbin Zuo, Hongying Liu 0001, Huadeng Wang, Lingqi Zeng, Ningning Tang, Fanhua Shang, Jingjing Deng 0001 |
Pattern Recognit. | 2 |
| 2026 | PathFusion-Net: A Rough Path Theory-Based Deep Learning Model for ECG Arrhythmia ClassificationabstractThis study introduces a novel electrocardiogram (ECG) arrhythmia classification model, PathFusion-Net, which integrates Rough Path Theory with deep learning technologies. The model combines Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Path Signatures, and Path Development to extract spatial morphological features from ECG images and multi-order temporal representations from ECG signals. By adopting an inter-patient split paradigm, our approach more closely reflects real-world clinical diagnostic settings compared to intra-patient methods. The model demonstrates state-of-the-art overall classification performance on both the MIT-BIH Arrhythmia Database and a private clinical dataset, achieving 94.7% and 95.1% accuracy, respectively, under the AAMI four-class standard with an inter-patient split paradigm. On the MIT-BIH dataset, the proposed method attains competitive precision and recall across multiple arrhythmia types, including 95.2% /87.9% for ventricular ectopic beats (V) and 75.7% /92.3% for supraventricular ectopic beats (S), indicating balanced performance across clinically diverse categories. This research highlights the potential of Rough Path Theory in time-series analysis and offers a novel deep learning framework for automated early detection and monitoring of ECG arrhythmias. The code used in this study is available at: https://github.com/Rand2AI/PathFusion-Net. Tianlong Feng, Qingchen Li, Yongzhi Liao, Di Lu 0001, Jianqin Zhao, Hao Ni 0001, Hongying Liu 0001, Jingjing Deng 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2025 | Unsupervised Degradation Representation Aware Transform for Real-World Blind Image Super-ResolutionabstractBlind image super-resolution (blind SR) aims to restore a high-resolution (HR) image from a low-resolution (LR) image with unknown degradation. Many existing methods explicitly estimate degradation information from various LR images. However, in most cases, image degradations are independent of image content. Their estimations may be influenced by the image content resulting in inaccuracy. Unlike existing works, we design a dual-encoder for degradation representation (DEDR) to preclude the influence of image content from LR images. This benefits in extracting the intrinsic degradation representation more accurately. To the best of our knowledge, this paper is the first work that estimates degradation representation through filtering out image content. Based on the degradation representation extracted by DEDR, we present a novel framework, named degradation representation aware transform network (DRAT) for blind SR. We propose global degradation aware (GDA) blocks to propagate degradation information across spatial and channel dimensions, in which a degradation representation transform module (DRT) is introduced to render features degradation-aware, thereby enhancing the restoration of LR images. Extensive experiments are conducted on three benchmark datasets (including Gaussian 8, DIV2KRK, and real-world datasets) under large scaling factors with complex degradations. The experimental results demonstrate that DRAT surpasses state-of-the-art supervised kernel estimation and unsupervised degradation representation methods. Hongying Liu 0001, Chaowei Fang, Fanhua Shang, Yuanyuan Liu 0001, Dongmei Jiang |
AAAI | 2 |
| 2025 | Beyond Background Shift: Rethinking Instance Replay in Continual Semantic SegmentationabstractIn this work, we focus on continual semantic segmentation (CSS), where segmentation networks are required to continuously learn new classes without erasing knowledge of previously learned ones. Although storing images of old classes and directly incorporating them into the training of new models has proven effective in mitigating catastrophic forgetting in classification tasks, this strategy presents notable limitations in CSS. Specifically, the stored and new images with partial category annotations leads to confusion between unannotated categories and the background, complicating model fitting. To tackle this issue, this paper proposes a novel Enhanced Instance Replay (EIR) method, which not only preserves knowledge of old classes while simultaneously eliminating background confusion by instance storage of old classes, but also mitigates background shifts in the new images by integrating stored instances with new images. By effectively resolving background shifts in both stored and new images, EIR alleviates catastrophic forgetting in the CSS task, thereby enhancing the model’s capacity for CSS. Experimental results validate the efficacy of our approach, which significantly outperforms state-of-the-art CSS methods. The code is available at https://github.com/YikeYin97/EIR. Hongmei Yin, Tingliang Feng, Fan Lyu, Fanhua Shang, Hongying Liu 0001, Wei Feng 0005 |
CVPR | 5 |
| 2025 | Improving Generalization in Federated Learning with Highly Heterogeneous Data via Momentum-Based Stochastic Controlled Weight AveragingabstractFor federated learning (FL) algorithms such as FedSAM, their generalization capability is crucial for real-word applications. In this paper, we revisit the generalization problem in FL and investigate the impact of data heterogeneity on FL generalization. We find that FedSAM usually performs worse than FedAvg in the case of highly heterogeneous data, and thus propose a novel and effective federated learning algorithm with Stochastic Weight Averaging (called \texttt{FedSWA}), which aims to find flatter minima in the setting of highly heterogeneous data. Moreover, we introduce a new momentum-based stochastic controlled weight averaging FL algorithm (\texttt{FedMoSWA}), which is designed to better align local and global models.
Theoretically, we provide both convergence analysis and generalization bounds for \texttt{FedSWA} and \texttt{FedMoSWA}. We also prove that the optimization and generalization errors of \texttt{FedMoSWA} are smaller than those of their counterparts, including FedSAM and its variants. Empirically, experimental results on CIFAR10/100 and Tiny ImageNet demonstrate the superiority of the proposed algorithms compared to their counterparts. Junkang Liu, Yuanyuan Liu 0001, Fanhua Shang, Hongying Liu 0001, Wei Feng 0005 |
ICML | 4 |
| 2025 | Consistency of Local and Global Flatness for Federated Learning
Junkang Liu, Fanhua Shang, Hongying Liu 0001, Yuanyuan Liu 0001 |
ACM Multimedia | 4 |
| 2025 | Tight High-Probability Bounds for Nonconvex Heavy-Tailed Scenario under Weaker AssumptionsabstractGradient clipping is increasingly important in centralized learning (CL) and federated learning (FL). Many works focus on its optimization properties under strong assumptions involving Gaussian noise and standard smoothness. However, practical machine learning tasks often only satisfy weaker conditions, such as heavy-tailed noise and $(L_0, L_1)$-smoothness. To bridge this gap, we propose a high-probability analysis for clipped Stochastic Gradient Descent (SGD) under these weaker assumptions. Our findings show a better convergence rate than existing ones can be achieved, and our high-probability analysis does not rely on the bounded gradient assumption. Moreover, we extend our analysis to FL, where a gap remains between expected and high-probability convergence, which the naive clipped SGD cannot bridge. Thus, we design a new \underline{Fed}erated \underline{C}lipped \underline{B}atched \underline{G}radient (FedCBG) algorithm, and prove the convergence and generalization bounds with high probability for the first time. Our analysis reveals the trade-offs between the optimization and generalization performance. Extensive experiments demonstrate that \methodname{} can generalize better to unseen client distributions than state-of-the-art baselines. Weixin An, Yuanyuan Liu 0001, Fanhua Shang, Junkang Liu, Hongying Liu 0001 |
NeurIPS | 6 |
| 2025 | QBasicVSR: Temporal Awareness Adaptation Quantization for Video Super-ResolutionabstractWhile model quantization has become pivotal for deploying super-resolution (SR) networks on mobile devices, existing works focus on quantization methods only for image super-resolution. Different from image SR quantization, the temporal error propagation, shared temporal parameterization, and temporal metric mismatch significantly degrade the quantization performance of a video SR model. To address these issues, we propose the first quantization method, QBasicVSR, for video super-resolution. A novel temporal awareness adaptation post-training quantization (PTQ) framework for video super-resolution with the flow-gradient video bit adaptation and temporal shared layer bit adaptation is presented. Moreover, we put forward a novel fine-tuning method for VSR with the supervision of the full-precision model. Our method achieves extraordinary performance with state-of-the-art efficient VSR approaches, delivering up to $\times$200 faster processing speed while utilizing only 1/8 of the GPU resources. Additionally, extensive experiments demonstrate that the proposed method significantly outperforms existing PTQ algorithms on various datasets. For instance, it attains a 2.53 dB increase on the UDM10 benchmark when quantizing BasicVSR to 4-bit with 100 unlabeled video clips. The code and models will be released on GitHub. Fanhua Shang, Hongying Liu 0001, Liang Wang 0001, Wei Feng 0005, Yanming Hui |
NeurIPS | 3 |
| 2025 | Distillation guided deep unfolding network with frequency hierarchical regularization for low-dose CT image denoising
Hongying Liu 0001, Yuanyuan Liu 0001, Fanhua Shang, Licheng Jiao |
Neurocomputing | 2 |
| 2025 | Semi-Supervised Remote Sensing Imagery Scene Classification Based on Probabilistic SelectionabstractTo address low-quality pseudo-labels and class imbalance in semi-supervised remote sensing scene classification, this letter proposes a novel Dynamic Adaptive Reweighting and Pseudo-labeling method (DARP) based on probabilistic random selection. First, we present a semi-supervised recursive learning strategy with adaptive category threshold and probabilistic random selection approach to solve the problems of sample diversity and imbalanced distribution of pseudo labeled sample categories. Second, we design a image category weight determination method based on the number and accuracy of image categories, and assign larger weights to classes with lower accuracy increasing the probability of these class samples being selected to improve overall accuracy. Experiments were conducted on the publicly available datasets UCM and NWPU-RESISC45, and the accuracies reached 95.87% and 88.86%, using only five labeled data for each category, which is an improvement of 5.71% and 1.56% over existing semi-supervised methods. Lixin Ye, Fanhua Shang, Hongying Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Completed Feature Disentanglement Learning for Multimodal MRIs AnalysisabstractMultimodal MRIs play a crucial role in clinical diagnosis and treatment. Feature disentanglement (FD)-based methods, aiming at learning superior feature representations for multimodal data analysis, have achieved significant success in multimodal learning (MML). Typically, existing FD-based methods separate multimodal data into modality-shared and modality-specific features, and employ concatenation or attention mechanisms to integrate these features. However, our preliminary experiments indicate that these methods could lead to a loss of shared information among subsets of modalities when the inputs contain more than two modalities, and such information is critical for prediction accuracy. Furthermore, these methods do not adequately interpret the relationships between the decoupled features at the fusion stage. To address these limitations, we propose a novel Complete Feature Disentanglement (CFD) strategy that recovers the lost information during feature decoupling. Specifically, the CFD strategy not only identifies modality-shared and modality-specific features, but also decouples shared features among subsets of multimodal inputs, termed as modality-partial-shared features. We further introduce a new Dynamic Mixture-of-Experts Fusion (DMF) module that dynamically integrates these decoupled features, by explicitly learning the local-global relationships among the features. The effectiveness of our approach is validated through classification tasks on three multimodal MRI datasets. Extensive experimental results demonstrate that our approach outperforms other state-of-the-art MML methods with obvious margins, showcasing its superior performance. Tianling Liu, Hongying Liu 0001, Fanhua Shang, Lequan Yu, Tong Han |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | DEs-Inspired Accelerated Unfolded Linearized ADMM Networks for Inverse ProblemsabstractMany research works have shown that the traditional alternating direction multiplier methods (ADMMs) can be better understood by continuous-time differential equations (DEs). On the other hand, many unfolded algorithms directly inherit the traditional iterations to build deep networks. Although they achieve superior practical performance and a faster convergence rate than traditional counterparts, there is a lack of clear insight into unfolded network structures. Thus, we attempt to explore the unfolded linearized ADMM (LADMM) from the perspective of DEs, and design more efficient unfolded networks. First, by proposing an unfolded Euler LADMM scheme and inspired by the trapezoid discretization, we design a new more accurate Trapezoid LADMM scheme. For the convenience of implementation, we provide its explicit version via a prediction-correction strategy. Then, to expand the representation space of unfolded networks, we design an accelerated variant of our Euler LADMM scheme, which can be interpreted as second-order DEs with stronger representation capabilities. To fully explore this representation space, we designed an accelerated Trapezoid LADMM scheme. To the best of our knowledge, this is the first work to explore a comprehensive connection with theoretical guarantees between unfolded ADMMs and first- (second-) order DEs. Finally, we instantiate our schemes as (A-)ELADMM and (A-)TLADMM with the proximal operators, and (A-)ELADMM-Net and (A-)TLADMM-Net with convolutional neural networks (CNNs). Extensive inverse problem experiments show that our Trapezoid LADMM schemes perform better than well-known methods. Weixin An, Yuanyuan Liu 0001, Fanhua Shang, Hongying Liu 0001, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | SAVSR: Arbitrary-Scale Video Super-Resolution via a Learned Scale-Adaptive NetworkabstractDeep learning-based video super-resolution (VSR) networks have gained significant performance improvements in recent years. However, existing VSR networks can only support a fixed integer scale super-resolution task, and when we want to perform VSR at multiple scales, we need to train several models. This implementation certainly increases the consumption of computational and storage resources, which limits the application scenarios of VSR techniques. In this paper, we propose a novel Scale-adaptive Arbitrary-scale Video Super-Resolution network (SAVSR), which is the first work focusing on spatial VSR at arbitrary scales including both non-integer and asymmetric scales. We also present an omni-dimensional scale-attention convolution, which dynamically adapts according to the scale of the input to extract inter-frame features with stronger representational power. Moreover, the proposed spatio-temporal adaptive arbitrary-scale upsampling performs VSR tasks using both temporal features and scale information. And we design an iterative bi-directional architecture for implicit feature alignment. Experiments at various scales on the benchmark datasets show that the proposed SAVSR outperforms state-of-the-art (SOTA) methods at non-integer and asymmetric scales. The source code is available at https://github.com/Weepingchestnut/SAVSR. Zekun Li 0014, Hongying Liu 0001, Fanhua Shang, Yuanyuan Liu 0001, Wei Feng 0005 |
AAAI | 2 |
| 2024 | FedBCGD: Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
Junkang Liu, Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001, Yuangang Li 0002, YunXiang Gong |
ACM Multimedia | 4 |
| 2024 | Robust and Faster Zeroth-Order Minimax Optimization: Complexity and ApplicationsabstractMany zeroth-order (ZO) optimization algorithms have been developed to solve nonconvex minimax problems in machine learning and computer vision areas. However, existing ZO minimax algorithms have high complexity and rely on some strict restrictive conditions for ZO estimations. To address these issues, we design a new unified ZO gradient descent extragradient ascent (ZO-GDEGA) algorithm, which reduces the overall complexity to $\mathcal{O}(d\epsilon^{-6})$ to find an $\epsilon$-stationary point of the function $\psi$ for nonconvex-concave (NC-C) problems, where $d$ is the variable dimension. To the best of our knowledge, ZO-GDEGA is the first ZO algorithm with complexity guarantees to solve stochastic NC-C problems. Moreover, ZO-GDEGA requires weaker conditions on the ZO estimations and achieves more robust theoretical results. As a by-product, ZO-GDEGA has advantages on the condition number for the NC-strongly concave case. Experimentally, ZO-GDEGA can generate more effective poisoning attack data with an average accuracy reduction of 5\%. The improved AUC performance also verifies the robustness of gradient estimations. Weixin An, Yuanyuan Liu 0001, Fanhua Shang, Hongying Liu 0001 |
NeurIPS | 4 |
| 2024 | A single frame and multi-frame joint network for 360-degree panorama video super-resolution
Hongying Liu 0001, Wanhao Ma, Zhubo Ruan, Chaowei Fang, Fanhua Shang, Yuanyuan Liu 0001, Chaoli Wang 0001, Dongmei Jiang |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Gradient Correction for White-Box Adversarial AttacksabstractDeep neural networks (DNNs) play key roles in various artificial intelligence applications such as image classification and object recognition. However, a growing number of studies have shown that there exist adversarial examples in DNNs, which are almost imperceptibly different from the original samples but can greatly change the output of DNNs. Recently, many white-box attack algorithms have been proposed, and most of the algorithms concentrate on how to make the best use of gradients per iteration to improve adversarial performance. In this article, we focus on the properties of the widely used activation function, rectified linear unit (ReLU), and find that there exist two phenomena (i.e., wrong blocking and over transmission) misguiding the calculation of gradients for ReLU during backpropagation. Both issues enlarge the difference between the predicted changes of the loss function from gradients and corresponding actual changes and misguide the optimized direction, which results in larger perturbations. Therefore, we propose a universal gradient correction adversarial example generation method, called ADV-ReLU, to enhance the performance of gradient-based white-box attack algorithms such as fast gradient signed method (FGSM), iterative FGSM (I-FGSM), momentum I-FGSM (MI-FGSM), and variance tuning MI-FGSM (VMI-FGSM). Through backpropagation, our approach calculates the gradient of the loss function with respect to the network input, maps the values to scores, and selects a part of them to update the misguided gradients. Comprehensive experimental results on ImageNet and CIFAR10 demonstrate that our ADV-ReLU can be easily integrated into many state-of-the-art gradient-based white-box attack algorithms, as well as transferred to black-box attacks, to further decrease perturbations measured in the -norm. Hongying Liu 0001, Zhijin Ge, Fanhua Shang, Yuanyuan Liu 0001, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Adaptive Non-Local Generative Adversarial Networks for Low-Dose CT Image DenoisingabstractLow-dose computed tomography (CT) has been widely used in medical diagnosis and treatment. Many deep networks have been proposed for low-dose CT denoising. The local receptive field of the convolution affects the network performance. For different input images, conventional neural networks always adopt a fixed number of channels which limits the performance of deep networks. To address these problems, we propose a channel-adaptive convolution and patch selection (CAPS) module to enhance the feature extraction of our network. CAPS enables our network to adaptively adjust the number of channels according to different inputs. Moreover, the concatenation of patches can expand the receptive field globally, so the shallow layer of our network can extract more global information. To further ensure the clarity of denoised images, we present a new wavelet loss function to the generator of our generative adversarial network. Compared with state-of-the-art methods, our network can obtain superior denoising results. Hongying Liu 0001, Fanhua Shang, Yuanyuan Liu 0001 |
ICASSP | 2 |
| 2023 | Improving the Transferability of Adversarial Examples with Arbitrary Style TransferabstractDeep neural networks are vulnerable to adversarial examples crafted by applying human-imperceptible perturbations on clean inputs. Although many attack methods can achieve high success rates in the white-box setting, they also exhibit weak transferability in the black-box setting. Recently, various methods have been proposed to improve adversarial transferability, in which the input transformation is one of the most effective methods. In this work, we notice that existing input transformation-based works mainly adopt the transformed data in the same domain for augmentation. Inspired by domain generalization, we aim to further improve the transferability using the data augmented from different domains. Specifically, a style transfer network can alter the distribution of low-level visual features in an image while preserving semantic content for humans. Hence, we propose a novel attack method named Style Transfer Method (STM) that utilizes a proposed arbitrary style transfer network to transform the images into different domains. To avoid inconsistent semantic information of stylized images for the classification network, we fine-tune the style transfer network and mix up the generated images added by random noise with the original images to maintain semantic consistency and boost input diversity. Extensive experimental results on the ImageNet-compatible dataset show that our proposed method can significantly improve the adversarial transferability on either normally trained models or adversarially trained models than state-of-the-art input transformation-based attacks. Code is available at: https://github.com/Zhijin-Ge/STM. Zhijin Ge, Fanhua Shang, Hongying Liu 0001, Yuanyuan Liu 0001, Wei Feng 0005, Xiaosen Wang |
ACM Multimedia | 3 |
| 2023 | A Single-Loop Accelerated Extra-Gradient Difference Algorithm with Improved Complexity Bounds for Constrained Minimax OptimizationabstractIn this paper, we propose a novel extra-gradient difference acceleration algorithm for solving constrained nonconvex-nonconcave (NC-NC) minimax problems. In particular, we design a new extra-gradient difference step to obtain an important quasi-cocoercivity property, which plays a key role to significantly improve the convergence rate in the constrained NC-NC setting without additional structural assumption. Then momentum acceleration is also introduced into our dual accelerating update step. Moreover, we prove that, to find an $\epsilon$-stationary point of the function $f$, our algorithm attains the complexity $\mathcal{O}(\epsilon^{-2})$ in the constrained NC-NC setting, while the best-known complexity bound is $\widetilde{\mathcal{O}}(\epsilon^{-4})$, where $\widetilde{\mathcal{O}}(\cdot)$ hides logarithmic factors compared to $\mathcal{O}(\cdot)$. As the special cases of the constrained NC-NC setting, our algorithm can also obtain the same complexity $\mathcal{O}(\epsilon^{-2})$ for both the nonconvex-concave (NC-C) and convex-nonconcave (C-NC) cases, while the best-known complexity bounds are $\widetilde{\mathcal{O}}(\epsilon^{-2.5})$ for the NC-C case and $\widetilde{\mathcal{O}}(\epsilon^{-4})$ for the C-NC case. For fair comparison with existing algorithms, we also analyze the complexity bound to find $\epsilon$-stationary point of the primal function $\phi$ for the constrained NC-C problem, which shows that our algorithm can improve the complexity bound from $\widetilde{\mathcal{O}}(\epsilon^{-3})$ to $\mathcal{O}(\epsilon^{-2})$. To the best of our knowledge, this is the first time that the proposed algorithm improves the best-known complexity bounds from $\mathcal{O}(\epsilon^{-4})$ and $\widetilde{\mathcal{O}}(\epsilon^{-3})$ to $\mathcal{O}(\epsilon^{-2})$ in both the NC-NC and NC-C settings. Yuanyuan Liu 0001, Fanhua Shang, Weixin An, Hongying Liu 0001, Zhouchen Lin |
NeurIPS | 5 |
| 2023 | Boosting Adversarial Transferability by Achieving Flat Local MaximaabstractTransfer-based attack adopts the adversarial examples generated on the surrogate model to attack various models, making it applicable in the physical world and attracting increasing interest. Recently, various adversarial attacks have emerged to boost adversarial transferability from different perspectives. In this work, inspired by the observation that flat local minima are correlated with good generalization, we assume and empirically validate that adversarial examples at a flat local region tend to have good transferability by introducing a penalized gradient norm to the original loss function. Since directly optimizing the gradient regularization norm is computationally expensive and intractable for generating adversarial examples, we propose an approximation optimization method to simplify the gradient update of the objective function. Specifically, we randomly sample an example and adopt a first-order procedure to approximate the curvature of the second-order Hessian matrix, which makes computing more efficient by interpolating two Jacobian matrices. Meanwhile, in order to obtain a more stable gradient direction, we randomly sample multiple examples and average the gradients of these examples to reduce the variance due to random sampling during the iterative process. Extensive experimental results on the ImageNet-compatible dataset show that the proposed method can generate adversarial examples at flat local regions, and significantly improve the adversarial transferability on either normally trained models or adversarially trained models than the state-of-the-art attacks. Our codes are available at: https://github.com/Trustworthy-AI-Group/PGN. Zhijin Ge, Xiaosen Wang, Hongying Liu 0001, Fanhua Shang, Yuanyuan Liu 0001 |
NeurIPS | 3 |
| 2022 | HNO: High-Order Numerical Architecture for ODE-Inspired Deep Unfolding NetworksabstractRecently, deep unfolding networks (DUNs) based on optimization algorithms have received increasing attention, and their high efficiency has been confirmed by many experimental and theoretical results. Since this type of networks combines model-based traditional optimization algorithms, they have high interpretability. In addition, ordinary differential equations (ODEs) are often used to explain deep neural networks, and provide some inspiration for designing innovative network models. In this paper, we transform DUNs into first-order ODE forms, and propose a high-order numerical architecture for ODE-inspired deep unfolding networks. To the best of our knowledge, this is the first work to establish the relationship between DUNs and ODEs. Moreover, we take two representative DUNs as examples, apply our architecture to them and design novel DUNs. In theory, we prove the existence, uniqueness of the solution and convergence of the proposed network, and also prove that our network obtains a fast linear convergence rate. Extensive experiments verify the effectiveness and advantages of our architecture. Lin Kong, Wei Sun 0049, Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001 |
AAAI | 5 |
| 2022 | Kill a Bird with Two Stones: Closing the Convergence Gaps in Non-Strongly Convex Optimization by Directly Accelerated SVRG with Double Compensation and SnapshotsabstractRecently, some accelerated stochastic variance reduction algorithms such as Katyusha and ASVRG-ADMM achieve faster convergence than non-accelerated methods such as SVRG and SVRG-ADMM. However, there are still some gaps between the oracle complexities and their lower bounds. To fill in these gaps, this paper proposes a novel Directly Accelerated stochastic Variance reductIon (DAVIS) algorithm with two Snapshots for non-strongly convex (non-SC) unconstrained problems. Our theoretical results show that DAVIS achieves the optimal convergence rate O(1/(nS^2)) and optimal gradient complexity O(n+\sqrt{nL/\epsilon}), which is identical to its lower bound. To the best of our knowledge, this is the first directly accelerated algorithm that attains the optimal lower bound and improves the convergence rate from O(1/S^2) to O(1/(nS^2)). Moreover, we extend DAVIS and theoretical results to non-SC problems with a structured regularizer, and prove that the proposed algorithm with double-snapshots also attains the optimal convergence rate O(1/(nS)) and optimal oracle complexity O(n+L/\epsilon) for such problems, and it is at least a factor n/S faster than existing accelerated stochastic algorithms, where n\gg S in general. Yuanyuan Liu 0001, Fanhua Shang, Weixin An, Hongying Liu 0001, Zhouchen Lin |
ICML | 4 |
| 2022 | PWPROP: A Progressive Weighted Adaptive Method for Training Deep Neural NetworksabstractIn recent years, adaptive optimization methods for deep learning have attracted considerable attention. AMSGRAD indicates that the adaptive methods may be hard to converge to optimal solutions of some convex problems due to the divergence of its adaptive learning rate as in ADAM. However, we find that AMSGRAD may generalize worse than ADAM for some deep learning tasks. We first show that AMSGRAD may not find a flat minimum. So how can we design an optimization method to find a flat minimum with low training loss? Few works focus on this important problem. We propose a novel progressive weighted adaptive optimization algorithm, called PWPROP, with fewer hyperparameters than its counterparts such as ADAM. By intuitively constructing a “sharp-flat minima” model, we show that how different second-order estimates affect the ability to escape a sharp minimum. Moreover, we also prove that PWPROP can address the non-convergence issue of ADAM and has a sublinear convergence rate for non-convex problems. Extensive experimental results show that PWPROP is effective and suitable for various deep learning architectures such as Transformer, and achieves state-of-the-art results. Dong Wang 0004, Huatian Zhang 0001, Fanhua Shang, Hongying Liu 0001, Yuanyuan Liu 0001, Shengmei Shen |
ICTAI | 5 |
| 2022 | A Numerical DEs Perspective on Unfolded Linearized ADMM Networks for Inverse ProblemsabstractMany research works show that the continuous-time Differential Equations (DEs) allow for a better understanding of traditional Alternating Direction Multiplier Methods (ADMMs). And many unfolded algorithms directly inherit the traditional iterations to build deep networks. Although they obtain a faster convergence rate and superior practical performance, there is a lack of an appropriate explanation of the unfolded network architectures. Thus, we attempt to explore the connection between the existing unfolded Linearized ADMM (LADMM) and numerical DEs, and propose efficient unfolded network design schemes. First, we present an unfolded Euler LADMM scheme as a by-product, which originates from the Euler method for solving first-order DEs. Then inspired by the trapezoid method in numerical DEs, we design a new more effective network scheme, called unfolded Trapezoid LADMM scheme. Moreover, we analyze that the Trapezoid LADMM scheme has higher precision than the Euler LADMM scheme. To the best of our knowledge, this is the first work to explore the connection between unfolded ADMMs and numerical DEs with theoretical guarantees. Finally, we instantiate our Euler LADMM and Trapezoid LADMM schemes into ELADMM and TLADMM with the proximal operators, and ELADMM-Net and TLADMM-Net with convolutional neural networks. And extensive experiments show that our algorithms are competitive with state-of-the-art methods. Weixin An, Yingjie Yue, Yuanyuan Liu 0001, Fanhua Shang, Hongying Liu 0001 |
ACM Multimedia | 5 |
| 2022 | Balanced Gradient Penalty Improves Deep Long-Tailed LearningabstractIn recent years, deep learning has achieved a great success in various image recognition tasks. However, the long-tailed setting over a semantic class plays a leading role in real-world applications. Common methods focus on optimization on balanced distribution or naive models. Few works explore long-tailed learning from a deep learning-based generalization perspective. The loss landscape on long-tailed learning is first investigated in this work. Empirical results show that sharpness-aware optimizers work not well on long-tailed learning. Because they do not take class priors into consideration, and they fail to improve performance of few-shot classes. To better guide the network and explicitly alleviate sharpness without extra computational burden, we develop a universal Balanced Gradient Penalty (BGP) method. Surprisingly, our BGP method does not need the detailed class priors and preserves privacy. Our new algorithm BGP, as a regularization loss, can achieve the state-of-the-art results on various image datasets (i.e., CIFAR-LT, ImageNet-LT and iNaturalist-2018) in the settings of different imbalance ratios. Dong Wang 0004, Liangji Fang, Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001 |
ACM Multimedia | 6 |
| 2022 | Loopless Variance Reduced Stochastic ADMM for Equality Constrained Problems in IoT ApplicationsabstractThe alternating direction method of multipliers (ADMMs) is an efficient optimization method for solving equality constrained problems in Internet of Things (IoT) applications. Recently, several stochastic variance reduced ADMM algorithms (e.g., SVRG-ADMM) have made exciting progress, such as linear convergence for strongly convex (SC) problems. However, SVRG-ADMM and its variants have an outer loop where the full gradient at the snapshot is computed, and their outer loop contains an inner loop, in which a large number of variance reduced gradients are estimated from random samples. This loopy design makes these methods more complex to analyze and determine the inner loop length, which must be proportional to the condition number to achieve best convergence, and is often set to$\mathcal {O}(n)$as a suboptimal choice, where$n$is the number of samples. To tackle these issues, we propose an efficient loopless variance reduced stochastic ADMM algorithm, called LVR-SADMM. In our LVR-SADMM, we remove the outer loop and replace it with a biased coin-flip, in which we update the snapshot with a small probability to trigger the full gradient computation. Moreover, we also theoretically analyze the convergence property of LVR-SADMM, which shows that it enjoys a fast linear convergence rate for SC problems. In particular, we also present an accelerated loopless SVRG-ADMM (LAVR-SADMM) method for both SC and non-SC problems. Various experimental results on many real-world data sets verify that the proposed methods can achieve an average speedup of$2\times $in the SC case and$5\times $in the non-SC case over their loopy counterparts, respectively. Yuanyuan Liu 0001, Jiacheng Geng, Fanhua Shang, Weixin An, Hongying Liu 0001 |
IEEE Internet Things J. | 5 |
| 2022 | Laplacian Smoothing Stochastic ADMMs With Differential Privacy GuaranteesabstractMany machine learning tasks such as structured sparse coding and multi-task learning can be converted into an equality constrained optimization problem. The stochastic alternating direction method of multipliers (SADMM) is a popular algorithm to solve such large-scale problems, and has been successfully used in many real-world applications. However, existing SADMMs fail to take into consideration an important issue in their designs, i.e., protecting sensitive information. To address this challenging issue, this paper proposes a novel differential privacy stochastic ADMM framework for solving equality constrained machine learning problems. In particular, to further lift the utility in privacy-preserving equality constrained optimization, a Laplacian smoothing operation is also introduced into our differential privacy ADMM framework, and it can smooth out the Gaussian noise used in the Gaussian mechanism. Then we propose an efficient differentially private variance reduced stochastic ADMM (DP-VRADMM) algorithm with Laplacian smoothing for both strongly convex and general convex objectives. As a by-product, we also present a new differentially private stochastic ADMM algorithm with DP guarantees. In theory, we provide both private guarantees and utility guarantees for the proposed algorithms, which show that Laplacian smoothing can improve the utility bounds of our algorithms. Experimental results on real-world datasets verify our theoretical results and the effectiveness of our algorithms. Yuanyuan Liu 0001, Jiacheng Geng, Fanhua Shang, Weixin An, Hongying Liu 0001, Wei Feng 0005 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Asynchronous Parallel, Sparse Approximated SVRG for High-Dimensional Machine LearningabstractWith the increasing of the data size and the development of multi-core computers, asynchronous parallel stochastic optimization algorithms such as KroMagnon have gained significant attention. In this paper, we propose a new Sparse approximation and asynchronous parallel Stochastic Variance Reduced Gradient (SSVRG) method for sparse and high-dimensional machine learning problems. Unlike standard SVRG and its asynchronous parallel variant, KroMagnon, the snapshot point of SSVRG is set to the average of all the iterates in the previous epoch, which allows it to take much larger learning rates and also makes it more robust to the choice of learning rates. In particular, we use the sparse approximation of the popular SVRG estimator to perform completely sparse updates at all iterations. Therefore, SSVRG has a much lower per-iteration computational cost than its dense counterpart, SVRG++, and is very friendly to asynchronous parallel implementation. Moreover, we provide the convergence guarantees of SSVRG for both strongly convex and non-strongly convex problems, while existing asynchronous algorithms (e.g., KroMagnon and ASAGA) only have convergence guarantees for strongly convex problems. Finally, we extend SSVRG to non-smooth and asynchronous parallel settings. Numerical experimental results demonstrate that SSVRG converges significantly faster than the state-of-the-art asynchronous parallel methods, e.g., KroMagnon, and is usually more than three orders of magnitude faster than SVRG++. Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Efficient Gradient Support Pursuit With Less Hard Thresholding for Cardinality-Constrained LearningabstractRecently, stochastic hard thresholding (HT) optimization methods [e.g., stochastic variance reduced gradient hard thresholding (SVRGHT)] are becoming more attractive for solving large-scale sparsity/rank-constrained problems. However, they have much higher HT oracle complexities, especially for high-dimensional data or large-scale matrices. To address this issue and inspired by the well-known Gradient Support Pursuit (GraSP) method, this article proposes a new Relaxed Gradient Support Pursuit (RGraSP) framework. Unlike GraSP, RGraSP only requires to yield an approximation solution at each iteration. Based on the property of RGraSP, we also present an efficient stochastic variance reduction-gradient support pursuit algorithm and its fast version (called stochastic variance reduced gradient support pursuit (SVRGSP+). We prove that the gradient oracle complexity of both our algorithms is two times less than that of SVRGHT. In particular, their HT complexity is about$\kappa _{\widehat {s}}$times less than that of SVRGHT, where$\kappa _{\widehat {s}}$is the restricted condition number. Moreover, we prove that our algorithms enjoy fast linear convergence to an approximately global optimum, and also present an asynchronous parallel variant to deal with very high-dimensional and sparse data. Experimental results on both synthetic and real-world datasets show that our algorithms yield superior results than the state-of-the-art gradient HT methods. Fanhua Shang, Bingkun Wei, Hongying Liu 0001, Yuanyuan Liu 0001, Pan Zhou 0002, Maoguo Gong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Learned Extragradient ISTA with Interpretable Residual Structures for Sparse CodingabstractRecently, the study on learned iterative shrinkage thresholding algorithm (LISTA) has attracted increasing attentions. A large number of experiments as well as some theories have proved the high efficiency of LISTA for solving sparse coding problems. However, existing LISTA methods are all serial connection. To address this issue, we propose a novel extragradient based LISTA (ELISTA), which has a residual structure and theoretical guarantees. Moreover, most LISTA methods use the soft thresholding function, which has been found to cause a large estimation bias. Therefore, we propose a thresholding function for ELISTA instead of soft thresholding. From a theoretical perspective, we prove that our method attains linear convergence. Through ablation experiments, the improvements of our method on the network structure and the thresholding function are verified in practice. Extensive empirical results verify the advantages of our method. Yangyang Li 0001, Lin Kong, Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001, Zhouchen Lin |
AAAI | 5 |
| 2021 | Large Motion Video Super-Resolution with Dual Subnet and Multi-Stage Communicated UpsamplingabstractVideo super-resolution (VSR) aims at restoring a video in low-resolution (LR) and improving it to higher-resolution (HR). Due to the characteristics of video tasks, it is very important that motion information among frames should be well concerned, summarized and utilized for guidance in a VSR algorithm. Especially, when a video contains large motion, conventional methods easily bring incoherent results or artifacts. In this paper, we propose a novel deep neural network with Dual Subnet and Multi-stage Communicated Upsampling (DSMC) for super-resolution of videos with large motion. We design a new module named U-shaped residual dense network with 3D convolution (U3D-RDN) for fine implicit motion estimation and motion compensation (MEMC) as well as coarse spatial feature extraction. And we present a new Multi-Stage Communicated Upsampling (MSCU) module to make full use of the intermediate results of upsampling for guiding the VSR. Moreover, a novel dual subnet is devised to aid the training of our DSMC, whose dual loss helps to reduce the solution space as well as enhance the generalization ability. Our experimental results confirm that our method achieves superior performance on videos with large motion compared to state-of-the-art methods. Hongying Liu 0001, Zhubo Ruan, Fanhua Shang, Yuanyuan Liu 0001 |
AAAI | 1 |
| 2021 | Behavior Mimics Distribution: Combining Individual and Group Behaviors for Federated LearningabstractFederated Learning (FL) has become an active and promising distributed machine learning paradigm. As a result of statistical heterogeneity, recent studies clearly show that the performance of popular FL methods (e.g., FedAvg) deteriorates dramatically due to the client drift caused by local updates. This paper proposes a novel Federated Learning algorithm (called IGFL), which leverages both Individual and Group behaviors to mimic distribution, thereby improving the ability to deal with heterogeneity. Unlike existing FL methods, our IGFL can be applied to both client and server optimization. As a by-product, we propose a new attention-based federated learning in the server optimization of IGFL. To the best of our knowledge, this is the first time to incorporate attention mechanisms into federated optimization. We conduct extensive experiments and show that IGFL can significantly improve the performance of existing federated learning methods. Especially when the distributions of data among individuals are diverse, IGFL can improve the classification accuracy by about 13% compared with prior baselines. Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001 |
IJCAI | 4 |
| 2021 | Progressive Semantic Matching for Video-Text RetrievalabstractCross-modal retrieval between texts and videos is important yet challenging. Until recently, previous works in this domain typically rely on learning a common space to match the text and video, but it is difficult to match due to the semantic gap between videos and texts. Although some methods employ coarse-to-fine or multi-expert networks to encode one or more common spaces for easier matching, they almost directly optimize one matching space, which is challenging, because of the huge semantic gap between different modalities. To address this issue, we aim at narrowing semantic gap by a progressive learning process with a coarse-to-fine architecture, and propose a novel Progressive Semantic Matching (PSM) method. We first construct a multilevel encoding network for videos and texts, and design some auxiliary common spaces, which are mapped by the outputs of encoders in different levels. Then all the common spaces are jointly trained end to end. In this way, the model can effectively encode videos and texts into a fusion common space by a progressive paradigm. Experimental results on three video-text datasets (i.e., MSR-VTT, TIGF and MSVD) demonstrate the advantages of our PSM, which achieves significant performance improvement compared with state-of-the-art approaches. Hongying Liu 0001, Ruyi Luo, Fanhua Shang, Mantang Niu, Yuanyuan Liu 0001 |
ACM Multimedia | 1 |
| 2021 | Principal component analysis in the stochastic differential privacy modelabstractIn this paper, we study the differentially private Principal Component Analysis (PCA) problem in stochastic optimization settings. We first propose a new stochastic gradient perturbation PCA mechanism (DP-SPCA) for the calculation of the right singular subspace to achieve $(\epsilon,\delta)$-differential privacy. For achieving a better utility guarantee and performance, we then present a new differential privacy stochastic variance reduction mechanism (DP-VRPCA) with gradient perturbation for PCA. To the best of our knowledge, this is the first work of stochastic gradient perturbation for $(\epsilon,\delta)$-differentially private PCA. We also compare the proposed algorithms with existing state-of-the-art methods, and experiments on real-world datasets and on classification tasks confirm the improved theoretical guarantees of our algorithms. Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001 |
UAI | 5 |
| 2021 | Accelerated Variance Reduction Stochastic ADMM for Large-Scale Machine LearningabstractRecently, many stochastic variance reduced alternating direction methods of multipliers (ADMMs) (e.g., SAG-ADMM and SVRG-ADMM) have made exciting progress such as linear convergence rate for strongly convex (SC) problems. However, their best-known convergence rate for non-strongly convex (non-SC) problems is$\mathcal {O}(1/T)$as opposed to$\mathcal {O}(1/T^2)$of accelerated deterministic algorithms, where$T$is the number of iterations. Thus, there remains a gap in the convergence rates of existing stochastic ADMM and deterministic algorithms. To bridge this gap, we introduce a new momentum acceleration trick into stochastic variance reduced ADMM, and propose a novel accelerated SVRG-ADMM method (called ASVRG-ADMM) for the machine learning problems with the constraint$Ax + By = c$. Then we design a linearized proximal update rule and a simple proximal one for the two classes of ADMM-style problems with$B = \tau I$and$B\ne \tau I$, respectively, where$I$is an identity matrix and$\tau$is an arbitrary bounded constant. Note that our linearized proximal update rule can avoid solving sub-problems iteratively. Moreover, we prove that ASVRG-ADMM converges linearly for SC problems. In particular, ASVRG-ADMM improves the convergence rate from$\mathcal {O}(1/T)$to$\mathcal {O}(1/T^2)$for non-SC problems. Finally, we apply ASVRG-ADMM to various machine learning problems, e.g., graph-guided fused Lasso, graph-guided logistic regression, graph-guided SVM, generalized graph-guided fused Lasso and multi-task learning, and show that ASVRG-ADMM consistently converges faster than the state-of-the-art methods. Yuanyuan Liu 0001, Fanhua Shang, Hongying Liu 0001, Lin Kong, Licheng Jiao, Zhouchen Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Differentially Private ADMM Algorithms for Machine LearningabstractIn this paper, we study efficient differentially private alternating direction methods of multipliers (ADMM) via gradient perturbation for many centralized machine learning problems. For smooth convex loss functions with (non)-smooth regularization, we propose the first differentially private ADMM (DP-ADMM) algorithm with the performance guarantee of (ϵ,δ)-differential privacy ((ϵ,δ)-DP). From the viewpoint of theoretical analysis, we use the Gaussian mechanism and the conversion relationship between Rényi Differential Privacy (RDP) and DP to perform a comprehensive privacy analysis for our algorithm. Then we establish a new criterion to prove the convergence of the proposed algorithms including DP-ADMM. We also give the utility analysis of our DP-ADMM. Moreover, we propose a new accelerated DP-ADMM (DP-AccADMM) algorithm with the Nesterov’s acceleration technique. Finally, we conduct numerical experiments on many real-world datasets to show the privacy-utility tradeoff of the two proposed algorithms, and all the comparative analysis shows that DP-AccADMM converges faster and has a better utility than DP-ADMM, when the privacy budget ϵ is larger than a threshold. Fanhua Shang, Yuanyuan Liu 0001, Hongying Liu 0001, Longjie Shen, Maoguo Gong |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | VR-SGD: A Simple Stochastic Variance Reduction Method for Machine LearningabstractIn this paper, we propose a simple variant of the original SVRG, called variance reduced stochastic gradient descent (VR-SGD). Unlike the choices of snapshot and starting points in SVRG and its proximal variant, Prox-SVRG, the two vectors of VR-SGD are set to the average and last iterate of the previous epoch, respectively. The settings allow us to use much larger learning rates, and also make our convergence analysis more challenging. We also design two different update rules for smooth and nonsmooth objective functions, respectively, which means that VR-SGD can tackle non-smooth and/or non-strongly convex problems directly without any reduction techniques. Moreover, we analyze the convergence properties of VR-SGD for strongly convex problems, which show that VR-SGD attains linear convergence. Different from most algorithms that have no convergence guarantees for nonstrongly convex problems, we also provide the convergence guarantees of VR-SGD for this case, and empirically verify that VR-SGD with varying learning rates achieves similar performance to its momentum accelerated variant that has the optimal convergence rate O(1=T2). Finally, we apply VR-SGD to solve various machine learning problems, such as convex and non-convex empirical risk minimization, and leading eigenvalue computation. Experimental results show that VR-SGD converges significantly faster than SVRG and Prox-SVRG, and usually outperforms state-of-the-art accelerated methods, e.g., Katyusha. Fanhua Shang, Kaiwen Zhou 0001, Hongying Liu 0001, James Cheng, Ivor W. Tsang, Lijun Zhang 0005, Dacheng Tao, Licheng Jiao |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Sparse Manifold-Regularized Neural Networks for Polarimetric SAR Terrain ClassificationabstractIn this article, a new deep neural network based on sparse filtering and manifold regularization (DSMR) is proposed for feature extraction and classification of polarimetric synthetic aperture radar (PolSAR) data. DSMR uses a novel deep neural network (DNN) to automatically learn features from raw SAR data. During preprocessing, the spatial information between pixels on PolSAR images is exploited to weight each data sample. Then, in the pretraining and fine-tuning, DSMR uses the population sparsity and the lifetime sparsity (dual sparsity) to learn the global features and preserves the local structure of data by neighborhood-based manifold regularization. The dual sparsity only needs to tune a few parameters, and the manifold regularization cuts down the number of training samples. Experimental results on synthesized and real PolSAR data sets from different SAR systems show that DSMR can improve classification accuracy compared with conventional DNNs, even for data sets with a large angle of incidence. Hongying Liu 0001, Fanhua Shang, Shuyuan Yang 0001, Maoguo Gong, Tianwen Zhu, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Loopless Semi-Stochastic Gradient Descent with Less Hard Thresholding for Sparse LearningabstractStochastic gradient hard thresholding methods have recently been shown to work favorably for solving large-scale empirical risk minimization problems under sparsity constraints. Many stochastic hard thresholding methods (e.g., SVRG-HT) conduct a full gradient update with a constant frequency and perform a hard thresholding operation at each iteration, which leads to a high computational complexity especially for high-dimensional and sparse problems. To be more efficient in large-scale datasets, we propose an efficient single-layer semi-stochastic gradient hard thresholding (LSSG-HT) method. The proposed algorithm updates full gradient with a given probability p and reduces lots of hard thresholding operations by setting frequency m, which reduces hard thresholding complexity in theory to O(κ_s/młog(1/ε)) compared with O(κ_słog(1/ε)) of SVRG-HT. We prove that our algorithm can converge to an optimal solution with a linear convergence rate. Furthermore, we also present an asynchronous parallel variant of LSSG-HT. Numerical experimental results demonstrate that the efficiency of our algorithms with comparison against the state-of-the-art algorithms. Bingkun Wei, Fanhua Shang, Hongying Liu 0001 |
CIKM | 4 |
| 2019 | Fast Semisupervised Classification Using Histogram-Based Density Estimation for Large-Scale Polarimetric SAR DataabstractIn order to obtain high classification accuracy and reduce time consumption for large-scale polarimetric synthetic-aperture radar (PolSAR) data. In this letter, we propose a fast semisupervised classification algorithm using histogram-based density estimation (called FSHDE). First, a noniterative collaborative training using our proposed Wishart-clustering selection strategy is designed to expand the labeled sample set from unlabeled samples. Second, a fast feature mapping based on histogram density estimation is employed to reliably capture the interaction of nonlinear features. Third, submodular optimization is used to select optimal subspace features to reduce feature correlation. Experimental results on synthetic and real PolSAR data indicate that FSHDE greatly reduces the time consumption and improves the accuracy for terrain classification compared with the state-of-the-art methods. Hongying Liu 0001, Feixiang Wang, Shuyuan Yang 0001, Biao Hou, Licheng Jiao, Ri Yang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Hyperspectral image classification based on stacked marginal discriminative autoencoderabstractIn this paper, a novel stacked marginal discriminative autoencoder (SMDAE) method is proposed for hyperspectral image classification. It uses a deep neural network to learn discriminative features from hyperspectral images automatically. In hyperspectral images, the collection of training samples is difficult. When the number of training samples is not enough, these training samples are difficult to estimate the statistical distribution of hyperspectral images accurately. In order to solve the small sample problem and improve the classification performance of the autoencoder, the marginal samples are selected through the distribution characteristics of samples. The marginal samples are searched based on k nearest neighbors between different classes. These samples are used to fine-tune the SMDAE network. The experimental results show that the proposed SMDAE method can achieve satisfying performance under small training set. Jie Feng 0003, Liguo Liu, Xiangrong Zhang, Rongfang Wang, Hongying Liu 0001 |
IGARSS | 5 |
| 2017 | Fast Classification for Large Polarimetric SAR Data Based on Refined Spatial-Anchor GraphabstractThe graph model-based semisupervised machine learning is well established. However, its computational complexity is still high in terms of the time consumption especially for large data. In this letter, we propose a fast semisupervised classification algorithm using the recently presented spatial-anchor graph for a large polarimetric synthetic aperture radar (Pol-SAR) data, named as Fast Spatial-Anchor Graph (FSAG) based algorithm. Based on an initial superpixel segmentation on the PolSAR image, the homogenous regions are obtained. The border pixels are reassigned to the most similar superpixel according to majority voting and distance measurement. Then, feature vectors are weighted within local homogenous regions. The refined spatial-anchor graph is constructed with these regions, and the semisupervised classification is conducted. Experimental results on synthesized and real PolSAR data indicate that the proposed FSAG greatly reduces time consumption and maintains the accuracy for terrain classifications compared with state-of-the-art graph-based approaches. Hongying Liu 0001, Shuyuan Yang 0001, Shuiping Gou, Puhua Chen, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Unsupervised saliency-guided SAR image change detection
Yaoguo Zheng, Licheng Jiao, Hongying Liu 0001, Xiangrong Zhang, Biao Hou, Shuang Wang 0001 |
Pattern Recognit. | 3 |
| 2017 | Deep Fully Convolutional Network-Based Spatial Distribution Prediction for Hyperspectral Image ClassificationabstractMost of the existing spatial-spectral-based hyperspectral image classification (HSIC) methods mainly extract the spatial-spectral information by combining the pixels in a small neighborhood or aggregating the statistical and morphological characteristics. However, those strategies can only generate shallow appearance features with limited representative ability for classes with high interclass similarity and spatial diversity and therefore reduce the classification accuracy. To this end, we present a novel HSIC framework, named deep multiscale spatial-spectral feature extraction algorithm, which focuses on learning effective discriminant features for HSIC. First, the well pretrained deep fully convolutional network based on VGG-verydeep-16 is introduced to excavate the potential deep multiscale spatial structural information in the proposed hyperspectral imaging framework. Then, the spectral feature and the deep multiscale spatial feature are fused by adopting the weighted fusion method. Finally, the fusion feature is put into a generic classifier to obtain the pixelwise classification. Compared with the existing spectral-spatial-based classification techniques, the proposed method provides the state-of-the-art performance and is much more effective, especially for images with high nonlinear distribution and spatial diversity. Licheng Jiao, Miaomiao Liang, Huan Chen 0006, Shuyuan Yang 0001, Hongying Liu 0001, Xianghai Cao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | Superpixel-Based Multiple Local CNN for Panchromatic and Multispectral Image ClassificationabstractRecently, very high resolution (VHR) panchromatic and multispectral (MS) remote-sensing images can be acquired easily. However, it is still a challenging task to fuse and classify these VHR images. Generally, there are two ways for the fusion and classification of panchromatic and MS images. One way is to use a panchromatic image to sharpen an MS image, and then classify a pan-sharpened MS image. Another way is to extract features from panchromatic and MS images, respectively, and then combine these features for classification. In this paper, we propose a superpixel-based multiple local convolution neural network (SML-CNN) model for panchromatic and MS images classification. In order to reduce the amount of input data for the CNN, we extend simple linear iterative clustering algorithm for segmenting MS images and generating superpixels. Superpixels are taken as the basic analysis unit instead of pixels. To make full advantage of the spatial-spectral and environment information of superpixels, a superpixel-based multiple local regions joint representation method is proposed. Then, an SML-CNN model is established to extract an efficient joint feature representation. A softmax layer is used to classify these features learned by multiple local CNN into different categories. Finally, in order to eliminate the adverse effects on the classification results within and between superpixels, we propose a multi-information modification strategy that combines the detailed information and semantic information to improve the classification performance. Experiments on the classification of Vancouver and Xi’an panchromatic and MS image data sets have demonstrated the effectiveness of the proposed approach. Wei Zhao 0014, Licheng Jiao, Wenping Ma 0001, Jiaqi Zhao 0001, Jin Zhao 0002, Hongying Liu 0001, Xianghai Cao, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2016 | Terrain classification with Polarimetric SAR based on Deep Sparse Filtering NetworkabstractA new method for Polarimetric Synthetic Aperture Radar (PolSAR) terrain classification based on Deep Sparse Filtering Network (DSFN) is proposed in this paper. It uses a novel deep learning network to learn features from the input raw data automatically. And the spatial information between pixels on PolSAR image is combined into the input data. Moreover, unlike the conventional deep networks, the DSFN only needs to tune very few parameters during pre-training and fine-tuning. A real PolSAR data is used to verify the proposed method. Experimental results show that the proposed DSFN is efficient with less parameters and effectively improves the classification accuracy compared with conventional deep networks. Hongying Liu 0001, Qiang Min, Jin Zhao 0002, Shuyuan Yang 0001, Biao Hou, Jie Feng 0003, Licheng Jiao |
IGARSS | 1 |
| 2016 | Fast semi-supervised classification based on parallel auction graph for polarimetric SAR dataabstractAlthough the graph-based machine learning has received considerable attention in the remote sensing area and it has been widely used for terrain classification, the construction of graph in most existing algorithms still takes large memory and plenty of computational time especially for large Polarimetric Synthetic Aperture Radar (PolSAR) data. Addressing these issues, we propose a fast semi-supervised classification method based on parallel auction graph in this paper. The spatial relation between pixels is firstly preprocessed using the superpixel segmentation. Then we divide the PolSAR data into multiple groups, and each of them is used to construct a sparse auction graph. The semi-supervised classification is performed parallel on those graphs. Experimental results on simulated and real PolSAR data demonstrate its efficiency and effectiveness compared with existing methods. Hongying Liu 0001, Xing Xing, Shigang Wang 0001, Zhixi Feng, Erlei Zhang, Shuyuan Yang 0001, Biao Hou, Licheng Jiao |
IGARSS | 1 |
| 2016 | New classifier based on compressed dictionary and LS-SVM
Licheng Jiao, Hongying Liu 0001, Shuyuan Yang 0001 |
Neurocomputing | 3 |
| 2016 | SAR image target recognition via Complementary Spatial Pyramid Coding
Shaona Wang, Licheng Jiao, Shuyuan Yang 0001, Hongying Liu 0001 |
Neurocomputing | 4 |
| 2016 | Weighted multifeature hyperspectral image classification via kernel joint sparse representation
Erlei Zhang, Xiangrong Zhang, Licheng Jiao, Hongying Liu 0001, Shuang Wang 0001, Biao Hou |
Neurocomputing | 4 |
| 2016 | Unsupervised High-Level Feature Extraction of SAR Imagery With Structured Sparsity Priors and Incremental Dictionary LearningabstractSparse representation is an effective model for high-level feature extraction, and the dictionary is critical, since it can provide a sparse and discriminative feature for image classification. However, the traditional sparse model with ℓ1- norm is unstable and ignores spatial context dependence. Furthermore, the traditional off-line dictionary learning is less efficient. In this letter, a high-level feature extraction approach is proposed, in which structured sparsity priors are imposed on the sparse representation to exploit the context dependence and an incremental structured dictionary learning method is proposed to exploit the inherent structures of a dictionary. The experiment results on unsupervised synthetic aperture radar imagery classification show that the structured priors improve classification performance and the proposed algorithm is more efficient in dictionary learning compared with existing works. Jiawei Chen 0001, Licheng Jiao, Wenping Ma 0001, Hongying Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2016 | Hierarchical semantic model and scattering mechanism based PolSAR image classification
Fang Liu 0001, Junfei Shi, Licheng Jiao, Hongying Liu 0001, Shuyuan Yang 0001, Jie Wu 0016, Hongxia Hao, Jialing Yuan |
Pattern Recognit. | 4 |
| 2016 | Classification and saliency detection by semi-supervised low-rank representation
Miaoyun Zhao, Licheng Jiao, Wenping Ma 0001, Hongying Liu 0001, Shuyuan Yang 0001 |
Pattern Recognit. | 4 |
| 2016 | Multiple Kernel Learning Based on Discriminative Kernel Clustering for Hyperspectral Band SelectionabstractIn hyperspectral images, band selection plays a crucial role for land-cover classification. Multiple kernel learning (MKL) is a popular feature selection method by selecting the relevant features and classifying the images simultaneously. Unfortunately, a large number of spectral bands in hyperspectral images result in excessive kernels, which limit the application of MKL. To address this problem, a novel MKL method based on discriminative kernel clustering (DKC) is proposed. In the proposed method, a discriminative kernel alignment (KA) (DKA) is defined. Traditional KA measures kernel similarity independently of the current classification task. Compared with KA, DKA measures the similarity of discriminative information by introducing the comparison of intraclass and interclass similarities. It can evaluate both kernel redundancy and kernel synergy for classification. Then, DKA-based affinity-propagation clustering is devised to reduce the kernel scale and retain the kernels having high discrimination and low redundancy for classification. Additionally, an analysis of necessity for DKC in hyperspectral band selection is provided by empirical Rademacher complexity. Experimental results on several hyperspectral images demonstrate the effectiveness of the proposed band selection method in terms of classification performance and computation efficiency. Jie Feng 0003, Licheng Jiao, Tao Sun 0007, Hongying Liu 0001, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Semi-supervised classification based on anchor-spatial graph for large polarimetric SAR dataabstractRecently a few works of semi-supervised learning methods based on graph have been proposed for remote sensing. The common idea of these methods are that they build a graph using the samples of the image. Most of their time complexity is relatively large, and they ignore the spatial information of the image, which leads to unsatisfactory classification results. this paper proposes a novel semi-supervised classification method based on anchor-spatial graph for large PolSAR data. Firstly the unsupervised Wishart clustering is performed to select representative samples, which served as anchors according to the least distance between samples. Then an anchor graph is built using the selected anchors according to the multiple features of the samples. And it is further combined with the spatial information of the samples to construct an anchor-spatial graph. Finally the class information from small quantities of labeled samples propagates to the unlabeled ones. Experimental results show that the proposed method has a low time complexity compared with existing works and it could effectively cut down the processing time for large PolSAR data meanwhile keeps the classification accuracy. Hongying Liu 0001, Dexiang Zhu, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 1 |
| 2015 | A Three-Component Fisher-Based Feature Weighting Method for Supervised PolSAR Image ClassificationabstractThis letter presents a feature weighting method for polarimetric synthetic aperture radar (PolSAR) image classification. Appropriate feature weighting is essential for obtaining accurate classifications but so far has remained an open research problem. We propose in this letter a supervised three-component feature weighting method based on the Fisher linear discriminant. Fisher linear discriminant method is used to calculate a coefficient for each feature. Then, these coefficients are modified according to a three-component scattering power decomposition model, combining both physical and statistical scattering characteristics to adapt them for the particular scattering mechanisms inherent in PolSAR data and assigned to the coherency matrix to enhance the discriminating ability of the features. Freeman decomposition and Wishart classifier are used to classify the PolSAR image. The effectiveness of the proposed method is demonstrated by experiments NASA/JPL AIRSAR L-band and CSA Radarsat-2 C-band PolSAR images of the San Francisco area. Bo Chen 0001, Shuang Wang 0001, Licheng Jiao, Rustam Stolkin, Hongying Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2015 | Fast Multifeature Joint Sparse Representation for Hyperspectral Image ClassificationabstractSince hyperspectral images (HSIs) usually have complex content and chaotic background, multiple kinds of features would be helpful for the classification task. Recently, representation-based methods with multifeature combination learning have been proposed. However, multifeature learning and the extended contextual information require much more computational burden, particularly for a large-scale dictionary case. In this letter, we propose a fast joint sparse representation classification method with multifeature combination learning for hyperspectral imagery. Once getting several complementary features (spectral, shape, and texture), the proposed model simultaneously acquires a representation vector for each kind of feature and imposes the joint sparsity ℓrow,0-norm regularization on the representation coefficients. The regularization can enforce the coefficients to share a common sparsity pattern, which preserves the crossfeature information. A new version of the simultaneous orthogonal matching pursuit is presented to solve the aforementioned problem because of its optimization with strong convergence guarantee and efficiency. Moreover, to further improve the classification performance, we incorporate contextual neighborhood information of the image into each kind of feature. Compared with state-of-the-art algorithms, it has been proved that the proposed algorithm with much less memory requirements performs tens to hundreds of times faster than those on real HSIs, while providing the same (or even better) accuracy. Erlei Zhang, Xiangrong Zhang, Hongying Liu 0001, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | A Resample-Based SVA Algorithm for Sidelobe Reduction of SAR/ISAR Imagery With Noninteger Nyquist Sampling RateabstractA resample-based spatial variant apodization (SVA) algorithm for sidelobe reduction was studied for synthetic aperture radar (SAR) and inverse SAR (ISAR) imagery with a noninteger Nyquist sampling rate. The weighting function of every sample in the image domain was calculated with the sample and two adjacent noninteger samples. The noninteger samples were obtained by interpolation in the image domain using sinc function. With the proper selection of two noninteger samples, the monotonic property of the weighting function on each side of the sampling point was preserved. The unequivocal determination of sidelobe suppression was achieved for noninteger Nyquist sampled (NINS) SAR and ISAR imagery. In addition, the lower and upper boundaries of the weighting function under the cosine-on-pedestal condition were extended for further sidelobe suppression and main lobe sharpening. The algorithm was implemented and applied to NINS imagery that is simulated. The algorithm was then assessed for acquired SAR and ISAR images. Improved results have been qualitatively and quantitatively achieved in sidelobe suppression and main lobe sharping in comparison with an existing algorithm. Shuang Wang 0001, Biao Hou, Yong Wang 0011, Hongying Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2014 | Unsupervised classification of polarimetric SAR images integrating color featuresabstractIn conventional terrain classification for the polarimetric SAR (POLSAR) images, color features are rarely involved unless in one recent supervised work. Unlike that work, the color features are exploited for the unsupervised classification in this paper. Firstly, based on the polarimetric decomposition of the POLSAR data, the common color spaces, such as RGB, HSI, and CIELab are calculated. The color feature is quantitatively selected from these color spaces by introducing the color entropy. Then together with the spatial information, extended scattering power entropy and the copolarized ratio, the adaptive Mean-shift algorithm is used to segment the POLSAR image. Finally, the segments are merged according to the Wishart distance measurement. The experiments using AIRSAR L-band POLSAR data indicate that the proposed method has better discriminative ability for urban areas and for boundary preservation compared with existing works. Hongying Liu 0001, Shuang Wang 0001, Biao Hou, Shuyuan Yang 0001, Junfei Shi, Licheng Jiao |
IGARSS | 1 |
| 2014 | Improve the performance of co-training by committee with refinement of class probability estimations
Shuang Wang 0001, Linsheng Wu, Licheng Jiao, Hongying Liu 0001 |
Neurocomputing | 4 |
| 2014 | Semi-supervised classification via kernel low-rank representation graph
Shuyuan Yang 0001, Zhixi Feng, Hongying Liu 0001, Licheng Jiao |
Knowl. Based Syst. | 4 |
| 2011 | Electromagnetic Analysis Enhancement with Signal Processing Techniques (Poster)
Hongying Liu 0001, Yukiyasu Tsunoo, Satoshi Goto |
ACISP | 1 |