EDBT 2026 Demo / reviewers in the wild / expert
Guoqiang Wu
dblp:98/4857
· DBLP profile ↗
32ranked-venue papers
11as first author
26since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 10 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Principled efficient triple-weighting for AUC-oriented imbalanced covariate shift
Yan Zhang 0145, Guoqiang Wu, Yilong Yin |
Expert Syst. Appl. | 5 |
| 2026 | LightRNN: A lightweight recurrent neural network for edge computing
Weihao Xia 0002, Huachuan Wang, Junlong Ma, Guoqiang Wu, Yongling Wu, James Lo |
Knowl. Based Syst. | 5 |
| 2026 | Diffusion classifier-driven reward for offline preference-based reinforcement learning
Teng Pang, Bingzheng Wang, Guoqiang Wu, Yilong Yin |
Pattern Recognit. | 3 |
| 2025 | Towards Macro-AUC Oriented Imbalanced Multi-Label Continual LearningabstractIn Continual Learning (CL), while existing work primarily focuses on the multi-class classification task, there has been limited research on Multi-Label Learning (MLL). In practice, MLL datasets are often class-imbalanced, making it inherently challenging, a problem that is even more acute in CL. Due to its sensitivity to imbalance, Macro-AUC is an appropriate and widely used measure in MLL. However, there is no research to optimize Macro-AUC in MLCL specifically. To fill this gap, in this paper, we propose a new memory replay-based method to tackle the imbalance issue for Macro-AUC-oriented MLCL. Specifically, inspired by recent theory work, we propose a new Reweighted Label-Distribution-Aware Margin (RLDAM) loss. Furthermore, to be compatible with the RLDAM loss, a new memory-updating strategy named Weight Retain Updating (WRU) is proposed to maintain the numbers of positive and negative instances of the original dataset in memory. Theoretically, we provide superior generalization analyses of the RLDAM-based algorithm in terms of Macro-AUC, separately in batch MLL and MLCL settings. This is the first work to offer theoretical generalization analyses in MLCL to our knowledge. Finally, a series of experimental results illustrate the effectiveness of our method over several baselines. Yan Zhang 0145, Guoqiang Wu, Bingzheng Wang, Teng Pang, Haoliang Sun, Yilong Yin |
AAAI | 2 |
| 2025 | N3C: Towards Replay-based Novelty Continual Clustering with Class-OverlappingabstractDeep clustering has excelled in batch settings, but little work has addressed the more practical and challenging continual clustering (CC) with shifting data distributions. Additionally, class-overlapping, also a challenging issue, where classes recur across tasks, is common in real-world scenarios. In this paper, we introduce a new framework for CC with class-overlapping, integrating OOD detection to distinguish between old and new classes and a two-step deep clustering process: contrastive learning for feature representation and rehearsal-based learning to retain previous knowledge. We also propose a memory-updating strategy for handling unsupervised data. Experiments validate our approach, examining factors like OOD detection, class-overlapping levels, etc. This work advances continual clustering toward real-world applications. Yan Zhang 0145, Guoqiang Wu, Bingzheng Wang, Teng Pang, Yilong Yin |
ICASSP | 2 |
| 2025 | A Theory for Conditional Generative Modeling on Multiple Data SourcesabstractThe success of large generative models has driven a paradigm shift, leveraging massive multi-source data to enhance model capabilities. However, the interaction among these sources remains theoretically underexplored. This paper provides a first attempt to fill this gap by rigorously analyzing multi-source training in conditional generative modeling, where each condition represents a distinct data source.
Specifically, we establish a general distribution estimation error bound in average total variation distance for conditional maximum likelihood estimation based on the bracketing number.
Our result shows that when source distributions share certain similarities and the model is expressive enough, multi-source training guarantees a sharper bound than single-source training.
We further instantiate the general theory on conditional Gaussian estimation and deep generative models including autoregressive and flexible energy-based models, by characterizing their bracketing numbers.
The results highlight that the number of sources and similarity among source distributions improve the advantage of multi-source training.
Simulations and real-world experiments validate our theory. Rongzhen Wang, Chenyu Zheng, Chongxuan Li, Guoqiang Wu |
ICML | 5 |
| 2025 | Learning to Drift in Extreme Turning with Active Exploration and Gaussian Process Based MPCabstractExtreme cornering in racing often leads to large sideslip angles, presenting a significant challenge for vehicle control. Conventional vehicle controllers struggle to manage this scenario, necessitating the use of a drifting controller. However, the large sideslip angle in drift conditions introduces model mismatch, which in turn affects control precision. To address this issue, we propose a model correction drift controller that integrates Model Predictive Control (MPC) with Gaussian Process Regression (GPR). GPR is employed to correct vehicle model mismatches during both drift equilibrium solving and the MPC optimization process. Additionally, the variance from GPR is utilized to actively explore different cornering drifting velocities, aiming to minimize trajectory tracking errors. The proposed algorithm is validated through simulations on the Simulink-Carsim platform and experiments with a 1:10 scale RC vehicle. In the simulation, the average lateral error with GPR is reduced by 52.8% compared to the non-GPR case. Incorporating exploration further decreases this error by 27.1%. The velocity tracking Root Mean Square Error (RMSE) also decreases by 10.6% with exploration. In the RC car experiment, the average lateral error with GPR is 36.7% lower, and exploration further leads to a 29.0% reduction. Moreover, the velocity tracking RMSE decreases by 7.2% with the inclusion of exploration. Guoqiang Wu, Wangjia Weng, Zhouheng Li, Yonghao Fu, Lei Xie 0007 |
IV | 1 |
| 2025 | RehearMixup: Improving rehearsal-based continual learning
Yan Zhang 0145, Kaiyuan Qi, Guoqiang Wu, Yilong Yin |
Neurocomputing | 4 |
| 2025 | QFAE: Q-Function guided Action Exploration for offline deep reinforcement learning
Teng Pang, Guoqiang Wu, Yan Zhang 0145, Bingzheng Wang, Yilong Yin |
Pattern Recognit. | 2 |
| 2024 | DiffAIL: Diffusion Adversarial Imitation LearningabstractImitation learning aims to solve the problem of defining reward functions in real-world decision-making tasks. The current popular approach is the Adversarial Imitation Learning (AIL) framework, which matches expert state-action occupancy measures to obtain a surrogate reward for forward reinforcement learning. However, the traditional discriminator is a simple binary classifier and doesn't learn an accurate distribution, which may result in failing to identify expert-level state-action pairs induced by the policy interacting with the environment. To address this issue, we propose a method named diffusion adversarial imitation learning (DiffAIL), which introduces the diffusion model into the AIL framework. Specifically, DiffAIL models the state-action pairs as unconditional diffusion models and uses diffusion loss as part of the discriminator's learning objective, which enables the discriminator to capture better expert demonstrations and improve generalization. Experimentally, the results show that our method achieves state-of-the-art performance and significantly surpasses expert demonstration on two benchmark tasks, including the standard state-action setting and state-only settings. Bingzheng Wang, Guoqiang Wu, Teng Pang, Yan Zhang 0145, Yilong Yin |
AAAI | 2 |
| 2024 | Lower Bounds of Uniform Stability in Gradient-Based Bilevel Algorithms for Hyperparameter OptimizationabstractGradient-based bilevel programming leverages unrolling differentiation (UD) or implicit function theorem (IFT) to solve hyperparameter optimization (HO) problems, and is proven effective and scalable in practice.
To understand their generalization behavior, existing works establish upper bounds on the uniform stability of these algorithms, while their tightness is still unclear.
To this end, this paper attempts to establish stability lower bounds for UD-based and IFT-based algorithms.
A central technical challenge arises from the dependency of each outer-level update on the concurrent stage of inner optimization in bilevel programming.
To address this problem, we introduce lower-bounded expansion properties to characterize the instability in update rules which can serve as general tools for lower-bound analysis.
These properties guarantee the hyperparameter divergence at the outer level and the Lipschitz constant of inner output at the inner level in the context of HO.
Guided by these insights, we construct a quadratic example that yields tight lower bounds for the UD-based algorithm and meaningful bounds for a representative IFT-based algorithm.
Our tight result indicates that uniform stability has reached its limit in stability analysis for the UD-based algorithm. Rongzhen Wang, Chenyu Zheng, Guoqiang Wu, Xu Min, Jun Zhou 0011, Chongxuan Li |
NeurIPS | 3 |
| 2024 | On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and CapabilityabstractAutoregressively trained transformers have brought a profound revolution to the world, especially with their in-context learning (ICL) ability to address downstream tasks.
Recently, several studies suggest that transformers learn a mesa-optimizer during autoregressive (AR) pretraining to implement ICL. Namely, the forward pass of the trained transformer is equivalent to optimizing an inner objective function in-context.
However, whether the practical non-convex training dynamics will converge to the ideal mesa-optimizer is still unclear.
Towards filling this gap, we investigate the non-convex dynamics of a one-layer linear causal self-attention model autoregressively trained by gradient flow, where the sequences are generated by an AR process $x_{t+1} = W x_t$. First, under a certain condition of data distribution, we prove that an autoregressively trained transformer learns $W$ by implementing one step of gradient descent to minimize an ordinary least squares (OLS) problem in-context. It then applies the learned $\widehat{W}$ for next-token prediction, thereby verifying the mesa-optimization hypothesis. Next, under the same data conditions, we explore the capability limitations of the obtained mesa-optimizer. We show that a stronger assumption related to the moments of data is the sufficient and necessary condition that the learned mesa-optimizer recovers the distribution. Besides, we conduct exploratory analyses beyond the first data condition
and prove that generally, the trained transformer will not perform vanilla gradient descent for the OLS problem. Finally, our simulation results verify the theoretical results. Chenyu Zheng, Wei Huang 0034, Rongzhen Wang, Guoqiang Wu, Jun Zhu 0001, Chongxuan Li |
NeurIPS | 4 |
| 2024 | Enhanced fatigue resistance of ferroelectric Al0.65Sc0.35N deposited by physical vapor deposition
Danyang Yao, Ruiqing Wang, Xu Ran, Jiuren Zhou, Qikun Wang, Guoqiang Wu, Genquan Han |
Sci. China Inf. Sci. | 9 |
| 2024 | Feature subset selection algorithm based on symmetric uncertainty and interaction factor
Xiangyuan Gu, Guoqiang Wu |
Multim. Tools Appl. | 3 |
| 2023 | Can Infinitely Wide Deep Nets Help Small-data Multi-label Learning?
Guoqiang Wu |
ACML | 1 |
| 2023 | Towards Understanding Generalization of Macro-AUC in Multi-label LearningabstractMacro-AUC is the arithmetic mean of the class-wise AUCs in multi-label learning and is commonly used in practice. However, its theoretical understanding is far lacking. Toward solving it, we characterize the generalization properties of various learning algorithms based on the corresponding surrogate losses w.r.t. Macro-AUC. We theoretically identify a critical factor of the dataset affecting the generalization bounds: the label-wise class imbalance. Our results on the imbalance-aware error bounds show that the widely-used univariate loss-based algorithm is more sensitive to the label-wise class imbalance than the proposed pairwise and reweighted loss-based ones, which probably implies its worse performance. Moreover, empirical results on various datasets corroborate our theory findings. To establish it, technically, we propose a new (and more general) McDiarmid-type concentration inequality, which may be of independent interest. Guoqiang Wu, Chongxuan Li, Yilong Yin |
ICML | 1 |
| 2023 | Revisiting Discriminative vs. Generative Classifiers: Theory and ImplicationsabstractA large-scale deep model pre-trained on massive labeled or unlabeled data transfers well to downstream tasks. Linear evaluation freezes parameters in the pre-trained model and trains a linear classifier separately, which is efficient and attractive for transfer. However, little work has investigated the classifier in linear evaluation except for the default logistic regression. Inspired by the statistical efficiency of naive Bayes, the paper revisits the classical topic on discriminative vs. generative classifiers. Theoretically, the paper considers the surrogate loss instead of the zero-one loss in analyses and generalizes the classical results from binary cases to multiclass ones. We show that, under mild assumptions, multiclass naive Bayes requires $O(\log n)$ samples to approach its asymptotic error while the corresponding multiclass logistic regression requires $O(n)$ samples, where $n$ is the feature dimension. To establish it, we present a multiclass $\mathcal{H}$-consistency bound framework and an explicit bound for logistic loss, which are of independent interests. Simulation results on a mixture of Gaussian validate our theoretical findings. Experiments on various pre-trained deep vision models show that naive Bayes consistently converges faster as the number of data increases. Besides, naive Bayes shows promise in few-shot cases and we observe the "two regimes” phenomenon in pre-trained supervised models. Our code is available at https://github.com/ML-GSAI/Revisiting-Dis-vs-Gen-Classifiers. Chenyu Zheng, Guoqiang Wu, Fan Bao, Yue Cao 0001, Chongxuan Li, Jun Zhu 0001 |
ICML | 2 |
| 2023 | Toward Understanding Generative Data AugmentationabstractGenerative data augmentation, which scales datasets by obtaining fake labeled examples from a trained conditional generative model, boosts classification performance in various learning tasks including (semi-)supervised learning, few-shot learning, and adversarially robust learning. However, little work has theoretically investigated the effect of generative data augmentation. To fill this gap, we establish a general stability bound in this not independently and identically
distributed (non-i.i.d.) setting, where the learned distribution is dependent on the original train set and generally not the same as the true distribution. Our theoretical result includes the divergence between the learned distribution and the true distribution. It shows that generative data augmentation can enjoy a faster learning rate when the order of divergence term is $o(\max\left( \log(m)\beta_m, 1 / \sqrt{m})\right)$, where $m$ is the train set size and $\beta_m$ is the corresponding stability constant. We further specify the learning setup to the Gaussian mixture model and generative adversarial nets. We prove that in both cases, though generative data augmentation does not enjoy a faster learning rate, it can improve the learning guarantees at a constant level when the train set is small, which is significant when the awful overfitting occurs. Simulation results on the Gaussian mixture model and empirical results on generative adversarial nets support our theoretical conclusions. Chenyu Zheng, Guoqiang Wu, Chongxuan Li |
NeurIPS | 2 |
| 2023 | Makeup transfer: A reviewabstractAbstract Makeup transfer (MT) aims to transfer the makeup style from a given reference makeup face image to a source image while preserving face identity and background information. In recent years, MT has attracted the attention of many scholars, and it has a wide range of application prospects and research value. Since then, many methods have been proposed to accomplish MT, most of which are based on Generative Adversarial Network methods. A taxonomy of existing algorithms in the field of MT is first proposed. Then, evaluation methods are proposed, existing methods are analysed, and existing datasets are introduced. This paper finally discusses the current problems in the field of MT and the trend of future research. Feng He 0008, Kai Bai, Yixin Zong, Yimai Jing, Guoqiang Wu, Chen Wang 0026 |
IET Comput. Vis. | 6 |
| 2023 | ACGAN: Age-compensated makeup transfer based on homologous continuity generative adversarial network modelabstractAbstract The authors focus on the makeup transformation problem, which refers to the transfer of makeup from a reference face to a source face image while maintaining the source makeup‐free face image. In recent years, makeup transformation has become a hot issue and a lot of research has been conducted on this basis, but there are some limitations in the existing methods, mainly due to the lack of consideration of age factor, which makes the final generated face makeup images appear not natural and lack appearance attractiveness. In order to further solve this problem, an age‐compensated makeup transformation framework based on homology continuity is proposed. In order to achieve a stable and controllable age‐compensation effect, the authors design a new coding module that can map the face makeup semantic vector into the higher feature space and achieve age compensation by adjusting the direction of the semantic vector. Finally, in order to comprehensively evaluate the effectiveness of the authors’ proposed method, a large number of qualitative and quantitative experiments have been conducted, and the experimental results show that the authors’ proposed framework outperforms existing methods. Guoqiang Wu, Feng He 0008, Yimai Jing, Xin Ning 0001, Chen Wang 0026, Bo Jin 0018 |
IET Comput. Vis. | 1 |
| 2023 | A smoothing Group Lasso based interval type-2 fuzzy neural network for simultaneous feature selection and system identification
Tao Gao 0003, Chen Wang 0026, Guoqiang Wu, Xin Ning 0001, Xiao Bai 0001, Jian Wang 0010 |
Knowl. Based Syst. | 4 |
| 2023 | Three-dimensional Softmax Mechanism Guided Bidirectional GRU Networks for Hyperspectral Remote Sensing Image Classification
Guoqiang Wu, Xin Ning 0001, Luyang Hou, Feng He 0008, Hengmin Zhang, Achyut Shankar |
Signal Process. | 1 |
| 2022 | Piezoelectric Micromachined Ultrasonic Transducer Array-Based Electronic Stethoscope for Internet of Medical ThingsabstractA piezoelectric micromachined ultrasonic transducer (PMUT) array-based smart electronic stethoscope is presented for Internet of Medical Things (IoMT). For the first time, an electronic stethoscope based on a high-performance PMUT array is demonstrated to record mechanoacoustic cardiopulmonary signals from the body. This electronic stethoscope can continuously capture the cardiopulmonary signals and enable the telemedicine as well as the digital smart medical becoming a reality. The measured cardiopulmonary signals are collected using the PMUT array and transmitted to a mobile terminal via Bluetooth. Then, primary diagnosis can be implemented in real time using a customized APP in the terminal. The PMUT array was designed with high acoustic pressure sensitivity, high noise resolution, and wide operation bandwidth. These crucial parameters enabled simultaneous monitoring of multiple health parameters associated with the cardiopulmonary system, namely, respiratory rate, heart rate, lung sounds, and heart sounds. A feasibility test was implemented on patients with preexisting medical condition, in order to detect weak mechanoacoustic signals related to tachyarrhythmia and asthma. This article demonstrates the potential use of smart electronic stethoscope in healthcare applications, such as the assessment of heart and lung diseases, which provides an innovative auscultation technology and brings great benefits to IoMT. Licheng Jia, Chongbin Liu, Yongjie Gao, Chengliang Sun, Guoqiang Wu |
IEEE Internet Things J. | 8 |
| 2021 | Stability and Generalization of Bilevel Programming in Hyperparameter OptimizationabstractThe (gradient-based) bilevel programming framework is widely used in hyperparameter optimization and has achieved excellent performance empirically. Previous theoretical work mainly focuses on its optimization properties, while leaving the analysis on generalization largely open. This paper attempts to address the issue by presenting an expectation bound w.r.t. the validation set based on uniform stability. Our results can explain some mysterious behaviours of the bilevel programming in practice, for instance, overfitting to the validation set. We also present an expectation bound for the classical cross-validation algorithm. Our results suggest that gradient-based algorithms can be better than cross-validation under certain conditions in a theoretical perspective. Furthermore, we prove that regularization terms in both the outer and inner levels can relieve the overfitting problem in gradient-based algorithms. In experiments on feature learning and data reweighting for noisy labels, we corroborate our theoretical findings. Fan Bao, Guoqiang Wu, Chongxuan Li, Jun Zhu 0001, Bo Zhang 0010 |
NeurIPS | 2 |
| 2021 | On the Convergence of Prior-Guided Zeroth-Order Optimization AlgorithmsabstractZeroth-order (ZO) optimization is widely used to handle challenging tasks, such as query-based black-box adversarial attacks and reinforcement learning. Various attempts have been made to integrate prior information into the gradient estimation procedure based on finite differences, with promising empirical results. However, their convergence properties are not well understood. This paper makes an attempt to fill up this gap by analyzing the convergence of prior-guided ZO algorithms under a greedy descent framework with various gradient estimators. We provide a convergence guarantee for the prior-guided random gradient-free (PRGF) algorithms. Moreover, to further accelerate over greedy descent methods, we present a new accelerated random search (ARS) algorithm that incorporates prior information, together with a convergence analysis. Finally, our theoretical results are confirmed by experiments on several numerical benchmarks as well as adversarial attacks. Shuyu Cheng, Guoqiang Wu, Jun Zhu 0001 |
NeurIPS | 2 |
| 2021 | Rethinking and Reweighting the Univariate Losses for Multi-Label Ranking: Consistency and GeneralizationabstractThe (partial) ranking loss is a commonly used evaluation measure for multi-label classification, which is usually optimized with convex surrogates for computational efficiency. Prior theoretical efforts on multi-label ranking mainly focus on (Fisher) consistency analyses. However, there is a gap between existing theory and practice --- some inconsistent pairwise losses can lead to promising performance, while some consistent univariate losses usually have no clear superiority in practice. To take a step towards filling up this gap, this paper presents a systematic study from two complementary perspectives of consistency and generalization error bounds of learning algorithms. We theoretically find two key factors of the distribution (or dataset) that affect the learning guarantees of algorithms: the instance-wise class imbalance and the label size $c$. Specifically, in an extremely imbalanced case, the algorithm with the consistent univariate loss has an error bound of $O(c)$, while the one with the inconsistent pairwise loss depends on $O(\sqrt{c})$ as shown in prior work. This may shed light on the superior performance of pairwise methods in practice, where real datasets are usually highly imbalanced. Moreover, we present an inconsistent reweighted univariate loss-based algorithm that enjoys an error bound of $O(\sqrt{c})$ for promising performance as well as the computational efficiency of univariate losses. Finally, experimental results confirm our theoretical findings. Guoqiang Wu, Chongxuan Li, Kun Xu 0004, Jun Zhu 0001 |
NeurIPS | 1 |
| 2020 | Multi-label classification: do Hamming loss and subset accuracy really conflict with each other?abstractVarious evaluation measures have been developed for multi-label classification, including Hamming Loss (HL), Subset Accuracy (SA) and Ranking Loss (RL). However, there is a gap between empirical results and the existing theories: 1) an algorithm often empirically performs well on some measure(s) while poorly on others, while a formal theoretical analysis is lacking; and 2) in small label space cases, the algorithms optimizing HL often have comparable or even better performance on the SA measure than those optimizing SA directly, while existing theoretical results show that SA and HL are conflicting measures. This paper provides an attempt to fill up this gap by analyzing the learning guarantees of the corresponding learning algorithms on both SA and HL measures. We show that when a learning algorithm optimizes HL with its surrogate loss, it enjoys an error bound for the HL measure independent of $c$ (the number of labels), while the bound for the SA measure depends on at most $O(c)$. On the other hand, when directly optimizing SA with its surrogate loss, it has learning guarantees that depend on $O(\sqrt{c})$ for both HL and SA measures. This explains the observation that when the label space is not large, optimizing HL with its surrogate loss can have promising performance for SA. We further show that our techniques are applicable to analyze the learning guarantees of algorithms on other measures, such as RL. Finally, the theoretical analyses are supported by experimental results. Guoqiang Wu, Jun Zhu 0001 |
NeurIPS | 1 |
| 2020 | Joint Ranking SVM and Binary Relevance with robust Low-rank learning for multi-label classification
Guoqiang Wu, Ruobing Zheng, Yingjie Tian 0001, Dalian Liu |
Neural Networks | 1 |
| 2018 | Privileged Multi-Target Support Vector RegressionabstractMulti-target regression is the problem where each instance is associated with multiple continuous target outputs simultaneously. Its major challenges arise with jointly exploring the complex input-output relationships and inter-target correlations. One representative approach is to build many independent single-target Support Vector Regression (SVR) models for each output target which can capture complex input-output relationships via the kernel trick. However, it does not involve inter-target correlations to improve the performance. Meanwhile, there are also many regularization-based methods which mainly explore the linear inter-target correlations, e.g., a low-rank constraint on the parameter matrix. However, in practice, it might be restrictive to assume the targets to be linearly related, and allowing for nonlinear relationships is a challenge. Motivated by Learning Using Privileged Information (LUPI), we propose a novel privileged multi-target support vector regression (MT-PSVR) model which can jointly explore the complex input-output relationships and nonlinear inter-target correlations. It explicitly explores inter-target correlations by viewing other targets as privileged information when training each target model. Besides, it can naturally use the kernel trick to explore both the complex input-output relationships and nonlinear inter-target correlations. Experimental results on many benchmark datasets validate the effectiveness of our approach. Guoqiang Wu, Yingjie Tian 0001, Dalian Liu |
ICPR | 1 |
| 2018 | A unified framework implementing linear binary relevance for multi-label learning
Guoqiang Wu, Yingjie Tian 0001 |
Neurocomputing | 1 |
| 2018 | Cost-sensitive multi-label learning with positive and negative label pairwise correlations
Guoqiang Wu, Yingjie Tian 0001, Dalian Liu |
Neural Networks | 1 |
| 2017 | Stochastic gradient descent for large-scale linear nonparallel SVMabstractIn recent years, nonparallel support vector machine (NPSVM) is proposed as a nonparallel hyperplane classifier with superior performance than standard SVM and existing nonparallel classifiers such as the twin support vector machine (TWSVM). With the perfect theoretical underpinnings and great practical success, NPSVM has been used to dealing with the classification tasks on different scales. Tackling large-scale classification problem is a challenge yet significant work. Although large-scale linear NPSVM model has already been efficiently solved by the dual coordinate descent (DCD) algorithm or alternating direction method of multipliers (ADMM), we present a new strategy to solve the primal form of linear NPSVM different from existing work in this paper. Our algorithm is designed in the framework of the stochastic gradient descent (SGD), which is well suited to large-scale problem. Experiments are conducted on five large-scale data sets to confirm the effectiveness of our method. Jingjing Tang 0004, Yingjie Tian 0001, Guoqiang Wu, Dewei Li 0002 |
WI | 3 |