Pinghua Gong

dblp:13/9925 · DBLP profile ↗
← Back
24ranked-venue papers
11as first author
3since 2021 · last 2026
0009-0005-6889-6792ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 11 first-author · 1 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Recommender systems · 80% Data mining · 20%
Theoretical computer science
8 papers
Mathematical optimization · 95% Information theory · 4% Algorithms and data structures · 1%
Artificial intelligence
10 papers
Learning paradigms · 24% Optimization for machine learning · 16% Deep learning architectures and training · 14%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Smart cities and intelligent transportation · 78% Medical and health informatics · 22%

Topics — the 30 heaviest of 43, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems
e-commerce recommendation
1.012026
OxygenREC: An Instruction-Following Generative Framework for E-commerce Recommendation · SIGIR 2026
Recommender systems › large language model-based recommendation
instruction-following recommendation
1.012026
OxygenREC: An Instruction-Following Generative Framework for E-commerce Recommendation · SIGIR 2026
Mathematical optimization
sparse learning
0.742015
HONOR: Hybrid Optimization for NOn-convex Regularized problems · NIPS 2015
A Modified Orthant-Wise Limited Memory Quasi-Newton Method with Convergence Analysis · ICML 2015
A General Iterative Shrinkage and Thresholding Algorithm for Non-convex Regularized Optimization Problems · ICML (2) 2013
Mathematical optimization
nonconvex optimization
0.532015
HONOR: Hybrid Optimization for NOn-convex Regularized problems · NIPS 2015
A General Iterative Shrinkage and Thresholding Algorithm for Non-convex Regularized Optimization Problems · ICML (2) 2013
Multi-Stage Multi-Task Feature Learning · NIPS 2012
Machine learning › Learning paradigms
multi-task learning
0.532014
Efficient multi-task feature learning with calibration · KDD 2014
Multi-stage multi-task feature learning · J. Mach. Learn. Res. 2013
Robust multi-task feature learning · KDD 2012
Mathematical optimization
continuous optimization
0.422015
A Modified Orthant-Wise Limited Memory Quasi-Newton Method with Convergence Analysis · ICML 2015
Efficient Multi-Stage Conjugate Gradient for Trust Region Step · AAAI 2012
Machine learning › Learning paradigms › multi-task learning
multi-task feature learning
0.322014
Efficient multi-task feature learning with calibration · KDD 2014
Robust multi-task feature learning · KDD 2012
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatio-temporal neural networks
0.312018
Deep Multi-View Spatial-Temporal Network for Taxi Demand Prediction · AAAI 2018
Smart cities and intelligent transportation › demand prediction
taxi demand prediction
0.312018
Deep Multi-View Spatial-Temporal Network for Taxi Demand Prediction · AAAI 2018
Machine learning › Optimization for machine learning
constrained optimization
0.312026
OCP: Orthogonal Constrained Projection for Sparse Scaling in Industrial Commodity Recommendation · SIGIR 2026
Natural language and speech › Language models and text generation
instruction following
0.312026
OxygenREC: An Instruction-Following Generative Framework for E-commerce Recommendation · SIGIR 2026
Machine learning › Deep learning architectures and training › regularization
orthogonal constraint
0.312026
OCP: Orthogonal Constrained Projection for Sparse Scaling in Industrial Commodity Recommendation · SIGIR 2026
Smart cities and intelligent transportation › spatio-temporal prediction
destination prediction
0.312017
A Taxi Order Dispatch Model based On Combinatorial Optimization · KDD 2017
Data mining
clustering
0.212016
Robust Convex Clustering Analysis · ICDM 2016
Data mining › clustering
convex clustering
0.212016
Robust Convex Clustering Analysis · ICDM 2016
Data mining › clustering
robust clustering
0.212016
Robust Convex Clustering Analysis · ICDM 2016
Machine learning › Optimization for machine learning › second-order optimization
quasi-newton method
0.212015
HONOR: Hybrid Optimization for NOn-convex Regularized problems · NIPS 2015
Mathematical optimization › regularization › sparse regularization
l1-regularized optimization
0.212015
A Modified Orthant-Wise Limited Memory Quasi-Newton Method with Convergence Analysis · ICML 2015
Mathematical optimization › numerical computation › numerical optimization › second-order methods › newton's method
quasi-newton method
0.212015
A Modified Orthant-Wise Limited Memory Quasi-Newton Method with Convergence Analysis · ICML 2015
Mathematical optimization › continuous optimization
composite optimization
0.212014
Efficient multi-task feature learning with calibration · KDD 2014
Mathematical optimization › constrained optimization › duality theory
dual optimization
0.212014
Efficient multi-task feature learning with calibration · KDD 2014
Machine learning › Deep learning architectures and training
regularization
0.212013
Multi-stage multi-task feature learning · J. Mach. Learn. Res. 2013
Machine learning › Learning theory › statistical estimation
estimation error bounds
0.112012
Multi-Stage Multi-Task Feature Learning · NIPS 2012
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
parameter estimation
0.112012
Multi-Stage Multi-Task Feature Learning · NIPS 2012
Machine learning › Trustworthy machine learning
robustness
0.112012
Robust multi-task feature learning · KDD 2012
Mathematical optimization › iterative methods
conjugate gradient method
0.112012
Efficient Multi-Stage Conjugate Gradient for Trust Region Step · AAAI 2012
Mathematical optimization › numerical computation › numerical optimization › second-order methods
trust region methods
0.112012
Efficient Multi-Stage Conjugate Gradient for Trust Region Step · AAAI 2012
Mathematical optimization › regularization › sparse regularization
l1 regularization
0.112011
A Fast Dual Projected Newton Method for l1-Regularized Least Squares · IJCAI 2011
Mathematical optimization › regularization
l1-regularized least squares
0.112011
A Fast Dual Projected Newton Method for l1-Regularized Least Squares · IJCAI 2011
Information theory › signal processing
sparse representation
0.112011
A Fast Dual Projected Newton Method for l1-Regularized Least Squares · IJCAI 2011

Methods — techniques the papers use, named apart from their topics

orthogonal constrained projection · 2.0generative recommendation · 2.0multi-view spatial-temporal network · 0.7LSTM · 0.7CNN · 0.7truncated l1 loss · 0.5stochastic dual coordinate ascent · 0.5gradient boosting · 0.5censored regression · 0.5dual optimization · 0.4convergence analysis · 0.4line search · 0.4barzilai-borwein rule · 0.4gradient descent · 0.4combinatorial optimization · 0.3bayesian framework · 0.3block coordinate descent · 0.2quasi-newton · 0.2
YearPublicationVenuePosition
2026 OCP: Orthogonal Constrained Projection for Sparse Scaling in Industrial Commodity Recommendation
Beilin Xu, Boheng Tan, Yuefeng Sun, Rite Bo, Yaqiang Zang, Pinghua Gong
SIGIR9
2026 OxygenREC: An Instruction-Following Generative Framework for E-commerce Recommendation
Qingyang Li 0001, Yanchen Qiao, Ziyang Ji, Xiangyu Qian, Yanlong Zang, Weijie Ding, Yaqiang Zang, Pinghua Gong
SIGIR14
2023 A survey on machine learning from few samples
Jiang Lu, Pinghua Gong, Jieping Ye, Jianwei Zhang 0001, Changshui Zhang
Pattern Recognit.2
2019 Hexagon-Based Convolutional Neural Network for Supply-Demand Forecasting of Ride-Sourcing Services
abstract
Ride-sourcing services are becoming an increasingly popular transportation mode in cities all over the world. With real-time information from both drivers and passengers, the ride-sourcing platform can reduce matching frictions and improve efficiencies by surge pricing, optimal vehicle-trip assignment, and proactive ridesplitting strategies. An important foundation of these strategies is the short-term supply-demand forecasting. In this paper, we tackle the problem of predicting the short-term supply-demand gap of ride-sourcing services. In contrast to the previous studies that partitioned a city area into numerous square lattices, we partition the city area into various regular hexagon lattices, which is motivated by the fact that hexagonal segmentation has an unambiguous neighborhood definition, smaller edge-to-area ratio, and isotropy. To capture the spatio-temporal characteristics in a hexagonal manner, we propose three hexagon-based convolutional neural networks (H-CNN), both the input and output of which are numerous local hexagon maps. Moreover, a hexagon-based ensemble mechanism is developed to enhance the prediction performance. Validated by a 3-week real-world ride-sourcing dataset in Guangzhou, China, the H-CNN models are found to significantly outperform the benchmark algorithms in terms of accuracy and robustness. Our approaches can be further extended to a broad range of spatio-temporal forecasting problems in the domain of shared mobility and urban computing.
Jintao Ke, Hai Yang 0003, Xiqun Chen, Yitian Jia, Pinghua Gong, Jieping Ye
IEEE Trans. Intell. Transp. Syst.6
2018 Deep Multi-View Spatial-Temporal Network for Taxi Demand Prediction
abstract
Taxi demand prediction is an important building block to enabling intelligent transportation systems in a smart city. An accurate prediction model can help the city pre-allocate resources to meet travel demand and to reduce empty taxis on streets which waste energy and worsen the traffic congestion. With the increasing popularity of taxi requesting services such as Uber and Didi Chuxing (in China), we are able to collect large-scale taxi demand data continuously. How to utilize such big data to improve the demand prediction is an interesting and critical real-world problem. Traditional demand prediction methods mostly rely on time series forecasting techniques, which fail to model the complex non-linear spatial and temporal relations. Recent advances in deep learning have shown superior performance on traditionally challenging tasks such as image classification by learning the complex features and correlations from large-scale data. This breakthrough has inspired researchers to explore deep learning techniques on traffic prediction problems. However, existing methods on traffic prediction have only considered spatial relation (e.g., using CNN) or temporal relation (e.g., using LSTM) independently. We propose a Deep Multi-View Spatial-Temporal Network (DMVST-Net) framework to model both spatial and temporal relations. Specifically, our proposed model consists of three views: temporal view (modeling correlations between future demand values with near time points via LSTM), spatial view (modeling local spatial correlation via local CNN), and semantic view (modeling correlations among regions sharing similar temporal patterns). Experiments on large-scale real taxi demand data demonstrate effectiveness of our approach over state-of-the-art methods.
Huaxiu Yao, Fei Wu 0007, Jintao Ke, Xianfeng Tang, Yitian Jia, Pinghua Gong, Jieping Ye, Zhenhui Li
AAAI7
2017 A Taxi Order Dispatch Model based On Combinatorial Optimization
abstract
Taxi-booking apps have been very popular all over the world as they provide convenience such as fast response time to the users. The key component of a taxi-booking app is the dispatch system which aims to provide optimal matches between drivers and riders. Traditional dispatch systems sequentially dispatch taxis to riders and aim to maximize the driver acceptance rate for each individual order. However, the traditional systems may lead to a low global success rate, which degrades the rider experience when using the app. In this paper, we propose a novel system that attempts to optimally dispatch taxis to serve multiple bookings. The proposed system aims to maximize the global success rate, thus it optimizes the overall travel efficiency, leading to enhanced user experience. To further enhance users' experience, we also propose a method to predict destinations of a user once the taxi-booking APP is started. The proposed method employs the Bayesian framework to model the distribution of a user's destination based on his/her travel histories.
Lingyu Zhang 0001, Yue Min, Guobin Wu 0001, Pengcheng Feng, Pinghua Gong, Jieping Ye
KDD7
2017 Computational Drug Discovery with Dyadic Positive-Unlabeled Learning
abstract
Computational Drug Discovery, which uses computational techniques to facilitate and improve the drug discovery process, has aroused considerable interests in recent years. Drug Repositioning (DR) and Drug-Drug Interaction (DDI) prediction are two key problems in drug discovery and many computational techniques have been proposed for them in the last decade. Although these two problems have mostly been researched separately in the past, both DR and DDI can be formulated as the problem of detecting positive interactions between data entities (DR is between drug and disease, and DDI is between pairwise drugs). The challenge in both problems is that we can only observe a very small portion of positive interactions. In this paper, we propose a novel framework called Dyadic Positive-Unlabeled learning (DyPU) to solve the problem of detecting positive interactions. DyPU forces positive data pairs to rank higher than the average score of unlabeled data pairs. Moreover, we also derive the dual formulation of the proposed method with the rectifier scoring function and we show that the associated non-trivial proximal operator admits a closed form solution. Extensive experiments are conducted on real drug data sets and the results show that our method achieves superior performance comparing with the state-of-the-art.
Yashu Liu 0001, Ping Zhang 0016, Pinghua Gong, Fei Wang 0001, Guoliang Xue, Jieping Ye
SDM4
2016 Robust Convex Clustering Analysis
abstract
Clustering is an unsupervised learning approach that explores data and seeks groups of similar objects. Many classical clustering models such as k-means and DBSCAN are based on heuristics algorithms and suffer from local optimal solutions and numerical instability. Recently convex clustering has received increasing attentions, which leverages the sparsity inducing norms and enjoys many attractive theoretical properties. However, convex clustering is based on Euclidean distance and is thus not robust against outlier features. Since the outlier features are very common especially when dimensionality is high, the vulnerability has greatly limited the applicability of convex clustering to analyze many real-world datasets. In this paper, we address the challenge by proposing a novel robust convex clustering method that simultaneously performs convex clustering and identifies outlier features. Specifically, the proposed method learns to decompose the data matrix into a clustering structure component and a group sparse component that captures feature outliers. We develop a block coordinate descent algorithm which iteratively performs convex clustering after outliers features are identified and eliminated. We also propose an efficient algorithm for solving the convex clustering by exploiting the structures on its dual problem. Moreover, to further illustrate the statistical stability, we present the theoretical performance bound of the proposed clustering method. Empirical studies on synthetic data and real-world data demonstrate that the proposed robust convex clustering can detect feature outliers as well as improve cluster quality.
Pinghua Gong, Shiyu Chang, Thomas S. Huang
ICDM2
2016 Predict Risk of Relapse for Patients with Multiple Stages of Treatment of Depression
abstract
Depression is a serious mood disorder afflicting millions of people around the globe. Medications of different types and with different effects on neural activity have been developed for its treatments during the past few decades. Due to the heterogeneity of the disorder, many patients cannot achieve symptomatic remission from a single clinical trial. Instead they need multiple clinical trials to achieve remission, resulting in a multiple stage treatment pattern. Furthermore those who indeed achieve symptom remission are still faced with substantial risk of relapse. One promising approach to predicting the risk of relapse is censored regression. Traditional censored regression typically applies only to situations in which the exact time of event of interest is known. However, follow-up studies that track the patients' relapse status can only provide an interval of time during which relapse occurs. The exact time of relapse is usually unknown. In this paper, we present a censored regression approach with a truncated $l_1$ loss function that can handle the uncertainty of relapse time. Based on this general loss function, we develop a gradient boosting algorithm and a stochastic dual coordinate ascent algorithm when the hypothesis in the loss function is represented as (1) an ensemble of decision trees and (2) a linear combination of covariates, respectively. As an extension of our linear model, a multi-stage linear approach is further proposed to harness the data collected from multiple stages of treatment. We evaluate the proposed algorithms using a real-world clinical trial dataset. Results show that our methods outperform the well-known Cox proportional hazard model. In addition, the risk factors identified by our multi-stage linear model not only corroborate findings from recent research but also yield some new insights into how to develop effective measures for prevention of relapse among patients after their initial remission from the acute treatment stage.
Zhi Nie, Pinghua Gong, Jieping Ye
KDD2
2016 Absolute Fused Lasso and Its Application to Genome-Wide Association Studies
abstract
In many real-world applications, the samples/features acquired are in spatial or temporal order. In such cases, the magnitudes of adjacent samples/features are typically close to each other. Meanwhile, in the high-dimensional scenario, identifying the most relevant samples/features is also desired. In this paper, we consider a regularized model which can simultaneously identify important features and group similar features together. The model is based on a penalty called Absolute Fused Lasso (AFL). The AFL penalty encourages sparsity in the coefficients as well as their successive differences of absolute values' i.e., local constancy of the coefficient components in absolute values. Due to the non-convexity of AFL, it is challenging to develop efficient algorithms to solve the optimization problem. To this end, we employ the Difference of Convex functions (DC) programming to optimize the proposed non-convex problem. At each DC iteration, we adopt the proximal algorithm to solve a convex regularized sub-problem. One of the major contributions of this paper is to develop a highly efficient algorithm to compute the proximal operator. Empirical studies on both synthetic and real-world data sets from Genome-Wide Association Studies demonstrate the efficiency and effectiveness of the proposed approach in simultaneous identifying important features and grouping similar features.
Tao Yang 0016, Jun Liu 0003, Pinghua Gong, Ruiwen Zhang, Xiaotong Shen, Jieping Ye
KDD3
2016 Multi-document summarization via group sparse learning
Ruifang He, Jiliang Tang, Pinghua Gong, Qinghua Hu, Bo Wang 0011
Inf. Sci.3
2015 A Modified Orthant-Wise Limited Memory Quasi-Newton Method with Convergence Analysis
abstract
The Orthant-Wise Limited memory Quasi-Newton (OWL-QN) method has been demonstrated to be very effective in solving the \ell_1-regularized sparse learning problem. OWL-QN extends the L-BFGS from solving unconstrained smooth optimization problems to \ell_1-regularized (non-smooth) sparse learning problems. At each iteration, OWL-QN does not involve any \ell_1-regularized quadratic optimization subproblem and only requires matrix-vector multiplications without an explicit use of the (inverse) Hessian matrix, which enables OWL-QN to tackle large-scale problems efficiently. Although many empirical studies have demonstrated that OWL-QN works quite well in practice, several recent papers point out that the existing convergence proof of OWL-QN is flawed and a rigorous convergence analysis for OWL-QN still remains to be established. In this paper, we propose a modified Orthant-Wise Limited memory Quasi-Newton (mOWL-QN) algorithm by slightly modifying the OWL-QN algorithm. As the main technical contribution of this paper, we establish a rigorous convergence proof for the mOWL-QN algorithm. To the best of our knowledge, our work fills the theoretical gap by providing the first rigorous convergence proof for the OWL-QN-type algorithm on solving \ell_1-regularized sparse learning problems. We also provide empirical studies to show that mOWL-QN works well and is as efficient as OWL-QN.
Pinghua Gong, Jieping Ye
ICML1
2015 HONOR: Hybrid Optimization for NOn-convex Regularized problems
abstract
Recent years have witnessed the superiority of non-convex sparse learning formulations over their convex counterparts in both theory and practice. However, due to the non-convexity and non-smoothness of the regularizer, how to efficiently solve the non-convex optimization problem for large-scale data is still quite challenging. In this paper, we propose an efficient \underline{H}ybrid \underline{O}ptimization algorithm for \underline{NO}n convex \underline{R}egularized problems (HONOR). Specifically, we develop a hybrid scheme which effectively integrates a Quasi-Newton (QN) step and a Gradient Descent (GD) step. Our contributions are as follows: (1) HONOR incorporates the second-order information to greatly speed up the convergence, while it avoids solving a regularized quadratic programming and only involves matrix-vector multiplications without explicitly forming the inverse Hessian matrix. (2) We establish a rigorous convergence analysis for HONOR, which shows that convergence is guaranteed even for non-convex problems, while it is typically challenging to analyze the convergence for non-convex problems. (3) We conduct empirical studies on large-scale data sets and results demonstrate that HONOR converges significantly faster than state-of-the-art algorithms.
Pinghua Gong, Jieping Ye
NIPS1
2014 Efficient multi-task feature learning with calibration
abstract
Multi-task feature learning has been proposed to improve the generalization performance by learning the shared features among multiple related tasks and it has been successfully applied to many real-world problems in machine learning, data mining, computer vision and bioinformatics. Most existing multi-task feature learning models simply assume a common noise level for all tasks, which may not be the case in real applications. Recently, a Calibrated Multivariate Regression (CMR) model has been proposed, which calibrates different tasks with respect to their noise levels and achieves superior prediction performance over the non-calibrated one. A major challenge is how to solve the CMR model efficiently as it is formulated as a composite optimization problem consisting of two non-smooth terms. In this paper, we propose a variant of the calibrated multi-task feature learning formulation by including a squared norm regularizer. We show that the dual problem of the proposed formulation is a smooth optimization problem with a piecewise sphere constraint. The simplicity of the dual problem enables us to develop fast dual optimization algorithms with low per-iteration cost. We also provide a detailed convergence analysis for the proposed dual optimization algorithm. Empirical studies demonstrate that, the dual optimization algorithm quickly converges and it is much more efficient than the primal optimization algorithm. Moreover, the calibrated multi-task feature learning algorithms with and without the squared norm regularizer achieve similar prediction performance and both outperform the non-calibrated ones. Thus, the proposed variant not only enables us to develop fast optimization algorithms, but also keeps the superior prediction performance of the calibrated multi-task feature learning over the non-calibrated one.
Pinghua Gong, Wei Fan 0001, Jieping Ye
KDD1
2013 Efficient blind separation of reflection layers with nonparametric transformations
abstract
Superimposed images are very common when taking photos behind glass. We address the reflection separation problem using multiple superimposed images photographed in different viewpoints. With viewpoints changing, the reflected scenes could contain arbitrarily complicated variations between mixtures, like human's motions or other nonrigid motions. In this article, we propose a moderate hypothesis to tackle the reflected scenes' arbitrary variations as well as the parametric transformations of transmitted scenes. To rapidly recover high-quality image layers, we propose an Efficient Superimposition Recovering Algorithm (ESRA) by extending the framework of accelerated gradient method. Our recovering method has good converging performance and is more than 30 times faster than state-of-the-art methods. Experimental results on synthetic and real world images demonstrate that our method is promising.
Han Li 0005, Kun Gai, Pinghua Gong, Changshui Zhang
ICASSP3
2013 A General Iterative Shrinkage and Thresholding Algorithm for Non-convex Regularized Optimization Problems
abstract
Non-convex sparsity-inducing penalties have recently received considerable attentions in sparse learning. Recent theoretical investigations have demonstrated their superiority over the convex counterparts in several sparse learning settings. However, solving the non-convex optimization problems associated with non-convex penalties remains a big challenge. A commonly used approach is the Multi-Stage (MS) convex relaxation (or DC programming), which relaxes the original non-convex problem to a sequence of convex problems. This approach is usually not very practical for large-scale problems because its computational cost is a multiple of solving a single convex problem. In this paper, we propose a General Iterative Shrinkage and Thresholding (GIST) algorithm to solve the nonconvex optimization problem for a large class of non-convex penalties. The GIST algorithm iteratively solves a proximal operator problem, which in turn has a closed-form solution for many commonly used penalties. At each outer iteration of the algorithm, we use a line search initialized by the Barzilai-Borwein (BB) rule that allows finding an appropriate step size quickly. The paper also presents a detailed convergence analysis of the GIST algorithm. The efficiency of the proposed algorithm is demonstrated by extensive experiments on large-scale data sets.
Pinghua Gong, Changshui Zhang, Zhaosong Lu, Jieping Ye
ICML (2)1
2013 Multi-stage multi-task feature learning
Pinghua Gong, Jieping Ye, Changshui Zhang
J. Mach. Learn. Res.1
2012 Efficient Multi-Stage Conjugate Gradient for Trust Region Step
abstract
The trust region step problem, by solving a sphere constrained quadratic programming, plays a critical role in the trust region Newton method. In this paper, we propose an efficient Multi-Stage Conjugate Gradient (MSCG) algorithm to compute the trust region step in a multi-stage manner. Specifically, when the iterative solution is in the interior of the sphere, we perform the conjugate gradient procedure. Otherwise, we perform a gradient descent procedure which points to the inner of the sphere and can make the next iterative solution be a interior point. Subsequently, we proceed with the conjugate gradient procedure again. We repeat the above procedures until convergence. We also present a theoretical analysis which shows that the MSCG algorithm converges. Moreover, the proposed MSCG algorithm can generate a solution in any prescribed precision controlled by a tolerance parameter which is the only parameter we need. Experimental results on large-scale text data sets demonstrate our proposed MSCG algorithm has a faster convergence speed compared with the state-of-the-art algorithms.
Pinghua Gong, Changshui Zhang
AAAI1
2012 Learning similarity metric with SVM
abstract
In this paper, we show how to learn a good similarity metric for SVM classification. We present a novel approach to simultaneously learn a Mahalanobis similarity metric and an SVM classifier. Different from previous approaches, we optimize the Mahalanobis metric directly for minimizing the SVM classification error. Our formulation generalizes the traditional large margin principle used in standard SVM, that is, we maximize the margin-radius-ratio. The learned similarity metric significantly improves the classification performance of standard SVM. Empirical studies on real datasets show the proposed approach achieves higher or comparable classification accuracies compared with state-of-the-art similarity learning methods.
Xiaoqiang Zhu, Pinghua Gong, Changshui Zhang
IJCNN2
2012 Robust multi-task feature learning
abstract
Multi-task learning (MTL) aims to improve the performance of multiple related tasks by exploiting the intrinsic relationships among them. Recently, multi-task feature learning algorithms have received increasing attention and they have been successfully applied to many applications involving high-dimensional data. However, they assume that all tasks share a common set of features, which is too restrictive and may not hold in real-world applications, since outlier tasks often exist. In this paper, we propose a Robust MultiTask Feature Learning algorithm (rMTFL) which simultaneously captures a common set of features among relevant tasks and identifies outlier tasks. Specifically, we decompose the weight (model) matrix for all tasks into two components. We impose the well-known group Lasso penalty on row groups of the first component for capturing the shared features among relevant tasks. To simultaneously identify the outlier tasks, we impose the same group Lasso penalty but on column groups of the second component. We propose to employ the accelerated gradient descent to efficiently solve the optimization problem in rMTFL, and show that the proposed algorithm is scalable to large-size problems. In addition, we provide a detailed theoretical analysis on the proposed rMTFL formulation. Specifically, we present a theoretical bound to measure how well our proposed rMTFL approximates the true evaluation, and provide bounds to measure the error between the estimated weights of rMTFL and the underlying true weights. Moreover, by assuming that the underlying true weights are above the noise level, we present a sound theoretical result to show how to obtain the underlying true shared features and outlier tasks (sparsity patterns). Empirical studies on both synthetic and real-world data demonstrate that our proposed rMTFL is capable of simultaneously capturing shared features among tasks and identifying outlier tasks.
Pinghua Gong, Jieping Ye, Changshui Zhang
KDD1
2012 Multi-Stage Multi-Task Feature Learning
abstract
Multi-task sparse feature learning aims to improve the generalization performance by exploiting the shared features among tasks. It has been successfully applied to many applications including computer vision and biomedical informatics. Most of the existing multi-task sparse feature learning algorithms are formulated as a convex sparse regularization problem, which is usually suboptimal, due to its looseness for approximating an $\ell_0$-type regularizer. In this paper, we propose a non-convex formulation for multi-task sparse feature learning based on a novel regularizer. To solve the non-convex optimization problem, we propose a Multi-Stage Multi-Task Feature Learning (MSMTFL) algorithm. Moreover, we present a detailed theoretical analysis showing that MSMTFL achieves a better parameter estimation error bound than the convex formulation. Empirical studies on both synthetic and real-world data sets demonstrate the effectiveness of MSMTFL in comparison with the state of the art multi-task sparse feature learning algorithms.
Pinghua Gong, Jieping Ye, Changshui Zhang
NIPS1
2012 Efficient Nonnegative Matrix Factorization via projected Newton method
Pinghua Gong, Changshui Zhang
Pattern Recognit.1
2011 A Fast Dual Projected Newton Method for l1-Regularized Least Squares
abstract
L1-regularized least squares, with the ability of discovering sparse representations, is quite prevalent in the field of machine learning, statistics and signal processing. In this paper, we propose a novel algorithm called Dual Projected Newton Method (DPNM) to solve the l1-regularized least squares problem. In DPNM, we first derive a new dual problem as a box constrained quadratic programming. Then, a projected Newton method is utilized to solve the dual problem, achieving a quadratic convergence rate. Moreover, we propose to utilize some practical techniques, thus it greatly reduces the computational cost and makes DPNM more efficient. Experimental results on six real-world data sets indicate that DPNM is very efficient for solving the l1-regularized least squares problem, by comparing it with state of the art methods.
Pinghua Gong, Changshui Zhang
IJCAI1
2011 Efficient Euclidean projections via Piecewise Root Finding and its application in gradient projection
Pinghua Gong, Kun Gai, Changshui Zhang
Neurocomputing1