Guo-Xun Yuan

dblp:86/9364 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
5 papers
Mathematical optimization · 100%
Artificial intelligence
3 papers
Optimization for machine learning · 83% Learning theory · 17%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization › regularization › regularized optimization
l1-regularized logistic regression
0.322012
An Improved GLMNET for L1-regularized Logistic Regression · J. Mach. Learn. Res. 2012
An improved GLMNET for l1-regularized logistic regression · KDD 2011
Mathematical optimization › continuous optimization
convex optimization
0.322012
An Improved GLMNET for L1-regularized Logistic Regression · J. Mach. Learn. Res. 2012
A Comparison of Optimization Methods and Software for Large-scale L1-regularized Linear Classification · J. Mach. Learn. Res. 2010
Machine learning › Optimization for machine learning › large-scale optimization
large-scale linear classification
0.112012
Recent Advances of Large-Scale Linear Classification · Proc. IEEE 2012
Machine learning › Optimization for machine learning › large-scale optimization
large-scale SVM training
0.112012
Scalable Training of Sparse Linear SVMs · ICDM 2012
Machine learning › Learning theory › classification
linear classification
0.112012
Recent Advances of Large-Scale Linear Classification · Proc. IEEE 2012
Machine learning › Optimization for machine learning › sparse learning
sparse linear SVM
0.112012
Scalable Training of Sparse Linear SVMs · ICDM 2012
Visualization and visual analytics › uncertainty visualization
uncertainty propagation
0.112012
Visualizing Flow of Uncertainty through Analytical Processes · IEEE Trans. Vis. Comput. Graph. 2012
Visualization and visual analytics
uncertainty visualization
0.112012
Visualizing Flow of Uncertainty through Analytical Processes · IEEE Trans. Vis. Comput. Graph. 2012
Mathematical optimization
continuous optimization
0.112012
Recent Advances of Large-Scale Linear Classification · Proc. IEEE 2012
Mathematical optimization
optimization for machine learning
0.112012
Recent Advances of Large-Scale Linear Classification · Proc. IEEE 2012
Machine learning › Optimization for machine learning
coordinate descent
0.112011
An improved GLMNET for l1-regularized logistic regression · KDD 2011
Machine learning › Optimization for machine learning › second-order optimization
newton-type methods
0.112011
An improved GLMNET for l1-regularized logistic regression · KDD 2011
Mathematical optimization › regularization
regularized optimization
0.112011
An improved GLMNET for l1-regularized logistic regression · KDD 2011
Data mining › predictive modeling
classification
0.012012
Recent Advances of Large-Scale Linear Classification · Proc. IEEE 2012
Data mining
pattern mining
0.012012
Recent Advances of Large-Scale Linear Classification · Proc. IEEE 2012
Visualization and visual analytics › visual analytics
visual analytics workflow
0.012012
Visualizing Flow of Uncertainty through Analytical Processes · IEEE Trans. Vis. Comput. Graph. 2012
Mathematical optimization › primal-dual method
alternating direction method
0.012012
Scalable Training of Sparse Linear SVMs · ICDM 2012

Methods — techniques the papers use, named apart from their topics

coordinate descent · 0.5soft-thresholding · 0.3dual alternating direction method · 0.3distributed optimization · 0.3newton method · 0.2elastic net · 0.2uncertainty flow metaphor · 0.1quadratic approximation · 0.1GLMNET · 0.1
YearPublicationVenuePosition
2018 Naive Parallelization of Coordinate Descent Methods and an Application on Multi-core L1-regularized Classification
abstract
It is well known that a direct parallelization of sequential optimization methods (e.g., coordinate descent and stochastic gradient methods) is often not effective. The reason is that at each iteration, the number of operations may be too small. We point out that this common understanding may not be true if the algorithm sequentially accesses the data in a feature-wise manner. For almost all real-world sparse sets we have examined, some features are much denser than others. Thus a direct parallelization of loops in a sequential method may result in excellent speedup. This approach possesses an advantage of retaining all convergence results because the algorithm is not changed at all. We apply this idea on coordinate descent (CD) methods, which are effective single-thread technique for L1-regularized classification. Further, an investigation on the shrinking technique commonly used to remove some features in the training process shows that this technique helps the parallelization of CD methods. Experiments indicate that a naive parallelization achieves better speedup than existing methods that laboriously modify the algorithm to achieve parallelism. Though a bit ironic, we conclude that the naive parallelization of the CD method is a highly competitive and robust multi-core implementation for L1-regularized classification.
Yong Zhuang, Yu-Chin Juan, Guo-Xun Yuan, Chih-Jen Lin
CIKM3
2012 Scalable Training of Sparse Linear SVMs
abstract
Sparse linear support vector machines have been widely applied to variable selection in many applications. For large data, managing the cost of training a sparse model with good predication performance is an essential topic. In this work, we propose a scalable training algorithm for large-scale data with millions of examples and features. We develop a dual alternating direction method for solving L1-regularized linear SVMs. The learning procedure simply involves quadratic programming in the same form as the standard SVM dual, followed by a soft-thresholding operation. The proposed training algorithm possesses two favorable properties. First, it is a decomposable algorithm by which a large problem can be reduced to small ones. Second, the sparsity of intermediate solutions is maintained throughout the training process. It naturally promotes the solution sparsity by soft-thresholding. We demonstrate that, by experiments, our method outperforms state-of-the-art approaches on large-scale benchmark data sets. We also show that it is well suited for training large sparse models on a distributed system.
Guo-Xun Yuan, Kwan-Liu Ma
ICDM1
2012 An Improved GLMNET for L1-regularized Logistic Regression
Guo-Xun Yuan, Chia-Hua Ho, Chih-Jen Lin
J. Mach. Learn. Res.1
2012 Recent Advances of Large-Scale Linear Classification
abstract
Linear classification is a useful tool in machine learning and data mining. For some data in a rich dimensional space, the performance (i.e., testing accuracy) of linear classifiers has shown to be close to that of nonlinear classifiers such as kernel methods, but training and testing speed is much faster. Recently, many research works have developed efficient optimization methods to construct linear classifiers and applied them to some large-scale applications. In this paper, we give a comprehensive survey on the recent development of this active research area.
Guo-Xun Yuan, Chia-Hua Ho, Chih-Jen Lin
Proc. IEEE1
2012 Visualizing Flow of Uncertainty through Analytical Processes
abstract
Uncertainty can arise in any stage of a visual analytics process, especially in data-intensive applications with a sequence of data transformations. Additionally, throughout the process of multidimensional, multivariate data analysis, uncertainty due to data transformation and integration may split, merge, increase, or decrease. This dynamic characteristic along with other features of uncertainty pose a great challenge to effective uncertainty-aware visualization. This paper presents a new framework for modeling uncertainty and characterizing the evolution of the uncertainty information through analytical processes. Based on the framework, we have designed a visual metaphor called uncertainty flow to visually and intuitively summarize how uncertainty information propagates over the whole analysis pipeline. Our system allows analysts to interact with and analyze the uncertainty information at different levels of detail. Three experiments were conducted to demonstrate the effectiveness and intuitiveness of our design.
Yingcai Wu, Guo-Xun Yuan, Kwan-Liu Ma
IEEE Trans. Vis. Comput. Graph.2
2011 An improved GLMNET for l1-regularized logistic regression
abstract
GLMNET proposed by Friedman et al. is an algorithm for generalized linear models with elastic net. It has been widely applied to solve L1-regularized logistic regression. However, recent experiments indicated that the existing GLMNET implementation may not be stable for large-scale problems. In this paper, we propose an improved GLMNET to address some theoretical and implementation issues. In particular, as a Newton-type method, GLMNET achieves fast local convergence, but may fail to quickly obtain a useful solution. By a careful design to adjust the effort for each iteration, our method is efficient regardless of loosely or strictly solving the optimization problem. Experiments demonstrate that the improved GLMNET is more efficient than a state-of-the-art coordinate descent method.
Guo-Xun Yuan, Chia-Hua Ho, Chih-Jen Lin
KDD1
2010 A Comparison of Optimization Methods and Software for Large-scale L1-regularized Linear Classification
Guo-Xun Yuan, Kai-Wei Chang 0001, Cho-Jui Hsieh, Chih-Jen Lin
J. Mach. Learn. Res.1