Young Woong Park

dblp:156/1532 · DBLP profile ↗
← Back
15ranked-venue papers
10as first author
5since 2021 · last 2025
0000-0003-0722-7729ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 4 since 2021Theory of computation · 6 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
YearPublicationVenuePosition
2025 Integrated subset selection and bandwidth estimation algorithm for geographically weighted regression
Young Woong Park
Pattern Recognit.2
2024 A systematic approach for learning imbalanced data: enhancing zero-inflated models through boosting
abstract
Abstract In this paper, we propose systematic approaches for learning imbalanced data based on a two-regime process: regime 0, which generates excess zeros (majority class), and regime 1, which contributes to generating an outcome of one (minority class). The proposed model contains two latent equations: a split probit (logit) equation in the first stage and an ordinary probit (logit) equation in the second stage. Because boosting improves the accuracy of prediction versus using a single classifier, we combined a boosting strategy with the two-regime process. Thus, we developed the zero-inflated probit boost (ZIPBoost) and zero-inflated logit boost (ZILBoost) methods. We show that the weight functions of ZIPBoost have the desired properties for good predictive performance. Like AdaBoost, the weight functions upweight misclassified examples and downweight correctly classified examples. We show that the weight functions of ZILBoost have similar properties to those of LogitBoost. The algorithm will focus more on examples that are hard to classify in the next iteration, resulting in improved prediction accuracy. We provide the relative performance of ZIPBoost and ZILBoost, which rely on the excess kurtosis of the data distribution. Furthermore, we show the convergence and time complexity of our proposed methods. We demonstrate the performance of our proposed methods using a Monte Carlo simulation, mergers and acquisitions (M&A) data application, and imbalanced datasets from the Keel repository. The results of the experiments show that our proposed methods yield better prediction accuracy compared to other learning algorithms.
Yeasung Jeong, Young Woong Park, Sumin Han
Mach. Learn.3
2024 Discordance minimization-based imputation algorithms for missing values in rating data
Young Woong Park, Jinhak Kim
Mach. Learn.1
2024 Observation weights matching approach for causal inference
Sumin Han, Hyeoncheol Baik, Yeasung Jeong, Young Woong Park
Pattern Recognit.5
2021 Optimization for L1-Norm Error Fitting via Data Aggregation
abstract
We propose a data aggregation-based algorithm with monotonic convergence to a global optimum for a generalized version of the L1-norm error fitting model with an assumption of the fitting function. The proposed algorithm generalizes the recent algorithm in the literature, aggregate and iterative disaggregate (AID), which selectively solves three specific L1-norm error fitting problems. With the proposed algorithm, any L1-norm error fitting model can be solved optimally if it follows the form of the L1-norm error fitting problem and if the fitting function satisfies the assumption. The proposed algorithm can also solve multidimensional fitting problems with arbitrary constraints on the fitting coefficients matrix. The generalized problem includes popular models, such as regression and the orthogonal Procrustes problem. The results of the computational experiment show that the proposed algorithms are faster than the state-of-the-art benchmarks for L1-norm regression subset selection and L1-norm regression over a sphere. Furthermore, the relative performance of the proposed algorithm improves as data size increases.
Young Woong Park
INFORMS J. Comput.1
2020 MILP Models for Complex System Reliability Redundancy Allocation with Mixed Components
abstract
The redundancy allocation problem (RAP) aims to find an optimal allocation of redundant components subject to resource constraints. In this paper, mixed integer linear programming (MILP) models and MILP-based algorithms are proposed for complex system reliability redundancy allocation problem with mixed components, where the system have bridges or interconnecting subsystems and each subsystem can have mixed types of components. Unlike the other algorithms in the literature, the proposed MILP models view the problem from a different point of view and approximate the nonconvex nonlinear system reliability function of a complex system using random samples. The solution to the MILP converges to the optimal solution of the original problem as sample size increases. In addition, data aggregation-based algorithms are proposed to improve the solution time and quality based on the proposed MILP models. A computational experiment shows that the proposed models and algorithms converge to the optimal or best-known solution as sample size increases. The proposed algorithms outperform popular metaheuristic algorithms in the literature.
Young Woong Park
INFORMS J. Comput.1
2020 Subset selection for multiple linear regression via optimization
Young Woong Park, Diego Klabjan
J. Glob. Optim.1
2020 A mathematical programming approach for integrated multiple linear regression subset selection and validation
Seokhyun Chung, Young Woong Park, Taesu Cheong
Pattern Recognit.2
2018 Three iteratively reweighted least squares algorithms for L1 -norm principal component analysis
Young Woong Park, Diego Klabjan
Knowl. Inf. Syst.1
2017 Algorithms for Generalized Clusterwise Linear Regression
abstract
Clusterwise linear regression (CLR), a clustering problem intertwined with regression, finds clusters of entities such that the overall sum of squared errors from regressions performed over these clusters is minimized, where each cluster may have different variances. We generalize the CLR problem by allowing each entity to have more than one observation and refer to this as generalized CLR. We propose an exact mathematical programming-based approach relying on column generation, a column generation–based heuristic algorithm that clusters predefined groups of entities, a metaheuristic genetic algorithm with adapted Lloyd’s algorithm for K-means clustering, a two-stage approach, and a modified algorithm of Späth [Späth (1979) Algorithm 39 clusterwise linear regression. Comput. 22(4):367–373] for solving generalized CLR. We examine the performance of our algorithms on a stock-keeping unit (SKU)-clustering problem employed in forecasting halo and cannibalization effects in promotions using real-world retail data from a large supermarket chain. In the SKU clustering problem, the retailer needs to cluster SKUs based on their seasonal effects in response to promotions. The seasonal effects result from regressions with predictors being promotion mechanisms and seasonal dummies performed over clusters generated. We compare the performance of all proposed algorithms for the SKU problem with real-world and synthetic data. The online supplement is available at https://doi.org/10.1287/ijoc.2016.0729 .
Young Woong Park, Diego Klabjan, Loren Williams
INFORMS J. Comput.1
2017 Bayesian Network Learning via Topological Order
abstract
We propose a mixed integer programming (MIP) model and iterative algorithms based on topological orders to solve optimization problems with acyclic constraints on a directed graph. The proposed MIP model has a significantly lower number of constraints compared to popular MIP models based on cycle elimination constraints and triangular inequalities. The proposed iterative algorithms use gradient descent and iterative reordering approaches, respectively, for searching topological orders. A computational experiment is presented for the Gaussian Bayesian network learning problem, an optimization problem minimizing the sum of squared errors of regression models with L1 penalty over a feature network with application of gene network inference in bioinformatics.
Young Woong Park, Diego Klabjan
J. Mach. Learn. Res.1
2016 Iteratively Reweighted Least Squares Algorithms for L1-Norm Principal Component Analysis
abstract
Principal component analysis (PCA) is often used to reduce the dimension of data by selecting a few orthonormal vectors that explain most of the variance structure of the data. L1 PCA uses the L1 norm to measure error, whereas the conventional PCA uses the L2 norm. For the L1 PCA problem minimizing the fitting error of the reconstructed data, we propose an exact reweighted and an approximate algorithm based on iteratively reweighted least squares. We provide convergence analyses, and compare their performance against benchmark algorithms in the literature. The computational experiment shows that the proposed algorithms consistently perform best.
Young Woong Park, Diego Klabjan
ICDM1
2016 An aggregate and iterative disaggregate algorithm with proven optimality in machine learning
Young Woong Park, Diego Klabjan
Mach. Learn.1
2015 Lot sizing with minimum order quantity
Young Woong Park, Diego Klabjan
Discret. Appl. Math.1
2015 Lifting for mixed integer programs with variable upper bounds
Sergey Shebalov, Young Woong Park, Diego Klabjan
Discret. Appl. Math.2