David Hong

dblp:195/1106 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2027
0000-0003-4698-4175ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2027 Certificate for orthogonal equivalence of real polynomials by polynomial-weighted principal component analysis
Martin Helmer, David Hong, Hoon Hong
J. Symb. Comput.2
2025 Optimal Sample Acquisition for Optimally Weighted PCA From Heterogeneous Quality Sources
abstract
Modern high-dimensional datasets are often formed by acquiring samples from multiple sources having heterogeneous quality, i.e., some sources are noisier than others. Collecting data in this manner raises the following natural question: what is the best way to collect the data (i.e., how many samples should be acquired from each source) given constraints (e.g., on time or energy)? In general, the answer depends on what analysis is to be performed. In this paper, we study the foundational signal processing task of estimating underlying low-dimensional principal components. Since the resulting dataset will be high-dimensional and will have heteroscedastic noise, we focus on the recently proposed optimally weighted PCA, which is designed specifically for this setting. We develop an efficient method for designing sample acquisitions that optimize the asymptotic performance of optimally weighted PCA given resource constraints, and we illustrate the proposed method through various case studies.
David Hong, Laura Balzano
IEEE Signal Process. Lett.1
2023 Provable Tradeoffs in Adversarially Robust Classification
abstract
It is well known that machine learning methods can be vulnerable to adversarially-chosen perturbations of their inputs. Despite significant progress in the area, foundational open problems remain. In this paper, we address several key questions. We derive exact and approximate Bayes-optimal robust classifiers for the important setting of two- and three-class Gaussian classification problems with arbitrary imbalance, for$\ell _{2}$and$\ell _{\infty} $adversaries. In contrast to classical Bayes-optimal classifiers, determining the optimal decisions here cannot be made pointwise and new theoretical approaches are needed. We develop and leverage new tools, including recent breakthroughs from probability theory on robust isoperimetry, which, to our knowledge, have not yet been used in the area. Our results reveal fundamental tradeoffs between standard and robust accuracy that grow when data is imbalanced. We also show further results, including an analysis of classification calibration for convex losses in certain models, and finite sample rates for the robust risk.
Edgar Dobriban, Seyed Hamed Hassani, David Hong, Alexander Robey
IEEE Trans. Inf. Theory3
2019 Convolutional Analysis Operator Learning: Dependence on Training Data
abstract
Convolutional analysis operator learning (CAOL) enables the unsupervised training of (hierarchical) convolutional sparsifying operators or autoencoders from large datasets. One can use many training images for CAOL, but a precise understanding of the impact of doing so has remained an open question. This letter presents a series of results that lend insight into the impact of dataset size on the filter update in CAOL. The first result is a general deterministic bound on errors in the estimated filters, and is followed by a bound on the expected errors as the number of training samples increases. The second result provides a high probability analogue. The bounds depend on properties of the training data, and we investigate their empirical values with real data. Taken together, these results provide evidence for the potential benefit of using more training data in CAOL.
Il Yong Chun, David Hong, Ben Adcock, Jeffrey A. Fessler
IEEE Signal Process. Lett.2