VLDB 2026 Research / reviewers in the wild / expert
Jiacheng Zhuo
dblp:198/0672
· DBLP profile ↗
8ranked-venue papers
2as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Learning theory · 39% Optimization for machine learning · 33% Image recognition and object detection · 10% | |
| Computer graphics and multimedia
2 papers |
Geometric modeling and processing · 75% Computational fabrication · 25% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 21 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › computational learning theory
computational-statistical gap |
0.8 | 1 | 2024 | On the Computational and Statistical Complexity of Over-parameterized Matrix Sensing · J. Mach. Learn. Res. 2024 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.8 | 1 | 2024 | Improving Computational Complexity in Statistical Models with Local Curvature Information · ICML 2024 |
Machine learning › Learning theory
local curvature |
0.8 | 1 | 2024 | Improving Computational Complexity in Statistical Models with Local Curvature Information · ICML 2024 |
Machine learning › Optimization for machine learning › low-rank optimization
matrix sensing |
0.8 | 1 | 2024 | On the Computational and Statistical Complexity of Over-parameterized Matrix Sensing · J. Mach. Learn. Res. 2024 |
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
normalized gradient descent |
0.8 | 1 | 2024 | Improving Computational Complexity in Statistical Models with Local Curvature Information · ICML 2024 |
Machine learning › Learning theory › statistical learning theory
statistical complexity |
0.8 | 1 | 2024 | On the Computational and Statistical Complexity of Over-parameterized Matrix Sensing · J. Mach. Learn. Res. 2024 |
Machine learning › Learning theory
statistical estimation |
0.8 | 1 | 2024 | Improving Computational Complexity in Statistical Models with Local Curvature Information · ICML 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
pretext task |
0.5 | 1 | 2021 | Predicting What You Already Know Helps: Provable Self-Supervised Learning · NeurIPS 2021 |
Machine learning › Learning theory
sample complexity |
0.5 | 1 | 2021 | Predicting What You Already Know Helps: Provable Self-Supervised Learning · NeurIPS 2021 |
Computer vision › 3D vision
3d reconstruction |
0.4 | 1 | 2019 | K-Best Transformation Synchronization · ICCV 2019 |
Computer vision › Image recognition and object detection › object detection
bottom-up detection |
0.4 | 1 | 2019 | Bottom-Up Object Detection by Grouping Extreme and Center Points · CVPR 2019 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.4 | 1 | 2019 | Bottom-Up Object Detection by Grouping Extreme and Center Points · CVPR 2019 |
Computer vision › Face, body and person analysis › human pose estimation
keypoint grouping |
0.4 | 1 | 2019 | Bottom-Up Object Detection by Grouping Extreme and Center Points · CVPR 2019 |
Machine learning › Optimization for machine learning
low-rank optimization |
0.4 | 1 | 2019 | Primal-Dual Block Generalized Frank-Wolfe · NeurIPS 2019 |
Computer vision › Image recognition and object detection
object detection |
0.4 | 1 | 2019 | Bottom-Up Object Detection by Grouping Extreme and Center Points · CVPR 2019 |
Machine learning › Optimization for machine learning › sparse learning
sparse optimization |
0.4 | 1 | 2019 | Primal-Dual Block Generalized Frank-Wolfe · NeurIPS 2019 |
Geometric modeling and processing
shape analysis |
0.4 | 1 | 2019 | K-Best Transformation Synchronization · ICCV 2019 |
Geometric modeling and processing › shape analysis
symmetry detection |
0.4 | 1 | 2019 | K-Best Transformation Synchronization · ICCV 2019 |
Mathematical optimization
continuous optimization |
0.4 | 1 | 2019 | Primal-Dual Block Generalized Frank-Wolfe · NeurIPS 2019 |
Mathematical optimization
frank-wolfe algorithm |
0.4 | 1 | 2019 | Primal-Dual Block Generalized Frank-Wolfe · NeurIPS 2019 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2019 | Primal-Dual Block Generalized Frank-Wolfe · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
transformation clustering · 0.8population-to-sample analysis · 0.8low-rank matrix sensing · 0.8iterative optimization · 0.8hessian eigenvalue scaling · 0.8frank-wolfe · 0.8factorized gradient descent · 0.8linear probing · 0.5conditional independence · 0.5variational algorithm · 0.4primal-dual method · 0.4linear convergence analysis · 0.4keypoint estimation networks · 0.4curl-free unit vector field optimization · 0.4branched cover · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improving Computational Complexity in Statistical Models with Local Curvature InformationabstractIt is known that when the statistical models are singular, i.e., the Fisher information matrix at the true parameter is degenerate, the fixed step-size gradient descent algorithm takes polynomial number of steps in terms of the sample size $n$ to converge to a final statistical radius around the true parameter, which can be unsatisfactory for the practical application. To further improve that computational complexity, we consider utilizing the local curvature information for parameter estimation. Even though there is a rich literature in using the local curvature information for optimization, the statistical rate of these methods in statistical models, to the best of our knowledge, has not been studied rigorously. The major challenge of this problem is due to the non-convex nature of sample loss function. To shed light on these problems, we specifically study the normalized gradient descent (NormGD) algorithm, a variant of gradient descent algorithm whose step size is scaled by the maximum eigenvalue of the Hessian matrix of the empirical loss function, and deal with the aforementioned issue with a population-to-sample analysis. When the population loss function is homogeneous, the NormGD iterates reach a final statistical radius around the true parameter after a logarithmic number of iterations in terms of $n$. Therefore, for fixed dimension $d$, the NormGD algorithm achieves the optimal computational complexity $\mathcal{O}(n)$ to reach the final statistical radius, which is cheaper than the complexity $\mathcal{O}(n^{\tau})$ of the fixed step-size gradient descent algorithm for some $\tau > 1$. Pedram Akbarian, Tongzheng Ren, Jiacheng Zhuo, Sujay Sanghavi, Nhat Ho |
ICML | 3 |
| 2024 | On the Computational and Statistical Complexity of Over-parameterized Matrix SensingabstractWe consider solving the low-rank matrix sensing problem with the Factorized Gradient Descent (FGD) method when the specified rank is larger than the true rank. We refer to this as over-parameterized matrix sensing. If the ground truth signal $\mathbf{X}^* \in \mathbb{R}^{d \times d}$ is of rank $r$, but we try to recover it using $\mathbf{F} \mathbf{F}^\top$ where $\mathbf{F} \in \mathbb{R}^{d \times k}$ and $k>r$, the existing statistical analysis either no longer holds or produces a vacuous statistical error upper bound (infinity) due to the flat local curvature of the loss function around the global maxima. By decomposing the factorized matrix $\mathbf{F}$ into separate column spaces to capture the impact of using $k > r$, we show that $\left\| {\mathbf{F}_t \mathbf{F}_t - \mathbf{X}^*} \right\|_F^2$ converges sub-linearly to a statistical error of $\tilde{\mathcal{O}} (k d \sigma^2/n)$ after $\tilde{\mathcal{O}}(\frac{\sigma_{r}}{\sigma}\sqrt{\frac{n}{d}})$ iterations, where $\mathbf{F}_t$ is the output of FGD after $t$ iterations, $\sigma^2$ is the variance of the observation noise, $\sigma_{r}$ is the $r$-th largest eigenvalue of $\mathbf{X}^*$, and $n$ is the number of samples. With a precise characterization of the convergence behavior and the statistical error, our results, therefore, offer a comprehensive picture of the statistical and computational complexity if we solve the over-parameterized matrix sensing problem with FGD. Jiacheng Zhuo, Jeongyeol Kwon, Nhat Ho, Constantine Caramanis |
J. Mach. Learn. Res. | 1 |
| 2021 | Predicting What You Already Know Helps: Provable Self-Supervised LearningabstractSelf-supervised representation learning solves auxiliary prediction tasks (known as pretext tasks), that do not require labeled data, to learn semantic representations. These pretext tasks are created solely using the input features, such as predicting a missing image patch, recovering the color channels of an image from context, or predicting missing words, yet predicting this \textit{known} information helps in learning representations effective for downstream prediction tasks. This paper posits a mechanism based on approximate conditional independence to formalize how solving certain pretext tasks can learn representations that provably decrease the sample complexity of downstream supervised tasks. Formally, we quantify how the approximate independence between the components of the pretext task (conditional on the label and latent variables) allows us to learn representations that can solve the downstream task with drastically reduced sample complexity by just training a linear layer on top of the learned representation. Jason D. Lee, Nikunj Saunshi, Jiacheng Zhuo |
NeurIPS | 4 |
| 2020 | Communication-Efficient Asynchronous Stochastic Frank-Wolfe over Nuclear-norm BallsabstractLarge-scale machine learning training suffers from two prior challenges, specifically for nuclear-norm constrained problems with distributed systems: the synchronization slowdown due to the straggling workers, and high communication costs. In this work, we propose an asynchronous Stochastic Frank Wolfe (SFW-asyn) method, which, for the first time, solves the two problems simultaneously, while successfully maintaining the same convergence rate as the vanilla SFW. We implement our algorithm in python (with MPI) to run on Amazon EC2, and demonstrate that SFW-asyn yields speed-ups almost linear to the number of machines compared to the vanilla SFW. Jiacheng Zhuo, Alexandros G. Dimakis, Constantine Caramanis |
AISTATS | 1 |
| 2019 | Bottom-Up Object Detection by Grouping Extreme and Center PointsabstractWith the advent of deep learning, object detection drifted from a bottom-up to a top-down recognition problem. State of the art algorithms enumerate a near-exhaustive list of object locations and classify each into: object or not. In this paper, we show that bottom-up approaches still perform competitively. We detect four extreme points (top-most, left-most, bottom-most, right-most) and one center point of objects using a standard keypoint estimation network. We group the five keypoints into a bounding box if they are geometrically aligned. Object detection is then a purely appearance-based keypoint estimation problem, without region classification or implicit feature learning. The proposed method performs on-par with the state-of-the-art region based detection methods, with a bounding box AP of 43.7% on COCO test-dev. In addition, our estimated extreme points directly span a coarse octagonal mask, with a COCO Mask AP of 18.9%, much better than the Mask AP of vanilla bounding boxes. Extreme point guided segmentation further improves this to 34.6% Mask AP. Xingyi Zhou, Jiacheng Zhuo, Philipp Krähenbühl |
CVPR | 2 |
| 2019 | K-Best Transformation SynchronizationabstractIn this paper, we introduce the problem of K-best transformation synchronization for the purpose of multiple scan matching. Given noisy pair-wise transformations computed between a subset of depth scan pairs, K-best transformation synchronization seeks to output multiple consistent relative transformations. This problem naturally arises in many geometry reconstruction applications, where the underlying object possesses self-symmetry. For approximately symmetric or even non-symmetric objects, K-best solutions offer an intermediate presentation for recovering the underlying single-best solution. We introduce a simple yet robust iterative algorithm for K-best transformation synchronization, which alternates between transformation propagation and transformation clustering. We present theoretical guarantees on the robust and exact recoveries of our algorithm. Experimental results demonstrate the advantage of our approach against state-of-the-art transformation synchronization techniques on both synthetic and real datasets. Yifan Sun 0007, Jiacheng Zhuo, Arnav Mohan, Qixing Huang |
ICCV | 2 |
| 2019 | Primal-Dual Block Generalized Frank-WolfeabstractWe propose a generalized variant of Frank-Wolfe algorithm for solving a class of sparse/low-rank optimization problems. Our formulation includes Elastic Net, regularized SVMs and phase retrieval as special cases. The proposed Primal-Dual Block Generalized Frank-Wolfe algorithm reduces the per-iteration cost while maintaining linear convergence rate. The per iteration cost of our method depends on the structural complexity of the solution (i.e. sparsity/low-rank) instead of the ambient dimension. We empirically show that our algorithm outperforms the state-of-the-art methods on (multi-class) classification tasks. Jiacheng Zhuo, Constantine Caramanis, Inderjit S. Dhillon, Alexandros G. Dimakis |
NeurIPS | 2 |
| 2019 | Weaving geodesic foliationsabstractWe study discrete geodesic foliations of surfaces---foliations whose leaves are all approximately geodesic curves---and develop several new variational algorithms for computing such foliations. Our key insight is a relaxation of vector field integrability in the discrete setting, which allows us to optimize for curl-free unit vector fields that remain well-defined near singularities and robustly recover a scalar function whose gradient is well aligned to these fields. We then connect the physics governing surfaces woven out of thin ribbons to the geometry of geodesic foliations, and present a design and fabrication pipeline for approximating surfaces of arbitrary geometry and topology by triaxially-woven structures, where the ribbon layout is determined by a geodesic foliation on a sixfold branched cover of the input surface. We validate the effectiveness of our pipeline on a variety of simulated and fabricated woven designs, including an example for readers to try at home. Josh Vekhter, Jiacheng Zhuo, Luisa F. Gil Fandino, Qixing Huang, Etienne Vouga |
ACM Trans. Graph. | 2 |