VLDB 2026 Research / reviewers in the wild / expert
Rachel Ward
dblp:364/6945
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 36% Optimization for machine learning · 21% Learning paradigms · 18% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 50% Algorithms and data structures · 50% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › gradient-based optimization
accelerated gradient methods |
0.8 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Algorithms and data structures › numerical linear algebra
matrix factorization |
0.8 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Mathematical optimization
nonconvex optimization |
0.8 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.7 | 1 | 2023 | Cluster-aware Semi-supervised Learning: Relational Knowledge Distillation Provably Learns Clustering · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
relational knowledge distillation |
0.7 | 1 | 2023 | Cluster-aware Semi-supervised Learning: Relational Knowledge Distillation Provably Learns Clustering · NeurIPS 2023 |
Machine learning › Learning paradigms
semi-supervised learning |
0.7 | 1 | 2023 | Cluster-aware Semi-supervised Learning: Relational Knowledge Distillation Provably Learns Clustering · NeurIPS 2023 |
Machine learning › Graph learning › graph clustering
spectral clustering |
0.7 | 1 | 2023 | Cluster-aware Semi-supervised Learning: Relational Knowledge Distillation Provably Learns Clustering · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks |
0.2 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
unbalanced initialization · 1.5nesterov's accelerated gradient · 1.5gradient descent · 1.5spectral clustering · 0.7data augmentation consistency regularization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural NetworksabstractWe study the convergence rate of first-order methods for rectangular matrix factorization, which is a canonical nonconvex optimization problem. Specifically, given a rank-$r$ matrix $\mathbf{A}\in\mathbb{R}^{m\times n}$, we prove that gradient descent (GD) can find a pair of $\epsilon$-optimal solutions $\mathbf{X}_T\in\mathbb{R}^{m\times d}$ and $\mathbf{Y}_T\in\mathbb{R}^{n\times d}$, where $d\geq r$, satisfying $\lVert\mathbf{X}_T\mathbf{Y}_T^\top-\mathbf{A}\rVert_F\leq\epsilon\lVert\mathbf{A}\rVert_F$ in $T=O(\kappa^2\log\frac{1}{\epsilon})$ iterations with high probability, where $\kappa$ denotes the condition number of $\mathbf{A}$. Furthermore, we prove that Nesterov's accelerated gradient (NAG) attains an iteration complexity of $O(\kappa\log\frac{1}{\epsilon})$, which is the best-known bound of first-order methods for rectangular matrix factorization. Different from small balanced random initialization in the existing literature, we adopt an unbalanced initialization, where $\mathbf{X}_0$ is large and $\mathbf{Y}_0$ is $0$. Moreover, our initialization and analysis can be further extended to linear neural networks, where we prove that NAG can also attain an accelerated linear convergence rate. In particular, we only require the width of the network to be greater than or equal to the rank of the output label matrix. In contrast, previous results achieving the same rate require excessive widths that additionally depend on the condition number and the rank of the input data matrix. Zhenghao Xu, Yuqing Wang 0005, Tuo Zhao, Rachel Ward, Molei Tao |
NeurIPS | 4 |
| 2023 | Cluster-aware Semi-supervised Learning: Relational Knowledge Distillation Provably Learns ClusteringabstractDespite the empirical success and practical significance of (relational) knowledge distillation that matches (the relations of) features between teacher and student models, the corresponding theoretical interpretations remain limited for various knowledge distillation paradigms. In this work, we take an initial step toward a theoretical understanding of relational knowledge distillation (RKD), with a focus on semi-supervised classification problems. We start by casting RKD as spectral clustering on a population-induced graph unveiled by a teacher model. Via a notion of clustering error that quantifies the discrepancy between the predicted and ground truth clusterings, we illustrate that RKD over the population provably leads to low clustering error. Moreover, we provide a sample complexity bound for RKD with limited unlabeled samples. For semi-supervised learning, we further demonstrate the label efficiency of RKD through a general framework of cluster-aware semi-supervised learning that assumes low clustering errors. Finally, by unifying data augmentation consistency regularization into this cluster-aware framework, we show that despite the common effect of learning accurate clusterings, RKD facilitates a "global" perspective through spectral clustering, whereas consistency regularization focuses on a "local" perspective via expansion. Yijun Dong, Rachel Ward |
NeurIPS | 4 |