EDBT 2026 Demo / reviewers in the wild / expert
Yishun Lu
dblp:415/2781
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Optimization for machine learning · 75% Deep learning architectures and training · 25% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
1.0 | 1 | 2026 | Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026 |
Machine learning › Deep learning architectures and training › training optimization
large-batch training |
1.0 | 1 | 2026 | Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026 |
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
natural gradient descent |
1.0 | 1 | 2026 | Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026 |
Machine learning › Optimization for machine learning
second-order optimization |
1.0 | 1 | 2026 | Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
variance-aware update · 1.0kronecker-factored approximate curvature · 1.0fisher-orthogonal projection · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch TrainingabstractModern GPUs are equipped with large amounts of high-bandwidth memory, enabling them to support mini-batch sizes of up to tens of thousands of training samples. However, most existing optimizers struggle to perform effectively at such a large batch size. As batch size increases, gradient noise decreases due to averaging over many samples, limiting the ability of first-order methods to escape sharp or suboptimal minima and reach the global minimum. Meanwhile, second-order methods like the natural gradient with Kronecker-Factored Approximate Curvature (KFAC) often require excessively high damping to remain stable at large batch sizes. This high damping effectively ``washes out" the curvature information that gives these methods their advantage, reducing their performance to that of simple gradient descent. In this paper, we introduce Fisher-Orthogonal Projection (FOP), a novel technique that restores the effectiveness of the second-order method at very large batch sizes, enabling scalable training with improved generalization and faster convergence. FOP constructs a variance-aware update direction by leveraging gradients from two sub-batches, enhancing the average gradient with a component of the gradient difference that is orthogonal to the average under the Fisher-metric. Through extensive benchmarks, we show that FOP accelerates convergence by ×1.2–1.3 over K-FAC and ×1.5–1.7 over SGD/AdamW at the same moderate batch sizes, while at extreme scales it achieves up to a ×7.5 speedup. Unlike other methods, FOP maintains small-batch accuracy when scaling to extremely large batch sizes. Moreover, it reduces Top-1 error by 2.3–3.3% on long-tailed CIFAR benchmarks, demonstrating robust generalization under severe class imbalance. Our lightweight, geometry-aware use of intra-batch variance makes natural-gradient optimization practical on modern data-centre GPUs. FOP is open-source and pip-installable, which can be integrated into existing training code with a single line and no extra configuration. Yishun Lu, Wes Armour |
AAAI | 1 |