Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yishun Lu

dblp:415/2781 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Optimization for machine learning · 75% Deep learning architectures and training · 25%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
1.012026
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026
Machine learning › Deep learning architectures and training › training optimization
large-batch training
1.012026
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
natural gradient descent
1.012026
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026
Machine learning › Optimization for machine learning
second-order optimization
1.012026
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training · AAAI 2026

Methods — techniques the papers use, named apart from their topics

variance-aware update · 1.0kronecker-factored approximate curvature · 1.0fisher-orthogonal projection · 1.0
YearPublicationVenuePosition
2026 Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
abstract
Modern GPUs are equipped with large amounts of high-bandwidth memory, enabling them to support mini-batch sizes of up to tens of thousands of training samples. However, most existing optimizers struggle to perform effectively at such a large batch size. As batch size increases, gradient noise decreases due to averaging over many samples, limiting the ability of first-order methods to escape sharp or suboptimal minima and reach the global minimum. Meanwhile, second-order methods like the natural gradient with Kronecker-Factored Approximate Curvature (KFAC) often require excessively high damping to remain stable at large batch sizes. This high damping effectively ``washes out" the curvature information that gives these methods their advantage, reducing their performance to that of simple gradient descent. In this paper, we introduce Fisher-Orthogonal Projection (FOP), a novel technique that restores the effectiveness of the second-order method at very large batch sizes, enabling scalable training with improved generalization and faster convergence. FOP constructs a variance-aware update direction by leveraging gradients from two sub-batches, enhancing the average gradient with a component of the gradient difference that is orthogonal to the average under the Fisher-metric. Through extensive benchmarks, we show that FOP accelerates convergence by ×1.2–1.3 over K-FAC and ×1.5–1.7 over SGD/AdamW at the same moderate batch sizes, while at extreme scales it achieves up to a ×7.5 speedup. Unlike other methods, FOP maintains small-batch accuracy when scaling to extremely large batch sizes. Moreover, it reduces Top-1 error by 2.3–3.3% on long-tailed CIFAR benchmarks, demonstrating robust generalization under severe class imbalance. Our lightweight, geometry-aware use of intra-batch variance makes natural-gradient optimization practical on modern data-centre GPUs. FOP is open-source and pip-installable, which can be integrated into existing training code with a single line and no extra configuration.
Yishun Lu, Wes Armour
AAAI1