Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Itai Kreisler

dblp:347/9893 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 35% Optimization for machine learning · 27% Efficient and distributed learning · 20%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization › stochastic optimization
stochastic convex optimization
1.822026
Clipping the Price of Adaptivity at the Tail · COLT 2026
Accelerated Parameter-Free Stochastic Optimization · COLT 2024
Machine learning › Optimization for machine learning
adaptive algorithm
1.012026
Clipping the Price of Adaptivity at the Tail · COLT 2026
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
adaptive optimization
1.012026
Clipping the Price of Adaptivity at the Tail · COLT 2026
Machine learning › Efficient and distributed learning
model acceleration
0.812024
Accelerated Parameter-Free Stochastic Optimization · COLT 2024
Mathematical optimization › online optimization
parameter-free optimization
0.812024
Accelerated Parameter-Free Stochastic Optimization · COLT 2024
Machine learning › Deep learning architectures and training › training dynamics
edge of stability
0.712023
Gradient Descent Monotonically Decreases the Sharpness of Gradient Flow Solutions in Scalar Networks and Beyond · ICML 2023
Machine learning › Deep learning architectures and training › training dynamics
gradient descent dynamics
0.712023
Gradient Descent Monotonically Decreases the Sharpness of Gradient Flow Solutions in Scalar Networks and Beyond · ICML 2023

Methods — techniques the papers use, named apart from their topics

tail events · 2.0clipping · 2.0iterate stabilization · 1.5dog · 1.5UniXGrad · 1.5gradient flow analysis · 0.7gradient descent · 0.7
YearPublicationVenuePosition
2026 Clipping the Price of Adaptivity at the Tail
abstract
Adaptive stochastic convex optimization (SCO) methods face a fundamental “price of adaptivity” barrier: under the standard set of assumptions, they cannot efficiently adapt to large uncertainty in both the initial distance to optimality and the Lipschitz constant. We circumvent this barrier by requiring a small amount of additional structure common to many learning problems. Specifically, we assume that the objective decomposes into a model and a loss function, enabling us to intervene by modifying the model’s output before it passes to the loss function. Under this assumption, we design a method that clips the learned model output in tail events where it deviates too much from the output of a fixed reference model. Our method matches the optimal bounds for known-parameter SCO up to logarithmic factors in the uncertainty in the distance and Lipschitz parameters, thus efficiently adapting to large uncertainty in both.
Itai Kreisler, Yair Carmon, Oliver Hinder
COLT1
2024 Accelerated Parameter-Free Stochastic Optimization
abstract
We propose a method that achieves near-optimal rates for \emph{smooth} stochastic convex optimization and requires essentially no prior knowledge of problem parameters. This improves on prior work which requires knowing at least the initial distance to optimality $d_0$. Our method, \textsc{U-DoG}, combines \textsc{UniXGrad} (Kavis et al., 2019) and \textsc{DoG} (Ivgi et al., 2023) with novel iterate stabilization techniques. It requires only loose bounds on $d_0$ and the noise magnitude, provides high probability guarantees under sub-Gaussian noise, and is also near-optimal in the non-smooth case. Our experiments show consistent, strong performance on convex problems and mixed results on neural network training.
Itai Kreisler, Maor Ivgi, Oliver Hinder, Yair Carmon
COLT1
2023 Gradient Descent Monotonically Decreases the Sharpness of Gradient Flow Solutions in Scalar Networks and Beyond
abstract
Recent research shows that when Gradient Descent (GD) is applied to neural networks, the loss almost never decreases monotonically. Instead, the loss oscillates as gradient descent converges to its “Edge of Stability” (EoS). Here, we find a quantity that does decrease monotonically throughout GD training: the sharpness attained by the gradient flow solution (GFS)—the solution that would be obtained if, from now until convergence, we train with an infinitesimal step size. Theoretically, we analyze scalar neural networks with the squared loss, perhaps the simplest setting where the EoS phenomena still occur. In this model, we prove that the GFS sharpness decreases monotonically. Using this result, we characterize settings where GD provably converges to the EoS in scalar networks. Empirically, we show that GD monotonically decreases the GFS sharpness in a squared regression model as well as practical neural network architectures.
Itai Kreisler, Mor Shpigel Nacson, Daniel Soudry, Yair Carmon
ICML1