Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Amnon Geifman

dblp:232/2462 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Learning theory · 32% Deep learning architectures and training · 23% Efficient and distributed learning · 21%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 50% Computational photography and imaging · 50%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
2.952024
Spectral Analysis of the Neural Tangent Kernel for Deep Residual Networks · J. Mach. Learn. Res. 2024
A Kernel Perspective of Skip Connections in Convolutional Networks · ICLR 2023
On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process Kernels · NeurIPS 2022
Machine learning › Efficient and distributed learning
model compression
1.722025
FFN Fusion: Rethinking Sequential Computation in Large Language Models · NeurIPS 2025
Puzzle: Distillation-Based NAS for Inference-Optimized LLMs · ICML 2025
Computer vision › 3D vision
structure from motion
1.232020
Averaging Essential and Fundamental Matrices in Collinear Camera Settings · CVPR 2020
Algebraic Characterization of Essential Matrices and Their Averaging in Multiview Settings · ICCV 2019
GPSfM: Global Projective SFM Using Algebraic Constraints on Multi-View Fundamental Matrices · CVPR 2019
Machine learning › Kernel, tree and ensemble methods
kernel methods
1.122023
A Kernel Perspective of Skip Connections in Convolutional Networks · ICLR 2023
On the Similarity between the Laplace and Neural Tangent Kernels · NeurIPS 2020
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
Puzzle: Distillation-Based NAS for Inference-Optimized LLMs · ICML 2025
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.912025
Puzzle: Distillation-Based NAS for Inference-Optimized LLMs · ICML 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
FFN Fusion: Rethinking Sequential Computation in Large Language Models · NeurIPS 2025
Machine learning › Deep learning architectures and training › convolutional neural network
residual network
0.812024
Spectral Analysis of the Neural Tangent Kernel for Deep Residual Networks · J. Mach. Learn. Res. 2024
Machine learning › Learning theory
spectral analysis
0.812024
Spectral Analysis of the Neural Tangent Kernel for Deep Residual Networks · J. Mach. Learn. Res. 2024
Machine learning › Deep learning architectures and training
convolutional neural network
0.712023
A Kernel Perspective of Skip Connections in Convolutional Networks · ICLR 2023
Machine learning › Deep learning architectures and training
skip connections
0.712023
A Kernel Perspective of Skip Connections in Convolutional Networks · ICLR 2023
Machine learning › Learning theory › inductive bias
spectral bias
0.612022
On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process Kernels · NeurIPS 2022
Computer vision › 3D vision
camera pose estimation
0.522020
Algebraic Characterization of Essential Matrices and Their Averaging in Multiview Settings · ICCV 2019
Averaging Essential and Fundamental Matrices in Collinear Camera Settings · CVPR 2020
Machine learning › Deep learning architectures and training › training dynamics
frequency bias
0.412020
Frequency Bias in Neural Networks for Input of Non-Uniform Density · ICML 2020
Machine learning › Deep learning architectures and training
training dynamics
0.412020
Frequency Bias in Neural Networks for Input of Non-Uniform Density · ICML 2020
Computer vision › 3D vision › structure from motion
projective structure from motion
0.412019
GPSfM: Global Projective SFM Using Algebraic Constraints on Multi-View Fundamental Matrices · CVPR 2019
Computational photography and imaging › camera geometry
fundamental matrix estimation
0.412019
GPSfM: Global Projective SFM Using Algebraic Constraints on Multi-View Fundamental Matrices · CVPR 2019
Geometric modeling and processing
multi-view geometry
0.412019
GPSfM: Global Projective SFM Using Algebraic Constraints on Multi-View Fundamental Matrices · CVPR 2019
Machine learning › Learning theory
generalization bounds
0.422024
Spectral Analysis of the Neural Tangent Kernel for Deep Residual Networks · J. Mach. Learn. Res. 2024
Frequency Bias in Neural Networks for Input of Non-Uniform Density · ICML 2020
Natural language and speech › Language models and text generation
large language model
0.312025
Puzzle: Distillation-Based NAS for Inference-Optimized LLMs · ICML 2025

Methods — techniques the papers use, named apart from their topics

spherical harmonics · 1.3neural tangent kernel · 1.2parallelization · 0.9mixed integer programming · 0.9feed-forward network fusion · 0.9blockwise local knowledge distillation · 0.9spectral analysis · 0.8eigenvalue and rank characterization · 0.8hierarchical factorizable kernels · 0.6spectral characterization · 0.4rank-constrained minimization · 0.4algebraic constraints · 0.4algebraic constraint · 0.4
YearPublicationVenuePosition
2025 Puzzle: Distillation-Based NAS for Inference-Optimized LLMs
abstract
Large language models (LLMs) offer remarkable capabilities, yet their high inference costs restrict wider adoption. While increasing parameter counts improves accuracy, it also broadens the gap between state-of-the-art capabilities and practical deployability. We present **Puzzle**, a hardware-aware framework that accelerates the inference of LLMs while preserving their capabilities. Using neural architecture search (NAS) at a large-scale, Puzzle optimizes models with tens of billions of parameters. Our approach utilizes blockwise local knowledge distillation (BLD) for parallel architecture exploration and employs mixed-integer programming for precise constraint optimization. We showcase our framework’s impact via Llama-3.1-Nemotron-51B-Instruct (Nemotron-51B) and Llama-3.3-Nemotron-49B, two publicly available models derived from Llama-70B-Instruct. Both models achieve a 2.17x inference throughput speedup, fitting on a single NVIDIA H100 GPU while retaining 98.4% of the original model's benchmark accuracies. These are the most accurate models supporting single H100 GPU inference with large batch sizes, despite training on 45B tokens at most, far fewer than the 15T used to train Llama-70B. Lastly, we show that lightweight alignment on these derived models allows them to surpass the parent model in specific capabilities. Our work establishes that powerful LLM models can be optimized for efficient deployment with only negligible loss in quality, underscoring that inference performance, not parameter count alone, should guide model selection.
Akhiad Bercovich, Tomer Ronen, Talor Abramovich, Nir Ailon, Nave Assaf, Mohammad Dabbah, Ido Galil, Amnon Geifman, Yonatan Geifman, Izhak Golan, Netanel Haber, Ehud Karpas, Roi Koren, Itay Levy, Pavlo Molchanov 0001, Shahar Mor, Zach Moshe, Najeeb Nabwani, Omri Puny, Ran Rubin, Itamar Schen, Ido Shahaf, Oren Tropp, Omer Ullman Argov, Ran Zilberstein, Ran El-Yaniv
ICML8
2025 FFN Fusion: Rethinking Sequential Computation in Large Language Models
abstract
We introduce \textit{FFN Fusion}, an architectural optimization technique that reduces sequential computation in large language models by identifying and exploiting natural opportunities for parallelization. Our key insight is that sequences of Feed-Forward Network (FFN) layers, particularly those remaining after the removal of specific attention layers, can often be parallelized with minimal accuracy impact. We develop a principled methodology for identifying and fusing such sequences, transforming them into parallel operations that significantly reduce inference latency while preserving model behavior. Applying these techniques to Llama-3.1-405B-Instruct, we create a 253B model (253B-Base), an efficient and soon-to-be publicly available model that achieves a 1.71$\times$ speedup in inference latency and 35$\times$ lower per-token cost while maintaining strong performance across benchmarks. Most intriguingly, we find that even full transformer blocks containing both attention and FFN layers can sometimes be parallelized, suggesting new directions for neural architecture design.
Akhiad Bercovich, Mohammad Dabbah, Omri Puny, Ido Galil, Amnon Geifman, Yonatan Geifman, Izik Golan, Ehud Karpas, Itay Levy, Zach Moshe, Najeeb Nabwani, Tomer Ronen, Itamar Schen, Ido Shahaf, Oren Tropp, Ran Zilberstein, Ran El-Yaniv
NeurIPS5
2024 Spectral Analysis of the Neural Tangent Kernel for Deep Residual Networks
abstract
Deep residual network architectures have been shown to achieve superior accuracy over classical feed-forward networks, yet their success is still not fully understood. Focusing on massively over-parameterized, fully connected residual networks with ReLU activation through their respective neural tangent kernels (ResNTK), we provide here a spectral analysis of these kernels. Specifically, we show that, much like NTK for fully connected networks (FC-NTK), for input distributed uniformly on the hypersphere $S^d$, the eigenvalues of ResNTK corresponding to their spherical harmonics eigenfunctions decay polynomially with frequency $k$ as $k^{-d}$. These in turn imply that the set of functions in their Reproducing Kernel Hilbert Space are identical to those of both FC-NTK as well as the standard Laplace kernel. Our spectral analysis allows us to highlight several additional properties of ResNTK, which depend on the choice of a hyper-parameter that balances between the skip and residual connections. Specifically, (1) with no bias, deep ResNTK is significantly biased toward even frequency functions; (2) unlike FC-NTK for deep networks, which is spiky and therefore yields poor generalization, ResNTK is stable and yields small generalization errors. We finally demonstrate these with experiments showing further that these phenomena arise in real networks.
Yuval Belfer, Amnon Geifman, Meirav Galun, Ronen Basri
J. Mach. Learn. Res.2
2023 A Kernel Perspective of Skip Connections in Convolutional Networks
Daniel Barzilai, Amnon Geifman, Meirav Galun, Ronen Basri
ICLR2
2022 On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process Kernels
abstract
We study the properties of various over-parameterized convolutional neural architectures through their respective Gaussian Process and Neural Tangent kernels. We prove that, with normalized multi-channel input and ReLU activation, the eigenfunctions of these kernels with the uniform measure are formed by products of spherical harmonics, defined over the channels of the different pixels. We next use hierarchical factorizable kernels to bound their respective eigenvalues. We show that the eigenvalues decay polynomially, quantify the rate of decay, and derive measures that reflect the composition of hierarchical features in these networks. Our theory provides a concrete quantitative characterization of the role of locality and hierarchy in the inductive bias of over-parameterized convolutional network architectures.
Amnon Geifman, Meirav Galun, David Jacobs 0001, Ronen Basri
NeurIPS1
2020 Averaging Essential and Fundamental Matrices in Collinear Camera Settings
abstract
Global methods to Structure from Motion have gained popularity in recent years. A significant drawback of global methods is their sensitivity to collinear camera settings. In this paper, we introduce an analysis and algorithms for averaging bifocal tensors (essential or fundamental matrices) when either subsets or all of the camera centers are collinear. We provide a complete spectral characterization of bifocal tensors in collinear scenarios and further propose two averaging algorithms. The first algorithm uses rank constrained minimization to recover camera matrices in fully collinear settings. The second algorithm enriches the set of possibly mixed collinear and non-collinear cameras with additional, ``virtual cameras," which are placed in general position, enabling the application of existing averaging methods to the enriched set of bifocal tensors. Our algorithms are shown to achieve state of the art results on various benchmarks that include autonomous car datasets and unordered image collections in both calibrated and unclibrated settings.
Amnon Geifman, Yoni Kasten, Meirav Galun, Ronen Basri
CVPR1
2020 Frequency Bias in Neural Networks for Input of Non-Uniform Density
abstract
Recent works have partly attributed the generalization ability of over-parameterized neural networks to frequency bias – networks trained with gradient descent on data drawn from a uniform distribution find a low frequency fit before high frequency ones. As realistic training sets are not drawn from a uniform distribution, we here use the Neural Tangent Kernel (NTK) model to explore the effect of variable density on training dynamics. Our results, which combine analytic and empirical observations, show that when learning a pure harmonic function of frequency $\kappa$, convergence at a point $x \in \S^{d-1}$ occurs in time $O(\kappa^d/p(x))$ where $p(x)$ denotes the local density at $x$. Specifically, for data in $\S^1$ we analytically derive the eigenfunctions of the kernel associated with the NTK for two-layer networks. We further prove convergence results for deep, fully connected networks with respect to the spectral decomposition of the NTK. Our empirical study highlights similarities and differences between deep and shallow networks in this model.
Ronen Basri, Meirav Galun, Amnon Geifman, David Jacobs 0001, Yoni Kasten, Shira Kritchman
ICML3
2020 On the Similarity between the Laplace and Neural Tangent Kernels
abstract
Recent theoretical work has shown that massively overparameterized neural networks are equivalent to kernel regressors that use Neural Tangent Kernels (NTKs). Experiments show that these kernel methods perform similarly to real neural networks. Here we show that NTK for fully connected networks with ReLU activation is closely related to the standard Laplace kernel. We show theoretically that for normalized data on the hypersphere both kernels have the same eigenfunctions and their eigenvalues decay polynomially at the same rate, implying that their Reproducing Kernel Hilbert Spaces (RKHS) include the same sets of functions. This means that both kernels give rise to classes of functions with the same smoothness properties. The two kernels differ for data off the hypersphere, but experiments indicate that when data is properly normalized these differences are not significant. Finally, we provide experiments on real data comparing NTK and the Laplace kernel, along with a larger class of $\gamma$-exponential kernels. We show that these perform almost identically. Our results suggest that much insight about neural networks can be obtained from analysis of the well-known Laplace kernel, which has a simple closed form.
Amnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun, David Jacobs 0001, Ronen Basri
NeurIPS1
2019 GPSfM: Global Projective SFM Using Algebraic Constraints on Multi-View Fundamental Matrices
abstract
This paper addresses the problem of recovering projective camera matrices from collections of fundamental matrices in multiview settings. We make two main contributions. First, given (2n) fundamental matrices computed for n images, we provide a complete algebraic characterization in the form of conditions that are both necessary and sufficient to enabling the recovery of camera matrices. These conditions are based on arranging the fundamental matrices as blocks in a single matrix, called the n-view fundamental matrix, and characterizing this matrix in terms of the signs of its eigenvalues and rank structures. Secondly, we propose a concrete algorithm for projective structure-formation that utilizes this characterization. Given a complete or partial collection of measured-fundamental matrices, our method seeks camera matrices that minimize a global algebraic error for the measured fundamental matrices. In contrast to existing methods, our optimization, without any initialization, produces a consistent set of fundamental matrices that corresponds to a unique set of cameras (up to a choice of projective frame). Our experiments indicate that our method achieves state of the art performance in both accuracy and running time.
Yoni Kasten, Amnon Geifman, Meirav Galun, Ronen Basri
CVPR2
2019 Algebraic Characterization of Essential Matrices and Their Averaging in Multiview Settings
abstract
Essential matrix averaging, i.e., the task of recovering camera locations and orientations in calibrated, multiview settings, is a first step in global approaches to Euclidean structure from motion. A common approach to essential matrix averaging is to separately solve for camera orientations and subsequently for camera positions. This paper presents a novel approach that solves simultaneously for both camera orientations and positions. We offer a complete characterization of the algebraic conditions that enable a unique Euclidean reconstruction of n cameras from a collection of (2n) essential matrices. We next use these conditions to formulate essential matrix averaging as a constrained optimization problem, allowing us to recover a consistent set of essential matrices given a (possibly partial) set of measured essential matrices computed independently for pairs of images. We finally use the recovered essential matrices to determine the global positions and orientations of the n cameras. We test our method on common SfM datasets, demonstrating high accuracy while maintaining efficiency and robustness, compared to existing methods.
Yoni Kasten, Amnon Geifman, Meirav Galun, Ronen Basri
ICCV2