Behnam Neyshabur

dblp:131/9898 · DBLP profile ↗
← Back
39ranked-venue papers
12as first author
18since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 11 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2023 REPAIR: REnormalizing Permuted Activations for Interpolation Repair
Keller Jordan, Hanie Sedghi, Olga Saukh, Rahim Entezari, Behnam Neyshabur
ICLR5
2023 Long Range Language Modeling via Gated State Spaces
Harsh Mehta, Ankit Gupta 0001, Ashok Cutkosky, Behnam Neyshabur
ICLR4
2022 Exploring the Limits of Large Scale Pre-training
Samira Abnar, Mostafa Dehghani 0001, Behnam Neyshabur, Hanie Sedghi
ICLR3
2022 The Role of Permutation Invariance in Linear Mode Connectivity of Neural Networks
Rahim Entezari, Hanie Sedghi, Olga Saukh, Behnam Neyshabur
ICLR4
2022 Leveraging unlabeled data to predict out-of-distribution performance
Sivaraman Balakrishnan, Zachary C. Lipton, Behnam Neyshabur, Hanie Sedghi
ICLR4
2022 A Loss Curvature Perspective on Training Instabilities of Deep Learning Models
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Reddy Kudugunta, Behnam Neyshabur, David Cardoze, George E. Dahl, Zachary Nado, Orhan Firat
ICLR5
2022 Data Scaling Laws in NMT: The Effect of Noise and Architecture
abstract
In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that the test loss of encoder-decoder transformer models scales as a power law in the number of training samples, with a dependence on the model size. Then, we systematically vary aspects of the training setup to understand how they impact the data scaling laws. In particular, we change the following (1) Architecture and task setup: We compare to a transformer-LSTM hybrid, and a decoder-only transformer with a language modeling loss (2) Noise level in the training distribution: We experiment with filtering, and adding iid synthetic noise. In all the above cases, we find that the data scaling exponents are minimally impacted, suggesting that marginally worse architectures or training data can be compensated for by adding more data. Lastly, we find that using back-translated data instead of parallel data, can significantly degrade the scaling exponent.
Yamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang 0006, Colin Cherry, Behnam Neyshabur, Orhan Firat
ICML6
2022 Revisiting Neural Scaling Laws in Language and Vision
abstract
The remarkable progress in deep learning in recent years is largely driven by improvements in scale, where bigger models are trained on larger datasets for longer schedules. To predict the benefit of scale empirically, we argue for a more rigorous methodology based on the extrapolation loss, instead of reporting the best-fitting (interpolating) parameters. We then present a recipe for estimating scaling law parameters reliably from learning curves. We demonstrate that it extrapolates more accurately than previous methods in a wide range of architecture families across several domains, including image classification, neural machine translation (NMT) and language modeling, in addition to tasks from the BIG-Bench evaluation benchmark. Finally, we release a benchmark dataset comprising of 90 evaluation tasks to facilitate research in this domain.
Ibrahim Alabdulmohsin, Behnam Neyshabur, Xiaohua Zhai
NeurIPS2
2022 Exploring Length Generalization in Large Language Models
abstract
The ability to extrapolate from short problem instances to longer ones is an important form of out-of-distribution generalization in reasoning tasks, and is crucial when learning from datasets where longer problem instances are rare. These include theorem proving, solving quantitative mathematics problems, and reading/summarizing novels. In this paper, we run careful empirical studies exploring the length generalization capabilities of transformer-based language models. We first establish that naively finetuning transformers on length generalization tasks shows significant generalization deficiencies independent of model scale. We then show that combining pretrained large language models' in-context learning abilities with scratchpad prompting (asking the model to output solution steps before producing an answer) results in a dramatic improvement in length generalization. We run careful failure analyses on each of the learning modalities and identify common sources of mistakes that highlight opportunities in equipping language models with the ability to generalize to longer problems.
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay V. Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, Behnam Neyshabur
NeurIPS10
2022 Block-Recurrent Transformers
abstract
We introduce the Block-Recurrent Transformer, which applies a transformer layer in a recurrent fashion along a sequence, and has linear complexity with respect to sequence length. Our recurrent cell operates on blocks of tokens rather than single tokens during training, and leverages parallel computation within a block in order to make efficient use of accelerator hardware. The cell itself is strikingly simple. It is merely a transformer layer: it uses self-attention and cross-attention to efficiently compute a recurrent function over a large set of state vectors and tokens. Our design was inspired in part by LSTM cells, and it uses LSTM-style gates, but it scales the typical LSTM cell up by several orders of magnitude. Our implementation of recurrence has the same cost in both computation time and parameter count as a conventional transformer layer, but offers dramatically improved perplexity in language modeling tasks over very long sequences. Our model out-performs a long-range Transformer XL baseline by a wide margin, while running twice as fast. We demonstrate its effectiveness on PG19 (books), arXiv papers, and GitHub source code. Our code has been released as open source.
DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, Behnam Neyshabur
NeurIPS5
2022 Solving Quantitative Reasoning Problems with Language Models
abstract
Language models have achieved remarkable performance on a wide range of tasks that require natural language understanding. Nevertheless, state-of-the-art models have generally struggled with tasks that require quantitative reasoning, such as solving mathematics, science, and engineering questions at the college level. To help close this gap, we introduce Minerva, a large language model pretrained on general natural language data and further trained on technical content. The model achieves strong performance in a variety of evaluations, including state-of-the-art performance on the MATH dataset. We also evaluate our model on over two hundred undergraduate-level problems in physics, biology, chemistry, economics, and other sciences that require quantitative reasoning, and find that the model can correctly answer nearly a quarter of them.
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V. Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, Vedant Misra
NeurIPS12
2021 Sharpness-aware Minimization for Efficiently Improving Generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam Neyshabur
ICLR4
2021 Are wider nets better given the same number of parameters?
Anna Golubeva, Guy Gur-Ari, Behnam Neyshabur
ICLR3
2021 Extreme Memorization via Scale of Initialization
Harsh Mehta, Ashok Cutkosky, Behnam Neyshabur
ICLR3
2021 Understanding the failure modes of out-of-distribution generalization
Vaishnavh Nagarajan, Anders Andreassen, Behnam Neyshabur
ICLR3
2021 The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers
Preetum Nakkiran, Behnam Neyshabur, Hanie Sedghi
ICLR2
2021 When Do Curricula Work?
Xiaoxia Wu, Ethan Dyer, Behnam Neyshabur
ICLR3
2021 Deep Learning Through the Lens of Example Difficulty
abstract
Existing work on understanding deep learning often employs measures that compress all data-dependent information into a few numbers. In this work, we adopt a perspective based on the role of individual examples. We introduce a measure of the computational difficulty of making a prediction for a given input: the (effective) prediction depth. Our extensive investigation reveals surprising yet simple relationships between the prediction depth of a given input and the model’s uncertainty, confidence, accuracy and speed of learning for that data point. We further categorize difficult examples into three interpretable groups, demonstrate how these groups are processed differently inside deep models and showcase how this understanding allows us to improve prediction accuracy. Insights from our study lead to a coherent view of a number of separately reported phenomena in the literature: early layers generalize while later layers memorize; early layers converge faster and networks learn easy data and simple functions first.
Robert J. N. Baldock, Hartmut Maennel, Behnam Neyshabur
NeurIPS3
2020 The intriguing role of module criticality in the generalization of deep networks
Niladri S. Chatterji, Behnam Neyshabur, Hanie Sedghi
ICLR2
2020 Fantastic Generalization Measures and Where to Find Them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, Samy Bengio
ICLR2
2020 Observational Overfitting in Reinforcement Learning
Xingyou Song, Yiding Jiang, Stephen Tu, Yilun Du, Behnam Neyshabur
ICLR5
2020 Towards Learning Convolutions from Scratch
abstract
Convolution is one of the most essential components of modern architectures used in computer vision. As machine learning moves towards reducing the expert bias and learning it from data, a natural next step seems to be learning convolution-like structures from scratch. This, however, has proven elusive. For example, current state-of-the-art architecture search algorithms use convolution as one of the existing modules rather than learning it from data. In an attempt to understand the inductive bias that gives rise to convolutions, we investigate minimum description length as a guiding principle and show that in some settings, it can indeed be indicative of the performance of architectures. To find architectures with small description length, we propose beta-LASSO, a simple variant of LASSO algorithm that, when applied on fully-connected networks for image classification tasks, learns architectures with local connections and achieves state-of-the-art accuracies for training fully-connected networks on CIFAR-10 (84.50%), CIFAR-100 (57.76%) and SVHN (93.84%) bridging the gap between fully-connected and convolutional networks.
Behnam Neyshabur
NeurIPS1
2020 What is being transferred in transfer learning?
abstract
One desired capability for machines is the ability to transfer their understanding of one domain to another domain where data is (usually) scarce. Despite ample adaptation of transfer learning in many deep learning applications, we yet do not understand what enables a successful transfer and which part of the network is responsible for that. In this paper, we provide new tools and analysis to address these fundamental questions. Through a series of analysis on transferring to block-shuffled images, we separate the effect of feature reuse from learning high-level statistics of data and show that some benefit of transfer learning comes from the latter. We present that when training from pre-trained weights, the model stays in the same basin in the loss landscape and different instances of such model are similar in feature space and close in parameter space.
Behnam Neyshabur, Hanie Sedghi, Chiyuan Zhang
NeurIPS1
2019 The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li 0005, Srinadh Bhojanapalli, Yann LeCun, Nathan Srebro
ICLR (Poster)1
2018 A PAC-Bayesian Approach to Spectrally-Normalized Margin Bounds for Neural Networks
Behnam Neyshabur, Srinadh Bhojanapalli, Nathan Srebro
ICLR (Poster)1
2018 Stronger Generalization Bounds for Deep Nets via a Compression Approach
abstract
Deep nets generalize well despite having more parameters than the number of training samples. Recent works try to give an explanation using PAC-Bayes and Margin-based analyses, but do not as yet result in sample complexity bounds better than naive parameter counting. The current paper shows generalization bounds that are orders of magnitude better in practice. These rely upon new succinct reparametrizations of the trained net — a compression that is explicit and efficient. These yield generalization bounds via a simple compression-based framework introduced here. Our results also provide some theoretical justification for widespread empirical success in compressing deep nets. Analysis of correctness of our compression relies upon some newly identified noise stability properties of trained deep nets, which are also experimentally verified. The study of these properties and resulting generalization bounds are also extended to convolutional nets, which had eluded earlier attempts on proving generalization.
Sanjeev Arora, Rong Ge 0001, Behnam Neyshabur, Yi Zhang 0074
ICML3
2018 Predicting protein-protein interactions through sequence-based deep learning
abstract
Motivation: High-throughput experimental techniques have produced a large amount of protein-protein interaction (PPI) data, but their coverage is still low and the PPI data is also very noisy. Computational prediction of PPIs can be used to discover new PPIs and identify errors in the experimental PPI data. Results: We present a novel deep learning framework, DPPI, to model and predict PPIs from sequence information alone. Our model efficiently applies a deep, Siamese-like convolutional neural network combined with random projection and data augmentation to predict PPIs, leveraging existing high-quality experimental PPI data and evolutionary information of a protein pair under prediction. Our experimental results show that DPPI outperforms the state-of-the-art methods on several benchmarks in terms of area under precision-recall curve (auPR), and computationally is more efficient. We also show that DPPI is able to predict homodimeric interactions where other methods fail to work accurately, and the effectiveness of DPPI in specific applications such as predicting cytokine-receptor binding affinities. Availability and implementation: Predicting protein-protein interactions through sequence-based deep learning): https://github.com/hashemifar/DPPI/. Supplementary information: Supplementary data are available at Bioinformatics online.
Somaye Hashemifar, Behnam Neyshabur, Aly Azeem Khan, Jinbo Xu
Bioinform.2
2017 Corralling a Band of Bandit Algorithms
abstract
We study the problem of combining multiple bandit algorithms (that is, online learning algorithms with partial feedback) with the goal of creating a master algorithm that performs almost as well as the best base algorithm \it if it were to be run on its own. The main challenge is that when run with a master, base algorithms unavoidably receive much less feedback and it is thus critical that the master not starve a base algorithm that might perform uncompetitively initially but would eventually outperform others if given enough feedback. We address this difficulty by devising a version of Online Mirror Descent with a special mirror map together with a sophisticated learning rate scheme. We show that this approach manages to achieve a more delicate balance between exploiting and exploring base algorithms than previous works yielding superior regret bounds. Our results are applicable to many settings, such as multi-armed bandits, contextual bandits, and convex bandits. As examples, we present two main applications. The first is to create an algorithm that enjoys worst-case robustness while at the same time performing much better when the environment is relatively easy. The second is to create an algorithm that works simultaneously under different assumptions of the environment, such as different priors or different loss structures.
Alekh Agarwal, Behnam Neyshabur, Robert E. Schapire
COLT3
2017 Implicit Regularization in Matrix Factorization
abstract
We study implicit regularization when optimizing an underdetermined quadratic objective over a matrix $X$ with gradient descent on a factorization of X. We conjecture and provide empirical and theoretical evidence that with small enough step sizes and initialization close enough to the origin, gradient descent on a full dimensional factorization converges to the minimum nuclear norm solution.
Suriya Gunasekar, Blake E. Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, Nathan Srebro
NIPS4
2017 Exploring Generalization in Deep Learning
abstract
With a goal of understanding what drives generalization in deep networks, we consider several recently suggested explanations, including norm-based control, sharpness and robustness. We study how these measures can ensure generalization, highlighting the importance of scale normalization, and making a connection between sharpness and PAC-Bayes theory. We then investigate how well the measures explain different observed phenomena.
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, Nathan Srebro
NIPS1
2016 Global Optimality of Local Search for Low Rank Matrix Recovery
abstract
We show that there are no spurious local minima in the non-convex factorized parametrization of low-rank matrix recovery from incoherent linear measurements. With noisy measurements we show all local minima are very close to a global optimum. Together with a curvature bound at saddle points, this yields a polynomial time global convergence guarantee for stochastic gradient descent {\em from random initialization}.
Srinadh Bhojanapalli, Behnam Neyshabur, Nathan Srebro
NIPS2
2016 Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations
abstract
We investigate the parameter-space geometry of recurrent neural networks (RNNs), and develop an adaptation of path-SGD optimization method, attuned to this geometry, that can learn plain RNNs with ReLU activations. On several datasets that require capturing long-term dependency structure, we show that path-SGD can significantly improve trainability of ReLU RNNs compared to RNNs trained with SGD, even with various recently suggested initialization schemes.
Behnam Neyshabur, Yuhuai Wu, Ruslan Salakhutdinov, Nathan Srebro
NIPS1
2015 Joint inference of tissue-specific networks with a scale free topology
abstract
High-throughput experimental techniques have produced an enormous number of gene expression profiles for various tissues of the human body. Tissue-specificity is a key component in reflecting the potentially different roles of proteins in diverse cell lineages. One way of understanding the tissue specificity is by reconstructing the tissue-specific co-expression networks (CENs) to analyze the correlation between genes. A few methods have been developed for estimating CENs, but it still remains challenging in terms of both accuracy and efficiency. In this paper we propose a new method, JointNet, for predicting tissue-specific co-expression networks. JointNet is exploiting the observation that, functionally related tissues have similar expression patterns and thus, similar networks. It uses different node penalties for hubs and non-hub nodes to accurately estimate the scale-free networks. Our experimental results show that the resulting tissue-specific CENs are accurate and that our method outperforms the current state of the art.
Somaye Hashemifar, Behnam Neyshabur, Jinbo Xu
BIBM2
2015 Norm-Based Capacity Control in Neural Networks
abstract
We investigate the capacity, convexity and characterization of a general family of norm-constrained feed-forward networks.
Behnam Neyshabur, Ryota Tomioka, Nathan Srebro
COLT1
2015 On Symmetric and Asymmetric LSHs for Inner Product Search
abstract
We consider the problem of designing locality sensitive hashes (LSH) for inner product similarity, and of the power of asymmetric hashes in this context. Shrivastava and Li (2014a) argue that there is no symmetric LSH for the problem and propose an asymmetric LSH based on different mappings for query and database points. However, we show there does exist a simple symmetric LSH that enjoys stronger guarantees and better empirical performance than the asymmetric LSH they suggest. We also show a variant of the settings where asymmetry is in-fact needed, but there a different asymmetric LSH is required.
Behnam Neyshabur, Nathan Srebro
ICML1
2015 Path-SGD: Path-Normalized Optimization in Deep Neural Networks
abstract
We revisit the choice of SGD for training deep neural networks by reconsidering the appropriate geometry in which to optimize the weights. We argue for a geometry invariant to rescaling of weights that does not affect the output of the network, and suggest Path-SGD, which is an approximate steepest descent method with respect to a path-wise regularizer related to max-norm regularization. Path-SGD is easy and efficient to implement and leads to empirical gains over SGD and AdaGrad.
Behnam Neyshabur, Ruslan Salakhutdinov, Nathan Srebro
NIPS1
2014 Clustering, Hamming Embedding, Generalized LSH and the Max Norm
Behnam Neyshabur, Yury Makarychev, Nathan Srebro
ALT1
2013 The Power of Asymmetry in Binary Hashing
abstract
When approximating binary similarity using the hamming distance between short binary hashes, we shown that even if the similarity is symmetric, we can have shorter and more accurate hashes by using two distinct code maps. I.e.~by approximating the similarity between $x$ and $x'$ as the hamming distance between $f(x)$ and $g(x')$, for two distinct binary codes $f,g$, rather than as the hamming distance between $f(x)$ and $f(x')$.
Behnam Neyshabur, Nathan Srebro, Ruslan Salakhutdinov, Yury Makarychev, Payman Yadollahpour
NIPS1
2013 NETAL: a new graph-based method for global alignment of protein-protein interaction networks
abstract
MOTIVATION: The interactions among proteins and the resulting networks of such interactions have a central role in cell biology. Aligning these networks gives us important information, such as conserved complexes and evolutionary relationships. Although there have been several publications on the global alignment of protein networks; however, none of proposed methods are able to produce a highly conserved and meaningful alignment. Moreover, time complexity of current algorithms makes them impossible to use for multiple alignment of several large networks together. RESULTS: We present a novel algorithm for the global alignment of protein-protein interaction networks. It uses a greedy method, based on the alignment scoring matrix, which is derived from both biological and topological information of input networks to find the best global network alignment. NETAL outperforms other global alignment methods in terms of several measurements, such as Edge Correctness, Largest Common Connected Subgraphs and the number of common Gene Ontology terms between aligned proteins. As the running time of NETAL is much less than other available methods, NETAL can be easily expanded to multiple alignment algorithm. Furthermore, NETAL overpowers all other existing algorithms in term of performance so that the short running time of NETAL allowed us to implement it as the first server for global alignment of protein-protein interaction networks. AVAILABILITY: Binaries supported on linux are freely available for download at http://www.bioinf.cs.ipm.ir/software/netal. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Behnam Neyshabur, Ahmadreza Khadem, Somaye Hashemifar, Seyed Shahriar Arab
Bioinform.1