Hossein Mobahi

dblp:94/1490 · DBLP profile ↗
← Back
29ranked-venue papers
11as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 9 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-authorSystems, architecture and hardware · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
abstract
Mihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen, Long Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Mihir Parmar, Palash Goyal, Yanfei Chen, Long T. Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang 0002, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi
EMNLP7
2025 Avoiding spurious sharpness minimization broadens applicability of SAM
abstract
Curvature regularization techniques like Sharpness Aware Minimization (SAM) have shown great promise in improving generalization on vision tasks. However, we find that SAM performs poorly in domains like natural language processing (NLP), often degrading performance — even with twice the compute budget. We investigate the discrepancy across domains and find that in the NLP setting, SAM is dominated by regularization of the logit statistics — instead of improving the geometry of the function itself. We use this observation to develop an alternative algorithm we call Functional SAM, which regularizes curvature only through modification of the statistics of the overall function implemented by the neural network, and avoids spurious minimization through logit manipulation. Furthermore, we argue that preconditioning the SAM perturbation also prevents spurious minimization, and when combined with Functional SAM, it gives further improvements. Our proposed algorithms show improved performance over AdamW and SAM baselines when trained for an equal number of steps, in both fixed-length and Chinchilla-style training settings, at various model scales (including billion-parameter scale). On the whole, our work highlights the importance of more precise characterizations of sharpness in broadening the applicability of curvature regularization to large language models (LLMs)
Sidak Pal Singh, Hossein Mobahi, Atish Agarwala, Yann N. Dauphin
ICML2
2024 On the Foundations of Shortcut Learning
abstract
Deep-learning models can extract a rich assortment of features from data. Which features a model uses depends not only on *predictivity*---how reliably a feature indicates training-set labels---but also on *availability*---how easily the feature can be extracted from inputs. The literature on shortcut learning has noted examples in which models privilege one feature over another, for example texture over shape and image backgrounds over foreground objects. Here, we test hypotheses about which input properties are more available to a model, and systematically study how predictivity and availability interact to shape models' feature use. We construct a minimal, explicit generative framework for synthesizing classification datasets with two latent features that vary in predictivity and in factors we hypothesize to relate to availability, and we quantify a model's shortcut bias---its over-reliance on the shortcut (more available, less predictive) feature at the expense of the core (less available, more predictive) feature. We find that linear models are relatively unbiased, but introducing a single hidden layer with ReLU or Tanh units yields a bias. Our empirical findings are consistent with a theoretical account based on Neural Tangent Kernels. Finally, we study how models used in practice trade off predictivity and availability in naturalistic datasets, discovering availability manipulations which increase models' degree of shortcut bias. Taken together, these findings suggest that the propensity to learn shortcut features is a fundamental characteristic of deep nonlinear architectures warranting systematic study given its role in shaping how models solve tasks.
Katherine L. Hermann, Hossein Mobahi, Thomas Fel, Michael C. Mozer
ICLR2
2024 Neglected Hessian component explains mysteries in sharpness regularization
abstract
Recent work has shown that methods that regularize second order information like SAM can improve generalization in deep learning. Seemingly similar methods like weight noise and gradient penalties often fail to provide such benefits. We investigate this inconsistency and reveal its connection to the the structure of the Hessian of the loss. Specifically, its decomposition into the positive semi-definite Gauss-Newton matrix and an indefinite matrix, which we call the Nonlinear Modeling Error (NME) matrix. Previous studies have largely overlooked the significance of the NME in their analysis for various reasons. However, we provide empirical and theoretical evidence that the NME is important to the performance of gradient penalties and explains their sensitivity to activation functions. We also provide evidence that the difference in regularization performance between gradient penalties and weight noise can be explained by the NME. Our findings emphasize the necessity of considering the NME in both experimental design and theoretical analysis for sharpness regularization.
Yann N. Dauphin, Atish Agarwala, Hossein Mobahi
NeurIPS3
2023 Sharpness-Aware Minimization Leads to Low-Rank Features
abstract
Sharpness-aware minimization (SAM) is a recently proposed method that minimizes the sharpness of the training loss of a neural network. While its generalization improvement is well-known and is the primary motivation, we uncover an additional intriguing effect of SAM: reduction of the feature rank which happens at different layers of a neural network. We show that this low-rank effect occurs very broadly: for different architectures such as fully-connected networks, convolutional networks, vision transformers and for different objectives such as regression, classification, language-image contrastive training. To better understand this phenomenon, we provide a mechanistic understanding of how low-rank features arise in a simple two-layer network. We observe that a significant number of activations gets entirely pruned by SAM which directly contributes to the rank reduction. We confirm this effect theoretically and check that it can also occur in deep networks, although the overall rank reduction mechanism can be more complex, especially for deep networks with pre-activation skip connections and self-attention layers.
Maksym Andriushchenko, Dara Bahri, Hossein Mobahi, Nicolas Flammarion
NeurIPS3
2023 On student-teacher deviations in distillation: does it pay to disobey?
abstract
Knowledge distillation (KD) has been widely used to improve the test accuracy of a "student" network, by training it to mimic the soft probabilities of a trained "teacher" network. Yet, it has been shown in recent work that, despite being trained to fit the teacher's probabilities, the student may not only significantly deviate from the teacher probabilities, but may also outdo than the teacher in performance. Our work aims to reconcile this seemingly paradoxical observation. Specifically, we characterize the precise nature of the student-teacher deviations, and argue how they _can_ co-occur with better generalization. First, through experiments on image and language data, we identify that these probability deviations correspond to the student systematically _exaggerating_ the confidence levels of the teacher. Next, we theoretically and empirically establish another form of exaggeration in some simple settings: KD exaggerates the implicit bias of gradient descent in converging faster along the top eigendirections of the data. Finally, we tie these two observations together: we demonstrate that the exaggerated bias of KD can simultaneously result in both (a) the exaggeration of confidence and (b) the improved generalization of the student, thus offering a resolution to the apparent paradox. Our analysis brings existing theory and practice closer by considering the role of gradient descent in KD and by demonstrating the exaggerated bias effect in both theoretical and empirical settings.
Vaishnavh Nagarajan, Aditya Krishna Menon, Srinadh Bhojanapalli, Hossein Mobahi, Sanjiv Kumar
NeurIPS4
2022 Sharpness-Aware Minimization Improves Language Model Generalization
abstract
The allure of superhuman-level capabilities has led to considerable interest in language models like GPT-3 and T5, wherein the research has, by and large, revolved around new model architectures, training tasks, and loss objectives, along with substantial engineering efforts to scale up model capacity and dataset size.Comparatively little work has been done to improve the generalization of these models through better optimization.In this work, we show that Sharpness-Aware Minimization (SAM), a recently proposed optimization procedure that encourages convergence to flatter minima, can substantially improve the generalization of language models without much computational overhead.We show that SAM is able to boost performance on SuperGLUE, GLUE, Web Questions, Natural Questions, Trivia QA, and TyDiQA, with particularly large gains when training data for these tasks is limited.
Dara Bahri, Hossein Mobahi, Yi Tay
ACL (1)2
2021 Sharpness-aware Minimization for Efficiently Improving Generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam Neyshabur
ICLR3
2021 A unifying view on implicit bias in training linear neural networks
Chulhee Yun, Shankar Krishnan, Hossein Mobahi
ICLR3
2020 Fantastic Generalization Measures and Where to Find Them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, Samy Bengio
ICLR3
2020 Self-Distillation Amplifies Regularization in Hilbert Space
abstract
Knowledge distillation introduced in the deep learning context is a method to transfer knowledge from one architecture to another. In particular, when the architectures are identical, this is called self-distillation. The idea is to feed in predictions of the trained model as new target values for retraining (and iterate this loop possibly a few times). It has been empirically observed that the self-distilled model often achieves higher accuracy on held out data. Why this happens, however, has been a mystery: the self-distillation dynamics does not receive any new information about the task and solely evolves by looping over training. To the best of our knowledge, there is no rigorous understanding of why this happens. This work provides the first theoretical analysis of self-distillation. We focus on fitting a nonlinear function to training data, where the model space is Hilbert space and fitting is subject to L2 regularization in this function space. We show that self-distillation iterations modify regularization by progressively limiting the number of basis functions that can be used to represent the solution. This implies (as we also verify empirically) that while a few rounds of self-distillation may reduce over-fitting, further rounds may lead to under-fitting and thus worse performance.
Hossein Mobahi, Mehrdad Farajtabar, Peter L. Bartlett
NeurIPS1
2019 Predicting the Generalization Gap in Deep Networks with Margin Distributions
Yiding Jiang, Dilip Krishnan, Hossein Mobahi, Samy Bengio
ICLR (Poster)3
2018 Large Margin Deep Networks for Classification
abstract
We present a formulation of deep learning that aims at producing a large margin classifier. The notion of \emc{margin}, minimum distance to a decision boundary, has served as the foundation of several theoretically profound and empirically successful results for both classification and regression tasks. However, most large margin algorithms are applicable only to shallow models with a preset feature representation; and conventional margin methods for neural networks only enforce margin at the output layer. Such methods are therefore not well suited for deep networks. In this work, we propose a novel loss function to impose a margin on any chosen set of layers of a deep network (including input and hidden layers). Our formulation allows choosing any $l_p$ norm ($p \geq 1$) on the metric measuring the margin. We demonstrate that the decision boundary obtained by our loss has nice properties compared to standard classification loss functions. Specifically, we show improved empirical results on the MNIST, CIFAR-10 and ImageNet datasets on multiple tasks: generalization from small training sets, corrupted labels, and robustness against adversarial perturbations. The resulting loss is general and complementary to existing data augmentation (such as random/adversarial input transform) and regularization techniques such as weight decay, dropout, and batch norm. \footnote{Code for the large margin loss function is released at \url{https://github.com/google-research/google-research/tree/master/large_margin}}
Gamaleldin F. Elsayed, Dilip Krishnan, Hossein Mobahi, Kevin Regan 0001, Samy Bengio
NeurIPS3
2017 Homotopy Analysis for Tensor PCA
abstract
Developing efficient and guaranteed nonconvex algorithms has been an important challenge in modern machine learning. Algorithms with good empirical performance such as stochastic gradient descent often lack theoretical guarantees. In this paper, we analyze the class of homotopy or continuation methods for global optimization of nonconvex functions. These methods start from an objective function that is efficient to optimize (e.g. convex), and progressively modify it to obtain the required objective, and the solutions are passed along the homotopy path. For the challenging problem of tensor PCA, we prove global convergence of the homotopy method in the “high noise” regime. The signal-to-noise requirement for our algorithm is tight in the sense that it matches the recovery guarantee for the \em best degree-$4$ sum-of-squares algorithm. In addition, we prove a phase transition along the homotopy path for tensor PCA. This allows us to simplify the homotopy method to a local search algorithm, viz., tensor power iterations, with a specific initialization and a noise injection procedure, while retaining the theoretical guarantees.
Anima Anandkumar, Rong Ge 0001, Hossein Mobahi
COLT4
2015 A Theoretical Analysis of Optimization by Gaussian Continuation
abstract
Optimization via continuation method is a widely used approach for solving nonconvex minimization problems. While this method generally does not provide a global minimum, empirically it often achieves a superior local minimum compared to alternative approaches such as gradient descent. However, theoretical analysis of this method is largely unavailable. Here, we provide a theoretical analysis that provides a bound on the endpoint solution of the continuation method. The derived bound depends on a problem specific characteristic that we refer to as optimization complexity. We show that this characteristic can be analytically computed when the objective function is expressed in some suitable basis functions. Our analysis combines elements of scale-space theory, regularization and differential equations.
Hossein Mobahi, John W. Fisher III
AAAI1
2015 The aperture problem for refractive motion
abstract
When viewed through a small aperture, a moving image provides incomplete information about the local motion. Only the component of motion along the local image gradient is constrained. In an essential part of optical flow algorithms, information must be aggregated from nearby image locations in order to estimate all components of motion. This limitation of local evidence for estimating optical flow is called “the aperture problem”. We pose and solve a generalization of the aperture problem for moving refractive elements. We consider a common setup in air flow imaging or telescope observation: a camera is viewing a static background, and an unknown refractive elements undergoing unknown motion between them. Then we are addressing this fundamental question: what does the local image motion tell us about the motion of refractive elements? We show that the information gleaned through a local aperture for this case is very different than that for optical flow. In optical flow, the movement of 1D structure already constrains the motion in a certain direction. However, we cannot infer any information about the refractive motion from the movement of 1D structure in the observed sequence, and can only recover one component of the motion from 2D structure. Results on both simulated and real sequences are shown to illustrate our theory.
Tianfan Xue, Hossein Mobahi, Frédo Durand, William T. Freeman
CVPR2
2015 Learning with a Wasserstein Loss
abstract
Learning to predict multi-label outputs is challenging, but in many problems there is a natural metric on the outputs that can be used to improve predictions. In this paper we develop a loss function for multi-label learning, based on the Wasserstein distance. The Wasserstein distance provides a natural notion of dissimilarity for probability measures. Although optimizing with respect to the exact Wasserstein distance is costly, recent work has described a regularized approximation that is efficiently computed. We describe an efficient learning algorithm based on this regularization, as well as a novel extension of the Wasserstein distance from probability measures to unnormalized measures. We also describe a statistical learning bound for the loss. The Wasserstein loss can encourage smoothness of the predictions with respect to a chosen metric on the output space. We demonstrate this property on a real-data tag prediction problem, using the Yahoo Flickr Creative Commons dataset, outperforming a baseline that doesn't use the metric.
Charlie Frogner, Chiyuan Zhang, Hossein Mobahi, Mauricio Araya-Polo, Tomaso A. Poggio
NIPS3
2014 A Compositional Model for Low-Dimensional Image Set Representation
abstract
Learning a low-dimensional representation of images is useful for various applications in graphics and computer vision. Existing solutions either require manually specified landmarks for corresponding points in the images, or are restricted to specific objects or shape deformations. This paper alleviates these limitations by imposing a specific model for generating images, the nested composition of color, shape, and appearance. We show that each component can be approximated by a low-dimensional subspace when the others are factored out. Our formulation allows for efficient learning and experiments show encouraging results.
Hossein Mobahi, Ce Liu 0001, William T. Freeman
CVPR1
2012 Seeing through the blur
abstract
This paper addresses the problem of image alignment using direct intensity-based methods for affine and homography transformations. Direct methods often employ scale-space smoothing (Gaussian blur) of the images to avoid local minima. Although, it is known that the isotropic blur used is not optimal for some motion models, the correct blur kernels have not been rigorously derived for motion models beyond translations. In this work, we derive blur kernels that result from smoothing the alignment objective function for some common motion models such as affine and homography. We show the derived kernels remove poor local minima and reach lower energy solutions in practice.
Hossein Mobahi, C. Lawrence Zitnick, Yi Ma 0001
CVPR1
2012 Toward a Practical Face Recognition System: Robust Alignment and Illumination by Sparse Representation
abstract
Many classic and contemporary face recognition algorithms work well on public data sets, but degrade sharply when they are used in a real recognition system. This is mostly due to the difficulty of simultaneously handling variations in illumination, image misalignment, and occlusion in the test image. We consider a scenario where the training images are well controlled and test images are only loosely controlled. We propose a conceptually simple face recognition system that achieves a high degree of robustness and stability to illumination variation, image misalignment, and partial occlusion. The system uses tools from sparse representation to align a test face image to a set of frontal training images. The region of attraction of our alignment algorithm is computed empirically for public face data sets such as Multi-PIE. We demonstrate how to capture a set of training images with enough illumination variation that they span test images taken under uncontrolled illumination. In order to evaluate how our algorithms work under practical testing conditions, we have implemented a complete face recognition system, including a projector-based training acquisition system. Our system can efficiently and effectively recognize faces under a variety of realistic conditions, using only frontal images under the proposed illuminations as training.
Andrew Wagner, John Wright 0001, Arvind Ganesh, Zihan Zhou 0001, Hossein Mobahi, Yi Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2011 Segmentation of Natural Images by Texture and Boundary Compression
Hossein Mobahi, Shankar R. Rao, Allen Y. Yang, S. Shankar Sastry, Yi Ma 0001
Int. J. Comput. Vis.1
2009 Natural Image Segmentation with Adaptive Texture and Boundary Encoding
Shankar R. Rao, Hossein Mobahi, Allen Y. Yang, S. Shankar Sastry, Yi Ma 0001
ACCV (1)2
2009 Face recognition with contiguous occlusion using markov random fields
abstract
Partially occluded faces are common in many applications of face recognition. While algorithms based on sparse representation have demonstrated promising results, they achieve their best performance on occlusions that are not spatially correlated (i.e. random pixel corruption). We show that such sparsity-based algorithms can be significantly improved by harnessing prior knowledge about the pixel error distribution. We show how a Markov Random Field model for spatial continuity of the occlusion can be integrated into the computation of a sparse representation of the test image with respect to the training images. Our algorithm efficiently and reliably identifies the corrupted regions and excludes them from the sparse representation. Extensive experiments on both laboratory and real-world datasets show that our algorithm tolerates much larger fractions and varieties of occlusion than current state-of-the-art algorithms.
Zihan Zhou 0001, Andrew Wagner, Hossein Mobahi, John Wright 0001, Yi Ma 0001
ICCV3
2009 Deep learning from temporal coherence in video
abstract
This work proposes a learning method for deep architectures that takes advantage of sequential data, in particular from the temporal coherence that naturally exists in unlabeled video recordings. That is, two successive frames are likely to contain the same object or objects. This coherence is used as a supervisory signal over the unlabeled data, and is used to improve the performance on a supervised task of interest. We demonstrate the effectiveness of this method on some pose invariant object and face recognition tasks. 1.
Hossein Mobahi, Ronan Collobert, Jason Weston
ICML1
2009 Data-driven image completion by image patch subspaces
abstract
We develop a new method for image completion on images with large missing regions. We assume that similar patches form low dimensional clusters in the image space where each cluster can be approximated by a (degenerate) Gaussian. We use sparse representation for subspace detection and then compute the most probable completion. Our results show almost no blurring or blocking effects. In addition, both the texture and structure of the missing regions look realistic to the human eye.
Hossein Mobahi, Shankar R. Rao, Yi Ma 0001
PCS1
2005 Concept Oriented Imitation Towards Verbal Human-Robot Interaction
abstract
Imitation equips robots with a simple and natural interface to learn new tasks. Although abstraction is a remarkable feature of imitation that discriminates it from mimicking, there has been no enough research on this dimension of imitation. Relational concepts are the simplest type of abstract concepts and can be an appropriate start point. These concepts may be learned by combining perceptual categorization and classical conditioning. The paper will first formalize relational concept learning within an imitative context. Internal modules of the learning agent are considered to be functions. We will prove that in this case the concept-motor mapping becomes one-to-one which simplifies learning. A learning algorithm for the model will be also proposed and evaluated in a phoneme acquisition experiment with a large number of highly overlapped samples.
Hossein Mobahi, Majid Nili Ahmadabadi, Babak Nadjar Araabi
ICRA1
2004 Peak stick RBF network for online system identification
abstract
In many practical problems of online system identification, the distribution of observed samples is uneven. For instance, at points where system is idle or changes slowly, the sample density increases and where system moves quickly, it is reduced. This generally results in performance degradation of learning. We will propose a new algorithm for training RBF networks that is particularly developed for online learning with uneven sample distribution. The basic idea is to find peaks and stick to them. Experiments show a notable improvement in convergence rate, settling of weights and error minimization.
Hossein Mobahi, Farrokh Janabi-Sharifi
IJCNN1
2004 Fast initialization of active contours
abstract
The field of robotics is currently undergoing a change toward creation of robots that can naturally interact with humans. For achieving this, interactive robots must be endowed with natural interfaces that can sense and respond in real-time. Vision can provide handy information for this purpose by detecting and tracking human limbs to analyze gestures, actions and even emotions. However, real-time processing of visual information is a challenging bottleneck. In this paper, we introduce a novel method, namely "self-organized contours", that can distinctly accelerate contour initialization, which is the slowest phase in visual tracking. Although the proposed method is general-purpose, it allows immediate initialization of active contours due to its similarity with snake structure. The proposed method is inspired from group behavior in insects and animals, particularly fishes.
Hossein Mobahi, Majid Nili Ahmadabadi, Babak Nadjar Araabi
IROS1
2003 Fuzzy perception, emotion and expression for interactive robots
abstract
Future robots need transparent interface that regular people can interpret, such as an emotional human-like face. Moreover, such robots must exhibit behaviors that are perceived believable and life-like. In this work we propose the use of fuzzy logic for effectively constructing the whole behavior system of these robots. This not only simplifies the design tasks, but also enriches human-robot interaction. The latter claim is justified by effortlessly generating intermediate and blend of emotions from a few basic emotions. Additionally, fuzzy motor commands yield smooth life-like motions and therefore improve believability.
Hossein Mobahi, Shahin Ansari
SMC1