Hrayr Harutyunyan

dblp:198/1465 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Efficient and distributed learning · 38% Learning theory · 18% Trustworthy machine learning · 9%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
parameter sharing
1.722025
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation · NeurIPS 2025
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA · ICLR 2025
Machine learning › Learning theory
generalization bounds
0.922021
Information-theoretic generalization bounds for black-box learning algorithms · NeurIPS 2021
Improving generalization by controlling label-noise information in neural network weights · ICML 2020
Machine learning › Efficient and distributed learning
adaptive computation
0.912025
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation · NeurIPS 2025
Natural language and speech › Language models and text generation
efficient language model
0.912025
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation · NeurIPS 2025
Machine learning › Efficient and distributed learning
inference acceleration
0.912025
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA · ICLR 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA · ICLR 2025
Machine learning › Deep learning architectures and training › transformer
recursive transformer
0.912025
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation · NeurIPS 2025
Machine learning › Efficient and distributed learning
data requirement estimation
0.712023
A Meta-Learning Approach to Predicting Performance and Data Requirements · CVPR 2023
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.712023
Supervision Complexity and its Role in Knowledge Distillation · ICLR 2023
Machine learning › Transfer learning and domain adaptation
meta-learning
0.712023
A Meta-Learning Approach to Predicting Performance and Data Requirements · CVPR 2023
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
performance prediction
0.712023
A Meta-Learning Approach to Predicting Performance and Data Requirements · CVPR 2023
Machine learning › Transfer learning and domain adaptation
domain generalization
0.612022
Failure Modes of Domain Generalization Algorithms · CVPR 2022
Machine learning › Trustworthy machine learning › Data-centric AI
data valuation
0.512021
Estimating informativeness of samples with Smooth Unique Information · ICLR 2021
Machine learning › Learning theory › generalization bounds
information-theoretic generalization bounds
0.512021
Information-theoretic generalization bounds for black-box learning algorithms · NeurIPS 2021
Machine learning › Learning theory
information-theoretic learning
0.512021
Estimating informativeness of samples with Smooth Unique Information · ICLR 2021
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.412020
Improving generalization by controlling label-noise information in neural network weights · ICML 2020
Machine learning › Learning theory › neural network theory
memorization and generalization
0.412020
Improving generalization by controlling label-noise information in neural network weights · ICML 2020
Machine learning › Representation and self-supervised learning
mutual information
0.412020
Improving generalization by controlling label-noise information in neural network weights · ICML 2020
Machine learning › Graph learning › graph neural network
graph convolutional network
0.412019
MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.412019
Fast structure learning with modular regularization · NeurIPS 2019
Machine learning › Graph learning
graph neural network
0.412019
MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.412019
Fast structure learning with modular regularization · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning
0.412019
Fast structure learning with modular regularization · NeurIPS 2019
Machine learning › Deep learning architectures and training
transformer
0.312025
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation · NeurIPS 2025
Machine learning › Learning paradigms › semi-supervised learning
graph-based semi-supervised learning
0.112019
MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing · ICML 2019
Machine learning › Learning paradigms
semi-supervised learning
0.112019
MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing · ICML 2019
Data mining › statistical analysis
covariance estimation
0.112019
Fast structure learning with modular regularization · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 1.5mutual information estimation · 1.0token-level routing · 0.9mixture of recursions · 0.9low-rank adaptation · 0.9early exiting · 0.9KV sharing · 0.9random forest · 0.7piecewise power law · 0.7meta-learning · 0.7structured latent factor models · 0.4information-theoretic measures · 0.4GLASSO · 0.4
YearPublicationVenuePosition
2025 Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
abstract
Large language models (LLMs) are expensive to deploy. Parameter sharing offers a possible path towards reducing their size and cost, but its effectiveness in modern LLMs remains fairly limited. In this work, we revisit "layer tying" as form of parameter sharing in Transformers, and introduce novel methods for converting existing LLMs into smaller "Recursive Transformers" that share parameters across layers, with minimal loss of performance. Here, our Recursive Transformers are efficiently initialized from standard pretrained Transformers, but only use a single block of unique layers that is then repeated multiple times in a loop. We further improve performance by introducing Relaxed Recursive Transformers that add flexibility to the layer tying constraint via depth-wise low-rank adaptation (LoRA) modules, yet still preserve the compactness of the overall model. We show that our recursive models (e.g., recursive Gemma 1B) outperform both similar-sized vanilla pretrained models (such as TinyLlama 1.1B and Pythia 1B) and knowledge distillation baselines---and can even recover most of the performance of the original "full-size" model (e.g., Gemma 2B with no shared parameters). Finally, we propose Continuous Depth-wise Batching, a promising new inference paradigm enabled by the Recursive Transformer when paired with early exiting. In a theoretical analysis, we show that this has the potential to lead to significant (2-3$\times$) gains in inference throughput.
Sangmin Bae, Adam Fisch, Hrayr Harutyunyan, Seungyeon Kim 0001, Tal Schuster
ICLR3
2025 Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
abstract
Scaling language models unlocks impressive capabilities, but the accompanying computational and memory demands make both training and deployment expensive. Existing efficiency efforts typically target either parameter sharing or adaptive computation, leaving open the question of how to attain both simultaneously. We introduce Mixture-of-Recursions (MoR), a unified framework that combines the two axes of efficiency inside a single Recursive Transformer. MoR reuses a shared stack of layers across recursion steps to achieve parameter efficiency, while lightweight routers enable adaptive token-level thinking by dynamically assigning different recursion depths to individual tokens. This allows MoR to focus quadratic attention computation only among tokens still active at a given recursion depth, further improving memory access efficiency by selectively caching only their key-value pairs. Beyond these core mechanisms, we also propose a KV sharing variant that reuses KV pairs from the first recursion, specifically designed to further decrease memory footprint. Across model scales ranging from 135M to 1.7B parameters, MoR forms a new Pareto frontier: at equal training FLOPs and smaller model sizes, it significantly lowers validation perplexity and improves few-shot accuracy, while delivering higher throughput compared with vanilla and existing recursive baselines.
Sangmin Bae, Reza Bayat, Sungnyun Kim, Jiyoun Ha, Tal Schuster, Adam Fisch, Hrayr Harutyunyan, Aaron C. Courville, Se-Young Yun
NeurIPS8
2023 A Meta-Learning Approach to Predicting Performance and Data Requirements
abstract
We propose an approach to estimate the number of samples required for a model to reach a target performance. We find that the power law, the de facto principle to estimate model performance, leads to a large error when using a small dataset (e.g., 5 samples per class) for extrapolation. This is because the log-performance error against the log-dataset size follows a nonlinear progression in the few-shot regime followed by a linear progression in the high-shot regime. We introduce a novel piecewise power law (PPL) that handles the two data regimes differently. To estimate the parameters of the PPL, we introduce a random forest regressor trained via meta learning that generalizes across classification/detection tasks, ResNet/ViT based architectures, and random/pre-trained initializations. The PPL improves the performance estimation on average by 37% across 16 classification and 33% across 10 detection datasets, compared to the power law. We further extend the PPL to provide a confidence bound and use it to limit the prediction horizon that reduces over-estimation of data by 76% on classification and 91% on detection datasets.
Achin Jain, Gurumurthy Swaminathan, Paolo Favaro, Hao Yang 0043, Avinash Ravichandran, Hrayr Harutyunyan, Alessandro Achille, Onkar Dabeer, Bernt Schiele, Ashwin Swaminathan, Stefano Soatto
CVPR6
2023 Supervision Complexity and its Role in Knowledge Distillation
Hrayr Harutyunyan, Ankit Singh Rawat, Aditya Krishna Menon, Seungyeon Kim 0001, Sanjiv Kumar
ICLR1
2022 Failure Modes of Domain Generalization Algorithms
abstract
Domain generalization algorithms use training data from multiple domains to learn models that generalize well to unseen domains. While recently proposed benchmarks demon-strate that most of the existing algorithms do not outperform simple baselines, the established evaluation methods fail to expose the impact of various factors that contribute to the poor performance. In this paper we propose an evaluation framework for domain generalization algorithms that allows decomposition of the error into components capturing distinct aspects of generalization. Inspired by the prevalence of algorithms based on the idea of domain-invariant representation learning, we extend the evaluation framework to capture various types of failures in achieving invariance. We show that the largest contributor to the generalization error varies across methods, datasets, regularization strengths and even training lengths. We observe two problems associated with the strategy of learning domain-invariant representations. On Colored MNIST, most domain generalization algorithms fail because they reach domain-invariance only on the training domains. On Camelyon-17, domain-invariance degrades the quality of representations on unseen domains. We hypothesize that focusing instead on tuning the classifier on top of a rich representation can be a promising direction.
Tigran Galstyan, Hrayr Harutyunyan, Hrant Khachatrian, Greg Ver Steeg, Aram Galstyan
CVPR2
2022 Formal limitations of sample-wise information-theoretic generalization bounds
abstract
Some of the tightest information-theoretic generalization bounds depend on the average information between the learned hypothesis and a single training example. However, these sample-wise bounds were derived only for expected generalization gap. We show that even for expected squared generalization gap no such sample-wise information-theoretic bounds exist. The same is true for PAC-Bayes and single-draw bounds. Remarkably, PAC-Bayes, single-draw and expected squared generalization gap bounds that depend on information in pairs of examples exist.
Hrayr Harutyunyan, Greg Ver Steeg, Aram Galstyan
ITW1
2021 Estimating informativeness of samples with Smooth Unique Information
Hrayr Harutyunyan, Alessandro Achille, Giovanni Paolini, Orchid Majumder, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto
ICLR1
2021 Information-theoretic generalization bounds for black-box learning algorithms
abstract
We derive information-theoretic generalization bounds for supervised learning algorithms based on the information contained in predictions rather than in the output of the training algorithm. These bounds improve over the existing information-theoretic bounds, are applicable to a wider range of algorithms, and solve two key challenges: (a) they give meaningful results for deterministic algorithms and (b) they are significantly easier to estimate. We show experimentally that the proposed bounds closely follow the generalization gap in practical scenarios for deep learning.
Hrayr Harutyunyan, Maxim Raginsky, Greg Ver Steeg, Aram Galstyan
NeurIPS1
2020 Improving generalization by controlling label-noise information in neural network weights
abstract
In the presence of noisy or incorrect labels, neural networks have the undesirable tendency to memorize information about the noise. Standard regularization techniques such as dropout, weight decay or data augmentation sometimes help, but do not prevent this behavior. If one considers neural network weights as random variables that depend on the data and stochasticity of training, the amount of memorized information can be quantified with the Shannon mutual information between weights and the vector of all training labels given inputs, $I(w; \mathbf{y} \mid \mathbf{x})$. We show that for any training algorithm, low values of this term correspond to reduction in memorization of label-noise and better generalization bounds. To obtain these low values, we propose training algorithms that employ an auxiliary network that predicts gradients in the final layers of a classifier without accessing labels. We illustrate the effectiveness of our approach on versions of MNIST, CIFAR-10, and CIFAR-100 corrupted with various noise models, and on a large-scale dataset Clothing1M that has noisy labels.
Hrayr Harutyunyan, Kyle Reing, Greg Ver Steeg, Aram Galstyan
ICML1
2019 MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing
abstract
Existing popular methods for semi-supervised learning with Graph Neural Networks (such as the Graph Convolutional Network) provably cannot learn a general class of neighborhood mixing relationships. To address this weakness, we propose a new model, MixHop, that can learn these relationships, including difference operators, by repeatedly mixing feature representations of neighbors at various distances. MixHop requires no additional memory or computational complexity, and outperforms on challenging baselines. In addition, we propose sparsity regularization that allows us to visualize how the network prioritizes neighborhood information across different graph datasets. Our analysis of the learned architectures reveals that neighborhood mixing varies per datasets.
Sami Abu-El-Haija, Bryan Perozzi, Amol Kapoor, Nazanin Alipourfard, Kristina Lerman, Hrayr Harutyunyan, Greg Ver Steeg, Aram Galstyan
ICML6
2019 Fast structure learning with modular regularization
abstract
Estimating graphical model structure from high-dimensional and undersampled data is a fundamental problem in many scientific fields. Existing approaches, such as GLASSO, latent variable GLASSO, and latent tree models, suffer from high computational complexity and may impose unrealistic sparsity priors in some cases. We introduce a novel method that leverages a newly discovered connection between information-theoretic measures and structured latent factor models to derive an optimization objective which encourages modular structures where each observed variable has a single latent parent. The proposed method has linear stepwise computational complexity w.r.t. the number of observed variables. Our experiments on synthetic data demonstrate that our approach is the only method that recovers modular structure better as the dimensionality increases. We also use our approach for estimating covariance structure for a number of real-world datasets and show that it consistently outperforms state-of-the-art estimators at a fraction of the computational cost. Finally, we apply the proposed method to high-resolution fMRI data (with more than 10^5 voxels) and show that it is capable of extracting meaningful patterns.
Greg Ver Steeg, Hrayr Harutyunyan, Daniel Moyer, Aram Galstyan
NeurIPS2