Ivan Kobyzev

dblp:234/1853 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-1934-4842ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021
YearPublicationVenuePosition
2026 BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMs
abstract
Post-training hybridization of large language models (LLMs) often replaces quadratic selfattention with sliding-window attention (SWA) to reduce KV cache usage and improve latency.Existing hybridization schemes are typically defined either at the layer level (e.g., interleaving) or at the head level via static rankings from local to global.Layer-level schemes ignore that local and global dependencies are routed through heads within the same layer, while static head-level rankings suffer from entanglement: a head's local/global behavior can change after hybridization.We propose BOSCH, Black-box Binary Optimization for Short-context Head Selection, a training-free method that formulates the problem as a Large Neighborhood Search and decomposes it into three subproblems: (i) layer-importance detection via small-budget black-box probes, (ii) adaptive per-layer SWA-ratio assignment based on these sensitivities, and (iii) grouped headlevel optimization within ratio buckets.Extensive experiments on 4 LLMs ranging from 1.7B to 30B parameters, across 4 SWA ratios, show that BOSCH consistently outperforms layerlevel heuristics and 6 strong static head-level methods, with larger gains at higher SWA ratios.Under continual pretraining, BOSCH recover original long-context performance faster and to a higher level.Analysis of the selected heads reveals substantial turnover for BOSCH across different SWA ratios, underscoring the importance of performing head-level selection for each target ratio rather than relying on fixed locality rankings.
Abbas Ghaddar, Ivan Kobyzev, Boxing Chen, Yufei Cui
ACL (1)2
2025 Integral Transformer: Denoising Attention, Not Too Much Not Too Little
abstract
Softmax self-attention often assigns disproportionate weight to semantically uninformative tokens such as special tokens and punctuation, a phenomenon known as attention noise.While recent methods like Cog Attention and the Differential Transformer have addressed this by introducing negative attention scores, they risk discarding useful information.In this paper, we propose the Integral Transformer, a novel selfattention mechanism that denoises attention by integrating signals sampled from the logit distribution.Our approach mitigates noise while preserving the contributions of special tokens critical for model performance.Extensive experiments demonstrate that our model outperforms vanilla, Cog, and Differential attention variants on well-established knowledge and reasoning language benchmarks.Moreover, our analysis reveals that employing vanilla self-attention in the lower Transformer layers enhances performance and that the Integral Transformer effectively balances attention distributions and reduces rank collapse in upper layers.
Ivan Kobyzev, Abbas Ghaddar, Dingtao Hu, Boxing Chen
EMNLP1
2025 ReGLA: Refining Gated Linear Attention
abstract
Peng Lu, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Peng Lu 0006, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais
NAACL (Long Papers)2
2023 Do we need Label Regularization to Fine-tune Pre-trained Language Models?
abstract
Ivan Kobyzev, Aref Jafari, Mehdi Rezagholizadeh, Tianda Li, Alan Do-Omri, Peng Lu, Pascal Poupart, Ali Ghodsi. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Ivan Kobyzev, Aref Jafari, Mehdi Rezagholizadeh, Tianda Li, Alan Do-Omri, Peng Lu 0006, Pascal Poupart, Ali Ghodsi 0001
EACL1
2023 DyLoRA: Parameter-Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank Adaptation
abstract
With the ever-growing size of pretrained models (PMs), fine-tuning them has become more expensive and resource-hungry.As a remedy, low-rank adapters (LoRA) keep the main pretrained weights of the model frozen and just introduce some learnable truncated SVD modules (so-called LoRA blocks) to the model.While LoRA blocks are parameter-efficient, they suffer from two major problems: first, the size of these blocks is fixed and cannot be modified after training (for example, if we need to change the rank of LoRA blocks, then we need to retrain them from scratch); second, optimizing their rank requires an exhaustive search and effort.In this work, we introduce a dynamic low-rank adaptation (DyLoRA) technique to address these two problems together.Our Dy-LoRA method trains LoRA blocks for a range of ranks instead of a single rank by sorting the representation learned by the adapter module at different ranks during training.We evaluate our solution on different natural language understanding (GLUE benchmark) and language generation tasks (E2E, DART and WebNLG) using different pretrained models such as RoBERTa and GPT with different sizes.Our results show that we can train dynamic search-free models with DyLoRA at least 4 to 7 times faster than LoRA without significantly compromising performance.Moreover, our models can perform consistently well on a much larger range of ranks compared to LoRA. 1
Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, Ali Ghodsi 0001
EACL3
2023 Efficient Classification of Long Documents via State-Space Models
abstract
Transformer-based models have achieved stateof-the-art performance on numerous NLP applications.However, long documents which are prevalent in real-world scenarios cannot be efficiently processed by transformers with the vanilla self-attention module due to their quadratic computation complexity and limited length extrapolation ability.Instead of tackling the computation difficulty for self-attention with sparse or hierarchical structures, in this paper, we investigate the use of State-Space Models (SSMs) for long document classification tasks.We conducted extensive experiments on six long document classification datasets, including binary, multi-class, and multi-label classification, comparing SSMs (with and without pre-training) to self-attention-based models.We also introduce the SSM-pooler model and demonstrate that it achieves comparable performance while being on average 36% more efficient.Additionally our method exhibits higher robustness to the input noise even in the extreme scenario of 40%. * Research done during internship in Huawei Noah's Ark Lab (Montreal).
Peng Lu 0006, Suyuchen Wang, Mehdi Rezagholizadeh, Bang Liu 0003, Ivan Kobyzev
EMNLP5
2022 Learning functions on multiple sets using multi-set transformers
abstract
We propose a general deep architecture for learning functions on multiple permutation-invariant sets. We also show how to generalize this architecture to sets of elements of any dimension by dimension equivariance. We demonstrate that our architecture is a universal approximator of these functions, and show superior results to existing methods on a variety of tasks including counting tasks, alignment tasks, distinguishability tasks and statistical distance measurements. This last task is quite important in Machine Learning. Although our approach is quite general, we demonstrate that it can generate approximate estimates of KL divergence and mutual information that are more accurate than previous techniques that are specifically designed to approximate those statistical distances.
Kira A. Selby, Ahmad Rashid, Ivan Kobyzev, Mehdi Rezagholizadeh, Pascal Poupart
UAI3
2021 Polarized-VAE: Proximity Based Disentangled Representation Learning for Text Generation
abstract
Vikash Balasubramanian, Ivan Kobyzev, Hareesh Bahuleyan, Ilya Shapiro, Olga Vechtomova. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Vikash Balasubramanian, Ivan Kobyzev, Hareesh Bahuleyan, Ilya Shapiro, Olga Vechtomova
EACL2
2021 Normalizing Flows: An Introduction and Review of Current Methods
abstract
Normalizing Flows are generative models which produce tractable distributions where both sampling and density evaluation can be efficient and exact. The goal of this survey article is to give a coherent and comprehensive review of the literature around the construction and use of Normalizing Flows for distribution learning. We aim to provide context and explanation of the models, review current state-of-the-art literature, and identify open questions and promising future directions.
Ivan Kobyzev, Simon Prince, Marcus A. Brubaker
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Tails of Lipschitz Triangular Flows
abstract
We investigate the ability of popular flow models to capture tail-properties of a target density by studying the increasing triangular maps used in these flow methods acting on a tractable source density. We show that the density quantile functions of the source and target density provide a precise characterization of the slope of transformation required to capture tails in a target density. We further show that any Lipschitz-continuous transport map acting on a source density will result in a density with similar tail properties as the source, highlighting the trade-off between the importance of choosing a complex source density and a sufficiently expressive transformation to capture desirable properties of a target density. Subsequently, we illustrate that flow models like Real-NVP, MAF, and Glow as implemented lack the ability to capture a distribution with non-Gaussian tails. We circumvent this problem by proposing tail-adaptive flows consisting of a source distribution that can be learned simultaneously with the triangular map to capture tail-properties of a target density. We perform several synthetic and real-world experiments to complement our theoretical findings.
Priyank Jaini, Ivan Kobyzev, Yaoliang Yu, Marcus A. Brubaker
ICML2
2020 Representation Learning for Dynamic Graphs: A Survey
abstract
Graphs arise naturally in many real-world applications including social networks, recommender systems, ontologies, biology, and computational finance. Traditionally, machine learning models for graphs have been mostly designed for static graphs. However, many applications involve evolving graphs. This introduces important challenges for learning and inference since nodes, attributes, and edges change over time. In this survey, we review the recent advances in representation learning for dynamic graphs, including dynamic knowledge graphs. We describe existing models from an encoder-decoder perspective, categorize these encoders and decoders based on the techniques they employ, and analyze the approaches in each category. We also review several prominent applications and widely used datasets and highlight directions for future research.
Mehran Kazemi, Rishab Goel, Kshitij Jain 0001, Ivan Kobyzev, Akshay Sethi, Peter Forsyth, Pascal Poupart
J. Mach. Learn. Res.4