VLDB 2026 Research / reviewers in the wild / expert
Ivan Kobyzev
dblp:234/1853
· DBLP profile ↗
11ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-1934-4842ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMsabstractPost-training hybridization of large language models (LLMs) often replaces quadratic selfattention with sliding-window attention (SWA) to reduce KV cache usage and improve latency.Existing hybridization schemes are typically defined either at the layer level (e.g., interleaving) or at the head level via static rankings from local to global.Layer-level schemes ignore that local and global dependencies are routed through heads within the same layer, while static head-level rankings suffer from entanglement: a head's local/global behavior can change after hybridization.We propose BOSCH, Black-box Binary Optimization for Short-context Head Selection, a training-free method that formulates the problem as a Large Neighborhood Search and decomposes it into three subproblems: (i) layer-importance detection via small-budget black-box probes, (ii) adaptive per-layer SWA-ratio assignment based on these sensitivities, and (iii) grouped headlevel optimization within ratio buckets.Extensive experiments on 4 LLMs ranging from 1.7B to 30B parameters, across 4 SWA ratios, show that BOSCH consistently outperforms layerlevel heuristics and 6 strong static head-level methods, with larger gains at higher SWA ratios.Under continual pretraining, BOSCH recover original long-context performance faster and to a higher level.Analysis of the selected heads reveals substantial turnover for BOSCH across different SWA ratios, underscoring the importance of performing head-level selection for each target ratio rather than relying on fixed locality rankings. Abbas Ghaddar, Ivan Kobyzev, Boxing Chen, Yufei Cui |
ACL (1) | 2 |
| 2025 | Integral Transformer: Denoising Attention, Not Too Much Not Too LittleabstractSoftmax self-attention often assigns disproportionate weight to semantically uninformative tokens such as special tokens and punctuation, a phenomenon known as attention noise.While recent methods like Cog Attention and the Differential Transformer have addressed this by introducing negative attention scores, they risk discarding useful information.In this paper, we propose the Integral Transformer, a novel selfattention mechanism that denoises attention by integrating signals sampled from the logit distribution.Our approach mitigates noise while preserving the contributions of special tokens critical for model performance.Extensive experiments demonstrate that our model outperforms vanilla, Cog, and Differential attention variants on well-established knowledge and reasoning language benchmarks.Moreover, our analysis reveals that employing vanilla self-attention in the lower Transformer layers enhances performance and that the Integral Transformer effectively balances attention distributions and reduces rank collapse in upper layers. Ivan Kobyzev, Abbas Ghaddar, Dingtao Hu, Boxing Chen |
EMNLP | 1 |
| 2025 | ReGLA: Refining Gated Linear AttentionabstractPeng Lu, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Peng Lu 0006, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais |
NAACL (Long Papers) | 2 |
| 2023 | Do we need Label Regularization to Fine-tune Pre-trained Language Models?abstractIvan Kobyzev, Aref Jafari, Mehdi Rezagholizadeh, Tianda Li, Alan Do-Omri, Peng Lu, Pascal Poupart, Ali Ghodsi. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Ivan Kobyzev, Aref Jafari, Mehdi Rezagholizadeh, Tianda Li, Alan Do-Omri, Peng Lu 0006, Pascal Poupart, Ali Ghodsi 0001 |
EACL | 1 |
| 2023 | DyLoRA: Parameter-Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank AdaptationabstractWith the ever-growing size of pretrained models (PMs), fine-tuning them has become more expensive and resource-hungry.As a remedy, low-rank adapters (LoRA) keep the main pretrained weights of the model frozen and just introduce some learnable truncated SVD modules (so-called LoRA blocks) to the model.While LoRA blocks are parameter-efficient, they suffer from two major problems: first, the size of these blocks is fixed and cannot be modified after training (for example, if we need to change the rank of LoRA blocks, then we need to retrain them from scratch); second, optimizing their rank requires an exhaustive search and effort.In this work, we introduce a dynamic low-rank adaptation (DyLoRA) technique to address these two problems together.Our Dy-LoRA method trains LoRA blocks for a range of ranks instead of a single rank by sorting the representation learned by the adapter module at different ranks during training.We evaluate our solution on different natural language understanding (GLUE benchmark) and language generation tasks (E2E, DART and WebNLG) using different pretrained models such as RoBERTa and GPT with different sizes.Our results show that we can train dynamic search-free models with DyLoRA at least 4 to 7 times faster than LoRA without significantly compromising performance.Moreover, our models can perform consistently well on a much larger range of ranks compared to LoRA. 1 Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, Ali Ghodsi 0001 |
EACL | 3 |
| 2023 | Efficient Classification of Long Documents via State-Space ModelsabstractTransformer-based models have achieved stateof-the-art performance on numerous NLP applications.However, long documents which are prevalent in real-world scenarios cannot be efficiently processed by transformers with the vanilla self-attention module due to their quadratic computation complexity and limited length extrapolation ability.Instead of tackling the computation difficulty for self-attention with sparse or hierarchical structures, in this paper, we investigate the use of State-Space Models (SSMs) for long document classification tasks.We conducted extensive experiments on six long document classification datasets, including binary, multi-class, and multi-label classification, comparing SSMs (with and without pre-training) to self-attention-based models.We also introduce the SSM-pooler model and demonstrate that it achieves comparable performance while being on average 36% more efficient.Additionally our method exhibits higher robustness to the input noise even in the extreme scenario of 40%. * Research done during internship in Huawei Noah's Ark Lab (Montreal). Peng Lu 0006, Suyuchen Wang, Mehdi Rezagholizadeh, Bang Liu 0003, Ivan Kobyzev |
EMNLP | 5 |
| 2022 | Learning functions on multiple sets using multi-set transformersabstractWe propose a general deep architecture for learning functions on multiple permutation-invariant sets. We also show how to generalize this architecture to sets of elements of any dimension by dimension equivariance. We demonstrate that our architecture is a universal approximator of these functions, and show superior results to existing methods on a variety of tasks including counting tasks, alignment tasks, distinguishability tasks and statistical distance measurements. This last task is quite important in Machine Learning. Although our approach is quite general, we demonstrate that it can generate approximate estimates of KL divergence and mutual information that are more accurate than previous techniques that are specifically designed to approximate those statistical distances. Kira A. Selby, Ahmad Rashid, Ivan Kobyzev, Mehdi Rezagholizadeh, Pascal Poupart |
UAI | 3 |
| 2021 | Polarized-VAE: Proximity Based Disentangled Representation Learning for Text GenerationabstractVikash Balasubramanian, Ivan Kobyzev, Hareesh Bahuleyan, Ilya Shapiro, Olga Vechtomova. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Vikash Balasubramanian, Ivan Kobyzev, Hareesh Bahuleyan, Ilya Shapiro, Olga Vechtomova |
EACL | 2 |
| 2021 | Normalizing Flows: An Introduction and Review of Current MethodsabstractNormalizing Flows are generative models which produce tractable distributions where both sampling and density evaluation can be efficient and exact. The goal of this survey article is to give a coherent and comprehensive review of the literature around the construction and use of Normalizing Flows for distribution learning. We aim to provide context and explanation of the models, review current state-of-the-art literature, and identify open questions and promising future directions. Ivan Kobyzev, Simon Prince, Marcus A. Brubaker |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Tails of Lipschitz Triangular FlowsabstractWe investigate the ability of popular flow models to capture tail-properties of a target density by studying the increasing triangular maps used in these flow methods acting on a tractable source density. We show that the density quantile functions of the source and target density provide a precise characterization of the slope of transformation required to capture tails in a target density. We further show that any Lipschitz-continuous transport map acting on a source density will result in a density with similar tail properties as the source, highlighting the trade-off between the importance of choosing a complex source density and a sufficiently expressive transformation to capture desirable properties of a target density. Subsequently, we illustrate that flow models like Real-NVP, MAF, and Glow as implemented lack the ability to capture a distribution with non-Gaussian tails. We circumvent this problem by proposing tail-adaptive flows consisting of a source distribution that can be learned simultaneously with the triangular map to capture tail-properties of a target density. We perform several synthetic and real-world experiments to complement our theoretical findings. Priyank Jaini, Ivan Kobyzev, Yaoliang Yu, Marcus A. Brubaker |
ICML | 2 |
| 2020 | Representation Learning for Dynamic Graphs: A SurveyabstractGraphs arise naturally in many real-world applications including social networks, recommender systems, ontologies, biology, and computational finance. Traditionally, machine learning models for graphs have been mostly designed for static graphs. However, many applications involve evolving graphs. This introduces important challenges for learning and inference since nodes, attributes, and edges change over time. In this survey, we review the recent advances in representation learning for dynamic graphs, including dynamic knowledge graphs. We describe existing models from an encoder-decoder perspective, categorize these encoders and decoders based on the techniques they employ, and analyze the approaches in each category. We also review several prominent applications and widely used datasets and highlight directions for future research. Mehran Kazemi, Rishab Goel, Kshitij Jain 0001, Ivan Kobyzev, Akshay Sethi, Peter Forsyth, Pascal Poupart |
J. Mach. Learn. Res. | 4 |