EDBT 2026 Demo / reviewers in the wild / expert
Boris Knyazev 0001
dblp:181/5675-1
· DBLP profile ↗
13ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0002-9484-1534ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 9 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Graph learning · 23% Efficient and distributed learning · 16% Deep learning architectures and training · 10% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph neural network |
2.5 | 4 | 2025 | Accelerating Training with Neuron Interaction and Nowcasting Networks · ICLR 2025 Graph Neural Networks for Learning Equivariant Representations of Neural Networks · ICLR 2024 Parameter Prediction for Unseen Deep Architectures · NeurIPS 2021 |
Machine learning › Efficient and distributed learning
parameter prediction |
1.2 | 2 | 2023 | Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models? · ICML 2023 Parameter Prediction for Unseen Deep Architectures · NeurIPS 2021 |
Computer vision › Segmentation and scene understanding
scene graph generation |
1.0 | 2 | 2021 | Context-aware Scene Graph Generation with Seq2Seq Transformers · ICCV 2021 Generative Compositional Augmentations for Scene Graph Prediction · ICCV 2021 |
Machine learning › Reinforcement learning › meta-reinforcement learning
learned update rules |
0.9 | 1 | 2025 | Accelerating Training with Neuron Interaction and Nowcasting Networks · ICLR 2025 |
Machine learning › Optimization for machine learning
neural network training acceleration |
0.9 | 1 | 2025 | Accelerating Training with Neuron Interaction and Nowcasting Networks · ICLR 2025 |
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation |
0.8 | 1 | 2024 | Graph Neural Networks for Learning Equivariant Representations of Neural Networks · ICLR 2024 |
Machine learning › Efficient and distributed learning
model initialization |
0.7 | 1 | 2023 | Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models? · ICML 2023 |
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pretrained model initialization |
0.7 | 1 | 2023 | Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models? · ICML 2023 |
Machine learning › Generative modeling
generative model evaluation |
0.6 | 1 | 2022 | On Evaluation Metrics for Graph Generative Models · ICLR 2022 |
Machine learning › Graph learning
graph generation |
0.6 | 1 | 2022 | On Evaluation Metrics for Graph Generative Models · ICLR 2022 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search › search space design
architecture encoding |
0.5 | 1 | 2021 | Parameter Prediction for Unseen Deep Architectures · NeurIPS 2021 |
Robotics › Robot manipulation
assembly |
0.5 | 1 | 2021 | Brick-by-Brick: Combinatorial Construction with Deep Reinforcement Learning · NeurIPS 2021 |
Natural language and speech › Language models and text generation
compositional generalization |
0.5 | 1 | 2021 | Generative Compositional Augmentations for Scene Graph Prediction · ICCV 2021 |
Machine learning › Generative modeling › generative adversarial network
conditional GAN |
0.5 | 1 | 2021 | Generative Compositional Augmentations for Scene Graph Prediction · ICCV 2021 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.5 | 1 | 2021 | Brick-by-Brick: Combinatorial Construction with Deep Reinforcement Learning · NeurIPS 2021 |
Machine learning › Deep learning architectures and training
hypernetwork |
0.5 | 1 | 2021 | Parameter Prediction for Unseen Deep Architectures · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › interpretability
attention analysis |
0.4 | 1 | 2019 | Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019 |
Machine learning › Graph learning › graph neural network
graph attention |
0.4 | 1 | 2019 | Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2019 | Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019 |
Machine learning › Deep learning architectures and training › symmetry-aware learning
permutation invariance |
0.2 | 1 | 2024 | Graph Neural Networks for Learning Equivariant Representations of Neural Networks · ICLR 2024 |
Computer vision › Segmentation and scene understanding
scene graph |
0.1 | 1 | 2021 | Generative Compositional Augmentations for Scene Graph Prediction · ICCV 2021 |
Computer vision › Vision and language
vision-language model |
0.1 | 1 | 2021 | Context-aware Scene Graph Generation with Seq2Seq Transformers · ICCV 2021 |
Machine learning › Graph learning
graph classification |
0.1 | 1 | 2019 | Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
graph neural network · 2.5transformer · 1.9weight nowcasting · 0.9parameter prediction · 0.7graph generative model · 0.6self-critical policy gradient · 0.5scene graph perturbation · 0.5reinforcement learning · 0.5monte carlo search · 0.5conditional generative adversarial network · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating Training with Neuron Interaction and Nowcasting NetworksabstractNeural network training can be accelerated when a learnable update rule is used in lieu of classic adaptive optimizers (e.g. Adam). However, learnable update rules can be costly and unstable to train and use. Recently, Jang et al. (2023) proposed a simpler approach to accelerate training based on weight nowcaster networks (WNNs). In their approach, Adam is used for most of the optimization steps and periodically, only every few steps, a WNN nowcasts (predicts near future) parameters. We improve WNNs by proposing neuron interaction and nowcasting (NiNo) networks. In contrast to WNNs, NiNo leverages neuron connectivity and graph neural networks to more accurately nowcast parameters. We further show that in some networks, such as Transformers, modeling neuron connectivity accurately is challenging. We address this and other limitations, which allows NiNo to accelerate Adam training by up to 50% in vision and language tasks. Boris Knyazev 0001, Abhinav Moudgil, Guillaume Lajoie, Eugene Belilovsky, Simon Lacoste-Julien |
ICLR | 1 |
| 2024 | Graph Neural Networks for Learning Equivariant Representations of Neural NetworksabstractNeural networks that process the parameters of other neural networks find applications in domains as diverse as classifying implicit neural representations, generating neural network weights, and predicting generalization errors. However, existing approaches either overlook the inherent permutation symmetry in the neural network or rely on intricate weight-sharing patterns to achieve equivariance, while ignoring the impact of the network architecture itself. In this work, we propose to represent neural networks as computational graphs of parameters, which allows us to harness powerful graph neural networks and transformers that preserve permutation symmetry. Consequently, our approach enables a single model to encode neural computational graphs with diverse architectures. We showcase the effectiveness of our method on a wide range of tasks, including classification and editing of implicit neural representations, predicting generalization performance, and learning to optimize, while consistently outperforming state-of-the-art methods. The source code is open-sourced at https://github.com/mkofinas/neural-graphs. Miltiadis Kofinas, Boris Knyazev 0001, Yunlu Chen, Gertjan J. Burghouts, Efstratios Gavves, Cees Snoek, David W. Zhang |
ICLR | 2 |
| 2023 | Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models?abstractPretraining a neural network on a large dataset is becoming a cornerstone in machine learning that is within the reach of only a few communities with large-resources. We aim at an ambitious goal of democratizing pretraining. Towards that goal, we train and release a single neural network that can predict high quality ImageNet parameters of other neural networks. By using predicted parameters for initialization we are able to boost training of diverse ImageNet models available in PyTorch. When transferred to other datasets, models initialized with predicted parameters also converge faster and reach competitive final performance. Boris Knyazev 0001, Doha Hwang, Simon Lacoste-Julien |
ICML | 1 |
| 2022 | On Evaluation Metrics for Graph Generative Models
Rylee Thompson, Boris Knyazev 0001, Elaheh Ghalebi, Jungtaek Kim 0001, Graham W. Taylor |
ICLR | 2 |
| 2021 | Generative Compositional Augmentations for Scene Graph PredictionabstractInferring objects and their relationships from an image in the form of a scene graph is useful in many applications at the intersection of vision and language. We consider a challenging problem of compositional generalization that emerges in this task due to a long tail data distribution. Current scene graph generation models are trained on a tiny fraction of the distribution corresponding to the most frequent compositions, e.g.. However, test images might contain zero- and few-shot compositions of objects and relationships, e.g.. Despite each of the object categories and the predicate (e.g. ‘on’) being frequent in the training data, the models often fail to properly understand such unseen or rare compositions. To improve generalization, it is natural to attempt increasing the diversity of the training distribution. However, in the graph domain this is non-trivial. To that end, we propose a method to synthesize rare yet plausible scene graphs by perturbing real ones. We then propose and empirically study a model based on conditional generative adversarial networks (GANs) that allows us to generate visual features of perturbed scene graphs and learn from them in a joint fashion. When evaluated on the Visual Genome dataset, our approach yields marginal, but consistent improvements in zero- and few-shot metrics. We analyze the limitations of our approach indicating promising directions for future research. Boris Knyazev 0001, Harm de Vries, Catalina Cangea, Graham W. Taylor, Aaron C. Courville, Eugene Belilovsky |
ICCV | 1 |
| 2021 | Context-aware Scene Graph Generation with Seq2Seq TransformersabstractScene graph generation is an important task in computer vision aimed at improving the semantic understanding of the visual world. In this task, the model needs to detect objects and predict visual relationships between them. Most of the existing models predict relationships in parallel assuming their independence. While there are different ways to capture these dependencies, we explore a conditional approach motivated by the sequence-to-sequence (Seq2Seq) formalism. Different from the previous research, our proposed model predicts visual relationships one at a time in an autoregressive manner by explicitly conditioning on the already predicted relationships. Drawing from translation models in NLP, we propose an encoder-decoder model built using Transformers where the encoder captures global context and long range interactions. The decoder then makes sequential predictions by conditioning on the scene graph constructed so far. In addition, we introduce a novel reinforcement learning-based training strategy tailored to Seq2Seq scene graph generation. By using a self-critical policy gradient training approach with Monte Carlo search we directly optimize for the (mean) recall metrics and bridge the gap between training and evaluation. Experimental results on two public benchmark datasets demonstrate that our Seq2Seq learning approach achieves strong empirical performance, outperforming previous state-of-the-art, while remaining efficient in terms of training and inference time. Full code for this work is available here: https://github.com/layer6ai-labs/SGG-Seq2Seq. Yichao Lu, Himanshu Rai, Boris Knyazev 0001, Guangwei Yu, Shashank Shekhar 0005, Graham W. Taylor, Maksims Volkovs |
ICCV | 4 |
| 2021 | Brick-by-Brick: Combinatorial Construction with Deep Reinforcement LearningabstractDiscovering a solution in a combinatorial space is prevalent in many real-world problems but it is also challenging due to diverse complex constraints and the vast number of possible combinations. To address such a problem, we introduce a novel formulation, combinatorial construction, which requires a building agent to assemble unit primitives (i.e., LEGO bricks) sequentially -- every connection between two bricks must follow a fixed rule, while no bricks mutually overlap. To construct a target object, we provide incomplete knowledge about the desired target (i.e., 2D images) instead of exact and explicit volumetric information to the agent. This problem requires a comprehensive understanding of partial information and long-term planning to append a brick sequentially, which leads us to employ reinforcement learning. The approach has to consider a variable-sized action space where a large number of invalid actions, which would cause overlap between bricks, exist. To resolve these issues, our model, dubbed Brick-by-Brick, adopts an action validity prediction network that efficiently filters invalid actions for an actor-critic network. We demonstrate that the proposed method successfully learns to construct an unseen object conditioned on a single image or multiple views of a target object. Hyunsoo Chung, Jungtaek Kim 0001, Boris Knyazev 0001, Jinhwi Lee, Graham W. Taylor, Jaesik Park, Minsu Cho |
NeurIPS | 3 |
| 2021 | Parameter Prediction for Unseen Deep ArchitecturesabstractDeep learning has been successful in automating the design of features in machine learning pipelines. However, the algorithms optimizing neural network parameters remain largely hand-designed and computationally inefficient. We study if we can use deep learning to directly predict these parameters by exploiting the past knowledge of training other networks. We introduce a large-scale dataset of diverse computational graphs of neural architectures - DeepNets-1M - and use it to explore parameter prediction on CIFAR-10 and ImageNet. By leveraging advances in graph neural networks, we propose a hypernetwork that can predict performant parameters in a single forward pass taking a fraction of a second, even on a CPU. The proposed model achieves surprisingly good performance on unseen and diverse networks. For example, it is able to predict all 24 million parameters of a ResNet-50 achieving a 60% accuracy on CIFAR-10. On ImageNet, top-5 accuracy of some of our networks approaches 50%. Our task along with the model and results can potentially lead to a new, more computationally efficient paradigm of training networks. Our model also learns a strong representation of neural architectures enabling their analysis. Boris Knyazev 0001, Michal Drozdzal, Graham W. Taylor, Adriana Romero-Soriano |
NeurIPS | 1 |
| 2020 | Graph Density-Aware Losses for Novel Compositions in Scene Graph Generation
Boris Knyazev 0001, Harm de Vries, Catalina Cangea, Graham W. Taylor, Aaron C. Courville, Eugene Belilovsky |
BMVC | 1 |
| 2019 | Image Classification with Hierarchical Multigraph Networks
Boris Knyazev 0001, Mohamed R. Amer, Graham W. Taylor |
BMVC | 1 |
| 2019 | Understanding Attention and Generalization in Graph Neural NetworksabstractWe aim to better understand attention over nodes in graph neural networks (GNNs) and identify factors influencing its effectiveness. We particularly focus on the ability of attention GNNs to generalize to larger, more complex or noisy graphs. Motivated by insights from the work on Graph Isomorphism Networks, we design simple graph reasoning tasks that allow us to study attention in a controlled environment. We find that under typical conditions the effect of attention is negligible or even harmful, but under certain conditions it provides an exceptional gain in performance of more than 60% in some of our classification tasks. Satisfying these conditions in practice is challenging and often requires optimal initialization or supervised training of attention. We propose an alternative recipe and train attention in a weakly-supervised fashion that approaches the performance of supervised models, and, compared to unsupervised models, improves results on several synthetic as well as real datasets. Source code and datasets are available at https://github.com/bknyaz/graphattentionpool. Boris Knyazev 0001, Graham W. Taylor, Mohamed R. Amer |
NeurIPS | 1 |
| 2018 | Leveraging Large Face Recognition Data for Emotion ClassificationabstractIn this paper we describe a solution to our entry for the emotion recognition challenge EmotiW 2017. We propose an ensemble of several models, which capture spatial and audio features from videos. Spatial features are captured by convolutional neural networks, pretrained on large face recognition datasets. We show that usage of strong industry-level face recognition networks increases the accuracy of emotion recognition. Using our ensemble we improve on the previous year's best result on the test set by about 1%, achieving a 60.03% classification accuracy without any use of visual temporal information, showing a top-2 result in this challenge. Boris Knyazev 0001, Roman Shvetsov, Natalia Efremova, Artem Kuharenko |
FG | 1 |
| 2017 | Recursive autoconvolution for unsupervised learning of convolutional neural networksabstractIn visual recognition tasks, such as image classification, unsupervised learning exploits cheap unlabeled data and can help to solve these tasks more efficiently. We show that the recursive autoconvolution operator, adopted from physics, boosts existing unsupervised methods by learning more discriminative filters. We take well established convolutional neural networks and train their filters layer-wise. In addition, based on previous works we design a network which extracts more than 600k features per sample, but with the total number of trainable parameters greatly reduced by introducing shared filters in higher layers. We evaluate our networks on the MNIST, CIFAR-10, CIFAR-100 and STL-10 image classification benchmarks and report several state of the art results among other unsupervised methods. Boris Knyazev 0001, Erhardt Barth, Thomas Martinetz |
IJCNN | 1 |