Boris Knyazev 0001

dblp:181/5675-1 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0002-9484-1534ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 9 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Graph learning · 23% Efficient and distributed learning · 16% Deep learning architectures and training · 10%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
2.542025
Accelerating Training with Neuron Interaction and Nowcasting Networks · ICLR 2025
Graph Neural Networks for Learning Equivariant Representations of Neural Networks · ICLR 2024
Parameter Prediction for Unseen Deep Architectures · NeurIPS 2021
Machine learning › Efficient and distributed learning
parameter prediction
1.222023
Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models? · ICML 2023
Parameter Prediction for Unseen Deep Architectures · NeurIPS 2021
Computer vision › Segmentation and scene understanding
scene graph generation
1.022021
Context-aware Scene Graph Generation with Seq2Seq Transformers · ICCV 2021
Generative Compositional Augmentations for Scene Graph Prediction · ICCV 2021
Machine learning › Reinforcement learning › meta-reinforcement learning
learned update rules
0.912025
Accelerating Training with Neuron Interaction and Nowcasting Networks · ICLR 2025
Machine learning › Optimization for machine learning
neural network training acceleration
0.912025
Accelerating Training with Neuron Interaction and Nowcasting Networks · ICLR 2025
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation
0.812024
Graph Neural Networks for Learning Equivariant Representations of Neural Networks · ICLR 2024
Machine learning › Efficient and distributed learning
model initialization
0.712023
Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models? · ICML 2023
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pretrained model initialization
0.712023
Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models? · ICML 2023
Machine learning › Generative modeling
generative model evaluation
0.612022
On Evaluation Metrics for Graph Generative Models · ICLR 2022
Machine learning › Graph learning
graph generation
0.612022
On Evaluation Metrics for Graph Generative Models · ICLR 2022
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search › search space design
architecture encoding
0.512021
Parameter Prediction for Unseen Deep Architectures · NeurIPS 2021
Robotics › Robot manipulation
assembly
0.512021
Brick-by-Brick: Combinatorial Construction with Deep Reinforcement Learning · NeurIPS 2021
Natural language and speech › Language models and text generation
compositional generalization
0.512021
Generative Compositional Augmentations for Scene Graph Prediction · ICCV 2021
Machine learning › Generative modeling › generative adversarial network
conditional GAN
0.512021
Generative Compositional Augmentations for Scene Graph Prediction · ICCV 2021
Machine learning › Reinforcement learning
deep reinforcement learning
0.512021
Brick-by-Brick: Combinatorial Construction with Deep Reinforcement Learning · NeurIPS 2021
Machine learning › Deep learning architectures and training
hypernetwork
0.512021
Parameter Prediction for Unseen Deep Architectures · NeurIPS 2021
Machine learning › Trustworthy machine learning › interpretability
attention analysis
0.412019
Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019
Machine learning › Graph learning › graph neural network
graph attention
0.412019
Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019
Machine learning › Trustworthy machine learning
interpretability
0.412019
Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019
Machine learning › Deep learning architectures and training › symmetry-aware learning
permutation invariance
0.212024
Graph Neural Networks for Learning Equivariant Representations of Neural Networks · ICLR 2024
Computer vision › Segmentation and scene understanding
scene graph
0.112021
Generative Compositional Augmentations for Scene Graph Prediction · ICCV 2021
Computer vision › Vision and language
vision-language model
0.112021
Context-aware Scene Graph Generation with Seq2Seq Transformers · ICCV 2021
Machine learning › Graph learning
graph classification
0.112019
Understanding Attention and Generalization in Graph Neural Networks · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

graph neural network · 2.5transformer · 1.9weight nowcasting · 0.9parameter prediction · 0.7graph generative model · 0.6self-critical policy gradient · 0.5scene graph perturbation · 0.5reinforcement learning · 0.5monte carlo search · 0.5conditional generative adversarial network · 0.5
YearPublicationVenuePosition
2025 Accelerating Training with Neuron Interaction and Nowcasting Networks
abstract
Neural network training can be accelerated when a learnable update rule is used in lieu of classic adaptive optimizers (e.g. Adam). However, learnable update rules can be costly and unstable to train and use. Recently, Jang et al. (2023) proposed a simpler approach to accelerate training based on weight nowcaster networks (WNNs). In their approach, Adam is used for most of the optimization steps and periodically, only every few steps, a WNN nowcasts (predicts near future) parameters. We improve WNNs by proposing neuron interaction and nowcasting (NiNo) networks. In contrast to WNNs, NiNo leverages neuron connectivity and graph neural networks to more accurately nowcast parameters. We further show that in some networks, such as Transformers, modeling neuron connectivity accurately is challenging. We address this and other limitations, which allows NiNo to accelerate Adam training by up to 50% in vision and language tasks.
Boris Knyazev 0001, Abhinav Moudgil, Guillaume Lajoie, Eugene Belilovsky, Simon Lacoste-Julien
ICLR1
2024 Graph Neural Networks for Learning Equivariant Representations of Neural Networks
abstract
Neural networks that process the parameters of other neural networks find applications in domains as diverse as classifying implicit neural representations, generating neural network weights, and predicting generalization errors. However, existing approaches either overlook the inherent permutation symmetry in the neural network or rely on intricate weight-sharing patterns to achieve equivariance, while ignoring the impact of the network architecture itself. In this work, we propose to represent neural networks as computational graphs of parameters, which allows us to harness powerful graph neural networks and transformers that preserve permutation symmetry. Consequently, our approach enables a single model to encode neural computational graphs with diverse architectures. We showcase the effectiveness of our method on a wide range of tasks, including classification and editing of implicit neural representations, predicting generalization performance, and learning to optimize, while consistently outperforming state-of-the-art methods. The source code is open-sourced at https://github.com/mkofinas/neural-graphs.
Miltiadis Kofinas, Boris Knyazev 0001, Yunlu Chen, Gertjan J. Burghouts, Efstratios Gavves, Cees Snoek, David W. Zhang
ICLR2
2023 Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models?
abstract
Pretraining a neural network on a large dataset is becoming a cornerstone in machine learning that is within the reach of only a few communities with large-resources. We aim at an ambitious goal of democratizing pretraining. Towards that goal, we train and release a single neural network that can predict high quality ImageNet parameters of other neural networks. By using predicted parameters for initialization we are able to boost training of diverse ImageNet models available in PyTorch. When transferred to other datasets, models initialized with predicted parameters also converge faster and reach competitive final performance.
Boris Knyazev 0001, Doha Hwang, Simon Lacoste-Julien
ICML1
2022 On Evaluation Metrics for Graph Generative Models
Rylee Thompson, Boris Knyazev 0001, Elaheh Ghalebi, Jungtaek Kim 0001, Graham W. Taylor
ICLR2
2021 Generative Compositional Augmentations for Scene Graph Prediction
abstract
Inferring objects and their relationships from an image in the form of a scene graph is useful in many applications at the intersection of vision and language. We consider a challenging problem of compositional generalization that emerges in this task due to a long tail data distribution. Current scene graph generation models are trained on a tiny fraction of the distribution corresponding to the most frequent compositions, e.g.. However, test images might contain zero- and few-shot compositions of objects and relationships, e.g.. Despite each of the object categories and the predicate (e.g. ‘on’) being frequent in the training data, the models often fail to properly understand such unseen or rare compositions. To improve generalization, it is natural to attempt increasing the diversity of the training distribution. However, in the graph domain this is non-trivial. To that end, we propose a method to synthesize rare yet plausible scene graphs by perturbing real ones. We then propose and empirically study a model based on conditional generative adversarial networks (GANs) that allows us to generate visual features of perturbed scene graphs and learn from them in a joint fashion. When evaluated on the Visual Genome dataset, our approach yields marginal, but consistent improvements in zero- and few-shot metrics. We analyze the limitations of our approach indicating promising directions for future research.
Boris Knyazev 0001, Harm de Vries, Catalina Cangea, Graham W. Taylor, Aaron C. Courville, Eugene Belilovsky
ICCV1
2021 Context-aware Scene Graph Generation with Seq2Seq Transformers
abstract
Scene graph generation is an important task in computer vision aimed at improving the semantic understanding of the visual world. In this task, the model needs to detect objects and predict visual relationships between them. Most of the existing models predict relationships in parallel assuming their independence. While there are different ways to capture these dependencies, we explore a conditional approach motivated by the sequence-to-sequence (Seq2Seq) formalism. Different from the previous research, our proposed model predicts visual relationships one at a time in an autoregressive manner by explicitly conditioning on the already predicted relationships. Drawing from translation models in NLP, we propose an encoder-decoder model built using Transformers where the encoder captures global context and long range interactions. The decoder then makes sequential predictions by conditioning on the scene graph constructed so far. In addition, we introduce a novel reinforcement learning-based training strategy tailored to Seq2Seq scene graph generation. By using a self-critical policy gradient training approach with Monte Carlo search we directly optimize for the (mean) recall metrics and bridge the gap between training and evaluation. Experimental results on two public benchmark datasets demonstrate that our Seq2Seq learning approach achieves strong empirical performance, outperforming previous state-of-the-art, while remaining efficient in terms of training and inference time. Full code for this work is available here: https://github.com/layer6ai-labs/SGG-Seq2Seq.
Yichao Lu, Himanshu Rai, Boris Knyazev 0001, Guangwei Yu, Shashank Shekhar 0005, Graham W. Taylor, Maksims Volkovs
ICCV4
2021 Brick-by-Brick: Combinatorial Construction with Deep Reinforcement Learning
abstract
Discovering a solution in a combinatorial space is prevalent in many real-world problems but it is also challenging due to diverse complex constraints and the vast number of possible combinations. To address such a problem, we introduce a novel formulation, combinatorial construction, which requires a building agent to assemble unit primitives (i.e., LEGO bricks) sequentially -- every connection between two bricks must follow a fixed rule, while no bricks mutually overlap. To construct a target object, we provide incomplete knowledge about the desired target (i.e., 2D images) instead of exact and explicit volumetric information to the agent. This problem requires a comprehensive understanding of partial information and long-term planning to append a brick sequentially, which leads us to employ reinforcement learning. The approach has to consider a variable-sized action space where a large number of invalid actions, which would cause overlap between bricks, exist. To resolve these issues, our model, dubbed Brick-by-Brick, adopts an action validity prediction network that efficiently filters invalid actions for an actor-critic network. We demonstrate that the proposed method successfully learns to construct an unseen object conditioned on a single image or multiple views of a target object.
Hyunsoo Chung, Jungtaek Kim 0001, Boris Knyazev 0001, Jinhwi Lee, Graham W. Taylor, Jaesik Park, Minsu Cho
NeurIPS3
2021 Parameter Prediction for Unseen Deep Architectures
abstract
Deep learning has been successful in automating the design of features in machine learning pipelines. However, the algorithms optimizing neural network parameters remain largely hand-designed and computationally inefficient. We study if we can use deep learning to directly predict these parameters by exploiting the past knowledge of training other networks. We introduce a large-scale dataset of diverse computational graphs of neural architectures - DeepNets-1M - and use it to explore parameter prediction on CIFAR-10 and ImageNet. By leveraging advances in graph neural networks, we propose a hypernetwork that can predict performant parameters in a single forward pass taking a fraction of a second, even on a CPU. The proposed model achieves surprisingly good performance on unseen and diverse networks. For example, it is able to predict all 24 million parameters of a ResNet-50 achieving a 60% accuracy on CIFAR-10. On ImageNet, top-5 accuracy of some of our networks approaches 50%. Our task along with the model and results can potentially lead to a new, more computationally efficient paradigm of training networks. Our model also learns a strong representation of neural architectures enabling their analysis.
Boris Knyazev 0001, Michal Drozdzal, Graham W. Taylor, Adriana Romero-Soriano
NeurIPS1
2020 Graph Density-Aware Losses for Novel Compositions in Scene Graph Generation
Boris Knyazev 0001, Harm de Vries, Catalina Cangea, Graham W. Taylor, Aaron C. Courville, Eugene Belilovsky
BMVC1
2019 Image Classification with Hierarchical Multigraph Networks
Boris Knyazev 0001, Mohamed R. Amer, Graham W. Taylor
BMVC1
2019 Understanding Attention and Generalization in Graph Neural Networks
abstract
We aim to better understand attention over nodes in graph neural networks (GNNs) and identify factors influencing its effectiveness. We particularly focus on the ability of attention GNNs to generalize to larger, more complex or noisy graphs. Motivated by insights from the work on Graph Isomorphism Networks, we design simple graph reasoning tasks that allow us to study attention in a controlled environment. We find that under typical conditions the effect of attention is negligible or even harmful, but under certain conditions it provides an exceptional gain in performance of more than 60% in some of our classification tasks. Satisfying these conditions in practice is challenging and often requires optimal initialization or supervised training of attention. We propose an alternative recipe and train attention in a weakly-supervised fashion that approaches the performance of supervised models, and, compared to unsupervised models, improves results on several synthetic as well as real datasets. Source code and datasets are available at https://github.com/bknyaz/graphattentionpool.
Boris Knyazev 0001, Graham W. Taylor, Mohamed R. Amer
NeurIPS1
2018 Leveraging Large Face Recognition Data for Emotion Classification
abstract
In this paper we describe a solution to our entry for the emotion recognition challenge EmotiW 2017. We propose an ensemble of several models, which capture spatial and audio features from videos. Spatial features are captured by convolutional neural networks, pretrained on large face recognition datasets. We show that usage of strong industry-level face recognition networks increases the accuracy of emotion recognition. Using our ensemble we improve on the previous year's best result on the test set by about 1%, achieving a 60.03% classification accuracy without any use of visual temporal information, showing a top-2 result in this challenge.
Boris Knyazev 0001, Roman Shvetsov, Natalia Efremova, Artem Kuharenko
FG1
2017 Recursive autoconvolution for unsupervised learning of convolutional neural networks
abstract
In visual recognition tasks, such as image classification, unsupervised learning exploits cheap unlabeled data and can help to solve these tasks more efficiently. We show that the recursive autoconvolution operator, adopted from physics, boosts existing unsupervised methods by learning more discriminative filters. We take well established convolutional neural networks and train their filters layer-wise. In addition, based on previous works we design a network which extracts more than 600k features per sample, but with the total number of trainable parameters greatly reduced by introducing shared filters in higher layers. We evaluate our networks on the MNIST, CIFAR-10, CIFAR-100 and STL-10 image classification benchmarks and report several state of the art results among other unsupervised methods.
Boris Knyazev 0001, Erhardt Barth, Thomas Martinetz
IJCNN1