EDBT 2026 Demo / reviewers in the wild / expert
Vyacheslav Kungurtsev
dblp:149/2722
· DBLP profile ↗
17ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-2229-8824ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Concord: Concept-Informed Diffusion for Dataset DistillationabstractDataset distillation (DD) has witnessed significant progress in creating small datasets that encapsulate rich information from large original ones. Particularly, methods based on generative priors show promising performance while maintaining computational efficiency and cross-architecture generalization. However, the generation process lacks explicit controllability for each sample. Previous distillation methods primarily match the real distribution from the perspective of the entire dataset, whereas overlooking concept completeness at the instance level. The missing or incorrectly represented object details cannot be efficiently compensated due to the constrained sample amount typical in DD settings. To this end, we propose incorporating the concept understanding of large language models (LLMs) to perform Concept-Informed Diffusion (Concord) for dataset distillation. Specifically, distinguishable and fine-grained concepts are retrieved based on category labels to inform the denoising process and refine essential object details. These concepts can be applied to any diffusion-based DD framework to enhance both the controllability and interpretability of the distilled image generation, without relying on pre-trained classifiers. We demonstrate the efficacy of Concord by achieving state-of-the-art performance on ImageNet-1K and subsets. Code is released at Concord. Jianyang Gu, Ruoxi Jia 0001, Saeed Vahidian, Vyacheslav Kungurtsev, Wei Jiang 0009, Yiran Chen 0001 |
WACV | 5 |
| 2026 | Contemporary data-driven innovations in peptide-based therapeutic designabstractIn recent years, peptides have grabbed significant attention across many fields, including pharmaceutical, biomedical, and biotechnological industries, owing to their notable biological activity, low toxicity, and high specificity. Naturally occurring peptides play crucial roles in handling different biological processes (cellular signaling, immune responses, and enzymatic functions, etc.), while laboratory-made synthetic peptides can be adopted for many applied industrial applications. Following the practical problems in large-scale peptide synthesis in the laboratory which demand high cost and time, recent developments have benefited from the best use of machine learning (ML), deep learning, active learning, reinforcement learning (RL), generative artificial intelligence (AI), and large language models (LLMs) to reduce the number of experiments. ML algorithms enable the prediction of the peptide structure-activity relationship, bioavailability, and other drug-like properties with high accuracy. The integration of AI with peptide-based therapeutics design marks a paradigm shift in drug discovery which was otherwise dominated by small organic molecules. This review comprehensively examines AI-driven methodologies, including classical ML approaches, deep generative models, RL, and LLMs, that overcome historical limitations in peptide design, such as structural flexibility, enzymatic degradation, and membrane impermeability. Recent advances in structure-aware algorithms and sequence-based frameworks have accelerated peptide-based therapeutic development across oncology, metabolic disorders, and infectious diseases. Despite challenges in data scarcity and validation gaps, the convergence of computational prediction with experimental automation promises clinical translation of AI-designed peptides in the near future. This review highlights the transformative potential of AI in ushering a new era of precision peptide therapeutics. Lipsa Priyadarsinee, Vyacheslav Kungurtsev, Vibhor Kumar, Bapi Chatterjee, G. Narahari Sastry, Natarajan Arul Murugan |
Briefings Bioinform. | 2 |
| 2025 | Binarizing Physics-Inspired GNNs for Combinatorial OptimizationabstractPhysics-inspired graph neural networks (PI-GNNs) have been utilized as an efficient unsupervised framework for relaxing combinatorial optimization problems encoded through a specific graph structure and loss, reflecting dependencies between the problem’s variables. While the framework has yielded promising results in various combinatorial problems, we show that the performance of PI-GNNs systematically plummets with an increasing density of the combinatorial problem graphs. Our analysis reveals an interesting phase transition in the PI-GNNs’ training dynamics, associated with degenerate solutions for the denser problems, highlighting a discrepancy between the relaxed, real-valued model outputs and the binary-valued problem solutions. To address the discrepancy, we propose principled alternatives to the naive strategy used in PI-GNNs by building on insights from fuzzy logic and binarized neural networks. Our experiments demonstrate that the portfolio of proposed methods significantly improves the performance of PI-GNNs in increasingly dense settings. Martin Krutsky, Gustav Sír, Vyacheslav Kungurtsev, Georgios Korpas |
ECAI | 3 |
| 2025 | Group Distributionally Robust Dataset Distillation with Risk MinimizationabstractDataset distillation (DD) has emerged as a widely adopted technique for crafting a synthetic dataset that captures the essential information of a training dataset, facilitating the training of accurate neural models. Its applications span various domains, including transfer learning, federated learning, and neural architecture search. The most popular methods for constructing the synthetic data rely on matching the convergence properties of training the model with the synthetic dataset and the training dataset. However, using the empirical loss as the criterion must be thought of as auxiliary in the same sense that the training set is an approximate substitute for the population distribution, and the latter is the data of interest. Yet despite its popularity, an aspect that remains unexplored is the relationship of DD to its generalization, particularly across uncommon subgroups. That is, how can we ensure that a model trained on the synthetic dataset performs well when faced with samples from regions with low population density? Here, the representativeness and coverage of the dataset become salient over the guaranteed training error at inference. Drawing inspiration from distributionally robust optimization, we introduce an algorithm that combines clustering with the minimization of a risk measure on the loss to conduct DD. We provide a theoretical rationale for our approach and demonstrate its effective generalization and robustness across subgroups through numerical experiments. Saeed Vahidian, Mingyu Wang 0004, Jianyang Gu, Vyacheslav Kungurtsev, Wei Jiang 0009, Yiran Chen 0001 |
ICLR | 4 |
| 2024 | Efficient Dataset Distillation via Minimax DiffusionabstractDataset distillation reduces the storage and computational consumption of training a network by generating a small surrogate dataset that encapsulates rich information of the original large-scale one. However, previous distillation methods heavily rely on the sample-wise iterative optimization scheme. As the images-per-class (IPC) setting or image resolution grows larger, the necessary computation will demand overwhelming time and resources. In this work, we intend to incorporate generative diffusion techniques for computing the surrogate dataset. Observing that key factors for constructing an effective surrogate dataset are representativeness and diversity, we design additional minimax criteria in the generative training to enhance these facets for the generated images of diffusion models. We present a theoretical model of the process as hierarchical diffusion control demonstrating the flexibility of the diffusion process to target these criteria without jeopardizing the faithfulness of the sample to the desired distribution. The proposed method achieves state-of-the-art validation performance while demanding much less computational resources. Under the 100-IPC setting on Image Woof, our method requires less than one-twentieth the distillation time of previous methods, yet yields even better performance. Source code and generated data are available in https://github.com/vimar-gu/MinimaxDiffusion. Jianyang Gu, Saeed Vahidian, Vyacheslav Kungurtsev, Wei Jiang 0009, Yang You 0001, Yiran Chen 0001 |
CVPR | 3 |
| 2024 | Unlocking the Potential of Federated Learning: The Symphony of Dataset Distillation via Deep Generative Latents
Yuqi Jia 0001, Saeed Vahidian, Jingwei Sun 0002, Vyacheslav Kungurtsev, Neil Zhenqiang Gong, Yiran Chen 0001 |
ECCV (78) | 5 |
| 2024 | Federated SGD with Local AsynchronyabstractParallel SGD in a shared-memory setting is oft-represented by the popular Hogwild! algorithm, in which lock-free updates are asynchronously performed by multiple computing processes. Unfortunately, scaling Hogwild! to distributed workers is largely unexplored. Specifically, it is unknown if any adaptation of Hogwild! to the popular decentralized multi-GPU setting offers any competitive speedup, either empirically or theoretically. In this work, we investigate the potential of decentralizing Hogwild! by incorporating simultaneously (a) asynchronous local gradient updates on the shared memory of GPUs, and (b) non-blocking asynchronous decentralized federated averaging. A naive direct implementation shows degradation in performance, arising from scheduling overheads and concurrent write conflicts on GPUs. To mitigate these drawbacks, we investigate and propose a new method, based on careful block selection rules, which update only portions of the parameter vectors. Our experiments show that the resulting decentralized training method exhibits improved throughput and competitive accuracy for standard image classification benchmarks on the CIFAR-10, CIFAR-100, and Imagenet datasets. On the theoretical side, we prove that our method guarantees sublinear ergodic convergence rates for non-convex objectives. Bapi Chatterjee, Vyacheslav Kungurtsev, Dan Alistarh |
ICDCS | 2 |
| 2024 | Towards Diverse Device Heterogeneous Federated Learning via Task Arithmetic Knowledge IntegrationabstractFederated Learning (FL) has emerged as a promising paradigm for collaborative machine learning, while preserving user data privacy. Despite its potential, standard FL algorithms lack support for diverse heterogeneous device prototypes, which vary significantly in model and dataset sizes---from small IoT devices to large workstations. This limitation is only partially addressed by existing knowledge distillation (KD) techniques, which often fail to transfer knowledge effectively across a broad spectrum of device prototypes with varied capabilities. This failure primarily stems from two issues: the dilution of informative logits from more capable devices by those from less capable ones, and the use of a single integrated logits as the distillation target across all devices, which neglects their individual learning capacities and and the unique contributions of each device. To address these challenges, we introduce TAKFL, a novel KD-based framework that treats the knowledge transfer from each device prototype's ensemble as a separate task, independently distilling each to preserve its unique contributions and avoid dilution. TAKFL also incorporates a KD-based self-regularization technique to mitigate the issues related to the noisy and unsupervised ensemble distillation process. To integrate the separately distilled knowledge, we introduce an adaptive task arithmetic knowledge integration process, allowing each student model to customize the knowledge integration for optimal performance. Additionally, we present theoretical results demonstrating the effectiveness of task arithmetic in transferring knowledge across heterogeneous device prototypes with varying capacities. Comprehensive evaluations of our method across both computer vision (CV) and natural language processing (NLP) tasks demonstrate that TAKFL achieves state-of-the-art results in a variety of datasets and settings, significantly outperforming existing KD-based methods. Our code is released at https://github.com/MMorafah/TAKFL and the project website is available at https://mmorafah.github.io/takflpage . Mahdi Morafah, Vyacheslav Kungurtsev, Hojin Chang, Chen Chen 0001, Bill Lin 0001 |
NeurIPS | 2 |
| 2023 | Efficient Distribution Similarity Identification in Clustered Federated Learning via Principal Angles between Client Data SubspacesabstractClustered federated learning (FL) has been shown to produce promising results by grouping clients into clusters. This is especially effective in scenarios where separate groups of clients have significant differences in the distributions of their local data. Existing clustered FL algorithms are essentially trying to group together clients with similar distributions so that clients in the same cluster can leverage each other's data to better perform federated learning. However, prior clustered FL algorithms attempt to learn these distribution similarities indirectly during training, which can be quite time consuming as many rounds of federated learning may be required until the formation of clusters is stabilized. In this paper, we propose a new approach to federated learning that directly aims to efficiently identify distribution similarities among clients by analyzing the principal angles between the client data subspaces. Each client applies a truncated singular value decomposition (SVD) step on its local data in a single-shot manner to derive a small set of principal vectors, which provides a signature that succinctly captures the main characteristics of the underlying distribution. This small set of principal vectors is provided to the server so that the server can directly identify distribution similarities among the clients to form clusters. This is achieved by comparing the similarities of the principal angles between the client data subspaces spanned by those principal vectors. The approach provides a simple, yet effective clustered FL framework that addresses a broad range of data heterogeneity issues beyond simpler forms of Non-IIDness like label skews. Our clustered FL approach also enables convergence guarantees for non-convex objectives. Saeed Vahidian, Mahdi Morafah, Weijia Wang 0002, Vyacheslav Kungurtsev, Chen Chen 0001, Mubarak Shah, Bill Lin 0001 |
AAAI | 4 |
| 2023 | When Do Curricula Work in Federated Learning?abstractAn oft-cited open problem of federated learning is the existence of data heterogeneity among clients. One pathway to understanding the drastic accuracy drop in federated learning is by scrutinizing the behavior of the clients’ deep models on data with different levels of "difficulty", which has been left unaddressed. In this paper, we investigate a different and rarely studied dimension of FL: ordered learning. Specifically, we aim to investigate how ordered learning principles can contribute to alleviating the heterogeneity effects in FL. We present theoretical analysis and conduct extensive empirical studies on the efficacy of orderings spanning three kinds of learning: curriculum, anti-curriculum, and random curriculum. We find that curriculum learning largely alleviates non-IIDness. Interestingly, the more disparate the data distributions across clients the more they benefit from ordered learning. We provide analysis explaining this phenomenon, specifically indicating how curriculum training appears to make the objective landscape progressively less convex, suggesting fast converging iterations at the beginning of the training procedure. We derive quantitative results of convergence for both convex and nonconvex objectives by modeling the curriculum training on federated devices as local SGD with locally biased stochastic gradients. Also, inspired by ordered learning, we propose a novel client selection technique that benefits from the real-world disparity in the clients. Our proposed approach to client selection has a synergic effect when applied together with ordered learning in FL. Saeed Vahidian, Sreevatsank Kadaveru, Woonjoon Baek, Weijia Wang 0002, Vyacheslav Kungurtsev, Chen Chen 0001, Mubarak Shah, Bill Lin 0001 |
ICCV | 5 |
| 2023 | Decentralized Bayesian learning with Metropolis-adjusted Hamiltonian Monte Carlo
Vyacheslav Kungurtsev, Adam D. Cobb, Tara Javidi, Brian Jalaian |
Mach. Learn. | 1 |
| 2022 | Mean-field Analysis of Piecewise Linear Solutions for Wide ReLU NetworksabstractUnderstanding the properties of neural networks trained via stochastic gradient descent (SGD) is at the heart of the theory of deep learning. In this work, we take a mean-field view, and consider a two-layer ReLU network trained via noisy-SGD for a univariate regularized regression problem. Our main result is that SGD with vanishingly small noise injected in the gradients is biased towards a simple solution: at convergence, the ReLU network implements a piecewise linear map of the inputs, and the number of “knot” points -- i.e., points where the tangent of the ReLU network estimator changes -- between two consecutive training inputs is at most three. In particular, as the number of neurons of the network grows, the SGD dynamics is captured by the solution of a gradient flow and, at convergence, the distribution of the weights approaches the unique minimizer of a related free energy, which has a Gibbs form. Our key technical contribution consists in the analysis of the estimator resulting from this minimizer: we show that its second derivative vanishes everywhere, except at some specific locations which represent the “knot” points. We also provide empirical evidence that knots at locations distinct from the data points might occur, as predicted by our theory. Alexander Shevchenko, Vyacheslav Kungurtsev, Marco Mondelli |
J. Mach. Learn. Res. | 2 |
| 2021 | Asynchronous Optimization Methods for Efficient Training of Deep Neural Networks with GuaranteesabstractAsynchronous distributed algorithms are a popular way to reduce synchronization costs in large-scale optimization, and in particular for neural network training. However, for nonsmooth and nonconvex objectives, few convergence guarantees exist beyond cases where closed-form proximal operator solutions are available. As training most popular deep neural networks corresponds to optimizing nonsmooth and nonconvex objectives, there is a pressing need for such convergence guarantees. In this paper, we analyze for the first time the convergence of stochastic asynchronous optimization for this general class of objectives. In particular, we focus on stochastic subgradient methods allowing for block variable partitioning, where the shared model is asynchronously updated by concurrent processes. To this end, we use a probabilistic model which captures key features of real asynchronous scheduling between concurrent processes. Under this model, we establish convergence with probability one to an invariant set for stochastic subgradient methods with momentum. From a practical perspective, one issue with the family of algorithms that we consider is that they are not efficiently supported by machine learning frameworks, which mostly focus on distributed data-parallel strategies. To address this, we propose a new implementation strategy for shared-memory based training of deep neural networks for a partitioned but shared model in single- and multi-GPU settings. Based on this implementation, we achieve on average1.2x speed-up in comparison to state-of-the-art training methods for popular image classification tasks, without compromising accuracy. Vyacheslav Kungurtsev, Malcolm Egan, Bapi Chatterjee, Dan Alistarh |
AAAI | 1 |
| 2021 | Elastic Consistency: A Practical Consistency Model for Distributed Stochastic Gradient DescentabstractOne key element behind the recent progress of machine learning has been the ability to train machine learning models in large-scale distributed shared-memory and message-passing environments. Most of these models are trained employing variants of stochastic gradient descent (SGD) based optimization, but most methods involve some type of consistency relaxation relative to sequential SGD, to mitigate its large communication or synchronization costs at scale. In this paper, we introduce a general consistency condition covering communication-reduced and asynchronous distributed SGD implementations. Our framework, called elastic consistency, decouples the system-specific aspects of the implementation from the SGD convergence requirements, giving a general way to obtain convergence bounds for a wide variety of distributed SGD methods used in practice. Elastic consistency can be used to re-derive or improve several previous convergence bounds in message-passing and shared-memory settings, but also to analyze new models and distribution schemes. As a direct application, we propose and analyze a new synchronization-avoiding scheduling scheme for distributed SGD, and show that it can be used to efficiently train deep convolutional models for image classification. Giorgi Nadiradze, Ilia Markov, Bapi Chatterjee, Vyacheslav Kungurtsev, Dan Alistarh |
AAAI | 4 |
| 2019 | Lifted Weight Learning of Markov Logic Networks RevisitedabstractWe study lifted weight learning of Markov logic networks. We show that there is an algorithm for maximum-likelihood learning of 2-variable Markov logic networks which runs in time polynomial in the domain size. Our results are based on existing lifted-inference algorithms and recent algorithmic results on computing maximum entropy distributions. Ondrej Kuzelka, Vyacheslav Kungurtsev |
AISTATS | 2 |
| 2017 | Asynchronous parallel nonconvex large-scale optimizationabstractWe propose a novel parallel asynchronous algorithmic framework for the minimization of the sum of a smooth (nonconvex) function and a convex (nonsmooth) regularizer. The framework hinges on Successive Convex Approximation (SCA) techniques and on a novel probabilistic model which describes in a unified way a variety of asynchronous settings in a more faithful and exhaustive way with respect to state-of-the-art models. Key features of our framework are: i) it accommodates inconsistent read, meaning that components of the variables may be written by some cores while being simultaneously read by others; ii) it covers in a unified way several existing methods; and iii) it accommodates a variety of parallel computing architectures. Almost sure convergence to stationary solutions is proved for the general case, and iteration complexity analysis is given for a specific version of our model. Numerical results show that our scheme outperforms existing asynchronous ones. Loris Cannelli, Francisco Facchinei, Vyacheslav Kungurtsev, Gesualdo Scutari |
ICASSP | 3 |
| 2017 | Capacity sensitivity in additive non-Gaussian noise channelsabstractIn this paper, a new framework based on the notion of capacity sensitivity is introduced to study the capacity of continuous memoryless point-to-point channels. The capacity sensitivity reflects how the capacity changes with small perturbations in any of the parameters describing the channel, even when the capacity is not available in closed-form. This includes perturbations of the cost constraints on the input distribution as well as on the channel distribution. The framework is based on continuity of the capacity, which is shown for a class of perturbations in the cost constraint and the channel distribution. The continuity then forms the foundation for obtaining bounds on the capacity sensitivity. As an illustration, the capacity sensitivity bound is applied to obtain scaling laws when the support of additive α-stable noise is truncated. Malcolm Egan, Samir Perlaza, Vyacheslav Kungurtsev |
ISIT | 3 |