EDBT 2026 Demo / reviewers in the wild / expert
Ivan V. Oseledets
dblp:56/7175
· DBLP profile ↗
71ranked-venue papers
1as first author
49since 2021 · last 2026
0000-0003-2071-2163ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 15 since 2021Databases, data management, data science and information retrieval · 10 · 6 since 2021Systems, architecture and hardware · 5 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse AutoencodersabstractRecent LLMs like DeepSeek-R1 have demonstrated state-of-the-art performance by integrating deep thinking and complex reasoning during generation. However, the internal mechanisms behind these reasoning processes remain unexplored. We observe reasoning LLMs consistently use vocabulary associated with human reasoning processes. We hypothesize these words correspond to specific reasoning moments within the models' internal mechanisms. To test this hypothesis, we employ Sparse Autoencoders (SAEs), a technique for sparse decomposition of neural network activations into human-interpretable features. We introduce ReasonScore, an automatic metric to identify active SAE features during these reasoning moments. We perform manual and automatic interpretation of the features detected by our metric, and find those with activation patterns matching uncertainty, exploratory thinking, and reflection. Through steering experiments, we demonstrate that amplifying these features increases performance on reasoning-intensive benchmarks (+2.2%) while producing longer reasoning traces (+20.5%). Using the model diffing technique, we provide evidence that these features are present only in models with reasoning capabilities. Our work provides the first step towards a mechanistic understanding of reasoning in LLMs. Andrey V. Galichin, Alexey Dontsov, Polina Druzhinina, Anton Razzhigaev, Oleg Rogov, Elena Tutubalina, Ivan V. Oseledets |
AAAI | 7 |
| 2026 | MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video UnderstandingabstractModern Video Large Language Models (VLLMs) often rely on uniform frame sampling for video understanding, but this approach frequently fails to capture critical information due to frame redundancy and variations in video content. We propose MaxInfo, the first training-free method based on the maximum volume principle, which is available in Fast and Slow versions and a Chunk-based version that selects and retains the most representative frames from a video. By maximizing the geometric volume formed by selected embeddings, MaxInfo ensures that the chosen frames cover the most informative regions of the embedding space, effectively reducing redundancy while preserving diversity. This method enhances the quality of input representations and improves long video comprehension performance across benchmarks. For instance, MaxInfo achieves a 3.28% improvement on LongVideoBench and a 6.4% improvement on EgoSchema for LLaVA-Video-7B. Moreover, MaxInfo boosts LongVideoBench performance by 3.47% on LLaVA-Video-72B and 3.44% on MiniCPM4.5. The approach is simple to implement and works with existing VLLMs without the need for additional training and very lower latency, making it a practical and effective alternative to traditional uniform sampling methods. Our code are available at https://github.com/FusionBrainLab/MaxInfo.git Pengyi Li 0003, Irina Abdullaeva, Alexander Gambashidze, Andrey Kuznetsov, Ivan V. Oseledets |
WACV | 5 |
| 2026 | Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers
Andrei Chertkov, Artem Basharin, Mikhail Saygin, Evgeny Frolov, Stanislav Straupe, Ivan V. Oseledets |
Neurocomputing | 6 |
| 2026 | Quasi-random physics-informed neural networks
Tianchi Yu, Ivan V. Oseledets |
Neurocomputing | 2 |
| 2026 | T-MLA: A targeted multiscale log-exponential attack framework for neural image compression
Nikolay I. Kalmykov, Razan Dibo, Kaiyu Shen, Zhonghan Xu, Anh Huy Phan 0001, Yipeng Liu 0001, Ivan V. Oseledets |
Inf. Sci. | 7 |
| 2026 | Sinc Kolmogorov-Arnold network and its application for solving PDEs with singularities
Tianchi Yu, Jingwei Qiu, Ivan V. Oseledets |
Neural Networks | 4 |
| 2026 | TERM Model: Tensor Ring Mixture Model for Density EstimationabstractProbabilistic modeling is a core challenge in statistical machine learning. Tensor-based probabilistic graph methods address interpretability and stability concerns encountered in neural network approaches and allow tractable inference (e.g., marginal inference and conditional inference). In this paper, we introduce tensor ring decomposition for density estimation, which reduces the number of permutation candidates compared to existing methods, while simultaneously enhancing expressive power and maintaining tractable inference. Different non-negative strategies for density function results in two variants: Born TRDE offers simpler inference and sampling but with slightly lower accuracy, while Energy TRDE, though more complex, achieves superior performance. Furthermore, a mixture model that incorporates multiple permutation candidates with adaptive weights is designed, resulting in increased expressive flexibility and comprehensiveness. Unlike existing methods that focus on finding a single optimal permutation, our approach, inspired by ensemble learning, demonstrates that combining multiple suboptimal permutations can yield superior results. Experiments demonstrate that the proposed approach excels in estimating probability density functions and sampling, capturing intricate details with competitive or superior performance compared to existing state-of-the-art (SOTA) tractable density methods. Ruituo Wu, Jiani Liu 0002, Bing Li 0002, Anh Huy Phan 0001, Ivan V. Oseledets, Ce Zhu, Yipeng Liu 0001 |
IEEE Trans. Big Data | 5 |
| 2025 | Certification of Speaker Recognition Models to Additive PerturbationsabstractSpeaker recognition technology is applied to various tasks, from personal virtual assistants to secure access systems. However, the robustness of these systems against adversarial attacks, particularly to additive perturbations, remains a significant challenge. In this paper, we pioneer applying robustness certification techniques to speaker recognition, initially developed for the image domain. Our work covers this gap by transferring and improving randomized smoothing certification techniques against norm-bounded additive perturbations for classification and few-shot learning tasks to speaker recognition. We demonstrate the effectiveness of these methods on VoxCeleb 1 and 2 datasets for several models. We expect this work to improve the robustness of voice biometrics and accelerate the research of certification methods in the audio domain. Dmitrii Korzh, Elvir Karimov, Mikhail Pautov, Oleg Rogov, Ivan V. Oseledets |
AAAI | 5 |
| 2025 | Exploring the Hidden Capacity of LLMs for One-Step Text GenerationabstractA recent study showed that large language models (LLMs) can reconstruct surprisingly long texts -up to thousands of tokens -via autoregressive generation from just one trained input embedding.In this work, we explore whether autoregressive decoding is essential for such reconstruction.We show that frozen LLMs can generate hundreds of accurate tokens in just one token-parallel forward pass, when provided with only two learned embeddings.This reveals a surprising and underexplored multi-token generation capability of autoregressive LLMs.We examine these embeddings and characterize the information they encode.We also empirically show that, although these representations are not unique for a given text, they form connected and local regions in embedding space -suggesting the potential to train a practical encoder.The existence of such representations hints that multi-token generation may be natively accessible in off-the-shelf LLMs via a learned input encoder, eliminating heavy retraining and helping to overcome the fundamental bottleneck of autoregressive decoding while reusing already-trained models. Gleb Mezentsev, Ivan V. Oseledets |
EMNLP | 2 |
| 2025 | Associative memory and dead neuronsabstractIn ``Large Associative Memory Problem in Neurobiology and Machine Learning,'' Dmitry Krotov and John Hopfield introduced a general technique for the systematic construction of neural ordinary differential equations with non-increasing energy or Lyapunov function. We study this energy function and identify that it is vulnerable to the problem of dead neurons. Each point in the state space where the neuron dies is contained in a non-compact region with constant energy. In these flat regions, energy function alone does not completely determine all degrees of freedom and, as a consequence, can not be used to analyze stability or find steady states or basins of attraction. We perform a direct analysis of the dynamical system and show how to resolve problems caused by flat directions corresponding to dead neurons: (i) all information about the state vector at a fixed point can be extracted from the energy and Hessian matrix (of Lagrange function), (ii) it is enough to analyze stability in the range of Hessian matrix, (iii) if steady state touching flat region is stable the whole flat region is the basin of attraction. The analysis of the Hessian matrix can be complicated for realistic architectures, so we show that for a slightly altered dynamical system (with the same structure of steady states), one can derive a diverse family of Lyapunov functions that do not have flat regions corresponding to dead neurons. In addition, these energy functions allow one to use Lagrange functions with Hessian matrices that are not necessarily positive definite and even consider architectures with non-symmetric feedforward and feedback connections. Vladimir Fanaskov, Ivan V. Oseledets |
ICLR | 2 |
| 2025 | AI Diagnostic Assistant (AIDA): A Predictive Model for Diagnoses from Health Records in Clinical Decision Support SystemsabstractClinical Decision Support Systems (CDSS) play an increasingly important role in medical diagnostics. We present AI Diagnostic Assistant (AIDA), a real-time predictive model designed to assist doctors in interpreting patient conditions. AIDA analyzes electronic health records (EHR), including medical history, laboratory results, and complaints, to suggest potential diagnoses from 95 common conditions before the doctor makes the final decision. The model acts as a verification and backup tool, ensuring that no critical details are overlooked. Trained on 1.5 million patient records and validated on a dataset curated by a panel of experts, AIDA proves trustworthy as a diagnosis-making assistant (87.7% accuracy compared to 91.7% accuracy among doctors). Integrated into a megapolis-wide CDSS, AIDA has assisted doctors in over 3 million real-world diagnoses to date. Dmitry Umerenkov, Alexandr Nesterov, Vladimir Shaposhnikov, Ruslan Abramov, Nikolay Romanenko, Vladimir Kokh, Marina Kirina, Anton Abrosimov, Dmitry V. Dylov, Ivan V. Oseledets |
IJCAI | 10 |
| 2025 | A case study of spatiotemporal forecasting techniques for weather forecasting
Shakir Showkat Sofi, Ivan V. Oseledets |
GeoInformatica | 2 |
| 2025 | Fast gradient-free activation maximization for neurons in spiking neural networks
Nikita Pospelov, Andrei Chertkov, Maxim Beketov, Ivan V. Oseledets, Konstantin Anokhin |
Neurocomputing | 4 |
| 2025 | GLiRA: Closed-Box Membership Inference Attack via Knowledge DistillationabstractWhile Deep Neural Networks demonstrate remarkable performance in practical tasks, they are vulnerable to membership inference attacks aimed at identifying whether a certain object belongs to the training dataset. To conduct a membership inference attack on a target model, an adversary has to train a set of shadow models and conduct a statistical test to determine the membership status of the particular input object. Usually, shadow models are trained without taking into account the target model; we argue that utilizing the predictions of the target model can guide the training process of the shadow model. To improve the efficiency of shadow model-based membership inference attacks, we propose GLiRA, a knowledge distillation-guided approach to membership inference attacks. We observe that the knowledge distillation significantly improves the efficiency of a likelihood ratio membership inference attack when the architecture of the target model is both known and unknown to an attacker. We evaluate the proposed method across multiple image classification datasets and model architectures and demonstrate that knowledge distillation-guided likelihood ratio attack outperforms the current state-of-the-art membership inference attacks in the majority of experimental settings. Andrey V. Galichin, Mikhail Pautov, Alexey Zhavoronkin, Oleg Rogov, Ivan V. Oseledets |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Your Transformer is Secretly LinearabstractAnton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova, Nikolai Gerasimenko, Ivan Oseledets, Denis Dimitrov, Andrey Kuznetsov. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Anton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova, Nikolai Gerasimenko, Ivan V. Oseledets, Denis Dimitrov, Andrey Kuznetsov |
ACL (1) | 5 |
| 2024 | RECE: Reduced Cross-Entropy Loss for Large-Catalogue Sequential RecommendersabstractScalability is a major challenge in modern recommender systems. In sequential recommendations, full Cross-Entropy (CE) loss achieves state-of-the-art recommendation quality but consumes excessive GPU memory with large item catalogs, limiting its practicality. Using a GPU-efficient locality-sensitive hashing-like algorithm for approximating large tensor of logits, this paper introduces a novel RECE (REduced Cross-Entropy) loss. RECE significantly reduces memory consumption while allowing one to enjoy the state-of-the-art performance of full CE loss. Experimental results on various datasets show that RECE cuts training peak memory usage by up to 12 times compared to existing methods while retaining or exceeding performance metrics of CE loss. The approach also opens up new possibilities for large-scale applications in other domains. Danil Gusak, Gleb Mezentsev, Ivan V. Oseledets, Evgeny Frolov |
CIKM | 3 |
| 2024 | General Lipschitz: Certified Robustness Against Resolvable Semantic Transformations via Transformation-Dependent Randomized SmoothingabstractRandomized smoothing is the state-of-the-art approach to constructing image classifiers that are provably robust against additive adversarial perturbations of bounded magnitude. However, it is more complicated to compute reasonable certificates against semantic transformations (e.g., image blurring, translation, gamma correction) and their compositions. In this work, we propose General Lipschitz (GL), a new flexible framework to certify neural networks against resolvable semantic transformations. Within the framework, we analyze transformation-dependent Lipschitz-continuity of smoothed classifiers w.r.t. transformation parameters and derive corresponding robustness certificates. To assess the effectiveness of the proposed approach, we evaluate it on different image classification datasets against several state-of-the-art certification methods. Dmitrii Korzh, Mikhail Pautov, Olga Tsymboi, Ivan V. Oseledets |
ECAI | 4 |
| 2024 | SparseGrad: A Selective Method for Efficient Fine-tuning of MLP LayersabstractViktoriia A. Chekalina, Anna Rudenko, Gleb Mezentsev, Aleksandr Mikhalev, Alexander Panchenko, Ivan Oseledets. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Viktoria Chekalina, Anna Rudenko, Gleb Mezentsev, Aleksandr Mikhalev, Alexander Panchenko, Ivan V. Oseledets |
EMNLP | 6 |
| 2024 | Neural operators meet conjugate gradients: The FCG-NO method for efficient PDE solvingabstractDeep learning solvers for partial differential equations typically have limited accuracy. We propose to overcome this problem by using them as preconditioners. More specifically, we apply discretization-invariant neural operators to learn preconditioners for the flexible conjugate gradient method (FCG). Architecture paired with novel loss function and training scheme allows for learning efficient preconditioners that can be used across different resolutions. On the theoretical side, FCG theory allows us to safely use nonlinear preconditioners that can be applied in $O(N)$ operations without constraining the form of the preconditioners matrix. To justify learning scheme components (the loss function and the way training data is collected) we perform several ablation studies. Numerical results indicate that our approach favorably compares with classical preconditioners and allows to reuse of preconditioners learned for lower resolution to the higher resolution data. Alexander Rudikov, Vladimir Fanaskov, Ekaterina A. Muravleva, Yuri M. Laevsky, Ivan V. Oseledets |
ICML | 5 |
| 2024 | Probabilistically Robust Watermarking of Neural Networks
Mikhail Pautov, Nikita Bogdanov, Stanislav Pyatkin, Oleg Rogov, Ivan V. Oseledets |
IJCAI | 5 |
| 2024 | Anatomical Positional Embeddings
Mikhail Goncharov, Valentin Samokhin, Eugenia Soboleva, Roman Sokolov, Boris Shirokikh, Mikhail Belyaev, Anvar Kurmukov, Ivan V. Oseledets |
MICCAI (10) | 8 |
| 2024 | Self-Attentive Sequential Recommendations with Hyperbolic RepresentationsabstractIn recent years, self-attentive sequential learning models have surpassed conventional collaborative filtering techniques in next-item recommendation tasks. However, Euclidean geometry utilized in these models may not be optimal for capturing a complex structure of behavioral data. Building on recent advances in the application of hyperbolic geometry to collaborative filtering tasks, we propose a novel approach that leverages hyperbolic geometry in the sequential learning setting. Our approach replaces final output of the Euclidean models with a linear predictor in the non-linear hyperbolic space, which increases the representational capacity and improves recommendation quality. Evgeny Frolov, Tatyana Matveeva, Leyla Mirvakhabova, Ivan V. Oseledets |
RecSys | 4 |
| 2024 | Scalable Cross-Entropy Loss for Sequential Recommendations with Large Item CatalogsabstractScalability issue plays a crucial role in productionizing modern recommender systems. Even lightweight architectures may suffer from high computational overload due to intermediate calculations, limiting their practicality in real-world applications. Specifically, applying full Cross-Entropy (CE) loss often yields state-of-the-art performance in terms of recommendations quality. Still, it suffers from excessive GPU memory utilization when dealing with large item catalogs. This paper introduces a novel Scalable Cross-Entropy (SCE) loss function in the sequential learning setup. It approximates the CE loss for datasets with large-size catalogs, enhancing both time efficiency and memory usage without compromising recommendations quality. Unlike traditional negative sampling methods, our approach utilizes a selective GPU-efficient computation strategy, focusing on the most informative elements of the catalog, particularly those most likely to be false positives. This is achieved by approximating the softmax distribution over a subset of the model outputs through the maximum inner product search. Experimental results on multiple datasets demonstrate the effectiveness of SCE in reducing peak memory usage by a factor of up to 100 compared to the alternatives, retaining or even exceeding their metrics values. The proposed approach also opens new perspectives for large-scale developments in different domains, such as large language models. Gleb Mezentsev, Danil Gusak, Ivan V. Oseledets, Evgeny Frolov |
RecSys | 3 |
| 2024 | Quantization of Large Language Models with an Overdetermined BasisabstractIn this paper, we introduce an algorithm for data quantization based on the principles of Kashin representation. This approach hinges on decomposing any given vector, matrix, or tensor into two factors. The first factor maintains a small infinity norm, while the second exhibits a similarly constrained norm when multiplied by an orthogonal matrix. Surprisingly, the entries of factors after decomposition are well-concentrated around several peaks, which allows us to efficiently replace them with corresponding centroids for quantization purposes. We study the theoretical properties of the proposed approach and rigorously evaluate our compression algorithm in the context of next-word prediction tasks, employing models like OPT of varying sizes. Our findings demonstrate that Kashin Quantization achieves competitive quality in model performance while ensuring superior data compression, marking a significant advancement in the field of data quantization. Daniil Merkulov, Daria Cherniuk, Alexander Rudikov, Ivan V. Oseledets, Ekaterina A. Muravleva, Aleksandr Mikhalev, Boris Kashin |
UAI | 4 |
| 2024 | Multiparticle Kalman filter for object localization in symmetric environments
Roman Korkin, Ivan V. Oseledets, Alexandr Katrutsa |
Expert Syst. Appl. | 2 |
| 2024 | Quantization Aware Factorization for Deep Neural Network CompressionabstractTensor decomposition of convolutional and fully-connected layers is an effective way to reduce parameters and FLOP in neural networks. Due to memory and power consumption limitations of mobile or embedded devices, the quantization step is usually necessary when pre-trained models are deployed. A conventional post-training quantization approach applied to networks with decomposed weights yields a drop in accuracy. This motivated us to develop an algorithm that finds tensor approximation directly with quantized factors and thus benefit from both compression techniques while keeping the prediction quality of the model. Namely, we propose to use Alternating Direction Method of Multipliers (ADMM) for Canonical Polyadic (CP) decomposition with factors whose elements lie on a specified quantization grid. We compress neural network weights with a devised algorithm and evaluate it’s prediction quality and performance. We compare our approach to state-of-the-art post-training quantization methods and demonstrate competitive results and high flexibility in achiving a desirable quality-performance tradeoff. Daria Cherniuk, Stanislav Abukhovich, Anh Huy Phan 0001, Ivan V. Oseledets, Andrzej Cichocki, Julia Gusak |
J. Artif. Intell. Res. | 4 |
| 2024 | Federated privacy-preserving collaborative filtering for on-device next app prediction
Albert Sayapin, Gleb Balitskiy, Daniel Bershatsky, Alexandr Katrutsa, Evgeny Frolov, Alexey A. Frolov, Ivan V. Oseledets, Vitaliy Kharin |
User Model. User Adapt. Interact. | 7 |
| 2023 | Understanding DDPM Latent Codes Through Optimal Transport
Valentin Khrulkov, Gleb V. Ryzhakov, Andrei Chertkov, Ivan V. Oseledets |
ICLR | 4 |
| 2023 | Constructive TT-representation of the tensors given as index interaction functions with applications
Gleb V. Ryzhakov, Ivan V. Oseledets |
ICLR | 2 |
| 2023 | General Covariance Data Augmentation for Neural PDE SolversabstractThe growing body of research shows how to replace classical partial differential equation (PDE) integrators with neural networks. The popular strategy is to generate the input-output pairs with a PDE solver, train the neural network in the regression setting, and use the trained model as a cheap surrogate for the solver. The bottleneck in this scheme is the number of expensive queries of a PDE solver needed to generate the dataset. To alleviate the problem, we propose a computationally cheap augmentation strategy based on general covariance and simple random coordinate transformations. Our approach relies on the fact that physical laws are independent of the coordinate choice, so the change in the coordinate system preserves the type of a parametric PDE and only changes PDE’s data (e.g., initial conditions, diffusion coefficient). For tried neural networks and partial differential equations, proposed augmentation improves test error by 23% on average. The worst observed result is a 17% increase in test error for multilayer perceptron, and the best case is a 80% decrease for dilated residual network. Vladimir Fanaskov, Tianchi Yu, Alexander Rudikov, Ivan V. Oseledets |
ICML | 4 |
| 2023 | Few-bit Backward: Quantized Gradients of Activation Functions for Memory Footprint ReductionabstractMemory footprint is one of the main limiting factors for large neural network training. In backpropagation, one needs to store the input to each operation in the computational graph. Every modern neural network model has quite a few pointwise nonlinearities in its architecture, and such operations induce additional memory costs that, as we show, can be significantly reduced by quantization of the gradients. We propose a systematic approach to compute optimal quantization of the retained gradients of the pointwise nonlinear functions with only a few bits per each element. We show that such approximation can be achieved by computing an optimal piecewise-constant approximation of the derivative of the activation function, which can be done by dynamic programming. The drop-in replacements are implemented for all popular nonlinearities and can be used in any existing pipeline. We confirm the memory reduction and the same convergence on several open benchmarks. Georgii S. Novikov, Daniel Bershatsky, Julia Gusak, Alex Shonenkov, Denis Dimitrov, Ivan V. Oseledets |
ICML | 6 |
| 2023 | PROTES: Probabilistic Optimization with Tensor SamplingabstractWe developed a new method PROTES for black-box optimization, which is based on the probabilistic sampling from a probability density function given in the low-parametric tensor train format. We tested it on complex multidimensional arrays and discretized multivariable functions taken, among others, from real-world applications, including unconstrained binary optimization and optimal control problems, for which the possible number of elements is up to $2^{1000}$. In numerical experiments, both on analytic model functions and on complex problems, PROTES outperforms popular discrete optimization methods (Particle Swarm Optimization, Covariance Matrix Adaptation, Differential Evolution, and others). Anastasia Batsheva, Andrei Chertkov, Gleb V. Ryzhakov, Ivan V. Oseledets |
NeurIPS | 4 |
| 2023 | Neural Harmonics: Bridging Spectral Embedding and Matrix Completion in Self-Supervised LearningabstractSelf-supervised methods received tremendous attention thanks to their seemingly heuristic approach to learning representations that respect the semantics of the data without any apparent supervision in the form of labels. A growing body of literature is already being published in an attempt to build a coherent and theoretically grounded understanding of the workings of a zoo of losses used in modern self-supervised representation learning methods.
In this paper, we attempt to provide an understanding from the perspective of a Laplace operator and connect the inductive bias stemming from the augmentation process to a low-rank matrix completion problem.
To this end, we leverage the results from low-rank matrix completion to provide theoretical analysis on the convergence of modern SSL methods and a key property that affects their downstream performance. Marina Munkhoeva, Ivan V. Oseledets |
NeurIPS | 2 |
| 2023 | Efficient GPT Model Pre-training using Tensor Train Matrix Representation
Viktoria Chekalina, Georgiy Novikov, Julia Gusak, Alexander Panchenko, Ivan V. Oseledets |
PACLIC | 5 |
| 2023 | An event-triggered iteratively reweighted convex optimization approach to multi-period portfolio selection
Filipp Skomorokhov, Jun Wang 0002, G. V. Ovchinnikov, Evgeny Burnaev, Ivan V. Oseledets |
Expert Syst. Appl. | 5 |
| 2023 | Fast cross tensor approximation for image and video completion
Salman Ahmadi-Asl, Maame G. Asante-Mensah, Andrzej Cichocki, Anh Huy Phan 0001, Ivan V. Oseledets, Jun Wang 0002 |
Signal Process. | 5 |
| 2022 | CC-CERT: A Probabilistic Approach to Certify General Robustness of Neural NetworksabstractIn safety-critical machine learning applications, it is crucial to defend models against adversarial attacks --- small modifications of the input that change the predictions. Besides rigorously studied $\ell_p$-bounded additive perturbations, semantic perturbations (e.g. rotation, translation) raise a serious concern on deploying ML systems in real-world. Therefore, it is important to provide provable guarantees for deep learning models against semantically meaningful input transformations. In this paper, we propose a new universal probabilistic certification approach based on Chernoff-Cramer bounds that can be used in general attack settings. We estimate the probability of a model to fail if the attack is sampled from a certain distribution. Our theoretical findings are supported by experimental results on different datasets. Mikhail Pautov, Nurislam Tursynbek, Marina Munkhoeva, Nikita Muravev, Aleksandr Petiushko, Ivan V. Oseledets |
AAAI | 6 |
| 2022 | T4DT: Tensorizing Time for Learning Temporal 3D Visual Data
Mikhail Usvyatsov, Rafael Ballester-Ripoll, Lina Bashaeva, Konrad Schindler, Gonzalo Ferrer 0001, Ivan V. Oseledets |
BMVC | 6 |
| 2022 | Hyperbolic Vision Transformers: Combining Improvements in Metric LearningabstractMetric learning aims to learn a highly discriminative model encouraging the embeddings of similar classes to be close in the chosen metrics and pushed apart for dissimilar ones. The common recipe is to use an encoder to extract embeddings and a distance-based loss function to match the representations - usually, the Euclidean distance is utilized. An emerging interest in learning hyperbolic data embeddings suggests that hyperbolic geometry can be beneficial for natural data. Following this line of work, we propose a new hyperbolic-based model for metric learning. At the core of our method is a vision transformer with output embeddings mapped to hyperbolic space. These embeddings are directly optimized using modified pairwise cross-entropy loss. We evaluate the proposed model with six different formulations on four datasets achieving the new state-of-the-art performance. The source code is available at https://github.com/htdt/hyp_metric. Aleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe, Ivan V. Oseledets |
CVPR | 5 |
| 2022 | Survey on Efficient Training of Large Neural NetworksabstractModern Deep Neural Networks (DNNs) require significant memory to store weight, activations, and other intermediate tensors during training. Hence, many models don’t fit one GPU device or can be trained using only a small per-GPU batch size. This survey provides a systematic overview of the approaches that enable more efficient DNNs training. We analyze techniques that save memory and make good use of computation and communication resources on architectures with a single or several GPUs. We summarize the main categories of strategies and compare strategies within and across categories. Along with approaches proposed in the literature, we discuss available implementations. Julia Gusak, Daria Cherniuk, Alena Shilova, Alexandr Katrutsa, Daniel Bershatsky, Xunyi Zhao, Lionel Eyraud-Dubois, Oleh Shliazhko, Denis Dimitrov, Ivan V. Oseledets, Olivier Beaumont |
IJCAI | 10 |
| 2022 | Smoothed Embeddings for Certified Few-Shot LearningabstractRandomized smoothing is considered to be the state-of-the-art provable defense against adversarial perturbations. However, it heavily exploits the fact that classifiers map input objects to class probabilities and do not focus on the ones that learn a metric space in which classification is performed by computing distances to embeddings of class prototypes. In this work, we extend randomized smoothing to few-shot learning models that map inputs to normalized embeddings. We provide analysis of the Lipschitz continuity of such models and derive a robustness certificate against $\ell_2$-bounded perturbations that may be useful in few-shot learning scenarios. Our theoretical results are confirmed by experiments on different datasets. Mikhail Pautov, Olesya Kuznetsova, Nurislam Tursynbek, Aleksandr Petiushko, Ivan V. Oseledets |
NeurIPS | 5 |
| 2022 | TTOpt: A Maximum Volume Quantized Tensor Train-based Optimization and its Application to Reinforcement LearningabstractWe present a novel procedure for optimization based on the combination of efficient quantized tensor train representation and a generalized maximum matrix volume principle.We demonstrate the applicability of the new Tensor Train Optimizer (TTOpt) method for various tasks, ranging from minimization of multidimensional functions to reinforcement learning.Our algorithm compares favorably to popular gradient-free methods and outperforms them by the number of function evaluations or execution time, often by a significant margin. Konstantin Sozykin, Andrei Chertkov, Roman Schutski, Anh Huy Phan 0001, Andrzej Cichocki, Ivan V. Oseledets |
NeurIPS | 6 |
| 2022 | Geometry-Inspired Top-k Adversarial PerturbationsabstractThe brittleness of deep image classifiers to small adversarial input perturbations has been extensively studied in the last several years. However, the main objective of existing perturbations is primarily limited to change the correctly predicted Top-1 class by an incorrect one, which does not intend to change the Top-k prediction. In many digital real-world scenarios Top-k prediction is more relevant. In this work, we propose a fast and accurate method of computing Top-k adversarial examples as a simple multi-objective optimization. We demonstrate its efficacy and performance by comparing it to other adversarial example crafting techniques. Moreover, based on this method, we propose Top-k Universal Adversarial Perturbations, image-agnostic tiny perturbations that cause the true class to be absent among the Top-k prediction for the majority of natural images. We experimentally show that our approach outperforms baseline methods and even improves existing techniques of finding Universal Adversarial Perturbations. Nurislam Tursynbek, Aleksandr Petiushko, Ivan V. Oseledets |
WACV | 3 |
| 2021 | Adversarial Turing Patterns from Cellular Automata
Nurislam Tursynbek, Ilya Vilkoviskiy, Maria Sindeeva, Ivan V. Oseledets |
AAAI | 4 |
| 2021 | Latent Transformations via NeuralODEs for GAN-based Image EditingabstractRecent advances in high-fidelity semantic image editing heavily rely on the presumably disentangled latent spaces of the state-of-the-art generative models, such as Style-GAN. Specifically, recent works show that it is possible to achieve decent controllability of attributes in face images via linear shifts along with latent directions. Several recent methods address the discovery of such directions, implicitly assuming that the state-of-the-art GANs learn the latent spaces with inherently linearly separable attribute distributions and semantic vector arithmetic properties.In our work, we show that nonlinear latent code manipulations realized as flows of a trainable Neural ODE are beneficial for many practical non-face image domains with more complex non-textured factors of variation. In particular, we investigate a large number of datasets with known attributes and demonstrate that certain attribute manipulations are challenging to obtain with linear shifts only. Valentin Khrulkov, Leyla Mirvakhabova, Ivan V. Oseledets, Artem Babenko |
ICCV | 3 |
| 2021 | Functional Space Analysis of Local GAN ConvergenceabstractRecent work demonstrated the benefits of studying continuous-time dynamics governing the GAN training. However, this dynamics is analyzed in the model parameter space, which results in finite-dimensional dynamical systems. We propose a novel perspective where we study the local dynamics of adversarial training in the general functional space and show how it can be represented as a system of partial differential equations. Thus, the convergence properties can be inferred from the eigenvalues of the resulting differential operator. We show that these eigenvalues can be efficiently estimated from the target dataset before training. Our perspective reveals several insights on the practical tricks commonly used to stabilize GANs, such as gradient penalty, data augmentation, and advanced integration schemes. As an immediate practical benefit, we demonstrate how one can a priori select an optimal data augmentation strategy for a particular generation task. Valentin Khrulkov, Artem Babenko, Ivan V. Oseledets |
ICML | 3 |
| 2021 | Tensor-train density estimationabstractEstimation of probability density function from samples is one of the central problems in statistics and machine learning. Modern neural network-based models can learn high dimensional distributions but have problems with hyperparameter selection and are often prone to instabilities during training and inference. We propose a new efficient tensor train-based model for density estimation (TTDE). Such density parametrization allows exact sampling, calculation of cumulative and marginal density functions, and partition function. It also has very intuitive hyperparameters. We develop an efficient non-adversarial training procedure for TTDE based on the Riemannian optimization. Experimental results demonstrate the competitive performance of the proposed method in density estimation and sampling tasks, while TTDE significantly outperforms competitors in training speed. Georgii S. Novikov, Maxim Panov, Ivan V. Oseledets |
UAI | 3 |
| 2021 | Dynamic Modeling of User Preferences for Stable RecommendationsabstractIn domains where users tend to develop long-term preferences that do not change too frequently, the stability of recommendations is an important factor of the perceived quality of a recommender system. In such cases, unstable recommendations may lead to poor personalization experience and distrust, driving users away from a recommendation service. We propose an incremental learning scheme that mitigates such problems via the dynamic modeling approach. It incorporates a generalized matrix form of a partial differential equation integrator that yields a dynamic low-rank approximation of time-dependent matrices representing user preferences. The scheme allows extending the famous PureSVD approach to time-aware settings and significantly improves its stability without sacrificing the accuracy in standard top-n recommendations tasks. Oluwafemi Olaleke, Ivan V. Oseledets, Evgeny Frolov |
UMAP | 2 |
| 2021 | FREDE: Anytime Graph EmbeddingsabstractLow-dimensional representations, or embeddings , of a graph's nodes facilitate several practical data science and data engineering tasks. As such embeddings rely, explicitly or implicitly, on a similarity measure among nodes, they require the computation of a quadratic similarity matrix, inducing a tradeoff between space complexity and embedding quality. To date, no graph embedding work combines (i) linear space complexity, (ii) a nonlinear transform as its basis, and (iii) nontrivial quality guarantees. In this paper we introduce FREDE ( FREquent Directions Embedding ), a graph embedding based on matrix sketching that combines those three desiderata. Starting out from the observation that embedding methods aim to preserve the covariance among the rows of a similarity matrix, FREDE iteratively improves on quality while individually processing rows of a nonlinearly transformed PPR similarity matrix derived from a state-of-the-art graph embedding method and provides, at any iteration , column-covariance approximation guarantees in due course almost indistinguishable from those of the optimal approximation by SVD. Our experimental evaluation on variably sized networks shows that FREDE performs almost as well as SVD and competitively against state-of-the-art embedding methods in diverse data science tasks, even when it is based on as little as 10% of node similarities. Anton Tsitsulin, Marina Munkhoeva, Davide Mottin, Panagiotis Karras, Ivan V. Oseledets, Emmanuel Müller |
Proc. VLDB Endow. | 5 |
| 2020 | Hyperbolic Image EmbeddingsabstractComputer vision tasks such as image classification, image retrieval, and few-shot learning are currently dominated by Euclidean and spherical embeddings so that the final decisions about class belongings or the degree of similarity are made using linear hyperplanes, Euclidean distances, or spherical geodesic distances (cosine similarity). In this work, we demonstrate that in many practical scenarios, hyperbolic embeddings provide a better alternative. Valentin Khrulkov, Leyla Mirvakhabova, Evgeniya Ustinova, Ivan V. Oseledets, Victor S. Lempitsky |
CVPR | 4 |
| 2020 | Stable Low-Rank Tensor Decomposition for Compression of Convolutional Neural Network
Anh Huy Phan 0001, Konstantin Sobolev, Konstantin Sozykin, Dmitry Ermilov, Julia Gusak, Petr Tichavský, Valeriy Glukhov, Ivan V. Oseledets, Andrzej Cichocki |
ECCV (29) | 8 |
| 2020 | The Shape of Data: Intrinsic Distance for Data Distributions
Anton Tsitsulin, Marina Munkhoeva, Davide Mottin, Panagiotis Karras, Alexander M. Bronstein, Ivan V. Oseledets, Emmanuel Müller |
ICLR | 6 |
| 2020 | TT-TSDF: Memory-Efficient TSDF with Low-Rank Tensor Train DecompositionabstractIn this paper we apply the low-rank Tensor Train decomposition for compression and operations on 3D objects and scenes represented by volumetric distance functions. Our study shows that not only it allows for a very efficient compression of the high-resolution TSDF maps (up to three orders of magnitude of the original memory footprint at resolution of 5123), but also allows to perform TSDF-Fusion directly in the low-rank form. This can potentially enable much more efficient 3D mapping on low-power mobile and consumer robot platforms. Alexey I. Boyko, Mikhail Matrosov, Ivan V. Oseledets, Dzmitry Tsetserukou, Gonzalo Ferrer 0001 |
IROS | 3 |
| 2020 | Interpolation Technique to Speed Up Gradients Propagation in Neural ODEsabstractWe propose a simple interpolation-based method for the efficient approximation of gradients in neural ODE models. We compare it with reverse dynamic method (known in literature as “adjoint method”) to train neural ODEs on classification, density estimation and inference approximation tasks. We also propose a theoretical justification of our approach using logarithmic norm formalism. As a result, our method allows faster model training than the reverse dynamic method what was confirmed and validated by extensive numerical experiments for several standard benchmarks. Talgat Daulbaev, Alexandr Katrutsa, Larisa Markeeva, Julia Gusak, Andrzej Cichocki, Ivan V. Oseledets |
NeurIPS | 6 |
| 2020 | Performance of Hyperbolic Geometry Models on Top-N Recommendation TasksabstractWe introduce a simple autoencoder based on hyperbolic geometry for solving standard collaborative filtering problem. In contrast to many modern deep learning techniques, we build our solution using only a single hidden layer. Remarkably, even with such a minimalistic approach, we not only outperform the Euclidean counterpart but also achieve a competitive performance with respect to the current state-of-the-art. We additionally explore the effects of space curvature on the quality of hyperbolic models and propose an efficient data-driven method for estimating its optimal value. Leyla Mirvakhabova, Evgeny Frolov, Valentin Khrulkov, Ivan V. Oseledets, Alexander Tuzhilin |
RecSys | 4 |
| 2020 | Tensor Train Decomposition on TensorFlow (T3F)abstractTensor Train decomposition is used across many branches of machine learning. We present T3F—a library for Tensor Train decomposition based on TensorFlow. T3F supports GPU execution, batch processing, automatic differentiation, and versatile functionality for the Riemannian optimization framework, which takes into account the underlying manifold structure to construct efficient optimization methods. The library makes it easier to implement machine learning papers that rely on the Tensor Train decomposition. T3F includes documentation, examples and 94% test coverage. Alexander Novikov 0001, Pavel Izmailov, Valentin Khrulkov, Michael Figurnov, Ivan V. Oseledets |
J. Mach. Learn. Res. | 5 |
| 2020 | Tensor Networks for Latent Variable Analysis: Higher Order Canonical Polyadic DecompositionabstractThe canonical polyadic decomposition (CPD) is a convenient and intuitive tool for tensor factorization; however, for higher order tensors, it often exhibits high computational cost and permutation of tensor entries, and these undesirable effects grow exponentially with the tensor order. Prior compression of tensor in-hand can reduce the computational cost of CPD, but this is only applicable when the rank R of the decomposition does not exceed the tensor dimensions. To resolve these issues, we present a novel method for CPD of higher order tensors, which rests upon a simple tensor network of representative inter-connected core tensors of orders not higher than 3. For rigor, we develop an exact conversion scheme from the core tensors to the factor matrices in CPD and an iterative algorithm of low complexity to estimate these factor matrices for the inexact case. Comprehensive simulations over a variety of scenarios support the proposed approach. Anh Huy Phan 0001, Andrzej Cichocki, Ivan V. Oseledets, Giuseppe Giovanni Calvi, Salman Ahmadi-Asl, Danilo P. Mandic |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Generalized Tensor Models for Recurrent Neural Networks
Valentin Khrulkov, Oleksii Hrinchuk, Ivan V. Oseledets |
ICLR (Poster) | 3 |
| 2019 | PROVEN: Verifying Robustness of Neural Networks with a Probabilistic ApproachabstractWe propose a novel framework PROVEN to \textbf{PRO}babilistically \textbf{VE}rify \textbf{N}eural network’s robustness with statistical guarantees. PROVEN provides probability certificates of neural network robustness when the input perturbation follow distributional characterization. Notably, PROVEN is derived from current state-of-the-art worst-case neural network robustness verification frameworks, and therefore it can provide probability certificates with little computational overhead on top of existing methods such as Fast-Lin, CROWN and CNN-Cert. Experiments on small and large MNIST and CIFAR neural network models demonstrate our probabilistic approach can tighten up robustness certificate to around $1.8 \times$ and $3.5 \times$ with at least a $99.99%$ confidence compared with the worst-case robustness certificate by CROWN and CNN-Cert. Lily Weng, Lam M. Nguyen, Mark S. Squillante, Akhilan Boopathy, Ivan V. Oseledets, Luca Daniel |
ICML | 6 |
| 2019 | IceVisionSet: lossless video dataset collected on Russian winter roads with traffic sign annotationsabstractAbility of autonomous vehicles to operate in complex dynamic environments requires, among other things, fast and accurate perception of surroundings, which includes recognition and tracking of traffic signs.For development and testing of modern sophisticated computer vision systems large and diverse datasets are of the major importance. To test the robustness of algorithms, image data with different moving speeds, camera settings, lighting and weather conditions are especially important.In this work we present a comprehensive, lifelike dataset of traffic sign images collected on the Russian winter roads in varying conditions, which include different weather, camera exposure, illumination and moving speeds. The dataset was annotated in accordance with the Russian traffic code. Annotation results and images are published under open CC BY 4.0 license and can be downloaded from the project website: http://oscar.skoltech.ru/. Artem L. Pavlov, Pavel A. Karpyshev, G. V. Ovchinnikov, Ivan V. Oseledets, Dzmitry Tsetserukou |
ICRA | 4 |
| 2019 | HybridSVD: when collaborative information is not enoughabstractWe propose a new hybrid algorithm that allows incorporating both user and item side information within the standard collaborative filtering technique. One of its key features is that it naturally extends a simple PureSVD approach and inherits its unique advantages, such as highly efficient Lanczos-based optimization procedure, simplified hyper-parameter tuning and a quick folding-in computation for generating recommendations instantly even in highly dynamic online environments. The algorithm utilizes a generalized formulation of the singular value decomposition, which adds flexibility to the solution and allows imposing the desired structure on its latent space. Conveniently, the resulting model also admits an efficient and straightforward solution for the cold start scenario. We evaluate our approach on a diverse set of datasets and show its superiority over similar classes of hybrid models. Evgeny Frolov, Ivan V. Oseledets |
RecSys | 2 |
| 2018 | Art of Singular Vectors and Universal Adversarial PerturbationsabstractVulnerability of Deep Neural Networks (DNNs) to adversarial attacks has been attracting a lot of attention in recent studies. It has been shown that for many state of the art DNNs performing image classification there exist universal adversarial perturbations - image-agnostic perturbations mere addition of which to natural images with high probability leads to their misclassification. In this work we propose a new algorithm for constructing such universal perturbations. Our approach is based on computing the so-called (p, q)-singular vectors of the Jacobian matrices of hidden layers of a network. Resulting perturbations present interesting visual patterns, and by using only 64 images we were able to construct universal perturbations with more than 60 % fooling rate on the dataset consisting of 50000 images. We also investigate a correlation between the maximal singular value of the Jacobian matrix and the fooling rate of the corresponding singular vector, and show that the constructed perturbations generalize across networks. Valentin Khrulkov, Ivan V. Oseledets |
CVPR | 2 |
| 2018 | Expressive power of recurrent neural networks
Valentin Khrulkov, Alexander Novikov 0001, Ivan V. Oseledets |
ICLR (Poster) | 3 |
| 2018 | Geometry Score: A Method For Comparing Generative Adversarial NetworksabstractOne of the biggest challenges in the research of generative adversarial networks (GANs) is assessing the quality of generated samples and detecting various levels of mode collapse. In this work, we construct a novel measure of performance of a GAN by comparing geometrical properties of the underlying data manifold and the generated one, which provides both qualitative and quantitative means for evaluation. Our algorithm can be applied to datasets of an arbitrary nature and is not limited to visual data. We test the obtained metric on various real-life models and datasets and demonstrate that our method provides new insights into properties of GANs. Valentin Khrulkov, Ivan V. Oseledets |
ICML | 2 |
| 2018 | AA-ICP: Iterative Closest Point with Anderson AccelerationabstractIterative Closest Point (ICP) is a widely used method for performing scan-matching and registration. Being simple and robust, this method is still computationally expensive and may be challenging to use in real-time applications with limited resources on mobile platforms. In this paper we propose a novel effective method for acceleration of ICP which does not require substantial modifications to the existing code. This method is based on an idea of Anderson acceleration which is an iterative procedure for finding a fixed point of contractive mapping. The latter is often faster than a standard Picard iteration, usually used in ICP implementations. We show that ICP, being a fixed point problem, can be significantly accelerated by this method enhanced by heuristics to improve overall robustness. We implement proposed approach into Point Cloud Library (PCL) and make it available online. Benchmarking on the real-world data fully supports our claims. Artem L. Pavlov, G. V. Ovchinnikov, Dmitry Yu. Derbyshev, Dzmitry Tsetserukou, Ivan V. Oseledets |
ICRA | 5 |
| 2018 | Quadrature-based features for kernel approximationabstractWe consider the problem of improving kernel approximation via randomized feature maps. These maps arise as Monte Carlo approximation to integral representations of kernel functions and scale up kernel methods for larger datasets. Based on an efficient numerical integration technique, we propose a unifying approach that reinterprets the previous random features methods and extends to better estimates of the kernel approximation. We derive the convergence behavior and conduct an extensive empirical study that supports our hypothesis. Marina Munkhoeva, Yermek Kapushev, Evgeny Burnaev, Ivan V. Oseledets |
NeurIPS | 4 |
| 2017 | Riemannian Optimization for Skip-Gram Negative SamplingabstractSkip-Gram Negative Sampling (SGNS) word embedding model, well known by its implementation in "word2vec" software, is usually optimized by stochastic gradient descent.However, the optimization of SGNS objective can be viewed as a problem of searching for a good matrix with the low-rank constraint.The most standard way to solve this type of problems is to apply Riemannian optimization framework to optimize the SGNS objective over the manifold of required low-rank matrices.In this paper, we propose an algorithm that optimizes SGNS objective using Riemannian optimization and demonstrates its superiority over popular competitors, such as the original method to train SGNS and SVD over SPPMI matrix. Alexander Fonarev, Oleksii Hrinchuk, Gleb Gusev, Pavel Serdyukov, Ivan V. Oseledets |
ACL (1) | 5 |
| 2016 | Efficient Rectangular Maximal-Volume Algorithm for Rating Elicitation in Collaborative FilteringabstractCold start problem in Collaborative Filtering can be solved by asking new users to rate a small seed set of representative items or by asking representative users to rate a new item. The question is how to build a seed set that can give enough preference information for making good recommendations. One of the most successful approaches, called Representative Based Matrix Factorization, is based on Maxvol algorithm. Unfortunately, this approach has one important limitation — a seed set of a particular size requires a rating matrix factorization of fixed rank that should coincide with that size. This is not necessarily optimal in the general case. In the current paper, we introduce a fast algorithm for an analytical generalization of this approach that we call Rectangular Maxvol. It allows the rank of factorization to be lower than the required size of the seed set. Moreover, the paper includes the theoretical analysis of the method's error, the complexity analysis of the existing methods and the comparison to the state-of-the-art approaches. Alexander Fonarev, Alexander Mikhalev, Pavel Serdyukov, Gleb Gusev, Ivan V. Oseledets |
ICDM | 5 |
| 2016 | Fifty Shades of Ratings: How to Benefit from a Negative Feedback in Top-N Recommendations TasksabstractConventional collaborative filtering techniques treat a top-n recommendations problem as a task of generating a list of the most relevant items. This formulation, however, disregards an opposite -- avoiding recommendations with completely irrelevant items. Due to that bias, standard algorithms, as well as commonly used evaluation metrics, become insensitive to negative feedback. In order to resolve this problem we propose to treat user feedback as a categorical variable and model it with users and items in a ternary way. We employ a third-order tensor factorization technique and implement a higher order folding-in method to support online recommendations. The method is equally sensitive to entire spectrum of user ratings and is able to accurately predict relevant items even from a negative only feedback. Our method may partially eliminate the need for complicated rating elicitation process as it provides means for personalized recommendations from the very beginning of an interaction with a recommender system. We also propose a modification of standard metrics which helps to reveal unwanted biases and account for sensitivity to a negative feedback. Our model achieves state-of-the-art quality in standard recommendation tasks while significantly outperforming other methods in the cold-start "no-positive-feedback" scenarios. Evgeny Frolov, Ivan V. Oseledets |
RecSys | 2 |
| 2015 | Enabling High-Dimensional Hierarchical Uncertainty Quantification by ANOVA and Tensor-Train DecompositionabstractHierarchical uncertainty quantification can reduce the computational cost of stochastic circuit simulation by employing spectral methods at different levels. This paper presents an efficient framework to simulate hierarchically some challenging stochastic circuits/systems that include high-dimensional subsystems. Due to the high parameter dimensionality, it is challenging to both extract surrogate models at the low level of the design hierarchy and to handle them in the high-level simulation. In this paper, we develop an efficient analysis of variance-based stochastic circuit/microelectromechanical systems simulator to efficiently extract the surrogate models at the low level. In order to avoid the curse of dimensionality, we employ tensor-train decomposition at the high level to construct the basis functions and Gauss quadrature points. As a demonstration, we verify our algorithm on a stochastic oscillator with four MEMS capacitors and 184 random parameters. This challenging example is efficiently simulated by our simulator at the cost of only 10min in MATLAB on a regular personal computer. Zheng Zhang 0005, Xiu Yang, Ivan V. Oseledets, George Em Karniadakis, Luca Daniel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Improved n-Term Karatsuba-Like Formulas in GF(2)abstractIt is well known that Chinese Remainder Theorem (CRT) can be used to construct efficient algorithms for multiplication of polynomials over GF(2). In this note, we show how to select an appropriate set of modulus polynomials to obtain minimal number of multiplications. Ivan V. Oseledets |
IEEE Trans. Computers | 1 |