Andrey Kuznetsov

dblp:50/11063 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 NoReGeo: Non-Reasoning Geometry Benchmark
abstract
We present NoReGeo, a novel benchmark designed to evaluate the intrinsic geometric understanding of large language models (LLMs) without relying on reasoning or algebraic computation. Unlike existing benchmarks that primarily assess models' proficiency in reasoning-based geometry-where solutions are derived using algebraic methods-NoReGeo focuses on evaluating whether LLMs can inherently encode spatial relationships and recognize geometric properties directly. Our benchmark comprises 2,500 trivial geometric problems spanning 25 categories, each carefully crafted to be solvable purely through native geometric understanding, assuming known object locations. We assess a range of state-of-the-art models on NoReGeo, including frontier models like GPT-4, observing that even the most advanced systems achieve an overall maximum of 65% accuracy in binary classification tasks. Further, our ablation experiments demonstrate that such geometric understanding does not emerge through fine-tuning alone, indicating that effective training for geometric comprehension requires a specialized approach from the outset. Our findings highlight a significant gap in current LLMs' ability to natively grasp geometric concepts, providing a foundation for future research toward models with true geometric cognition.
Irina Abdullaeva, Anton Vasiliuk, Elizaveta Goncharova, Temurbek Rahmatullaev, Zagorulko Ivan, Maxim Kurkin, Andrey Kuznetsov
AAAI7
2026 BREPS: Bounding-Box Robustness Evaluation of Promptable Segmentation
abstract
Promptable segmentation models such as SAM have established a powerful paradigm, enabling strong generalization to unseen objects and domains with minimal user input, including points, bounding boxes, and text prompts. Among these, bounding boxes stand out as particularly effective, often outperforming points while significantly reducing annotation costs. However, current training and evaluation protocols typically rely on synthetic prompts generated through simple heuristics, offering limited insight into real-world robustness. In this paper, we investigate the robustness of promptable segmentation models to natural variations in bounding box prompts. First, we conduct a controlled user study and collect thousands of real bounding box annotations. Our analysis reveals substantial variability in segmentation quality across users for the same model and instance, indicating that SAM-like models are highly sensitive to natural prompt noise. Then, since exhaustive testing of all possible user inputs is computationally prohibitive, we reformulate robustness evaluation as a white-box optimization problem over the bounding box prompt space. We introduce BREPS, a method for generating adversarial bounding boxes that minimize or maximize segmentation error while adhering to naturalness constraints. Finally, we benchmark state-of-the-art models across 10 datasets, spanning everyday scenes to medical imaging.
Andrey Moskalenko, Danil Kuznetsov, Irina Dudko, Anastasiia Iasakova, Nikita Boldyrev, Denis Shepelev, Andrei Spiridonov, Andrey Kuznetsov, Vlad Shakhuro
AAAI8
2026 T-LoRA: Single Image Diffusion Model Customization Without Overfitting
abstract
While diffusion model fine-tuning offers a powerful approach for customizing pre-trained models to generate specific objects, it frequently suffers from overfitting when training samples are limited, compromising both generalization capability and output diversity. This paper tackles the challenging yet most impactful task of adapting a diffusion model using just a single concept image, as single-image customization holds the greatest practical potential. We introduce T-LoRA, a Timestep-Dependent Low-Rank Adaptation framework specifically designed for diffusion model personalization. In our work we show that higher diffusion timesteps are more prone to overfitting than lower ones, necessitating a timestep-sensitive fine-tuning strategy. T-LoRA incorporates two key innovations: (1) a dynamic fine-tuning strategy that adjusts rank-constrained updates based on diffusion timesteps, and (2) a weight parametrization technique that ensures independence between adapter components through orthogonal initialization. Extensive experiments show that T-LoRA and its individual components outperform standard LoRA and other diffusion model personalization techniques. They achieve a superior balance between concept fidelity and text alignment, highlighting the potential of T-LoRA in data-limited and resource-constrained scenarios.
Vera Soboleva, Aibek Alanov, Andrey Kuznetsov, Konstantin Sobolev
AAAI3
2026 Fast and Accurate Fisher-Guided Quantization via Efficient Kronecker Factorization
abstract
Quantization has shown strong results in preserving model quality under compression.However, under aggressive bit-width reductions, even quantization may require additional information to prevent performance degradation.A natural source of it is the second-order curvature information, captured by the Hessian.Since the Hessian of the model layers is prohibitively large, direct computation is infeasible, making structured parameterizations and approximations crucial in practice.In this work, we propose an efficient Kroneckerfactored approximation yielding state-of-theart performance when integrated into existing quantization schemes.Evaluations on the LLaMA and Qwen model families show near-baseline quality at 4-bit compression and only a 5-6% degradation at 2-bit for models with 7-8B parameters.Moreover, our method substantially accelerates the most expensive component in second-order quantization -Hessian parameterization -achieving up to a 10× speedup over prior approaches.Quantized model checkpoints are available at https://huggingface.co/collections/ timo13113/fastkron-collection.
Viktoria Chekalina, Gerasin Timofey, Andrey Kuznetsov, Evgeny Frolov
ACL (1)3
2026 Feature Inversion as a Lens on Vision Encoders
abstract
Vision encoders power modern vision-only and vision-language systems, yet the geometry of their internal features remains opaque. In this work, we introduce a simple, general approach for vision latent analysis: reconstruct images from frozen encoder features and treat reconstructability as a proxy for retained information and feature organization. Concretely, we train a lightweight reconstructor to invert feature tensors and use it to compare various vision encoders — CLIP-based ViT, SigLIP, SAM, and InternViT. We rank models by the informativeness of their features and observe consistent gains with image-centric objectives and higher spatial resolution. Beyond measurement, controlled manipulations in feature space produce predictable pixel-level edits: orthogonal rotations (rather than spatial transformations) implement channel permutations and drive systematic color changes; linear contractions implement channel suppression; and a learned linear map enables plausible colorization of grayscale inputs. VLM-based experiments confirm that feature-space color swaps translate into semantic color changes in reconstructions. Our approach is encoder-agnostic in principle (demonstrated on ViT-based models), requires only access to features, and offers a practical diagnostic of what encoders remember, how that information is organized, and how it can be manipulated.
Eduard Allakhverdov, Dmitrii Tarasov, Elizaveta Goncharova, Andrey Kuznetsov
WACV4
2026 MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
abstract
Modern Video Large Language Models (VLLMs) often rely on uniform frame sampling for video understanding, but this approach frequently fails to capture critical information due to frame redundancy and variations in video content. We propose MaxInfo, the first training-free method based on the maximum volume principle, which is available in Fast and Slow versions and a Chunk-based version that selects and retains the most representative frames from a video. By maximizing the geometric volume formed by selected embeddings, MaxInfo ensures that the chosen frames cover the most informative regions of the embedding space, effectively reducing redundancy while preserving diversity. This method enhances the quality of input representations and improves long video comprehension performance across benchmarks. For instance, MaxInfo achieves a 3.28% improvement on LongVideoBench and a 6.4% improvement on EgoSchema for LLaVA-Video-7B. Moreover, MaxInfo boosts LongVideoBench performance by 3.47% on LLaVA-Video-72B and 3.44% on MiniCPM4.5. The approach is simple to implement and works with existing VLLMs without the need for additional training and very lower latency, making it a practical and effective alternative to traditional uniform sampling methods. Our code are available at https://github.com/FusionBrainLab/MaxInfo.git
Pengyi Li 0003, Irina Abdullaeva, Alexander Gambashidze, Andrey Kuznetsov, Ivan V. Oseledets
WACV4
2024 Your Transformer is Secretly Linear
abstract
Anton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova, Nikolai Gerasimenko, Ivan Oseledets, Denis Dimitrov, Andrey Kuznetsov. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Anton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova, Nikolai Gerasimenko, Ivan V. Oseledets, Denis Dimitrov, Andrey Kuznetsov
ACL (1)7
2019 Digital video forgery detection based on statistical features calculation
abstract
Fake digital information is distributed heavily nowadays using social networks, new and other information sources. Digital forgeries use may lead to an unexpected result and it is quite difficult to detect tampering with just an expert view. A lot of algorithms for digital image forgery detection exist, but video forgery detection is on its early development stage. We propose a new approach for digital video forgery detection, which is based on statistical features calculation on difference shift frames. We selected three types of features for research: CC-PEV, SPAM and MP-486. We also estimated the quality of several classification techniques to detect altered frames: RBF-based SVM, linear ensemble classifier and decision tree. The experimental results showed the best combination of feature and classification algorithms for the video forgery detection problem solution.
Andrey Kuznetsov
ICMV1
2018 Person reidentification on video surveillance data
abstract
Person reidentification is a very challenging problem nowadays because of a big amount of video surveillance systems used. The data from such systems is processed to analyze events or emergency situations, find specific people, etc. One of the ways of solving the problem of an area security is the development of person reidentification algorithms. In this paper we propose an algorithm for person reidentification based on RGB histogram features calculation. On the first stage HOG descriptor is selected to detect a person on an image. Then we used k-means++ clustering algorithm to remove background on a person image. Finally, Bayes and SVM classification methods were used for person reidentification. Experimental results showed that the proposed solution can be used for person reidentification with high precision (not less than 82%). To carry out research 3D People Surveillance Dataset was used.
Andrey Kuznetsov
ICMV1
2015 Approach to building a web-based expert system interface and its application for software provisioning in clouds
abstract
This paper focuses on a generalized approach to providing user interface to a web-based expert system (WBES).We examine MVC and MVP design patterns used traditionally to construct a web application user interface.In order to leverage the strength of the MVC/MVP design patterns we propose a special ontology representing a user communication domain.We describe a self-service networked infrastructure for automatic deployment of command line interface (CLI) applications.We demonstrate how to apply the proposed ontology for the design of a WBES aimed at supporting client software re-execution in clouds.In particular, we address the problems existing in the area of software development for music information retrieval algorithms implementation.
Evgeny Pyshkin, Andrey Kuznetsov
FedCSIS2
2015 An evaluation of popular hyperspectral images classification approaches
abstract
This work is devoted to the problem of the best hyperspectral images classification algorithm selection. The following algorithms are used for comparison: decision tree using full cross-validation; decision tree C 4.5; Bayesian classifier; maximum-likelihood method; MSE minimization classifier, including a special case – classification by conjugation; spectral angle classifier (for empirical mean and nearest neighbor), spectral mismatch classifier and support vector machine (SVM). There are used AVIRIS and SpecTIR hyperspectral images to conduct experiments.
Andrey Kuznetsov, Vladislav V. Myasnikov
ICMV1
2013 An Approach for Developing a Mobile Accessed Music Search Integration Platform
Marina Purgina, Andrey Kuznetsov, Evgeny Pyshkin
FedCSIS2