Ivan P. Yamshchikov

dblp:178/9094 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-3784-0671ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Theory of computation · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution
abstract
Efficiency and safety of Large Language Models (LLMs), among other factors, rely on the quality of tokenization.A good tokenizer not only improves inference speed and language understanding but also provides extra defense against jailbreak attacks and lowers the risk of hallucinations.In this work, we investigate the efficiency of code tokenization, in particular from the perspective of data source diversity.We demonstrate that code tokenizers are prone to producing unused, and thus under-trained, tokens due to the imbalance in repository and language diversity in the training data, as well as the dominance of source-specific, repetitive tokens that are often unusable in future inference.By modifying the BPE objective and introducing merge skipping, we implement different techniques under the name Source-Attributed BPE (SA-BPE) to regularize BPE training and minimize overfitting, thereby substantially reducing the number of under-trained tokens while maintaining the same inference procedure as with regular BPE.This provides an effective tool suitable for production use. pchizhov/sa-bpe
Pavel Chizhov, Egor Bogomolov, Ivan P. Yamshchikov
ACL (1)3
2026 app.build: A Production Framework for Scaling Agentic Prompt-to-App Generation with Environment Scaffolding
abstract
We present app.build (https://github.com/neondatabase/appdotbuild-agent), an open-source framework that improves LLM-based application generation through systematic validation and structured environments. Our approach combines multi-layered validation pipelines, stack-specific orchestration, and model-agnostic architecture, implemented across three reference stacks. Through evaluation on 30 generation tasks, we demonstrate that comprehensive validation achieves 73.3% viability rate with 30% reaching perfect quality scores, while open-weights models achieve 80.8% of closed-model performance when provided structured environments. The open-source framework has been adopted by the community, with over 3,000 applications generated to date. This work demonstrates that scaling reliable AI agents requires scaling environments, not just models -- providing empirical insights and complete reference implementations for production-oriented agent systems.
Evgenii Kniazev, Arseny Kravchenko, Igor Rekun, James Broadhead, Nikita Shamgunov, Pranav Kumar Sah, Pratik Nichite, Ivan P. Yamshchikov
SANER8
2025 ComicScene154: A Scene Dataset for Comic Analysis
abstract
Comics offer a compelling yet under-explored domain for computational narrative analysis, combining text and imagery in ways distinct from purely textual or audiovisual media.We introduce ComicScene154, a manually annotated dataset of scene-level narrative arcs derived from public-domain comic books spanning diverse genres.By conceptualizing comics as an abstraction for narrative-driven, multimodal data, we highlight their potential to inform broader research on multi-modal storytelling.To demonstrate the utility of Comic-Scene154, we present a scene segmentation baseline, providing an initial benchmark for future studies to build upon.Our results indicate that ComicScene154 constitutes a valuable resource for advancing computational methods in multimodal narrative understanding and expanding the scope of comic analysis within the Natural Language Processing community.
Sandro Paval, Pascal Meissner, Ivan P. Yamshchikov
EMNLP3
2025 Generalization potential of large language models
abstract
Abstract The rise of deep learning techniques and especially the advent of large language models (LLMs) intensified the discussions around possibilities that artificial intelligence with higher generalization capability entails. The range of opinions on the capabilities of LLMs is extremely broad: from equating language models with stochastic parrots to stating that they are already conscious. This paper represents an attempt to review LLM landscape in the context of their generalization capacity as an information theoretic property of those complex systems. We discuss the suggested theoretical explanations for generalization in LLMs and highlight possible mechanisms responsible for these generalization properties. Through an examination of existing literature and theoretical frameworks, we endeavor to provide insights into the mechanisms driving the generalization capacity of LLMs, thus contributing to a deeper understanding of their capabilities and limitations in natural language processing tasks.
Mikhail Budnikov, Anna Bykova, Ivan P. Yamshchikov
Neural Comput. Appl.3
2024 Vygotsky Distance: Measure for Benchmark Task Similarity
abstract
Evaluation plays a significant role in modern natural language processing. Most modern NLP benchmarks consist of arbitrary sets of tasks that neither guarantee any generalization potential for the model once applied outside the test set nor try to minimize the resource consumption needed for model evaluation. This paper presents a theoretical instrument and a practical algorithm to calculate similarity between benchmark tasks, we call this similarity measure “Vygotsky distance”. The core idea of this similarity measure is that it is based on relative performance of the “students” on a given task, rather that on the properties of the task itself. If two tasks are close to each other in terms of Vygotsky distance the models tend to have similar relative performance on them. Thus knowing Vygotsky distance between tasks one can significantly reduce the number of evaluation tasks while maintaining a high validation quality. Experiments on various benchmarks, including GLUE, SuperGLUE, CLUE, and RussianSuperGLUE, demonstrate that a vast majority of NLP benchmarks could be at least 40% smaller in terms of the tasks included. Most importantly, Vygotsky distance could also be used for the validation of new tasks thus increasing the generalization potential of the future NLP models.
Maxim K. Surkov, Ivan P. Yamshchikov
LREC/COLING2
2024 BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
abstract
Language models can greatly benefit from efficient tokenization.However, they still mostly utilize the classical Byte-Pair Encoding (BPE) algorithm, a simple and reliable method.BPE has been shown to cause such issues as undertrained tokens and sub-optimal compression that may affect the downstream performance.We introduce PickyBPE, a modified BPE algorithm that carries out vocabulary refinement during tokenizer training by removing merges that leave intermediate "junk" tokens.Our method improves vocabulary efficiency, eliminates under-trained tokens, and does not compromise text compression.Our experiments show that this method either improves downstream performance or does not harm it.
Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova, Ivan P. Yamshchikov
EMNLP4
2023 Rehabilitating Homeless: Dataset and Key Insights
abstract
This paper presents a large anonymized dataset of homelessness alongside insights into the data-driven rehabilitation of homeless people. The dataset was gathered by a large non-profit organization working on rehabilitating the homeless for twenty years. This is the first dataset that we know of that contains rich information on thousands of homeless individuals seeking rehabilitation. We show how data analysis can help to make the rehabilitation of homeless people more effective and successful. Thus, we hope this paper alerts the data science community to the problem of homelessness.
Anna Bykova, Nikolai Filippov, Ivan P. Yamshchikov
AAAI3
2023 Fine-tuning transformers: Vocabulary transfer
Vladislav D. Mosin, Igor Samenko, Borislav Kozlovskii, Alexey Tikhonov, Ivan P. Yamshchikov
Artif. Intell.5
2022 Moving Other Way: Exploring Word Mover Distance Extensions
abstract
The word mover's distance (WMD) is a popular semantic similarity metric for two texts. This position paper studies several possible extensions of WMD. We experiment with the frequency of words in the corpus as a weighting factor and the geometry of the word vector space. We validate possible extensions of WMD on six document classification datasets. Some proposed extensions show better results in terms of the k-nearest neighbor classification error than WMD.
Ilya S. Smirnov, Ivan P. Yamshchikov
COMPLEXIS2
2022 BERT in Plutarch's Shadows
abstract
The extensive surviving corpus of the ancient scholar Plutarch of Chaeronea (ca.45-120 CE) also contains several texts which, according to current scholarly opinion, did not originate with him and are therefore attributed to an anonymous author Pseudo-Plutarch.These include, in particular, the work Placita Philosophorum (Quotations and Opinions of the Ancient Philosophers), which is extremely important for the history of ancient philosophy.Little is known about the identity of that anonymous author and its relation to other authors from the same period.This paper presents a BERT language model for Ancient Greek.The model discovers previously unknown statistical properties relevant to these literary, philosophical, and historical problems and can shed new light on this authorship question.In particular, the Placita Philosophorum, together with one of the other Pseudo-Plutarch texts, shows similarities with the texts written by authors from an Alexandrian context (2nd/3rd century CE)."I do not need a friend who changes when I change and who nods when I nod; my shadow does that much better."(Plutarch, Quomodo adulator ab amico internoscatur 53b 10)
Ivan P. Yamshchikov, Alexey Tikhonov, Yorgos Pantis, Charlotte Schubert, Jürgen Jost
EMNLP1
2021 Style-transfer and Paraphrase: Looking for a Sensible Semantic Similarity Metric
abstract
The rapid development of such natural language processing tasks as style transfer, paraphrase, and machine translation often calls for the use of semantic similarity metrics. In recent years a lot of methods to measure the semantic similarity of two short texts were developed. This paper provides a comprehensive analysis for more than a dozen of such methods. Using a new dataset of fourteen thousand sentence pairs human-labeled according to their semantic similarity, we demonstrate that none of the metrics widely used in the literature is close enough to human judgment in these tasks. A number of recently proposed metrics provide comparable results, yet Word Mover Distance is shown to be the most reasonable solution to measure semantic similarity in reformulated texts at the moment.
Ivan P. Yamshchikov, Viacheslav Shibaev, Nikolay Khlebnikov, Alexey Tikhonov
AAAI1
2021 Artificial Neural Networks Jamming on the Beat
abstract
This paper addresses the issue of long-scale correlations that is characteristic for symbolic music and is a challenge for modern generative algorithms. It suggests a very simple workaround for this challenge, namely, generation of a drum pattern that could be further used as a foundation for melody generation. The paper presents a large dataset of drum patterns alongside with corresponding melodies. It explores two possible methods for drum pattern generation. Exploring a latent space of drum patterns one could generate new drum patterns with a given music style. Finally, the paper demonstrates that a simple artificial neural network could be trained to generate melodies corresponding with these drum patters used as inputs. Resulting system could be used for end-to-end generation of symbolic music with song-like structure and higher long-scale correlations between the notes.
Alexey Tikhonov, Ivan P. Yamshchikov
COMPLEXIS2
2020 Text Classification for Monolingual Political Manifestos with Words Out of Vocabulary
Arsenii Rasov, Ilya Obabkov, Eckehard Olbrich, Ivan P. Yamshchikov
COMPLEXIS4
2020 It Means More If It Sounds Good: Yet Another Hypotheses Concerning the Evolution of Polysemous Words
Ivan P. Yamshchikov, Nono S. C. Merleau, Igor Samenko, Jürgen Jost
COMPLEXIS1
2020 Paranoid Transformer: Reading Narrative of Madness as Computational Approach to Creativity
Yana Agafonova, Alexey Tikhonov, Ivan P. Yamshchikov
ICCC3
2020 Drum Beats and Where To Find Them: Sampling Drum Patterns from a Latent Space
Alexey Tikhonov, Ivan P. Yamshchikov
ICCC2
2019 Style Transfer for Texts: Retrain, Report Errors, Compare with Rewrites
abstract
Alexey Tikhonov, Viacheslav Shibaev, Aleksander Nagaev, Aigul Nugmanova, Ivan P. Yamshchikov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alexey Tikhonov, Viacheslav Shibaev, Aleksander Nagaev, Aigul Nugmanova, Ivan P. Yamshchikov
EMNLP/IJCNLP (1)5
2018 Elephants, Donkeys, and Colonel Blotto
abstract
This paper employs a novel method for the empirical analysis of political discourse and develops a model that demonstrates dynamics comparable with the empirical data. Applying a set of binary text classifiers based on convolutional neural networks, we label statements in the political programs of the Democratic and the Republican Party in the United States. Extending the framework of the Colonel Blotto game by a stochastic activation structure, we show that, under a simple learning rule, the simulated game exhibits dynamics that resemble the empirical data.
Ivan P. Yamshchikov, Sharwin Rezagholi
COMPLEXIS1
2018 Guess who? Multilingual Approach For The Automated Generation Of Author-Stylized Poetry
abstract
This paper addresses the problem of stylized text generation in a multilingual setup. A version of a language model based on a long short-term memory (LSTM) artificial neural network with extended phonetic and semantic embeddings is used for stylized poetry generation. The quality of the resulting poems generated by the network is estimated through bilingual evaluation understudy (BLEU), a survey and a new cross-entropy based metric that is suggested for the problems of such type. The experiments show that the proposed model consistently outperforms random sample and vanilla-LSTM baselines, humans also tend to associate machine generated texts with the target author.
Alexey Tikhonov, Ivan P. Yamshchikov
SLT2