VLDB 2026 Research / reviewers in the wild / expert
Ivan P. Yamshchikov
dblp:178/9094
· DBLP profile ↗
19ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-3784-0671ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Theory of computation · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Where Words Come: Efficient Regularization of Code Tokenizers Through Source AttributionabstractEfficiency and safety of Large Language Models (LLMs), among other factors, rely on the quality of tokenization.A good tokenizer not only improves inference speed and language understanding but also provides extra defense against jailbreak attacks and lowers the risk of hallucinations.In this work, we investigate the efficiency of code tokenization, in particular from the perspective of data source diversity.We demonstrate that code tokenizers are prone to producing unused, and thus under-trained, tokens due to the imbalance in repository and language diversity in the training data, as well as the dominance of source-specific, repetitive tokens that are often unusable in future inference.By modifying the BPE objective and introducing merge skipping, we implement different techniques under the name Source-Attributed BPE (SA-BPE) to regularize BPE training and minimize overfitting, thereby substantially reducing the number of under-trained tokens while maintaining the same inference procedure as with regular BPE.This provides an effective tool suitable for production use. pchizhov/sa-bpe Pavel Chizhov, Egor Bogomolov, Ivan P. Yamshchikov |
ACL (1) | 3 |
| 2026 | app.build: A Production Framework for Scaling Agentic Prompt-to-App Generation with Environment ScaffoldingabstractWe present app.build (https://github.com/neondatabase/appdotbuild-agent), an open-source framework that improves LLM-based application generation through systematic validation and structured environments. Our approach combines multi-layered validation pipelines, stack-specific orchestration, and model-agnostic architecture, implemented across three reference stacks. Through evaluation on 30 generation tasks, we demonstrate that comprehensive validation achieves 73.3% viability rate with 30% reaching perfect quality scores, while open-weights models achieve 80.8% of closed-model performance when provided structured environments. The open-source framework has been adopted by the community, with over 3,000 applications generated to date. This work demonstrates that scaling reliable AI agents requires scaling environments, not just models -- providing empirical insights and complete reference implementations for production-oriented agent systems. Evgenii Kniazev, Arseny Kravchenko, Igor Rekun, James Broadhead, Nikita Shamgunov, Pranav Kumar Sah, Pratik Nichite, Ivan P. Yamshchikov |
SANER | 8 |
| 2025 | ComicScene154: A Scene Dataset for Comic AnalysisabstractComics offer a compelling yet under-explored domain for computational narrative analysis, combining text and imagery in ways distinct from purely textual or audiovisual media.We introduce ComicScene154, a manually annotated dataset of scene-level narrative arcs derived from public-domain comic books spanning diverse genres.By conceptualizing comics as an abstraction for narrative-driven, multimodal data, we highlight their potential to inform broader research on multi-modal storytelling.To demonstrate the utility of Comic-Scene154, we present a scene segmentation baseline, providing an initial benchmark for future studies to build upon.Our results indicate that ComicScene154 constitutes a valuable resource for advancing computational methods in multimodal narrative understanding and expanding the scope of comic analysis within the Natural Language Processing community. Sandro Paval, Pascal Meissner, Ivan P. Yamshchikov |
EMNLP | 3 |
| 2025 | Generalization potential of large language modelsabstractAbstract The rise of deep learning techniques and especially the advent of large language models (LLMs) intensified the discussions around possibilities that artificial intelligence with higher generalization capability entails. The range of opinions on the capabilities of LLMs is extremely broad: from equating language models with stochastic parrots to stating that they are already conscious. This paper represents an attempt to review LLM landscape in the context of their generalization capacity as an information theoretic property of those complex systems. We discuss the suggested theoretical explanations for generalization in LLMs and highlight possible mechanisms responsible for these generalization properties. Through an examination of existing literature and theoretical frameworks, we endeavor to provide insights into the mechanisms driving the generalization capacity of LLMs, thus contributing to a deeper understanding of their capabilities and limitations in natural language processing tasks. Mikhail Budnikov, Anna Bykova, Ivan P. Yamshchikov |
Neural Comput. Appl. | 3 |
| 2024 | Vygotsky Distance: Measure for Benchmark Task SimilarityabstractEvaluation plays a significant role in modern natural language processing. Most modern NLP benchmarks consist of arbitrary sets of tasks that neither guarantee any generalization potential for the model once applied outside the test set nor try to minimize the resource consumption needed for model evaluation. This paper presents a theoretical instrument and a practical algorithm to calculate similarity between benchmark tasks, we call this similarity measure “Vygotsky distance”. The core idea of this similarity measure is that it is based on relative performance of the “students” on a given task, rather that on the properties of the task itself. If two tasks are close to each other in terms of Vygotsky distance the models tend to have similar relative performance on them. Thus knowing Vygotsky distance between tasks one can significantly reduce the number of evaluation tasks while maintaining a high validation quality. Experiments on various benchmarks, including GLUE, SuperGLUE, CLUE, and RussianSuperGLUE, demonstrate that a vast majority of NLP benchmarks could be at least 40% smaller in terms of the tasks included. Most importantly, Vygotsky distance could also be used for the validation of new tasks thus increasing the generalization potential of the future NLP models. Maxim K. Surkov, Ivan P. Yamshchikov |
LREC/COLING | 2 |
| 2024 | BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer TrainingabstractLanguage models can greatly benefit from efficient tokenization.However, they still mostly utilize the classical Byte-Pair Encoding (BPE) algorithm, a simple and reliable method.BPE has been shown to cause such issues as undertrained tokens and sub-optimal compression that may affect the downstream performance.We introduce PickyBPE, a modified BPE algorithm that carries out vocabulary refinement during tokenizer training by removing merges that leave intermediate "junk" tokens.Our method improves vocabulary efficiency, eliminates under-trained tokens, and does not compromise text compression.Our experiments show that this method either improves downstream performance or does not harm it. Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova, Ivan P. Yamshchikov |
EMNLP | 4 |
| 2023 | Rehabilitating Homeless: Dataset and Key InsightsabstractThis paper presents a large anonymized dataset of homelessness alongside insights into the data-driven rehabilitation of homeless people. The dataset was gathered by a large non-profit organization working on rehabilitating the homeless for twenty years. This is the first dataset that we know of that contains rich information on thousands of homeless individuals seeking rehabilitation. We show how data analysis can help to make the rehabilitation of homeless people more effective and successful. Thus, we hope this paper alerts the data science community to the problem of homelessness. Anna Bykova, Nikolai Filippov, Ivan P. Yamshchikov |
AAAI | 3 |
| 2023 | Fine-tuning transformers: Vocabulary transfer
Vladislav D. Mosin, Igor Samenko, Borislav Kozlovskii, Alexey Tikhonov, Ivan P. Yamshchikov |
Artif. Intell. | 5 |
| 2022 | Moving Other Way: Exploring Word Mover Distance ExtensionsabstractThe word mover's distance (WMD) is a popular semantic similarity metric for two texts. This position paper studies several possible extensions of WMD. We experiment with the frequency of words in the corpus as a weighting factor and the geometry of the word vector space. We validate possible extensions of WMD on six document classification datasets. Some proposed extensions show better results in terms of the k-nearest neighbor classification error than WMD. Ilya S. Smirnov, Ivan P. Yamshchikov |
COMPLEXIS | 2 |
| 2022 | BERT in Plutarch's ShadowsabstractThe extensive surviving corpus of the ancient scholar Plutarch of Chaeronea (ca.45-120 CE) also contains several texts which, according to current scholarly opinion, did not originate with him and are therefore attributed to an anonymous author Pseudo-Plutarch.These include, in particular, the work Placita Philosophorum (Quotations and Opinions of the Ancient Philosophers), which is extremely important for the history of ancient philosophy.Little is known about the identity of that anonymous author and its relation to other authors from the same period.This paper presents a BERT language model for Ancient Greek.The model discovers previously unknown statistical properties relevant to these literary, philosophical, and historical problems and can shed new light on this authorship question.In particular, the Placita Philosophorum, together with one of the other Pseudo-Plutarch texts, shows similarities with the texts written by authors from an Alexandrian context (2nd/3rd century CE)."I do not need a friend who changes when I change and who nods when I nod; my shadow does that much better."(Plutarch, Quomodo adulator ab amico internoscatur 53b 10) Ivan P. Yamshchikov, Alexey Tikhonov, Yorgos Pantis, Charlotte Schubert, Jürgen Jost |
EMNLP | 1 |
| 2021 | Style-transfer and Paraphrase: Looking for a Sensible Semantic Similarity MetricabstractThe rapid development of such natural language processing tasks as style transfer, paraphrase, and machine translation often calls for the use of semantic similarity metrics. In recent years a lot of methods to measure the semantic similarity of two short texts were developed. This paper provides a comprehensive analysis for more than a dozen of such methods. Using a new dataset of fourteen thousand sentence pairs human-labeled according to their semantic similarity, we demonstrate that none of the metrics widely used in the literature is close enough to human judgment in these tasks. A number of recently proposed metrics provide comparable results, yet Word Mover Distance is shown to be the most reasonable solution to measure semantic similarity in reformulated texts at the moment. Ivan P. Yamshchikov, Viacheslav Shibaev, Nikolay Khlebnikov, Alexey Tikhonov |
AAAI | 1 |
| 2021 | Artificial Neural Networks Jamming on the BeatabstractThis paper addresses the issue of long-scale correlations that is characteristic for symbolic music and is a challenge for modern generative algorithms. It suggests a very simple workaround for this challenge, namely, generation of a drum pattern that could be further used as a foundation for melody generation. The paper presents a large dataset of drum patterns alongside with corresponding melodies. It explores two possible methods for drum pattern generation. Exploring a latent space of drum patterns one could generate new drum patterns with a given music style. Finally, the paper demonstrates that a simple artificial neural network could be trained to generate melodies corresponding with these drum patters used as inputs. Resulting system could be used for end-to-end generation of symbolic music with song-like structure and higher long-scale correlations between the notes. Alexey Tikhonov, Ivan P. Yamshchikov |
COMPLEXIS | 2 |
| 2020 | Text Classification for Monolingual Political Manifestos with Words Out of Vocabulary
Arsenii Rasov, Ilya Obabkov, Eckehard Olbrich, Ivan P. Yamshchikov |
COMPLEXIS | 4 |
| 2020 | It Means More If It Sounds Good: Yet Another Hypotheses Concerning the Evolution of Polysemous Words
Ivan P. Yamshchikov, Nono S. C. Merleau, Igor Samenko, Jürgen Jost |
COMPLEXIS | 1 |
| 2020 | Paranoid Transformer: Reading Narrative of Madness as Computational Approach to Creativity
Yana Agafonova, Alexey Tikhonov, Ivan P. Yamshchikov |
ICCC | 3 |
| 2020 | Drum Beats and Where To Find Them: Sampling Drum Patterns from a Latent Space
Alexey Tikhonov, Ivan P. Yamshchikov |
ICCC | 2 |
| 2019 | Style Transfer for Texts: Retrain, Report Errors, Compare with RewritesabstractAlexey Tikhonov, Viacheslav Shibaev, Aleksander Nagaev, Aigul Nugmanova, Ivan P. Yamshchikov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Alexey Tikhonov, Viacheslav Shibaev, Aleksander Nagaev, Aigul Nugmanova, Ivan P. Yamshchikov |
EMNLP/IJCNLP (1) | 5 |
| 2018 | Elephants, Donkeys, and Colonel BlottoabstractThis paper employs a novel method for the empirical analysis of political discourse and develops a model that demonstrates dynamics comparable with the empirical data. Applying a set of binary text classifiers based on convolutional neural networks, we label statements in the political programs of the Democratic and the Republican Party in the United States. Extending the framework of the Colonel Blotto game by a stochastic activation structure, we show that, under a simple learning rule, the simulated game exhibits dynamics that resemble the empirical data. Ivan P. Yamshchikov, Sharwin Rezagholi |
COMPLEXIS | 1 |
| 2018 | Guess who? Multilingual Approach For The Automated Generation Of Author-Stylized PoetryabstractThis paper addresses the problem of stylized text generation in a multilingual setup. A version of a language model based on a long short-term memory (LSTM) artificial neural network with extended phonetic and semantic embeddings is used for stylized poetry generation. The quality of the resulting poems generated by the network is estimated through bilingual evaluation understudy (BLEU), a survey and a new cross-entropy based metric that is suggested for the problems of such type. The experiments show that the proposed model consistently outperforms random sample and vanilla-LSTM baselines, humans also tend to associate machine generated texts with the target author. Alexey Tikhonov, Ivan P. Yamshchikov |
SLT | 2 |