VLDB 2026 Research / reviewers in the wild / expert
Irina Piontkovskaya
dblp:211/7823
· DBLP profile ↗
14ranked-venue papers
0as first author
12since 2021 · last 2026
0009-0003-0299-5849ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RuBIN: A Russian Benchmark for Evaluating LLMs with Cultural Insights
Polina Lazukova, Irina Piontkovskaya |
LREC | 2 |
| 2025 | LAMeD: LLM-generated Annotations for Memory Leak DetectionabstractStatic analysis tools are widely used to detect software bugs and vulnerabilities but often struggle with scalability and efficiency in complex codebases. Traditional approaches rely on manually crafted annotations—labeling functions as sources or sinks—to track data flows, e.g., ensuring that allocated memory is eventually freed, and code analysis tools such as CodeQL, Infer, or Cooddy can use function specifications, but manual annotation is laborious and error-prone, especially for large or third-party libraries. We present LAMeD (LLM-generated Annotations for Memory leak Detection), a novel approach that leverages large language models (LLMs) to automatically generate function-specific annotations. When integrated with analyzers such as Cooddy, LAMeD significantly improves memory leak detection and reduces path explosion. We also suggest directions for extending LAMeD to broader code analysis. Ekaterina Shemetova, Ivan Smirnov, Anton Alekseev 0001, Ilya Shenbin, Alexey D. Rukhovich, Sergey I. Nikolenko, Vadim Lomshakov, Irina Piontkovskaya |
EASE | 8 |
| 2025 | Quantifying Logical Consistency in Transformers via Query-Key AlignmentabstractEduard Tulchinskii, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Eduard Tulchinskii, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov |
EMNLP | 5 |
| 2023 | GEC-DePenD: Non-Autoregressive Grammatical Error Correction with Decoupled Permutation and DecodingabstractKonstantin Yakovlev, Alexander Podolskiy, Andrey Bout, Sergey Nikolenko, Irina Piontkovskaya. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Konstantin Yakovlev, Alexander Podolskiy, Andrey Bout, Sergey I. Nikolenko, Irina Piontkovskaya |
ACL (1) | 5 |
| 2023 | Efficient Grammatical Error Correction Via Multi-Task Training and Optimized Training ScheduleabstractProgress in neural grammatical error correction (GEC) is hindered by the lack of annotated training data.Sufficient amounts of highquality manually annotated data are not available, so recent research has relied on generating synthetic data, pretraining on it, and then finetuning on real datasets; performance gains have been achieved either by ensembling or by using huge pretrained models such as XXL-T5 as the backbone.In this work, we explore an orthogonal direction: how to use available data more efficiently.First, we propose auxiliary tasks that exploit the alignment between the original and corrected sentences, such as predicting a sequence of corrections.We formulate each task as a sequence-to-sequence problem and perform multi-task training.Second, we discover that the order of datasets used for training and even individual instances within a dataset may have important effects on the final performance, so we set out to find the best training schedule.Together, these two ideas lead to significant improvements, producing results that improve state of the art with much smaller models; in particular, we outperform the best models based on T5-XXL (11B parameters) with a BART-based model (400M parameters). Andrey Bout, Alexander Podolskiy, Sergey I. Nikolenko, Irina Piontkovskaya |
EMNLP | 4 |
| 2023 | Topological Data Analysis for Speech ProcessingabstractInternational audience Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Daniil Cherniavskii, Serguei Barannikov, Irina Piontkovskaya, Sergey I. Nikolenko, Evgeny Burnaev |
INTERSPEECH | 6 |
| 2023 | Intrinsic Dimension Estimation for Robust Detection of AI-Generated TextsabstractRapidly increasing quality of AI-generated content makes it difficult to distinguish between human and AI-generated texts, which may lead to undesirable consequences for society. Therefore, it becomes increasingly important to study the properties of human texts that are invariant over text domains and various proficiency of human writers, can be easily calculated for any language, and can robustly separate natural and AI-generated texts regardless of the generation model and sampling method. In this work, we propose such an invariant of human texts, namely the intrinsic dimensionality of the manifold underlying the set of embeddings of a given text sample. We show that the average intrinsic dimensionality of fluent texts in natural language is hovering around the value $9$ for several alphabet-based languages and around $7$ for Chinese, while the average intrinsic dimensionality of AI-generated texts for each language is $\approx 1.5$ lower, with a clear statistical separation between human-generated and AI-generated distributions. This property allows us to build a score-based artificial text detector. The proposed detector's accuracy is stable over text domains, generator models, and human writer proficiency levels, outperforming SOTA detectors in model-agnostic and cross-domain scenarios by a significant margin. Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Daniil Cherniavskii, Sergey I. Nikolenko, Evgeny Burnaev, Serguei Barannikov, Irina Piontkovskaya |
NeurIPS | 8 |
| 2023 | Sinkhorn Transformations for Single-Query Postprocessing in Text-Video RetrievalabstractA recent trend in multimodal retrieval is related to postprocessing test set results via the dual-softmax loss (DSL). While this approach can bring significant improvements, it usually presumes that an entire matrix of test samples is available as DSL input. This work introduces a new postprocessing approach based on Sinkhorn transformations that outperforms DSL. Further, we propose a new postprocessing setting that does not require access to multiple test queries. We show that our approach can significantly improve the results of state of the art models such as CLIP4Clip, BLIP, X-CLIP, and DRL, thus achieving a new state-of-the-art on several standard text-video retrieval datasets both with access to the entire test set and in the single-query setting. Konstantin Yakovlev, Gregory Polyakov, Ilseyar Alimova, Alexander Podolskiy, Andrey Bout, Sergey I. Nikolenko, Irina Piontkovskaya |
SIGIR | 7 |
| 2022 | Ask Me Anything in Your Native LanguageabstractNikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, Valentin Malykh. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Nikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, Valentin Malykh |
NAACL-HLT | 3 |
| 2021 | Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain DetectionabstractReal-life applications, heavily relying on machine learning, such as dialog systems, demand for out-of-domain detection methods. Intent classification models should be equipped with a mechanism to distinguish seen intents from unseen ones so that the dialog agent is capable of rejecting the latter and avoiding undesired behavior. However, despite increasing attention paid to the task, the best practices for out-of-domain intent detection have not yet been fully established. This paper conducts a thorough comparison of out-of-domain intent detection methods. We prioritize the methods, not requiring access to out-of-domain data during training, gathering of which is extremely time- and labor-consuming due to lexical and stylistic variation of user utterances. We evaluate multiple contextual encoders and methods, proven to be efficient, on three common datasets for intent classification, expanded with out-of-domain utterances. Our main findings show that fine-tuning Transformer-based encoders on in-domain data leads to superior results. Mahalanobis distance, together with utterance representations, derived from Transformer-based encoders, outperform other methods by a wide margin(1-5% in terms of AUROC) and establish new state-of-the-art results for all datasets. The broader analysis shows that the reason for success lies in the fact that the fine-tuned Transformer is capable of constructing homogeneous representations of in-domain utterances, revealing geometrical disparity to out of domain utterances. In turn, the Mahalanobis distance captures this disparity easily. Alexander Podolskiy, Dmitry Lipin, Andrey Bout, Ekaterina Artemova, Irina Piontkovskaya |
AAAI | 5 |
| 2021 | Artificial Text Detection via Examining the Topology of Attention MapsabstractLaida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova, Serguei Barannikov, Alexander Bernstein, Irina Piontkovskaya, Dmitri Piontkovski, Evgeny Burnaev. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Laida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova, Serguei Barannikov, Alexander V. Bernstein, Irina Piontkovskaya, Dmitri Piontkovski, Evgeny Burnaev |
EMNLP (1) | 7 |
| 2021 | Single Example Can Improve Zero-Shot Data GenerationabstractSub-tasks of intent classification, such as robustness to distribution shift, adaptation to specific user groups and personalization, out-ofdomain detection, require extensive and flexible datasets for experiments and evaluation.As collecting such datasets is time-and laborconsuming, we propose to use text generation methods to gather datasets.The generator should be trained to generate utterances that belong to the given intent.We explore two approaches to generating task-oriented utterances.In the zero-shot approach, the model is trained to generate utterances from seen intents and is further used to generate utterances for intents unseen during training.In the oneshot approach, the model is presented with a single utterance from a test intent.We perform a thorough automatic, and human evaluation of the dataset generated utilizing two proposed approaches.Our results reveal that the attributes of the generated data are close to original test sets, collected via crowd-sourcing. Pavel Burnyshev, Valentin Malykh, Andrey Bout, Ekaterina Artemova, Irina Piontkovskaya |
INLG | 5 |
| 2020 | SumTitles: a Summarization Dataset with Low ExtractivenessabstractThe existing dialogue summarization corpora are significantly extractive.We introduce a methodology for dataset extractiveness evaluation and present a new low-extractive corpus of movie dialogues for abstractive text summarization along with baseline evaluation.The corpus contains 153k dialogues and consists of three parts: 1) automatically aligned subtitles, 2) automatically aligned scenes from scripts, and 3) manually aligned scenes from scripts.We also present an alignment algorithm which we use to construct the corpus. 1 Valentin Malykh, Konstantin Chernis, Ekaterina Artemova, Irina Piontkovskaya |
COLING | 4 |
| 2018 | Distributed Fine-tuning of Language Models on Private Data
Vadim Popov, Mikhail A. Kudinov, Irina Piontkovskaya, Petr Vytovtov, Alex Nevidomsky |
ICLR (Poster) | 3 |