Irina Piontkovskaya

dblp:211/7823 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
12since 2021 · last 2026
0009-0003-0299-5849ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RuBIN: A Russian Benchmark for Evaluating LLMs with Cultural Insights
Polina Lazukova, Irina Piontkovskaya
LREC2
2025 LAMeD: LLM-generated Annotations for Memory Leak Detection
abstract
Static analysis tools are widely used to detect software bugs and vulnerabilities but often struggle with scalability and efficiency in complex codebases. Traditional approaches rely on manually crafted annotations—labeling functions as sources or sinks—to track data flows, e.g., ensuring that allocated memory is eventually freed, and code analysis tools such as CodeQL, Infer, or Cooddy can use function specifications, but manual annotation is laborious and error-prone, especially for large or third-party libraries. We present LAMeD (LLM-generated Annotations for Memory leak Detection), a novel approach that leverages large language models (LLMs) to automatically generate function-specific annotations. When integrated with analyzers such as Cooddy, LAMeD significantly improves memory leak detection and reduces path explosion. We also suggest directions for extending LAMeD to broader code analysis.
Ekaterina Shemetova, Ivan Smirnov, Anton Alekseev 0001, Ilya Shenbin, Alexey D. Rukhovich, Sergey I. Nikolenko, Vadim Lomshakov, Irina Piontkovskaya
EASE8
2025 Quantifying Logical Consistency in Transformers via Query-Key Alignment
abstract
Eduard Tulchinskii, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Eduard Tulchinskii, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov
EMNLP5
2023 GEC-DePenD: Non-Autoregressive Grammatical Error Correction with Decoupled Permutation and Decoding
abstract
Konstantin Yakovlev, Alexander Podolskiy, Andrey Bout, Sergey Nikolenko, Irina Piontkovskaya. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Konstantin Yakovlev, Alexander Podolskiy, Andrey Bout, Sergey I. Nikolenko, Irina Piontkovskaya
ACL (1)5
2023 Efficient Grammatical Error Correction Via Multi-Task Training and Optimized Training Schedule
abstract
Progress in neural grammatical error correction (GEC) is hindered by the lack of annotated training data.Sufficient amounts of highquality manually annotated data are not available, so recent research has relied on generating synthetic data, pretraining on it, and then finetuning on real datasets; performance gains have been achieved either by ensembling or by using huge pretrained models such as XXL-T5 as the backbone.In this work, we explore an orthogonal direction: how to use available data more efficiently.First, we propose auxiliary tasks that exploit the alignment between the original and corrected sentences, such as predicting a sequence of corrections.We formulate each task as a sequence-to-sequence problem and perform multi-task training.Second, we discover that the order of datasets used for training and even individual instances within a dataset may have important effects on the final performance, so we set out to find the best training schedule.Together, these two ideas lead to significant improvements, producing results that improve state of the art with much smaller models; in particular, we outperform the best models based on T5-XXL (11B parameters) with a BART-based model (400M parameters).
Andrey Bout, Alexander Podolskiy, Sergey I. Nikolenko, Irina Piontkovskaya
EMNLP4
2023 Topological Data Analysis for Speech Processing
abstract
International audience
Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Daniil Cherniavskii, Serguei Barannikov, Irina Piontkovskaya, Sergey I. Nikolenko, Evgeny Burnaev
INTERSPEECH6
2023 Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts
abstract
Rapidly increasing quality of AI-generated content makes it difficult to distinguish between human and AI-generated texts, which may lead to undesirable consequences for society. Therefore, it becomes increasingly important to study the properties of human texts that are invariant over text domains and various proficiency of human writers, can be easily calculated for any language, and can robustly separate natural and AI-generated texts regardless of the generation model and sampling method. In this work, we propose such an invariant of human texts, namely the intrinsic dimensionality of the manifold underlying the set of embeddings of a given text sample. We show that the average intrinsic dimensionality of fluent texts in natural language is hovering around the value $9$ for several alphabet-based languages and around $7$ for Chinese, while the average intrinsic dimensionality of AI-generated texts for each language is $\approx 1.5$ lower, with a clear statistical separation between human-generated and AI-generated distributions. This property allows us to build a score-based artificial text detector. The proposed detector's accuracy is stable over text domains, generator models, and human writer proficiency levels, outperforming SOTA detectors in model-agnostic and cross-domain scenarios by a significant margin.
Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Daniil Cherniavskii, Sergey I. Nikolenko, Evgeny Burnaev, Serguei Barannikov, Irina Piontkovskaya
NeurIPS8
2023 Sinkhorn Transformations for Single-Query Postprocessing in Text-Video Retrieval
abstract
A recent trend in multimodal retrieval is related to postprocessing test set results via the dual-softmax loss (DSL). While this approach can bring significant improvements, it usually presumes that an entire matrix of test samples is available as DSL input. This work introduces a new postprocessing approach based on Sinkhorn transformations that outperforms DSL. Further, we propose a new postprocessing setting that does not require access to multiple test queries. We show that our approach can significantly improve the results of state of the art models such as CLIP4Clip, BLIP, X-CLIP, and DRL, thus achieving a new state-of-the-art on several standard text-video retrieval datasets both with access to the entire test set and in the single-query setting.
Konstantin Yakovlev, Gregory Polyakov, Ilseyar Alimova, Alexander Podolskiy, Andrey Bout, Sergey I. Nikolenko, Irina Piontkovskaya
SIGIR7
2022 Ask Me Anything in Your Native Language
abstract
Nikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, Valentin Malykh. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Nikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, Valentin Malykh
NAACL-HLT3
2021 Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain Detection
abstract
Real-life applications, heavily relying on machine learning, such as dialog systems, demand for out-of-domain detection methods. Intent classification models should be equipped with a mechanism to distinguish seen intents from unseen ones so that the dialog agent is capable of rejecting the latter and avoiding undesired behavior. However, despite increasing attention paid to the task, the best practices for out-of-domain intent detection have not yet been fully established. This paper conducts a thorough comparison of out-of-domain intent detection methods. We prioritize the methods, not requiring access to out-of-domain data during training, gathering of which is extremely time- and labor-consuming due to lexical and stylistic variation of user utterances. We evaluate multiple contextual encoders and methods, proven to be efficient, on three common datasets for intent classification, expanded with out-of-domain utterances. Our main findings show that fine-tuning Transformer-based encoders on in-domain data leads to superior results. Mahalanobis distance, together with utterance representations, derived from Transformer-based encoders, outperform other methods by a wide margin(1-5% in terms of AUROC) and establish new state-of-the-art results for all datasets. The broader analysis shows that the reason for success lies in the fact that the fine-tuned Transformer is capable of constructing homogeneous representations of in-domain utterances, revealing geometrical disparity to out of domain utterances. In turn, the Mahalanobis distance captures this disparity easily.
Alexander Podolskiy, Dmitry Lipin, Andrey Bout, Ekaterina Artemova, Irina Piontkovskaya
AAAI5
2021 Artificial Text Detection via Examining the Topology of Attention Maps
abstract
Laida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova, Serguei Barannikov, Alexander Bernstein, Irina Piontkovskaya, Dmitri Piontkovski, Evgeny Burnaev. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Laida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova, Serguei Barannikov, Alexander V. Bernstein, Irina Piontkovskaya, Dmitri Piontkovski, Evgeny Burnaev
EMNLP (1)7
2021 Single Example Can Improve Zero-Shot Data Generation
abstract
Sub-tasks of intent classification, such as robustness to distribution shift, adaptation to specific user groups and personalization, out-ofdomain detection, require extensive and flexible datasets for experiments and evaluation.As collecting such datasets is time-and laborconsuming, we propose to use text generation methods to gather datasets.The generator should be trained to generate utterances that belong to the given intent.We explore two approaches to generating task-oriented utterances.In the zero-shot approach, the model is trained to generate utterances from seen intents and is further used to generate utterances for intents unseen during training.In the oneshot approach, the model is presented with a single utterance from a test intent.We perform a thorough automatic, and human evaluation of the dataset generated utilizing two proposed approaches.Our results reveal that the attributes of the generated data are close to original test sets, collected via crowd-sourcing.
Pavel Burnyshev, Valentin Malykh, Andrey Bout, Ekaterina Artemova, Irina Piontkovskaya
INLG5
2020 SumTitles: a Summarization Dataset with Low Extractiveness
abstract
The existing dialogue summarization corpora are significantly extractive.We introduce a methodology for dataset extractiveness evaluation and present a new low-extractive corpus of movie dialogues for abstractive text summarization along with baseline evaluation.The corpus contains 153k dialogues and consists of three parts: 1) automatically aligned subtitles, 2) automatically aligned scenes from scripts, and 3) manually aligned scenes from scripts.We also present an alignment algorithm which we use to construct the corpus. 1
Valentin Malykh, Konstantin Chernis, Ekaterina Artemova, Irina Piontkovskaya
COLING4
2018 Distributed Fine-tuning of Language Models on Private Data
Vadim Popov, Mikhail A. Kudinov, Irina Piontkovskaya, Petr Vytovtov, Alex Nevidomsky
ICLR (Poster)3