VLDB 2026 Research / reviewers in the wild / expert
Dmitry I. Ilvovsky
dblp:140/3022 · also Dmitri I. Ilvovsky, Dmitry Ilvovsky
· DBLP profile ↗
9ranked-venue papers in the field
0as first author
8since 2021 · last 2026
0000-0002-5484-372XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 7Data Mining & Knowledge Discovery · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiteKV: A KV cache compression method for efficient inference on resource-constrained devices
Dmitry I. Ilvovsky |
Inf. Process. Manag. | 2 |
| 2025 | RuSemCor: A Word Sense Disambiguation corpus for RussianabstractWe present RuSemCor, an open Word Sense Disambiguation (WSD) corpus for Russian. The corpus was constructed by manually linking tokens from the OpenCorpora corpus to senses in the Russian wordnet RuWordNet. It consists of 869 documents with 121,710 tokens of which 51,588 are wordnet annotated. The resource is represented using the NIF, OLiA, OntoLex, and Global WordNet ontologies and integrated into the Linguistic Linked Open Data cloud. We used RuSemCor as a diagnostic benchmark to evaluate a range of WSD methods. Our experiments yielded three main findings. 1)~Generative LLMs substantially outperform traditional knowledge-based methods such as Personalized PageRank. 2)~Despite their strengths, generative LLMs do not surpass encoder-based models specifically trained for WSD. 3)~Incorporating lexical-semantic relations from RuWordNet produces mixed results: it enhances the performance of encoder-based models and leading LLMs like GPT-4, DeepSeek, and Mistral 24B, but tends to degrade accuracy for smaller generative models such as GPT-3 and Mistral 7B. The resource is distributed under the CC BY-SA open license and is available at: https://github.com/LLOD-Ru/rusemcor. Alexander Kirillovich, Ilia Karpov, Natalia V. Loukachevitch, Maksim Kulaev, Dmitry I. Ilvovsky |
CIKM | 5 |
| 2025 | Frozen in the Middle: Hidden States Remain Unchanged Across Intermediate Layers of Language ModelsabstractThis paper investigates the internal mechanisms of large language models (LLMs) through the lens of Mechanistic Interpretability (MI). We present novel findings on how information is processed and propagated within these models. Our key contributions include: (1) providing evidence for the localized nature of fact storage and information propagation from subject tokens; (2) introducing a new observation that hidden states remain largely unchanged across multiple middle layers, which we call the ''plateau'' phenomenon; and (3) developing a manually crafted diagnostic dataset of factual prompts. Our work complements and extends prior research on transformer information flow by demonstrating that, contrary to the prevailing assumption of sequential representation enrichment across layers, subject token states stabilize early and remain functionally static throughout multiple middle layers while containing all necessary information for the final prediction. These insights advance our understanding of how transformers process factual information and suggest a more complex pattern of layer specialization than previously identified. Pavel Tikhonov, Dmitry I. Ilvovsky |
CIKM | 2 |
| 2025 | Enhancing FEVER-Style Claim Fact-Checking Against Wikipedia: A Diagnostic Taxonomy and a Generative Framework
Anton Chernyavskiy, Dmitry I. Ilvovsky, Preslav Nakov |
ECIR (1) | 2 |
| 2024 | Truth-O-Meter: Handling Multiple Inconsistent Sources Repairing LLM HallucinationsabstractLarge Language Models (LLM) often produce text with incorrect facts and hallucinations. To address this issue, we developed a fact-checking system Truth-O-Meter12 which verifies LLM results on the Internet and other sources of information to detect wrong claims/facts and proposes corrections for them. NLP and reasoning techniques such as Abstract Meaning Representation and syntactic alignment are applied to match hallucinating sentences with truthful ones. To handle inconsistent sources while fact-checking, we rely on argumentation analysis in the form of defeasible logic programming, selecting the most authoritative source. Our evaluation shows that LLM content can be substantially improved for factual correctness and meaningfulness on an industrial scale. Boris A. Galitsky, Anton Chernyavskiy, Dmitry I. Ilvovsky |
SIGIR | 3 |
| 2021 | WhatTheWikiFact: Fact-Checking Claims Against WikipediaabstractThe rise of Internet has made it a major source of information. Unfortunately, not all information online is true, and thus a number of fact-checking initiatives have been launched, both manual and automatic, to deal with the problem. Here, we present our contribution in this regard: WhatTheWikiFact, a system for automatic claim verification using Wikipedia. The system can predict the veracity of an input claim, and it further shows the evidence it has retrieved as part of the verification process. It shows confidence scores and a list of relevant Wikipedia articles, together with detailed information about each article, including the phrase used to retrieve it, the most relevant sentences extracted from it and their stance with respect to the input claim, as well as the associated probabilities. The system supports several languages: Bulgarian, English, and Russian. Anton Chernyavskiy, Dmitry I. Ilvovsky, Preslav Nakov |
CIKM | 2 |
| 2021 | TatWordNet: A Linguistic Linked Open Data-Integrated WordNet Resource for TatarabstractWe present the first release of TatWordNet (http://wordnet.tatar), a wordnet resource for Tatar. TatWordNet has been constructed by the combination of the expand and the merge approaches. The synsets of TatWordNet have been compiled by: (i) the automatic conversion of concepts of TatThes, a socio-political Tatar; (ii) semi-automatic translation of synsets of RuWordNet, a wordnet resource for Russian with the followed manual verification and correction; (iii) manual translation of base RuWordNet synsets; (iv) and manual translation of the all hypernyms of the previously translated RuWordNet synsets. The currents version of TatWordNet contains 18,583 synsets, 36,540 lexical entries and 49,525 senses. The resource has been published to the Linguistic Linked Open Data cloud and interlinked with the Global WordNet Grid. Alexander Kirillovich, Marat Shaekhov, Alfiya M. Galieva, Olga Nevzorova, Dmitry I. Ilvovsky, Natalia V. Loukachevitch |
LDK | 5 |
| 2021 | Transformers: "The End of History" for Natural Language Processing?
Anton Chernyavskiy, Dmitry I. Ilvovsky, Preslav Nakov |
ECML/PKDD (3) | 2 |
| 2019 | On a Chatbot Conducting Virtual DialoguesabstractWe present a demo of the chatbot that delivers content in the form of virtual dialogues automatically produced from the plain texts extracted and selected from the documents. This virtual dialogue content is provided in the form of answers derived from the found and selected documents split into fragments, and questions are automatically generated for these answers. Boris A. Galitsky, Dmitry I. Ilvovsky |
CIKM | 2 |