VLDB 2026 Research / reviewers in the wild / expert
Denny Vrandecic
dblp:59/5618 · also Zdenko Vrandecic
· DBLP profile ↗
27ranked-venue papers
5as first author
1since 2021 · last 2026
0000-0002-9593-2294ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 18 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 1 first-authorArtificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 60% Transfer learning and domain adaptation · 30% Efficient and distributed learning · 9% | |
| Databases, data mining, and information retrieval
3 papers |
Data integration and cleaning · 51% Knowledge graphs · 49% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
1.0 | 1 | 2026 | Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis
fact-checking |
1.0 | 1 | 2026 | Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis
multilingual NLP |
1.0 | 1 | 2026 | Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model compression › lightweight neural network
small language models |
0.3 | 1 | 2026 | Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages · ACL (1) 2026 |
Data integration and cleaning
data mapping |
0.2 | 1 | 2016 | From Freebase to Wikidata: The Great Migration · WWW 2016 |
Knowledge graphs
knowledge graph construction |
0.2 | 1 | 2016 | From Freebase to Wikidata: The Great Migration · WWW 2016 |
Data integration and cleaning
data migration |
0.1 | 1 | 2016 | From Freebase to Wikidata: The Great Migration · WWW 2016 |
Data integration and cleaning › web data integration
data mashup |
0.1 | 1 | 2007 | The two cultures: mashing up web 2.0 and the semantic web · WWW 2007 |
Knowledge graphs
semantic web |
0.1 | 1 | 2007 | The two cultures: mashing up web 2.0 and the semantic web · WWW 2007 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology |
0.0 | 1 | 2007 | The two cultures: mashing up web 2.0 and the semantic web · WWW 2007 |
Methods — techniques the papers use, named apart from their topics
prompting · 1.0fine-tuning · 1.0semantic web technologies · 0.3ontology design · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource LanguagesabstractIn automated fact-checking (AFC), checkworthiness detection identifies claims requiring verification based on domain-specific criteria.On Wikipedia, this task instantiates as Citation Needed Detection (CND), which flags claims lacking supporting citations.However, existing research has largely overlooked lowerresource languages, and recent AFC pipelines rely on large language models (LLMs), which are inaccessible to low-resource organizations.We introduce MCN, a multilingual CND corpus spanning 18 languages across three resource levels, on which we conduct an extensive study of small decoder-based language models (SLMs).Our experiments show that SLMs fine-tuned with an encoder-style objective substantially outperform prompted LLMs across languages.We further present one of the first studies on cross-lingual CND, demonstrating that SLMs fine-tuned solely on English claims surpass LLMs, even with little to no target-language adaptation.Our findings have important implications for lower-resource Wikipedia communities and suggest that compact, task-specific models are preferable to LLMs for CND.We release all data and code at https://github.com/gerritq/mcnDataset Task Domain Languages (resource-level) Size NLP4IF 2021 (Shaar et al., 2021) CD/CWD Covid-19/Politics 3 (2 high, 1 medium) 1.3K-4K CitationNeeded (Redi et al., 2019) CWD Wikipedia 3 (3 high) 20K Kazemi et al. (2021) CD Covid-19, Politics 5 (2 high, 2 medium, 1 low) 5K Dutta et al. (2023) CD Politics 3 (2 high, 1 medium) 600-1.4KHalitaj and Zubiaga (2024) CWD Wikipedia 3 (2 high, 1 medium) 31K-1.1mCheckThat 2024 (Hasanain et al., 2024) CD/CWD Politics, Covid-19 3 (3 high) 1K-23K Baigutanova et al. (2026) CWD Wikipedia 5 (5 high) 20K MCN (ours) CWD Wikipedia 18 (8 high, 5 medium, 5 low) 2K-1.1m Gerrit Quaremba, Amy Rechkemmer, Elizabeth Black, Denny Vrandecic, Elena Simperl |
ACL (1) | 4 |
| 2020 | Introducing Lexical Masks: a New Representation of Lexical Entries for Better Evaluation and Exchange of LexiconsabstractThe evaluation and exchange of large lexicon databases remains a challenge in many NLP applications. Despite the existence of commonly accepted standards for the format and the features used in a lexicon, there is still a lack of precise and interoperable specification requirements about how lexical entries of a particular language should look like, both in terms of the numbers of forms and in terms of features associated with these forms. This paper presents the notion of “lexical masks”, a powerful tool used to evaluate and exchange lexicon databases in many languages. Bruno Cartoni, Daniel Calvelo Aros, Denny Vrandecic, Saran Lertpradit |
LREC | 3 |
| 2020 | Wiki-40B: Multilingual Language Model DatasetabstractWe propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families. With around 40 billion characters, we hope this new resource will accelerate the research of multilingual modeling. We train monolingual causal language models using a state-of-the-art model (Transformer-XL) establishing baselines for many languages. We also introduce the task of multilingual causal language modeling where we train our model on the combined text of 40+ languages from Wikipedia with different vocabulary sizes and evaluate on the languages individually. We released the cleaned-up text of 40+ Wikipedia language editions, the corresponding trained monolingual language models, and several multilingual language models with different fixed vocabulary sizes. Mandy Guo, Zihang Dai, Denny Vrandecic, Rami Al-Rfou |
LREC | 3 |
| 2019 | Describing Datasets in WikidataabstractWe propose to use Wikidata to provide metadata for datasets when the traditional approach via Schema.org is not feasible. We describe and discuss the proposal, and believe that the process described in this paper can help with increasing findability and accessibility of certain datasets. Denny Vrandecic |
eScience | 1 |
| 2019 | Wikidata: A large-scale collaborative ontological medical database
Houcemeddine Turki, Thomas Shafee, Mohamed Ali Hadj Taieb, Mohamed Benaouicha, Denny Vrandecic, Diptanshu Das, Helmi Hamdi |
J. Biomed. Informatics | 5 |
| 2017 | Using WikiData as a Multi-lingual Multi-dialectal Dictionary for Arabic DialectsabstractSince 2012, Wikidata has been developed as a freely accessible and community generated knowledge database that represents not only the name of each item (tree, famous person...) in all available languages but also to define links (like "component of", "born in", "instance of"...) between items. Nowadays, the output of Wikidata has become one of the greatest semantic web data in the world and has been proved to be useful to solve many currently existing problems in Computational Linguistics, in Medicine, and in many other fields. In this research work, we propose to convert Wikidata into a multilingual multi-dialectal dictionary for Arabic dialects and we explain how Wikidata as a multi-lingual multi-dialectal dictionary for Arabic dialects can be later used for the natural language processing of the varieties of Arabic by computational linguists and computer scientists. Houcemeddine Turki, Denny Vrandecic, Helmi Hamdi, Imed Adel |
AICCSA | 2 |
| 2016 | From Freebase to Wikidata: The Great MigrationabstractCollaborative knowledge bases that make their data freely available in a machine-readable form are central for the data strategy of many projects and organizations. The two major collaborative knowledge bases are Wikimedia's Wikidata and Google's Freebase. Due to the success of Wikidata, Google decided in 2014 to offer the content of Freebase to the Wikidata community. In this paper, we report on the ongoing transfer efforts and data mapping challenges, and provide an analysis of the effort so far. We describe the Primary Sources Tool, which aims to facilitate this and future data migrations. Throughout the migration, we have gained deep insights into both Wikidata and Freebase, and share and discuss detailed statistics on both knowledge bases. Thomas Pellissier Tanon, Denny Vrandecic, Sebastian Schaffert, Thomas Steiner, Lydia Pintscher |
WWW | 2 |
| 2014 | Introducing Wikidata to the Linked Data Web
Fredo Erxleben, Michael Günther 0002, Markus Krötzsch, Julian Alfredo Mendez, Denny Vrandecic |
ISWC (1) | 5 |
| 2012 | Capturing Common Knowledge about Tasks: Intelligent Assistance for To-Do ListsabstractAlthough to-do lists are a ubiquitous form of personal task management, there has been no work on intelligent assistance to automate, elaborate, or coordinate a user’s to-dos. Our research focuses on three aspects of intelligent assistance for to-dos. We investigated the use of intelligent agents to automate to-dos in an office setting. We collected a large corpus from users and developed a paraphrase-based approach to matching agent capabilities with to-dos. We also investigated to-dos for personal tasks and the kinds of assistance that can be offered to users by elaborating on them on the basis of substep knowledge extracted from the Web. Finally, we explored coordination of user tasks with other users through a to-do management application deployed in a popular social networking site. We discuss the emergence of Social Task Networks, which link users‘ tasks to their social network as well as to relevant resources on the Web. We show the benefits of using common sense knowledge to interpret and elaborate to-dos. Conversely, we also show that to-do lists are a valuable way to create repositories of common sense knowledge about tasks. Yolanda Gil, Varun Ratnakar, Timothy Chklovski, Paul Groth, Denny Vrandecic |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2011 | Wiki-Based Maturing of Process Descriptions
Frank Dengler, Denny Vrandecic |
BPM | 2 |
| 2011 | Zhi# - OWL Aware Compilation
Alexander Paar, Denny Vrandecic |
ESWC (2) | 2 |
| 2011 | Want world domination? win at risk!: matching to-do items with how-tos from the webabstractTo-Do lists are widely used for personal task management. We propose a novel approach to assist users in managing their To-Dos by matching them to How-To knowledge from the Web. We have implemented a system that, given a To-Do item, provides a number of possibly matching How-Tos, broken down into steps that can be used as new To-Do entries. Our implementation is in the form of a web service that can be easily integrated into existing To-Do applications. This can help users by providing them with an approach to tackle the To-Do by listing smaller, more actionable To-Dos. In this paper we present our implementation, an evaluation of the matching component over two sets of To-Do corpora with very different characteristics, and a discussion of the results. Denny Vrandecic, Yolanda Gil, Varun Ratnakar |
IUI | 1 |
| 2011 | Wikiing pro: semantic wiki-based process editorabstractRecently, a trend toward collaborative, user-centric, on-line process modeling can be observed. Unfortunately, current social software approaches mostly focus on the graphical development of processes and do not consider existing textual process description like HowTos or guidelines. We address this issue by combining graphical process modeling techniques with a wiki-based light-weight knowledge capturing approach and a background semantic knowledge base. Our approach enables the collaborative maturing of process descriptions with a graphical representation, formal semantic annotations, and natural language. By translating existing textual process descriptions into graphical descriptions and formal semantic annotations, we provide a holistic approach for collaborative process development that is designed to foster knowledge reuse and maturing within the system. Frank Dengler, Denny Vrandecic, Elena Simperl |
K-CAP | 2 |
| 2011 | Language resources extracted from WikipediaabstractWikipedia provides an interesting amount of text for more than hundred languages. This also includes languages where no reference corpora or other linguistic resources are easily available. We have extracted background language models built from the content of Wikipedia in various languages. The models generated from Simple and English Wikipedia are compared to language models derived from other established corpora. The differences between the models in regard to term coverage, term distribution and correlation are described and discussed. We provide access to the full dataset and create visualizations of the language models that can be used exploratory. The paper describes the newly released dataset for 33 languages, and the services that we provide on top of them. Denny Vrandecic, Philipp Sorg, Rudi Studer |
K-CAP | 1 |
| 2011 | Labels in the Web of Data
Basil Ell, Denny Vrandecic, Elena Simperl |
ISWC (1) | 2 |
| 2011 | Shortipedia aggregating and curating Semantic Web data
Denny Vrandecic, Varun Ratnakar, Markus Krötzsch, Yolanda Gil |
J. Web Semant. | 1 |
| 2009 | Tempus Fugit
Uta Lösch, Sebastian Rudolph, Denny Vrandecic, Rudi Studer |
ESWC | 3 |
| 2008 | Workshop on social web and knowledge management (SWKM2008)abstractThis paper provides an overview on the synergies between social web and knowledge managemen, topics, program committee members as well as summary of accepted papers for the SWKM2008 workshop. Peter Dolog, Markus Krötzsch, Sebastian Schaffert, Denny Vrandecic |
WWW | 4 |
| 2008 | The two cultures: Mashing up Web 2.0 and the Semantic Web
Anupriya Ankolekar, Markus Krötzsch, Thanh Tran 0001, Denny Vrandecic |
J. Web Semant. | 4 |
| 2007 | Learning Disjointness
Johanna Völker, Denny Vrandecic, York Sure-Vetter, Andreas Hotho |
ESWC | 2 |
| 2007 | How to Design Better Ontology Metrics
Denny Vrandecic, York Sure-Vetter |
ESWC | 1 |
| 2007 | The two cultures: mashing up web 2.0 and the semantic webabstractA common perception is that there are two competing visions for the future evolution of the Web: the Semantic Web and Web 2.0. A closer look, though, reveals that the core technologies and concerns of these two approaches are complementary and that each field can and must draw from the other's strengths. We believe that future web applications will retain the Web 2.0 focus on community and usability, while drawing on Semantic Web infrastructure to facilitate mashup-like information sharing. However, there are several open issues that must be addressed before such applications can become commonplace. In this paper, we outline a semantic weblogs scenario that illustrates the potential for combining Web 2.0 and Semantic Web technologies, while highlighting the unresolved issues that impede its realization. Nevertheless, we believe that the scenario can be realized in the short-term. We point to recent progress made in resolving each of the issues as well as future research directions for each of the communities. Anupriya Ankolekar, Markus Krötzsch, Thanh Tran 0001, Denny Vrandecic |
WWW | 4 |
| 2007 | Semantic Wikipedia
Markus Krötzsch, Denny Vrandecic, Max Völkel, Heiko Haller, Rudi Studer |
J. Web Semant. | 2 |
| 2006 | Semantic MediaWiki
Markus Krötzsch, Denny Vrandecic, Max Völkel |
ISWC | 2 |
| 2006 | Semantic WikipediaabstractWikipedia is the world's largest collaboratively edited source of encyclopaedic knowledge. But in spite of its utility, its contents are barely machine-interpretable. Structural knowledge, e.,g. about how concepts are interrelated, can neither be formally stated nor automatically processed. Also the wealth of numerical data is only available as plain text and thus can not be processed by its actual meaning.We provide an extension to be integrated in Wikipedia, that allows the typing of links between articles and the specification of typed data inside the articles in an easy-to-use manner.Enabling even casual users to participate in the creation of an open semantic knowledge base, Wikipedia has the chance to become a resource of semantic statements, hitherto unknown regarding size, scope, openness, and internationalisation. These semantic enhancements bring to Wikipedia benefits of today's semantic technologies: more specific ways of searching and browsing. Also, the RDF export, that gives direct access to the formalised knowledge, opens Wikipedia up to a wide range of external applications, that will be able to use it as a background knowledge base.In this paper, we present the design, implementation, and possible uses of this extension. Max Völkel, Markus Krötzsch, Denny Vrandecic, Heiko Haller, Rudi Studer |
WWW | 3 |
| 2005 | Resolution-Based Approximate Reasoning for OWL DL
Pascal Hitzler, Denny Vrandecic |
ISWC | 2 |
| 2005 | Automatic Evaluation of Ontologies (AEON)
Johanna Völker, Denny Vrandecic, York Sure-Vetter |
ISWC | 2 |