Denny Vrandecic

dblp:59/5618 · also Zdenko Vrandecic · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
1since 2021 · last 2026
0000-0002-9593-2294ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 1 first-authorArtificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Information extraction and text analysis · 60% Transfer learning and domain adaptation · 30% Efficient and distributed learning · 9%
Databases, data mining, and information retrieval
3 papers
Data integration and cleaning · 51% Knowledge graphs · 49%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
1.012026
Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages · ACL (1) 2026
Natural language and speech › Information extraction and text analysis
fact-checking
1.012026
Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages · ACL (1) 2026
Natural language and speech › Information extraction and text analysis
multilingual NLP
1.012026
Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages · ACL (1) 2026
Machine learning › Efficient and distributed learning › model compression › lightweight neural network
small language models
0.312026
Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages · ACL (1) 2026
Data integration and cleaning
data mapping
0.212016
From Freebase to Wikidata: The Great Migration · WWW 2016
Knowledge graphs
knowledge graph construction
0.212016
From Freebase to Wikidata: The Great Migration · WWW 2016
Data integration and cleaning
data migration
0.112016
From Freebase to Wikidata: The Great Migration · WWW 2016
Data integration and cleaning › web data integration
data mashup
0.112007
The two cultures: mashing up web 2.0 and the semantic web · WWW 2007
Knowledge graphs
semantic web
0.112007
The two cultures: mashing up web 2.0 and the semantic web · WWW 2007
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology
0.012007
The two cultures: mashing up web 2.0 and the semantic web · WWW 2007

Methods — techniques the papers use, named apart from their topics

prompting · 1.0fine-tuning · 1.0semantic web technologies · 0.3ontology design · 0.1
YearPublicationVenuePosition
2026 Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages
abstract
In automated fact-checking (AFC), checkworthiness detection identifies claims requiring verification based on domain-specific criteria.On Wikipedia, this task instantiates as Citation Needed Detection (CND), which flags claims lacking supporting citations.However, existing research has largely overlooked lowerresource languages, and recent AFC pipelines rely on large language models (LLMs), which are inaccessible to low-resource organizations.We introduce MCN, a multilingual CND corpus spanning 18 languages across three resource levels, on which we conduct an extensive study of small decoder-based language models (SLMs).Our experiments show that SLMs fine-tuned with an encoder-style objective substantially outperform prompted LLMs across languages.We further present one of the first studies on cross-lingual CND, demonstrating that SLMs fine-tuned solely on English claims surpass LLMs, even with little to no target-language adaptation.Our findings have important implications for lower-resource Wikipedia communities and suggest that compact, task-specific models are preferable to LLMs for CND.We release all data and code at https://github.com/gerritq/mcnDataset Task Domain Languages (resource-level) Size NLP4IF 2021 (Shaar et al., 2021) CD/CWD Covid-19/Politics 3 (2 high, 1 medium) 1.3K-4K CitationNeeded (Redi et al., 2019) CWD Wikipedia 3 (3 high) 20K Kazemi et al. (2021) CD Covid-19, Politics 5 (2 high, 2 medium, 1 low) 5K Dutta et al. (2023) CD Politics 3 (2 high, 1 medium) 600-1.4KHalitaj and Zubiaga (2024) CWD Wikipedia 3 (2 high, 1 medium) 31K-1.1mCheckThat 2024 (Hasanain et al., 2024) CD/CWD Politics, Covid-19 3 (3 high) 1K-23K Baigutanova et al. (2026) CWD Wikipedia 5 (5 high) 20K MCN (ours) CWD Wikipedia 18 (8 high, 5 medium, 5 low) 2K-1.1m
Gerrit Quaremba, Amy Rechkemmer, Elizabeth Black, Denny Vrandecic, Elena Simperl
ACL (1)4
2020 Introducing Lexical Masks: a New Representation of Lexical Entries for Better Evaluation and Exchange of Lexicons
abstract
The evaluation and exchange of large lexicon databases remains a challenge in many NLP applications. Despite the existence of commonly accepted standards for the format and the features used in a lexicon, there is still a lack of precise and interoperable specification requirements about how lexical entries of a particular language should look like, both in terms of the numbers of forms and in terms of features associated with these forms. This paper presents the notion of “lexical masks”, a powerful tool used to evaluate and exchange lexicon databases in many languages.
Bruno Cartoni, Daniel Calvelo Aros, Denny Vrandecic, Saran Lertpradit
LREC3
2020 Wiki-40B: Multilingual Language Model Dataset
abstract
We propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families. With around 40 billion characters, we hope this new resource will accelerate the research of multilingual modeling. We train monolingual causal language models using a state-of-the-art model (Transformer-XL) establishing baselines for many languages. We also introduce the task of multilingual causal language modeling where we train our model on the combined text of 40+ languages from Wikipedia with different vocabulary sizes and evaluate on the languages individually. We released the cleaned-up text of 40+ Wikipedia language editions, the corresponding trained monolingual language models, and several multilingual language models with different fixed vocabulary sizes.
Mandy Guo, Zihang Dai, Denny Vrandecic, Rami Al-Rfou
LREC3
2019 Describing Datasets in Wikidata
abstract
We propose to use Wikidata to provide metadata for datasets when the traditional approach via Schema.org is not feasible. We describe and discuss the proposal, and believe that the process described in this paper can help with increasing findability and accessibility of certain datasets.
Denny Vrandecic
eScience1
2019 Wikidata: A large-scale collaborative ontological medical database
Houcemeddine Turki, Thomas Shafee, Mohamed Ali Hadj Taieb, Mohamed Benaouicha, Denny Vrandecic, Diptanshu Das, Helmi Hamdi
J. Biomed. Informatics5
2017 Using WikiData as a Multi-lingual Multi-dialectal Dictionary for Arabic Dialects
abstract
Since 2012, Wikidata has been developed as a freely accessible and community generated knowledge database that represents not only the name of each item (tree, famous person...) in all available languages but also to define links (like "component of", "born in", "instance of"...) between items. Nowadays, the output of Wikidata has become one of the greatest semantic web data in the world and has been proved to be useful to solve many currently existing problems in Computational Linguistics, in Medicine, and in many other fields. In this research work, we propose to convert Wikidata into a multilingual multi-dialectal dictionary for Arabic dialects and we explain how Wikidata as a multi-lingual multi-dialectal dictionary for Arabic dialects can be later used for the natural language processing of the varieties of Arabic by computational linguists and computer scientists.
Houcemeddine Turki, Denny Vrandecic, Helmi Hamdi, Imed Adel
AICCSA2
2016 From Freebase to Wikidata: The Great Migration
abstract
Collaborative knowledge bases that make their data freely available in a machine-readable form are central for the data strategy of many projects and organizations. The two major collaborative knowledge bases are Wikimedia's Wikidata and Google's Freebase. Due to the success of Wikidata, Google decided in 2014 to offer the content of Freebase to the Wikidata community. In this paper, we report on the ongoing transfer efforts and data mapping challenges, and provide an analysis of the effort so far. We describe the Primary Sources Tool, which aims to facilitate this and future data migrations. Throughout the migration, we have gained deep insights into both Wikidata and Freebase, and share and discuss detailed statistics on both knowledge bases.
Thomas Pellissier Tanon, Denny Vrandecic, Sebastian Schaffert, Thomas Steiner, Lydia Pintscher
WWW2
2014 Introducing Wikidata to the Linked Data Web
Fredo Erxleben, Michael Günther 0002, Markus Krötzsch, Julian Alfredo Mendez, Denny Vrandecic
ISWC (1)5
2012 Capturing Common Knowledge about Tasks: Intelligent Assistance for To-Do Lists
abstract
Although to-do lists are a ubiquitous form of personal task management, there has been no work on intelligent assistance to automate, elaborate, or coordinate a user’s to-dos. Our research focuses on three aspects of intelligent assistance for to-dos. We investigated the use of intelligent agents to automate to-dos in an office setting. We collected a large corpus from users and developed a paraphrase-based approach to matching agent capabilities with to-dos. We also investigated to-dos for personal tasks and the kinds of assistance that can be offered to users by elaborating on them on the basis of substep knowledge extracted from the Web. Finally, we explored coordination of user tasks with other users through a to-do management application deployed in a popular social networking site. We discuss the emergence of Social Task Networks, which link users‘ tasks to their social network as well as to relevant resources on the Web. We show the benefits of using common sense knowledge to interpret and elaborate to-dos. Conversely, we also show that to-do lists are a valuable way to create repositories of common sense knowledge about tasks.
Yolanda Gil, Varun Ratnakar, Timothy Chklovski, Paul Groth, Denny Vrandecic
ACM Trans. Interact. Intell. Syst.5
2011 Wiki-Based Maturing of Process Descriptions
Frank Dengler, Denny Vrandecic
BPM2
2011 Zhi# - OWL Aware Compilation
Alexander Paar, Denny Vrandecic
ESWC (2)2
2011 Want world domination? win at risk!: matching to-do items with how-tos from the web
abstract
To-Do lists are widely used for personal task management. We propose a novel approach to assist users in managing their To-Dos by matching them to How-To knowledge from the Web. We have implemented a system that, given a To-Do item, provides a number of possibly matching How-Tos, broken down into steps that can be used as new To-Do entries. Our implementation is in the form of a web service that can be easily integrated into existing To-Do applications. This can help users by providing them with an approach to tackle the To-Do by listing smaller, more actionable To-Dos. In this paper we present our implementation, an evaluation of the matching component over two sets of To-Do corpora with very different characteristics, and a discussion of the results.
Denny Vrandecic, Yolanda Gil, Varun Ratnakar
IUI1
2011 Wikiing pro: semantic wiki-based process editor
abstract
Recently, a trend toward collaborative, user-centric, on-line process modeling can be observed. Unfortunately, current social software approaches mostly focus on the graphical development of processes and do not consider existing textual process description like HowTos or guidelines. We address this issue by combining graphical process modeling techniques with a wiki-based light-weight knowledge capturing approach and a background semantic knowledge base. Our approach enables the collaborative maturing of process descriptions with a graphical representation, formal semantic annotations, and natural language. By translating existing textual process descriptions into graphical descriptions and formal semantic annotations, we provide a holistic approach for collaborative process development that is designed to foster knowledge reuse and maturing within the system.
Frank Dengler, Denny Vrandecic, Elena Simperl
K-CAP2
2011 Language resources extracted from Wikipedia
abstract
Wikipedia provides an interesting amount of text for more than hundred languages. This also includes languages where no reference corpora or other linguistic resources are easily available. We have extracted background language models built from the content of Wikipedia in various languages. The models generated from Simple and English Wikipedia are compared to language models derived from other established corpora. The differences between the models in regard to term coverage, term distribution and correlation are described and discussed. We provide access to the full dataset and create visualizations of the language models that can be used exploratory. The paper describes the newly released dataset for 33 languages, and the services that we provide on top of them.
Denny Vrandecic, Philipp Sorg, Rudi Studer
K-CAP1
2011 Labels in the Web of Data
Basil Ell, Denny Vrandecic, Elena Simperl
ISWC (1)2
2011 Shortipedia aggregating and curating Semantic Web data
Denny Vrandecic, Varun Ratnakar, Markus Krötzsch, Yolanda Gil
J. Web Semant.1
2009 Tempus Fugit
Uta Lösch, Sebastian Rudolph, Denny Vrandecic, Rudi Studer
ESWC3
2008 Workshop on social web and knowledge management (SWKM2008)
abstract
This paper provides an overview on the synergies between social web and knowledge managemen, topics, program committee members as well as summary of accepted papers for the SWKM2008 workshop.
Peter Dolog, Markus Krötzsch, Sebastian Schaffert, Denny Vrandecic
WWW4
2008 The two cultures: Mashing up Web 2.0 and the Semantic Web
Anupriya Ankolekar, Markus Krötzsch, Thanh Tran 0001, Denny Vrandecic
J. Web Semant.4
2007 Learning Disjointness
Johanna Völker, Denny Vrandecic, York Sure-Vetter, Andreas Hotho
ESWC2
2007 How to Design Better Ontology Metrics
Denny Vrandecic, York Sure-Vetter
ESWC1
2007 The two cultures: mashing up web 2.0 and the semantic web
abstract
A common perception is that there are two competing visions for the future evolution of the Web: the Semantic Web and Web 2.0. A closer look, though, reveals that the core technologies and concerns of these two approaches are complementary and that each field can and must draw from the other's strengths. We believe that future web applications will retain the Web 2.0 focus on community and usability, while drawing on Semantic Web infrastructure to facilitate mashup-like information sharing. However, there are several open issues that must be addressed before such applications can become commonplace. In this paper, we outline a semantic weblogs scenario that illustrates the potential for combining Web 2.0 and Semantic Web technologies, while highlighting the unresolved issues that impede its realization. Nevertheless, we believe that the scenario can be realized in the short-term. We point to recent progress made in resolving each of the issues as well as future research directions for each of the communities.
Anupriya Ankolekar, Markus Krötzsch, Thanh Tran 0001, Denny Vrandecic
WWW4
2007 Semantic Wikipedia
Markus Krötzsch, Denny Vrandecic, Max Völkel, Heiko Haller, Rudi Studer
J. Web Semant.2
2006 Semantic MediaWiki
Markus Krötzsch, Denny Vrandecic, Max Völkel
ISWC2
2006 Semantic Wikipedia
abstract
Wikipedia is the world's largest collaboratively edited source of encyclopaedic knowledge. But in spite of its utility, its contents are barely machine-interpretable. Structural knowledge, e.,g. about how concepts are interrelated, can neither be formally stated nor automatically processed. Also the wealth of numerical data is only available as plain text and thus can not be processed by its actual meaning.We provide an extension to be integrated in Wikipedia, that allows the typing of links between articles and the specification of typed data inside the articles in an easy-to-use manner.Enabling even casual users to participate in the creation of an open semantic knowledge base, Wikipedia has the chance to become a resource of semantic statements, hitherto unknown regarding size, scope, openness, and internationalisation. These semantic enhancements bring to Wikipedia benefits of today's semantic technologies: more specific ways of searching and browsing. Also, the RDF export, that gives direct access to the formalised knowledge, opens Wikipedia up to a wide range of external applications, that will be able to use it as a background knowledge base.In this paper, we present the design, implementation, and possible uses of this extension.
Max Völkel, Markus Krötzsch, Denny Vrandecic, Heiko Haller, Rudi Studer
WWW3
2005 Resolution-Based Approximate Reasoning for OWL DL
Pascal Hitzler, Denny Vrandecic
ISWC2
2005 Automatic Evaluation of Ontologies (AEON)
Johanna Völker, Denny Vrandecic, York Sure-Vetter
ISWC2