EDBT 2026 Demo / reviewers in the wild / expert
Christian Chiarcos
dblp:90/6019
· DBLP profile ↗
18ranked-venue papers in the field
11as first author
9since 2021 · last 2025
0000-0002-4428-029XORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 18 (11 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MOOC on Linguistic Linked Data
Jorge Gracia, Slavko Zitnik, Maxim Ionov, Christian Chiarcos, Dagmar Gromann, Francesco Mambrini, Marco Passarotti, Armando Stellato, John P. McCrae, Gilles Sérasset, Andon Tchechmedjiev, Sara Carvalho, Penny Labropoulou, Rute Costa |
ESWC (2) | 4 |
| 2025 | Putting Low German on the Map (of Linguistic Linked Open Data)abstractWe describe the creation of a cross-dialectal lexical resource for Low German, a regional language spoken primarily in Germany and the Netherlands, based on the application of Linguistic Linked Open Data (LLOD) technologies. We argue that this approach is particularly well-suited for a language without a written standard, but with multiple, incompatible orthographies and considerable internal variation in phonology, spelling and grammar. A major hurdle in the preservation and documentation of and in the creation of educational materials (such as texts and dictionaries) for this variety is its internal degree of linguistic and orthographic variation, intensified by mutually exclusive influences from different national languages and their respective orthographies. We thus aim to provide a “digital Rosetta stone” to unify lexical materials from different dialects through linking dictionaries and mapping corresponding words without the need for a standardvariety. This involves two components, a mapping between different orthographies and phonological systems, and a technology for linking regional dictionaries maintained by different hosts and developed by or for different communities of speakers. Christian Chiarcos, Tabea Gröger, Christian Fäth |
LDK | 1 |
| 2023 | Towards a Conversational Web? A Benchmark for Analysing Semantic Change with Conversational Knowledge Bots and Linked Open Data
Florentina Armaselu, Elena Apostol, Christian Chiarcos, Anas Fahad Khan, Chaya Liebeskind, Barbara McGillivray, Ciprian-Octavian Truica, Andrius Utka, Giedre Valunaite Oleskeviciene |
LDK | 3 |
| 2023 | Crowdsourcing OLiA Annotation Models the Indirect Way
Christian Chiarcos |
LDK | 1 |
| 2023 | Validation of Language Agnostic Models for Discourse Marker Detection
Mariana Damova, Kostadin Mishev, Giedre Valunaite Oleskeviciene, Chaya Liebeskind, Purificação Silvano, Dimitar Trajanov, Ciprian-Octavian Truica, Elena Apostol, Christian Chiarcos, Anna Baczkowska |
LDK | 9 |
| 2023 | Towards ELTeC-LLOD: European Literary Text Collection Linguistic Linked Open Data
Ranka Stankovic, Christian Chiarcos, Milos Utvic, Olivera Kitanovic |
LDK | 2 |
| 2021 | Get! Mimetypes! Right! (Crazy New Idea)abstractThis paper identifies three technical requirements - availability of data, sustainable hosting and resolvable URIs for hosted data - as minimal pre-conditions for Linguistic Linked Open Data technology to develop towards a mature technological ecosystem that third party applications can build upon. While a critical amount of data is available (and it continues to grow), there does not seem to exist a hosting solution that combines the prospects of long-term availability with an unrestricted capability to support resolvable URIs. In particular, data hosting services do currently not allow data to be declared as RDF content by means of their media type (mime type), so that the capability of clients to recognize formats and to resolve URIs on that basis is severely limited. Christian Chiarcos |
LDK | 1 |
| 2021 | Linking Discourse Marker InventoriesabstractThe paper describes the first comprehensive edition of machine-readable discourse marker lexicons. Discourse markers such as and, because, but, though or thereafter are essential communicative signals in human conversation, as they indicate how an utterance relates to its communicative context. As much of this information is implicit or expressed differently in different languages, discourse parsing, context-adequate natural language generation and machine translation are considered particularly challenging aspects of Natural Language Processing. Providing this data in machine-readable, standard-compliant form will thus facilitate such technical tasks, and moreover, allow to explore techniques for translation inference to be applied to this particular group of lexical resources that was previously largely neglected in the context of Linguistic Linked (Open) Data. Christian Chiarcos, Maxim Ionov |
LDK | 1 |
| 2021 | An Ontology for CoNLL-RDF: Formal Data Structures for TSV Formats in Language TechnologyabstractIn language technology and language sciences, tab-separated values (TSV) represent a frequently used formalism to represent linguistically annotated natural language, often addressed as "CoNLL formats". A large number of such formats do exist, but although they share a number of common features, they are not interoperable, as different pieces of information are encoded differently in these dialects. CoNLL-RDF refers to a programming library and the associated data model that has been introduced to facilitate processing and transforming such TSV formats in a serialization-independent way. CoNLL-RDF represents CoNLL data, by means of RDF graphs and SPARQL update operations, but so far, without machine-readable semantics, with annotation properties created dynamically on the basis of a user-defined mapping from columns to labels. Current applications of CoNLL-RDF include linking between corpora and dictionaries [Mambrini and Passarotti, 2019] and knowledge graphs [Tamper et al., 2018], syntactic parsing of historical languages [Chiarcos et al., 2018; Chiarcos et al., 2018], the consolidation of syntactic and semantic annotations [Chiarcos and Fäth, 2019], a bridge between RDF corpora and a traditional corpus query language [Ionov et al., 2020], and language contact studies [Chiarcos et al., 2018]. We describe a novel extension of CoNLL-RDF, introducing a formal data model, formalized as an ontology. The ontology is a basis for linking RDF corpora with other Semantic Web resources, but more importantly, its application for transformation between different TSV formats is a major step for providing interoperability between CoNLL formats. Christian Chiarcos, Maxim Ionov, Luis Glaser, Christian Fäth |
LDK | 1 |
| 2020 | Leveraging Linguistic Linked Data for Cross-Lingual Model Transfer in the Pharmaceutical Domain
Jorge Gracia, Christian Fäth, Matthias Hartung, Maxim Ionov, Julia Bosque-Gil, Susana Veríssimo, Christian Chiarcos, Matthias Orlikowski |
ISWC (2) | 7 |
| 2019 | Automatic Detection of Language and Annotation Model Information in CoNLL CorporaabstractWe introduce AnnoHub, an on-going effort to automatically complement existing language resources with metadata about the languages they cover and the annotation schemes (tagsets) that they apply, to provide a web interface for their curation and evaluation by means of domain experts, and to publish them as a RDF dataset and as part of the (Linguistic) Linked Open Data (LLOD) cloud. In this paper, we focus on tabular formats with tab-separated values (TSV), a de-facto standard for annotated corpora as popularized as part of the CoNLL Shared Tasks. By extension, other formats for which a converter to CoNLL and/or TSV formats does exist, can be processed analoguously. We describe our implementation and its evaluation against a sample of 93 corpora from the Universal Dependencies, v.2.3. Frank Abromeit, Christian Chiarcos |
LDK | 2 |
| 2019 | Graph-Based Annotation Engineering: Towards a Gold Corpus for Role and Reference GrammarabstractThis paper describes the application of annotation engineering techniques for the construction of a corpus for Role and Reference Grammar (RRG). RRG is a semantics-oriented formalism for natural language syntax popular in comparative linguistics and linguistic typology, and predominantly applied for the description of non-European languages which are less-resourced in terms of natural language processing. Because of its cross-linguistic applicability and its conjoint treatment of syntax and semantics, RRG also represents a promising framework for research challenges within natural language processing. At the moment, however, these have not been explored as no RRG corpus data is publicly available. While RRG annotations cannot be easily derived from any single treebank in existence, we suggest that they can be reliably inferred from the intersection of syntactic and semantic annotations as represented by, for example, the Universal Dependencies (UD) and PropBank (PB), and we demonstrate this for the English Web Treebank, a 250,000 token corpus of various genres of English internet text. The resulting corpus is a gold corpus for future experiments in natural language processing in the sense that it is built on existing annotations which have been created manually. A technical challenge in this context is to align UD and PB annotations, to integrate them in a coherent manner, and to distribute and to combine their information on RRG constituent and operator projections. For this purpose, we describe a framework for flexible and scalable annotation engineering based on flexible, unconstrained graph transformations of sentence graphs by means of SPARQL Update. Christian Chiarcos, Christian Fäth |
LDK | 1 |
| 2019 | Ligt: An LLOD-Native Vocabulary for Representing Interlinear Glossed Text as RDFabstractThe paper introduces Ligt, a native RDF vocabulary for representing linguistic examples as text with interlinear glosses (IGT) in a linked data formalism. Interlinear glossing is a notation used in various fields of linguistics to provide readers with a way to understand linguistic phenomena and to provide corpus data when documenting endangered languages. This data is usually provided with morpheme-by-morpheme correspondence which is not supported by any established vocabularies for representing linguistic corpora or automated annotations. Interlinear Glossed Text can be stored and exchanged in several formats specifically designed for the purpose, but these differ in their designs and concepts, and they are tied to particular tools, so the reusability of the annotated data is limited. To improve interoperability and reusability, we propose to convert such glosses to a tool-independent representation well-suited for the Web of Data, i.e., a representation in RDF. Beyond establishing structural (format) interoperability by means of a common data representation, our approach also allows using shared vocabularies and terminology repositories available from the (Linguistic) Linked Open Data cloud. We describe the core vocabulary and the converters that use this vocabulary to convert IGT in a format of various widely-used tools into RDF. Ultimately, a Linked Data representation will facilitate the accessibility of language data from less-resourced language varieties within the (Linguistic) Linked Open Data cloud, as well as enable novel ways to access and integrate this information with (L)LOD dictionary data and other types of lexical-semantic resources. In a longer perspective, data currently only available through these formats will become more visible and reusable and contribute to the development of a truly multilingual (semantic) web. Christian Chiarcos, Maxim Ionov |
LDK | 1 |
| 2019 | CoNLL-Merge: Efficient Harmonization of Concurrent Tokenization and Textual VariationabstractThe proper detection of tokens in of running text represents the initial processing step in modular NLP pipelines. But strategies for defining these minimal units can differ, and conflicting analyses of the same text seriously limit the integration of subsequent linguistic annotations into a shared representation. As a solution, we introduce CoNLL Merge, a practical tool for harmonizing TSV-related data models, as they occur, e.g., in multi-layer corpora with non-sequential, concurrent tokenizations, but also in ensemble combinations in Natural Language Processing. CoNLL Merge works unsupervised, requires no manual intervention or external data sources, and comes with a flexible API for fully automated merging routines, validity and sanity checks. Users can chose from several merging strategies, and either preserve a reference tokenization (with possible losses of annotation granularity), create a common tokenization layer consisting of minimal shared subtokens (loss-less in terms of annotation granularity, destructive against a reference tokenization), or present tokenization clashes (loss-less and non-destructive, but introducing empty tokens as place-holders for unaligned elements). We demonstrate the applicability of the tool on two use cases from natural language processing and computational philology. Christian Chiarcos, Niko Schenk |
LDK | 1 |
| 2017 | CoNLL-RDF: Linked Corpora Done in an NLP-Friendly Way
Christian Chiarcos, Christian Fäth |
LDK | 1 |
| 2017 | LLODifying Linguistic Glosses
Christian Chiarcos, Maxim Ionov, Monika Rind-Pawlowski, Christian Fäth, Jesse Wichers Schreur, Irina Nevskaya |
LDK | 1 |
| 2017 | OnLiT: An Ontology for Linguistic TerminologyabstractUnderstanding the differences underlying the scope, usage and content of language data requires the provision of a clarifying terminological basis which is integrated in the metadata describing a particular language resource. While terminological resources such as the SIL Glossary of Linguistic Terms, ISOcat or the GOLD ontology provide a considerable amount of linguistic terms, their practical usage is limited to a look up of a defined term whose relation to other terms is unspecified or insufficient. Therefore, in this paper we propose an ontology for linguistic terminology, called OnLiT. It is a data model which can be used to represent linguistic terms and concepts in a semantically interrelated data structure and, thus, overcomes prevalent isolating definition-based term descriptions. OnLiT is based on the LiDo Glossary of Linguistic Terms and enables the creation of RDF datasets, that represent linguistic terms and their meanings within the whole or a subdomain of linguistics. Bettina Klimek, John P. McCrae, Christian Lehmann, Christian Chiarcos, Sebastian Hellmann 0001 |
LDK | 4 |
| 2012 | POWLA: Modeling Linguistic Corpora in OWL/DL
Christian Chiarcos |
ESWC | 1 |