Sylvia Melzer

dblp:24/3315 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-0144-5429ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Treating OCR Output as a Language (TOOL) - Improving OCR Output with Seq2Seq Translation
abstract
Optical Character Recognition (OCR) systems are frequently used to digitise text, but often produce noisy results, especially with historical, poor-quality or multilingual data.Despite advances in OCR technology, post-processing remains a significant bottleneck.We propose TOOL (Treating OCR Output as a Language), a new approach that understands OCR correction as a machine translation task.By treating noisy OCR text as a language in its own right, TOOL employs sequenceto-sequence models like Marian to translate it into clean, standardised text.This method is scalable, model-independent and language-flexible.We demonstrate this approach by translating "OCR German" to Standard German from around 1871 to the present day, improving accuracy at the token level by using matched training pairs of OCR output and base text.
Thomas Asselborn, Magnus Bender, Ralf Möller 0001, Sylvia Melzer
FedCSIS4
2023 EpiDoc Data Matching for Federated Information Retrieval in the Humanities
abstract
The importance of federated information retrieval (FIR) is growing in humanities research.Unlike traditional centralized information retrieval methods, where searches are conducted within a logically centralised collection of documents, FIR treats each information system as an independent source with its own unique characteristics.Searching these systems together as a centralised source results in lower precision in humanities research, even when the research data itself is structured and stored according to standardised guidelines such as EpiDoc, and requires the need to be able to trace the origin of records to avoid incorrect historical conclusions.Matching of queries against all data sets in each source is proving less effective.A global search index that enables traceable matching of key values deemed relevant would provide a more robust solution here.In this article, we propose a solution that introduces a novel EpiDoc data matching procedure, facilitating traceable FIR across distinct epigraphic sources.
Sylvia Melzer, Meike Klettke, Franziska Weise, Kaja Harter-Uibopuu, Ralf Möller 0001
FedCSIS1
2022 TEI-Based Interactive Critical Editions
Simon Schiff, Sylvia Melzer, Eva Wilden, Ralf Möller 0001
DAS2
2008 On Ontology Based Abduction for Text Interpretation
Irma Sofía Espinosa Peraldí, Atila Kaya, Sylvia Melzer, Ralf Möller 0001
CICLing3
2007 Towards a Media Interpretation Framework for the Semantic Web
abstract
We present a formal framework for media interpretation that leverages low-level information extraction to a higher level of abstraction in order to support semantics-based information retrieval for the Semantic Web. The overall goal of the framework is to provide high-level content descriptions of documents for maximizing precision and recall of semantics-based information retrieval.
Irma Sofía Espinosa Peraldí, Atila Kaya, Sylvia Melzer, Ralf Möller 0001, Michael Wessel
Web Intelligence3