Cerstin Mahlow

dblp:40/1283 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
3since 2021 · last 2025
0000-0003-0215-5551ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author
YearPublicationVenuePosition
2025 Session details: DocEng Demonstrations
Cerstin Mahlow
DocEng1
2025 Designing Visual Tools for Writing Process Analysis
abstract
Understanding how texts are produced is crucial not only for the development of theoretical models and writing strategies, but also for practical applications. However, the writing process itself--- including intermediate versions, copy-paste actions, input from co-authors or LLMs---remains invisible in the final text. This study addresses this gap by visualizing fine-grained keystroke logging data to capture both the product (final text) and the process (writer's actions) at sentence and text level. We design and implement custom JavaScript visualizations of linguistically processed keystroke logging data. Our pilot study examines data from nine students writing under identical conditions; we analyze temporal, spatial, and structural aspects of writing. The results reveal diverse, nonlinear writing strategies and suggest that individualized process visualizations can inform both document engineering and writing analytics. The novel visualization types we present demonstrate how process and product can be meaningfully integrated.
Cerstin Mahlow
DocEng1
2022 Academic writing and publishing beyond documents
abstract
Research on writing tools stopped in the late 1980s when Microsoft Word had achieved monopoly status. However, the development of the Web and the advent of mobile devices are increasingly rendering static print-like documents obsolete. In this vision paper we reflect on the impact of this development on scholarly writing and publishing. Academic publications increasingly include dynamic elements, e.g., code, data plots, and other visualizations, which clearly requires other tools for document production than traditional word processors. When the printed page no longer is the desired final product, content and form can be addressed explicitly and separately, thus emphasizing the structure of texts rather than the structure of documents. The resulting challenges have not yet been fully addressed by document engineering.
Cerstin Mahlow, Michael Piotrowski
DocEng1
2020 Swiss-AL: A Multilingual Swiss Web Corpus for Applied Linguistics
abstract
The Swiss Web Corpus for Applied Linguistics (Swiss-AL) is a multilingual (German, French, Italian) collection of texts from selected web sources. Unlike most other web corpora it is not intended for NLP purposes, but rather designed to support data-based and data-driven research on societal and political discourses in Switzerland. It currently contains 8 million texts (approx. 1.55 billion tokens), including news and specialist publications, governmental opinions, and parliamentary records, web sites of political parties, companies, and universities, statements from industry associations and NGOs, etc. A flexible processing pipeline using state-of-the-art components allows researchers in applied linguistics to create tailor-made subcorpora for studying discourse in a wide range of domains. So far, Swiss-AL has been used successfully in research on Swiss public discourses on energy and on antibiotic resistance.
Julia Krasselt, Philipp Dressen, Matthias Fluor, Cerstin Mahlow, Klaus Rothenhäusler, Maren Runte
LREC4
2016 C-WEP―Rich Annotated Collection of Writing Errors by Professionals
Cerstin Mahlow
LREC1
2012 A framework for retrieval and annotation in digital humanities using XQuery full text and update in BaseX
abstract
A key difference between traditional humanities research and the emerging field of digital humanities is that the latter aims to complement qualitative methods with quantitative data. In linguistics, this means the use of large corpora of text, which are usually annotated automatically using natural language processing tools. However, these tools do not exist for historical texts, so scholars have to work with unannotated data. We have developed a system for systematic, iterative exploration and annotation of historical text corpora, which relies on an XML database (BaseX) and in particular on the Full Text and Update facilities of XQuery.
Cerstin Mahlow, Christian Grün, Alexander Holupirek, Marc H. Scholl
ACM Symposium on Document Engineering1
2009 Linguistic editing support
abstract
Unlike programmers, authors only get very little support from their writing tools, i.e., their word processors and editors. Current editors are unaware of the objects and structures of natural languages and only offer character-based operations for manipulating text. Writers thus have to execute complex sequences of low-level functions to achieve their rhetoric or stylistic goals while composing. Software requiring long and complex sequences of operations causes users to make slips. In the case of editing and revising, these slips result in typical revision errors, such as sentences without a verb, agreement errors, or incorrect word order. In the LingURed project, we are developing language-aware editing functions to prevent errors. These functions operate on linguistic elements, not characters, thus shortening the command sequences writers have to execute. This paper describes the motivation and background of the LingURed project and shows some prototypical language-aware functions.
Michael Piotrowski, Cerstin Mahlow
ACM Symposium on Document Engineering2
2008 Linguistic Support for Revising and Editing
Cerstin Mahlow, Michael Piotrowski
CICLing1