Bernardo Stearns

dblp:229/4352 · DBLP profile ↗
← Back
3ranked-venue papers in the field
0as first author
3since 2021 · last 2025
0000-0001-9377-8572ORCID · corroborated

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 3
YearPublicationVenuePosition
2025 Cuaċ: Fast and Small Universal Representations of Corpora
abstract
The increasing size and diversity of corpora in natural language processing requires highly efficient processing frameworks. Building on the universal corpus format, Teanga, we present Cuaċ, a format for the compact representation of corpora. We describe this methodology based on short-string compression and indexing techniques and show that the files created with this methodology are similar to compressed human-readable serializations and can be further compressed using lossless compression. We also show that this introduces no computational penalty on the time to process files. This methodology aims to speed up natural language processing pipelines and is the basis for a fast database system for corpora.
John P. McCrae, Bernardo Stearns, Alamgir Munir Qazi, Shubhanker Banerjee, Atul Kr. Ojha
LDK2
2023 The Cardamom Workbench for Historical and Under-Resourced Languages
Adrian Doyle, Theodorus Fransen, Bernardo Stearns, John P. McCrae, Oksana Dereza, Priya Rani
LDK3
2023 A new learner language data set for the study of English for Specific Purposes at university
Cyriel Mallart, Nicolas Ballier, Jen-Yu Li, Andrew J. Simpkin, Bernardo Stearns, Rémi Venant, Thomas Gaillat
LDK5