Laurentiu Zoicas

dblp:266/0981 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 100%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational social science and digital humanities
historical linguistics
1.422024
Verba volant, scripta volant? Don't worry! There are computational solutions for protoword reconstruction · EMNLP 2024
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification · EMNLP 2023
Computational social science and digital humanities › computational linguistics
cognate identification
0.712023
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification · EMNLP 2023
Natural language and speech › Information extraction and text analysis
multilingual NLP
0.212023
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

machine learning · 1.3deep learning · 1.3sequence modeling · 0.8computational historical linguistics · 0.8
YearPublicationVenuePosition
2024 Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages
abstract
Identifying the type of relationship between words (cognates, borrowings, inherited) provides a deeper insight into the history of a language and allows for a better characterization of language relatedness. In this paper, we propose a computational approach for discriminating between cognates and borrowings, one of the most difficult tasks in historical linguistics. We compare the discriminative power of graphic and phonetic features and we analyze the underlying linguistic factors that prove relevant in the classification task. We perform experiments for pairs of languages in the Romance language family (French, Italian, Spanish, Portuguese, and Romanian), based on a comprehensive database of Romance cognates and borrowings. To our knowledge, this is one of the first attempts of this kind and the most comprehensive in terms of covered languages.
Liviu P. Dinu, Ana Sabina Uban, Ioan-Bogdan Iordache, Alina Maria Cristea, Simona Georgescu, Laurentiu Zoicas
LREC/COLING6
2024 Verba volant, scripta volant? Don't worry! There are computational solutions for protoword reconstruction
abstract
Liviu P Dinu, Ana Sabina Uban, Alina Maria Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Liviu P. Dinu, Ana Sabina Uban, Alina Maria Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas
EMNLP7
2023 RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification
abstract
The identification of cognates is a fundamental process in historical linguistics, on which any further research is based.Even though there are several cognate databases for Romance languages, they are rather scattered, incomplete, noisy, contain unreliable information, or have uncertain availability.In this paper we introduce a comprehensive database of Romance cognates and borrowings based on the etymological information provided by the dictionaries (the largest known database of this kind, in our best knowledge).We extract pairs of cognates between any two Romance languages by parsing electronic dictionaries of Romanian, Italian, Spanish, Portuguese and French.Based on this resource, we propose a strong benchmark for the automatic detection of cognates, by applying machine learning and deep learning based methods on any two pairs of Romance languages.We find that automatic identification of cognates is possible with accuracy averaging around 94% for the more difficult task formulations.
Liviu P. Dinu, Ana Sabina Uban, Alina Maria Cristea, Anca P. Dinu, Ioan-Bogdan Iordache, Simona Georgescu, Laurentiu Zoicas
EMNLP7
2020 Automatic Reconstruction of Missing Romanian Cognates and Unattested Latin Words
abstract
Producing related words is a key concern in historical linguistics. Given an input word, the task is to automatically produce either its proto-word, a cognate pair or a modern word derived from it. In this paper, we apply a method for producing related words based on sequence labeling, aiming to fill in the gaps in incomplete cognate sets in Romance languages with Latin etymology (producing Romanian cognates that are missing) and to reconstruct uncertified Latin words. We further investigate an ensemble-based aggregation for combining and re-ranking the word productions of multiple languages.
Alina Maria Cristea, Liviu P. Dinu, Laurentiu Zoicas
LREC3