Verginica Barbu Mititelu

dblp:22/5737 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0003-1945-2587ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 9 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 The Romanian Corpus Annotated with Multiword Expressions. PARSEME-Ro Version 2.0
Verginica Barbu Mititelu, Mihaela Cristescu, Elena Irimia, Carmen Mîrzea Vasile
LREC1
2026 PARSEME 2.0 Multilingual Corpus of Multiword Expressions
abstract
International audience
Agata Savary, Manon Scholivet, Carlos Ramisch, Takuya Nakamura, Eric Bilinski, Sara Stymne, Voula Giouli, Stella Markantonatou, Vasile Florian Pais, Maria Mitrofan, Louis Estève, Bruno Guillaume, Verginica Barbu Mititelu, Jaka Cibej, Roberto Díaz Hernández, Victoria Fendel, Polona Gantar, Olha Kanishcheva, Cvetana Krstev, Chaya Liebeskind, Irina Lobzhanidze, Aleksandra M. Markovic, Gunta Nespore-Berzkalne, Adriana S. Pagano, Mehrnoush Shamsfard, Ranka Stankovic, Vahideh Tajalli, Carole Tiberius, Aakanksha Padhye
LREC13
2023 Adopting Linguistic Linked Data Principles: Insights on Users' Experience
Verginica Barbu Mititelu, Maria Pia di Buono, Hugo Gonçalo Oliveira, Blerina Spahiu, Giedre Valunaite Oleskeviciene
LDK1
2023 The Romanian Wordnet in Linked Open Data Format
abstract
In this paper we present the standardization of the Romanian Wordnet by means of conversion to the Linked Open Data format.We describe the vocabularies used to encode data and metadata of this resource.The decisions made are in accordance with the characteristics of the Romanian Wordnet, which are the outcome of the development method, enrichment strategies and resources used for its creations.By interlinking with other resources, words in the Romanian Wordnet have now the pronunciation associated, as well as syntagmatic information, in the form of contexts of occurrences.
Elena Irimia, Verginica Barbu Mititelu
GWC2
2023 RoLEX: The development of an extended Romanian lexical dataset and its evaluation at predicting concurrent lexical information
abstract
Abstract In this article, we introduce an extended, freely available resource for the Romanian language, named RoLEX. The dataset was developed mainly for speech processing applications, yet its applicability extends beyond this domain. RoLEX includes over 330,000 curated entries with information regarding lemma, morphosyntactic description, syllabification, lexical stress and phonemic transcription. The process of selecting the list of word entries and semi-automatically annotating the complete lexical information associated with each of the entries is thoroughly described. The dataset’s inherent knowledge is then evaluated in a task of concurrent prediction of syllabification, lexical stress marking and phonemic transcription. The evaluation looked into several dataset design factors, such as the minimum viable number of entries for correct prediction, the optimisation of the minimum number of required entries through expert selection and the augmentation of the input with morphosyntactic information, as well as the influence of each task in the overall accuracy. The best results were obtained when the orthographic form of the entries was augmented with the complete morphosyntactic tags. A word error rate of 3.08% and a character error rate of 1.08% were obtained this way. We show that using a carefully selected subset of entries for training can result in a similar performance to the performance obtained by a larger set of randomly selected entries (twice as many). In terms of prediction complexity, the lexical stress marking posed most problems and accounts for around 60% of the errors in the predicted sequence.
Beáta Lorincz, Elena Irimia, Adriana Cornelia Stan, Verginica Barbu Mititelu
Nat. Lang. Eng.4
2022 Aligning the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs
abstract
We present here the efforts of aligning two language resources for Romanian: the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs: for each occurrence of those verbs in the treebank that were included as entries in the lexicon, a set of valence frames is automatically assigned, then manually validated by two linguists and, when necessary, corrected. Validating a valence frame also means semantically disambiguating the verb in the respective context. The validation is done by two linguists, on complementary datasets. However, a subset of verbs were validated by both annotators and Cohen’s κ is 0.87 for this subset. The alignment we have made also serves as a method of enhancing the quality of the two resources, as in the process we identify morpho-syntactic annotation mistakes, incomplete valence frames or missing ones. Information from each resource complements the information from the other, thus their value increases. The treebank and the lexicon are freely available, while the links discovered between them are also made available on GitHub.
Ana-Maria Barbu, Verginica Barbu Mititelu, Catalin Mititelu
LREC2
2022 Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources
abstract
This article presents the current outcomes of the CURLICAT CEF Telecom project, which aims to collect and deeply annotate a set of large corpora from selected domains. The CURLICAT corpus includes 7 monolingual corpora (Bulgarian, Croatian, Hungarian, Polish, Romanian, Slovak and Slovenian) containing selected samples from respective national corpora. These corpora are automatically tokenized, lemmatized and morphologically analysed and the named entities annotated. The annotations are uniformly provided for each language specific corpus while the common metadata schema is harmonised across the languages. Additionally, the corpora are annotated for IATE terms in all languages. The file format is CoNLL-U Plus format, containing the ten columns specific to the CoNLL-U format and three extra columns specific to our corpora as defined by Varádi et al. (2020). The CURLICAT corpora represent a rich and valuable source not just for training NMT models, but also for further studies and developments in machine learning, cross-lingual terminological data extraction and classification.
Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadic, Vanja Stefanec, Maciej Ogrodniczuk, Bartlomiej Niton, Piotr Pezik, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan, Dan Tufis, Radovan Garabík, Simon Krek, Andraz Repar
LREC9
2021 Semantic Analysis of Verb-Noun Derivation in Princeton WordNet
abstract
We present here the results of a morphosemantic analysis of the verb-noun pairs in the Princeton WordNet as reflected in the standoff file containing pairs annotated with a set of 14 semantic relations.We have automatically distinguished between zero-derivation and affixal derivation in the data and identified the affixes and manually checked the results.The data show that for each semantic relation an affix prevails in creating new words, although we cannot talk about their specificity with respect to such a relation.Moreover, certain pairs of verb-noun semantic primes are better represented for each semantic relation, and some semantic clusters (in the form of WordNet subtrees) take shape as a result.We thus employ a large-scale data-driven linguistically motivated analysis afforded by the rich derivational and morphosemantic description in WordNet to the end of capturing finer regularities in the process of derivation as represented in the semantic properties of the words involved and as reflected in the structure of the lexicon.
Verginica Barbu Mititelu, Svetlozara Leseva, Ivelina Stoyanova
GWC1
2020 The MARCELL Legislative Corpus
abstract
This article presents the current outcomes of the MARCELL CEF Telecom project aiming to collect and deeply annotate a large comparable corpus of legal documents. The MARCELL corpus includes 7 monolingual sub-corpora (Bulgarian, Croatian, Hungarian, Polish, Romanian, Slovak and Slovenian) containing the total body of respective national legislative documents. These sub-corpora are automatically sentence split, tokenized, lemmatized and morphologically and syntactically annotated. The monolingual sub-corpora are complemented by a thematically related parallel corpus (Croatian-English). The metadata and the annotations are uniformly provided for each language specific sub-corpus. Besides the standard morphosyntactic analysis plus named entity and dependency annotation, the corpus is enriched with the IATE and EUROVOC labels. The file format is CoNLL-U Plus Format, containing the ten columns specific to the CoNLL-U format and four extra columns specific to our corpora. The MARCELL corpora represents a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification.
Tamás Váradi, Svetla Koeva, Martin Yamalov, Marko Tadic, Bálint Sass, Bartlomiej Niton, Maciej Ogrodniczuk, Piotr Pezik, Verginica Barbu Mititelu, Radu Ion, Elena Irimia, Maria Mitrofan, Vasile Florian Pais, Dan Tufis, Radovan Garabík, Simon Krek, Andraz Repar, Matjaz Rihtar, Janez Brank
LREC9
2019 Evaluating the Wordnet and CoRoLa-based Word Embedding Vectors for Romanian as Resources in the Task of Microworlds Lexicon Expansion
abstract
Within a larger frame of facilitating human-robot interaction, we present here the creation of a core vocabulary to be learned by a robot.It is extracted from two tokenised and lemmatized scenarios pertaining to two imagined microworlds in which the robot is supposed to play an assistive role.We also evaluate two resources for their utility for expanding this vocabulary so as to better cope with the robot's communication needs.The language under study is Romanian and the resources used are the Romanian wordnet and word embedding vectors extracted from the large representative corpus of contemporary Romanian, CoRoLa.The evaluation is made for two situations: one in which the words are not semantically disambiguated before expanding the lexicon, and another one in which they are disambiguated with senses from the Romanian wordnet.The appropriateness of each resource is discussed.
Elena Irimia, Maria Mitrofan, Verginica Barbu Mititelu
GWC3
2019 Leaving No Stone Unturned When Identifying and Classifying Verbal Multiword Expressions in the Romanian Wordnet
abstract
We present here the enhancement of the Romanian wordnet with a new type of information, very useful in language processing, namely types of verbal multiword expressions.All verb literals made of two or more words are attached a label specific to the type of verbal multiword expression they correspond to.These labels were created in the PARSEME Cost Action and were used in the version 1.1 of the shared task they organized.The results of this annotation are compared to those obtained in the annotation of a Romanian news corpus with the same labels.Given the alignment of the Romanian wordnet to the Princeton WordNet, this type of annotation can be further used for drawing comparisons between equivalent verbal literals in various languages, provided that such information is annotated in the wordnets of the respective languages and their wordnets are aligned to Princeton WordNet, and thus to the Romanian wordnet.
Verginica Barbu Mititelu, Maria Mitrofan
GWC1
2018 Ensemble Romanian Dependency Parsing with Neural Networks
Radu Ion, Elena Irimia, Verginica Barbu Mititelu
LREC3
2018 The Reference Corpus of the Contemporary Romanian Language (CoRoLa)
Verginica Barbu Mititelu, Dan Tufis, Elena Irimia
LREC1
2018 Investigating English Affixes and their Productivity with Princeton WordNet
abstract
Such a rich language resource like Princeton WordNet, containing linguistic information of different types (semantic, lexical, syntactic, derivational, dialectal, etc.), is a thesaurus which is worth both being used in various language-enabled applications and being explored in order to study a language.In this paper we show how we used Princeton WordNet version 3.0 to study the English affixes.We extracted pairs of base-derived words and identified the affixes by means of which the derived words were created from their bases.We distinguished among four types of derivation depending on the type of overlapping between the senses of the base word and those of the derived word that are linked by derivational relations in Princeton WordNet.We studied the behaviour of affixes with respect to these derivation types.Drawing on these data, we inferred about their productivity.
Verginica Barbu Mititelu
GWC1
2016 The IPR-cleared Corpus of Contemporary Written and Spoken Romanian Language
Dan Tufis, Verginica Barbu Mititelu, Elena Irimia, Stefan Daniel Dumitrescu, Tiberiu Boros
LREC2
2014 CoRoLa ― The Reference Corpus of Contemporary Romanian Language
Verginica Barbu Mititelu, Elena Irimia, Dan Tufis
LREC1
2014 News about the Romanian Wordnet
abstract
There are more than 60 wordnets worldwide; the Romanian wordnet is among those that are maintained and further developed.Begun within the BalkaNet project and further enriched in various (application oriented) projects, it was used in word sense disambiguation, machine translation and question answering with promising results.We present here the latest qualitative and quantitative improvements of our lexical resource, special attention being paid to derivational relations, the latest statistics, as well as the development of an Application Programming Interface, meant to facilitate work with the wordnet, both for its further development purposes and for its use in applications.In the context of creating a common European research infrastruc
Verginica Barbu Mititelu, Stefan Daniel Dumitrescu, Dan Tufis
GWC1
2012 Adding Morpho-semantic Relations to the Romanian Wordnet
Verginica Barbu Mititelu
LREC1
2008 Annotation of WordNet Verbs with TimeML Event Classes
Georgiana Puscasu, Verginica Barbu Mititelu
LREC2
2006 Romanian Valence Dictionary in XML Format
Ana-Maria Barbu, Emil Ionescu, Verginica Barbu Mititelu
LREC3