Andrei-Marius Avram

dblp:248/7814 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0001-7045-038XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation
abstract
Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language processing, most of the datasets available for summarization are in English, posing challenges for developing summarization models in other languages. Thus, in this work, we introduce RoLargeSum, a novel large-scale summarization dataset for the Romanian language crawled from various publicly available news websites from Romania and the Republic of Moldova that were thoroughly cleaned to ensure a high-quality standard. RoLargeSum contains more than 615K news articles, together with their summaries, as well as their headlines, keywords, dialect, and other metadata that we found on the targeted websites. We further evaluated the performance of several BART variants and open-source large language models on RoLargeSum for benchmarking purposes. We manually evaluated the results of the best-performing system to gain insight into the potential pitfalls of this data set and future development.
Andrei-Marius Avram, Mircea Timpuriu, Andreea Iuga, Vlad-Cristian Matei, Iulian-Marius Taiatu, Tudor Gaina, Dumitru-Clementin Cercel, Mihaela-Claudia Cercel, Florin Pop
COLING1
2025 UniBERT: adversarial training for language-universal representations
abstract
Abstract This paper presents UniBERT, a compact multilingual language model that uses an innovative training framework that integrates three components: masked language modeling, adversarial training, and knowledge distillation. Pre-trained on a meticulously curated Wikipedia corpus spanning 107 languages, UniBERT is designed to reduce the computational demands of large-scale models while maintaining competitive performance across various natural language processing tasks. Comprehensive evaluations on four tasks, named entity recognition, natural language inference, question answering, and semantic textual similarity, demonstrate that our multilingual training strategy, enhanced by an adversarial objective, significantly improves cross-lingual generalization. Specifically, UniBERT models show an average relative improvement of 7.72% over traditional baselines, which achieved an average relative improvement of only 1.12%, and statistical analysis confirms the significance of these gains (p value = 0.0184). This work highlights the benefits of combining adversarial training and knowledge distillation to build robust and scalable language models, thus advancing the field of multilingual and cross-lingual natural language processing.
Andrei-Marius Avram, Marian Lupascu, Dumitru-Clementin Cercel, Ionut Mironica, Stefan Trausan-Matu
Neural Comput. Appl.1
2024 HistNERo: Historical Named Entity Recognition for the Romanian Language
Andrei-Marius Avram, Andreea Iuga, George-Vlad Manolache, Vlad-Cristian Matei, Razvan-Gabriel Miclius, Vlad-Andrei Muntean, Manuel-Petru Sorlescu, Dragos-Andrei Serban, Adrian-Dinu Urse, Vasile Florian Pais, Dumitru-Clementin Cercel
ICDAR (3)1
2023 TA-DA: Topic-Aware Domain Adaptation for Scientific Keyphrase Identification and Classification (Student Abstract)
abstract
Keyphrase identification and classification is a Natural Language Processing and Information Retrieval task that involves extracting relevant groups of words from a given text related to the main topic. In this work, we focus on extracting keyphrases from scientific documents. We introduce TA-DA, a Topic-Aware Domain Adaptation framework for keyphrase extraction that integrates Multi-Task Learning with Adversarial Training and Domain Adaptation. Our approach improves performance over baseline models by up to 5% in the exact match of the F1-score.
Razvan-Alexandru Smadu, George-Eduard Zaharia, Andrei-Marius Avram, Dumitru-Clementin Cercel, Mihai Dascalu, Florin Pop
AAAI3
2023 Adversarial Capsule Networks for Romanian Satire Detection and Sentiment Analysis
Sebastian-Vasile Echim, Razvan-Alexandru Smadu, Andrei-Marius Avram, Dumitru-Clementin Cercel, Florin Pop
NLDB3
2023 RoBERTweet: A BERT Language Model for Romanian Tweets
Iulian-Marius Taiatu, Andrei-Marius Avram, Dumitru-Clementin Cercel, Florin Pop
NLDB2
2022 Distilling the Knowledge of Romanian BERTs Using Multiple Teachers
abstract
Running large-scale pre-trained language models in computationally constrained environments remains a challenging problem yet to be addressed, while transfer learning from these models has become prevalent in Natural Language Processing tasks. Several solutions, including knowledge distillation, network quantization, or network pruning have been previously proposed; however, these approaches focus mostly on the English language, thus widening the gap when considering low-resource languages. In this work, we introduce three light and fast versions of distilled BERT models for the Romanian language: Distil-BERT-base-ro, Distil-RoBERT-base, and DistilMulti-BERT-base-ro. The first two models resulted from the individual distillation of knowledge from two base versions of Romanian BERTs available in literature, while the last one was obtained by distilling their ensemble. To our knowledge, this is the first attempt to create publicly available Romanian distilled BERT models, which were thoroughly evaluated on five tasks: part-of-speech tagging, named entity recognition, sentiment analysis, semantic textual similarity, and dialect identification. Our experimental results argue that the three distilled models offer performance comparable to their teachers, while being twice as fast on a GPU and ~35% smaller. In addition, we further test the similarity between the predictions of our students versus their teachers by measuring their label and probability loyalty, together with regression loyalty - a new metric introduced in this work.
Andrei-Marius Avram, Darius Catrina, Dumitru-Clementin Cercel, Mihai Dascalu, Traian Rebedea, Vasile Florian Pais, Dan Tufis
LREC1
2021 A Modular Approach for Romanian-English Speech Translation
Andrei-Marius Avram, Vasile Florian Pais, Dan Tufis
NLDB1
2020 Introducing RONEC - the Romanian Named Entity Corpus
abstract
We present RONEC - the Named Entity Corpus for the Romanian language. The corpus contains over 26000 entities in ~5000 annotated sentences, belonging to 16 distinct classes. The sentences have been extracted from a copy-right free newspaper, covering several styles. This corpus represents the first initiative in the Romanian language space specifically targeted for named entity recognition. It is available in BRAT and CoNLL-U Plus formats, and it is free to use and extend at github.com/dumitrescustefan/ronec
Stefan Daniel Dumitrescu, Andrei-Marius Avram
LREC2