Razvan-Alexandru Smadu

dblp:290/1947 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-5660-2225ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
abstract
Satire, irony, and sarcasm are techniques that are typically used humorously or critically, rather than deceptively; they can occasionally be mistaken for factual reporting, akin to fake news. These techniques can be applied at a more granular level, allowing satirical information to be incorporated into news articles. In this paper, we introduce the first sentence-level dataset for Romanian satire detection for news articles, called SeLeRoSa. The dataset comprises 13,873 manually annotated sentences spanning various domains, including social issues, IT, science, and movies. With the rise and recent progress of large language models (LLMs) in the natural language processing literature, LLMs have demonstrated enhanced capabilities to tackle various tasks in zero-shot settings. We evaluate multiple baseline models based on LLMs in both zero-shot and fine-tuning settings, as well as transformer-based models. Our findings reveal the current limitations of these models in the sentence-level satire detection task, paving the way for new research directions.
Razvan-Alexandru Smadu, Andreea Iuga, Dumitru-Clementin Cercel, Florin Pop
CIKM1
2024 Benchmarking Adversarial Robustness in Speech Emotion Recognition: Insights into Low-Resource Romanian and German Languages
abstract
Therapy, interviews, and emergency services assisted by artificial intelligence (AI) are applications where speech emotion recognition (SER) plays an essential role, for which performance and robustness are subject to improvement. Deep learning approaches have proven effective in SER; nevertheless, they can underperform when exposed to adversarial attacks. In this paper, we explore and enhance architectures, such as convolutional neural networks with long short-term memory (CNN-LSTM), AlexNet, VGG16, Convolutional Vision Transformer (CvT), Vision Transformer (ViT), and LeViT, by finding the suitable setup for SER models regarding speech processing, network hyperparameters, spectrogram augmentations, and adversarial examples. We apply our methodology to Romanian and German SER datasets and achieve state-of-the-art results, with 89.81% validation weighted accuracy and 98.09% average weighted accuracy on the trained models. Our highly robust models reach complete adversarial defense and up to 5.56% weighted accuracy improvement when attacked. We also show how adversarial attacks influence model behavior in SER through explainable AI techniques.
Sebastian-Vasile Echim, Razvan-Alexandru Smadu, Dumitru-Clementin Cercel
ECAI2
2024 Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain Setups
abstract
Complex Word Identification (CWI) is an essential step in the lexical simplification task and has recently become a task on its own.Some variations of this binary classification task have emerged, such as lexical complexity prediction (LCP) and complexity evaluation of multi-word expressions (MWE).Large language models (LLMs) recently became popular in the Natural Language Processing community because of their versatility and capability to solve unseen tasks in zero/few-shot settings.Our work investigates LLM usage, specifically open-source models such as Llama 2, Llama 3, and Vicuna v1.5, and closed-source, such as ChatGPT-3.5turboand GPT-4o, in the CWI, LCP, and MWE settings.We evaluate zero-shot, few-shot, and fine-tuning settings and show that LLMs struggle in certain conditions or achieve comparable results against existing methods.In addition, we provide some views on meta-learning combined with prompt learning.In the end, we conclude that the current state of LLMs cannot or barely outperform existing methods, which are usually much smaller.
Razvan-Alexandru Smadu, David-Gabriel Ion, Dumitru-Clementin Cercel, Florin Pop, Mihaela-Claudia Cercel
EMNLP1
2024 Enhancing Romanian Offensive Language Detection Through Knowledge Distillation, Multi-task Learning, and Data Augmentation
Vlad-Cristian Matei, Iulian-Marius Taiatu, Razvan-Alexandru Smadu, Dumitru-Clementin Cercel
NLDB (1)3
2024 A Cross-Lingual Meta-Learning Method Based on Domain Adaptation for Speech Emotion Recognition
David-Gabriel Ion, Razvan-Alexandru Smadu, Dumitru-Clementin Cercel, Florin Pop, Mihaela-Claudia Cercel
WISE (1)2
2023 TA-DA: Topic-Aware Domain Adaptation for Scientific Keyphrase Identification and Classification (Student Abstract)
abstract
Keyphrase identification and classification is a Natural Language Processing and Information Retrieval task that involves extracting relevant groups of words from a given text related to the main topic. In this work, we focus on extracting keyphrases from scientific documents. We introduce TA-DA, a Topic-Aware Domain Adaptation framework for keyphrase extraction that integrates Multi-Task Learning with Adversarial Training and Domain Adaptation. Our approach improves performance over baseline models by up to 5% in the exact match of the F1-score.
Razvan-Alexandru Smadu, George-Eduard Zaharia, Andrei-Marius Avram, Dumitru-Clementin Cercel, Mihai Dascalu, Florin Pop
AAAI1
2023 Adversarial Capsule Networks for Romanian Satire Detection and Sentiment Analysis
Sebastian-Vasile Echim, Razvan-Alexandru Smadu, Andrei-Marius Avram, Dumitru-Clementin Cercel, Florin Pop
NLDB2
2022 Domain Adaptation in Multilingual and Multi-Domain Monolingual Settings for Complex Word Identification
abstract
Complex word identification (CWI) is a cornerstone process towards proper text simplification.CWI is highly dependent on context, whereas its difficulty is augmented by the scarcity of available datasets which vary greatly in terms of domains and languages.As such, it becomes increasingly more difficult to develop a robust model that generalizes across a wide array of input examples.In this paper, we propose a novel training technique for the CWI task based on domain adaptation to improve the target character and context representations.This technique addresses the problem of working with multiple domains, inasmuch as it creates a way of smoothing the differences between the explored datasets.Moreover, we also propose a similar auxiliary task, namely text simplification, that can be used to complement lexical complexity prediction.Our model obtains a boost of up to 2.42% in terms of Pearson Correlation Coefficients in contrast to vanilla training techniques, when considering the CompLex from the Lexical Complexity Prediction 2021 dataset.At the same time, we obtain an increase of 3% in Pearson scores, while considering a cross-lingual setup relying on the Complex Word Identification 2018 dataset.In addition, our model yields state-ofthe-art results in terms of Mean Absolute Error.
George-Eduard Zaharia, Razvan-Alexandru Smadu, Dumitru-Clementin Cercel, Mihai Dascalu
ACL (1)2