Ana Sabina Uban

dblp:174/7148 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0003-2197-3947ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational social science and digital humanities · 87% Computing education · 13%
Artificial intelligence
2 papers
Information extraction and text analysis · 88% Machine translation · 12%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational social science and digital humanities
historical linguistics
1.422024
Verba volant, scripta volant? Don't worry! There are computational solutions for protoword reconstruction · EMNLP 2024
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification · EMNLP 2023
Natural language and speech › Information extraction and text analysis › multilingual NLP
cognate identification
0.912025
Friend or Foe? A Computational Investigation of Semantic False Friends across Romance Languages · EMNLP 2025
Natural language and speech › Information extraction and text analysis
lexical semantics
0.912025
Friend or Foe? A Computational Investigation of Semantic False Friends across Romance Languages · EMNLP 2025
Computational social science and digital humanities › computational linguistics
cognate identification
0.712023
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification · EMNLP 2023
Computational social science and digital humanities › science of science
peer review
0.412019
The Myth of Double-Blind Review Revisited: ACL vs. EMNLP · EMNLP/IJCNLP (1) 2019
Computing education
research ethics
0.412019
The Myth of Double-Blind Review Revisited: ACL vs. EMNLP · EMNLP/IJCNLP (1) 2019
Natural language and speech › Information extraction and text analysis
multilingual NLP
0.212023
RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification · EMNLP 2023
Computational social science and digital humanities
scientometrics
0.112019
The Myth of Double-Blind Review Revisited: ACL vs. EMNLP · EMNLP/IJCNLP (1) 2019

Methods — techniques the papers use, named apart from their topics

machine learning · 1.3deep learning · 1.3etymological dictionary analysis · 0.9sequence modeling · 0.8computational historical linguistics · 0.8statistical analysis · 0.4
YearPublicationVenuePosition
2025 Friend or Foe? A Computational Investigation of Semantic False Friends across Romance Languages
abstract
In this paper we present a comprehensive analysis of lexical semantic divergence between cognate words and borrowings in the Romance languages.We experiment with different algorithms for false friend detection including deceptive cognate and deceptive borrowings and correction and evaluate them systematically on cognate and borrowing pairs in the five Romance languages.We use the most complete and reliable dataset of cognate words and borrowings based on etymological dictionaries for the five main Romance languages (Italian, Spanish, Portuguese, French and Romanian) to extract deceptive cognates and borrowings automatically based on usage, and freely publish the lexicon of obtained true and deceptive cognate and borrowings in every Romance language pair.
Ana Sabina Uban, Liviu P. Dinu, Ioan-Bogdan Iordache, Simona Georgescu, Claudia Vlad
EMNLP1
2024 Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages
abstract
Identifying the type of relationship between words (cognates, borrowings, inherited) provides a deeper insight into the history of a language and allows for a better characterization of language relatedness. In this paper, we propose a computational approach for discriminating between cognates and borrowings, one of the most difficult tasks in historical linguistics. We compare the discriminative power of graphic and phonetic features and we analyze the underlying linguistic factors that prove relevant in the classification task. We perform experiments for pairs of languages in the Romance language family (French, Italian, Spanish, Portuguese, and Romanian), based on a comprehensive database of Romance cognates and borrowings. To our knowledge, this is one of the first attempts of this kind and the most comprehensive in terms of covered languages.
Liviu P. Dinu, Ana Sabina Uban, Ioan-Bogdan Iordache, Alina Maria Cristea, Simona Georgescu, Laurentiu Zoicas
LREC/COLING2
2024 Verba volant, scripta volant? Don't worry! There are computational solutions for protoword reconstruction
abstract
Liviu P Dinu, Ana Sabina Uban, Alina Maria Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Liviu P. Dinu, Ana Sabina Uban, Alina Maria Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas
EMNLP2
2023 RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification
abstract
The identification of cognates is a fundamental process in historical linguistics, on which any further research is based.Even though there are several cognate databases for Romance languages, they are rather scattered, incomplete, noisy, contain unreliable information, or have uncertain availability.In this paper we introduce a comprehensive database of Romance cognates and borrowings based on the etymological information provided by the dictionaries (the largest known database of this kind, in our best knowledge).We extract pairs of cognates between any two Romance languages by parsing electronic dictionaries of Romanian, Italian, Spanish, Portuguese and French.Based on this resource, we propose a strong benchmark for the automatic detection of cognates, by applying machine learning and deep learning based methods on any two pairs of Romance languages.We find that automatic identification of cognates is possible with accuracy averaging around 94% for the more difficult task formulations.
Liviu P. Dinu, Ana Sabina Uban, Alina Maria Cristea, Anca P. Dinu, Ioan-Bogdan Iordache, Simona Georgescu, Laurentiu Zoicas
EMNLP2
2022 Investigating the Relationship Between Romanian Financial News and Closing Prices from the Bucharest Stock Exchange
abstract
A new data set is gathered from a Romanian financial news website for the duration of four years. It is further refined to extract only information related to one company by selecting only paragraphs and even sentences that referred to it. The relation between the extracted sentiment scores of the texts and the stock prices from the corresponding dates is investigated using various approaches like the lexicon-based Vader tool, Financial BERT, as well as Transformer-based models. Automated translation is used, since some models could be only applied for texts in English. It is encouraging that all models, be that they are applied to Romanian or English texts, indicate a correlation between the sentiment scores and the increase or decrease of the stock closing prices.
Ioan-Bogdan Iordache, Ana Sabina Uban, Catalin Stoean, Liviu P. Dinu
LREC2
2022 Multi-Aspect Transfer Learning for Detecting Low Resource Mental Disorders on Social Media
abstract
Mental disorders are a serious and increasingly relevant public health issue. NLP methods have the potential to assist with automatic mental health disorder detection, but building annotated datasets for this task can be challenging; moreover, annotated data is very scarce for disorders other than depression. Understanding the commonalities between certain disorders is also important for clinicians who face the problem of shifting standards of diagnosis. We propose that transfer learning with linguistic features can be useful for approaching both the technical problem of improving mental disorder detection in the context of data scarcity, and the clinical problem of understanding the overlapping symptoms between certain disorders. In this paper, we target four disorders: depression, PTSD, anorexia and self-harm. We explore multi-aspect transfer learning for detecting mental disorders from social media texts, using deep learning models with multi-aspect representations of language (including multiple types of interpretable linguistic features). We explore different transfer learning strategies for cross-disorder and cross-platform transfer, and show that transfer learning can be effective for improving prediction performance for disorders where little annotated data is available. We offer insights into which linguistic features are the most useful vehicles for transferring knowledge, through ablation experiments, as well as error analysis.
Ana Sabina Uban, Berta Chulvi, Paolo Rosso
LREC1
2022 Detecting Early Signs of Depression in the Conversational Domain: The Role of Transfer Learning in Low-Resource Scenarios
Petr Lorenc, Ana Sabina Uban, Paolo Rosso, Jan Sedivý
NLDB2
2021 On the Explainability of Automatic Predictions of Mental Disorders from Social Media Data
Ana Sabina Uban, Berta Chulvi, Paolo Rosso
NLDB1
2021 An emotion and cognitive based analysis of mental health disorders from social media data
abstract
Mental disorders can severely affect quality of life, constitute a major predictive factor of suicide, and are usually underdiagnosed and undertreated. Early detection of signs of mental health problems is particularly important, since unattended, they can be life-threatening. This is why a deep understanding of the complex manifestations of mental disorder development is important. We present a study of mental disorders in social media, from different perspectives. We are interested in understanding whether monitoring language in social media could help with early detection of mental disorders, using computational methods. We developed deep learning models to learn linguistic markers of disorders, at different levels of the language (content, style, emotions), and further try to interpret the behavior of our models for a deeper understanding of mental disorder signs. We complement our prediction models with computational analyses grounded in theories from psychology related to cognitive styles and emotions, in order to understand to what extent it is possible to connect cognitive styles with the communication of emotions over time. The final goal is to distinguish between users diagnosed with a mental disorder and healthy users, in order to assist clinicians in diagnosing patients. We consider three different mental disorders, which we analyze separately and comparatively: depression, anorexia, and self-harm tendencies.
Ana Sabina Uban, Berta Chulvi, Paolo Rosso
Future Gener. Comput. Syst.1
2020 Automatically Building a Multilingual Lexicon of False Friends With No Supervision
abstract
Cognate words, defined as words in different languages which derive from a common etymon, can be useful for language learners, who can leverage the orthographical similarity of cognates to more easily understand a text in a foreign language. Deceptive cognates, or false friends, do not share the same meaning anymore; these can be instead deceiving and detrimental for language acquisition or text understanding in a foreign language. We use an automatic method of detecting false friends from a set of cognates, in a fully unsupervised fashion, based on cross-lingual word embeddings. We implement our method for English and five Romance languages, including a low-resource language (Romanian), and evaluate it against two different gold standards. The method can be extended easily to any language pair, requiring only large monolingual corpora for the involved languages and a small bilingual dictionary for the pair. We additionally propose a measure of “falseness” of a false friends pair. We publish freely the database of false friends in the six languages, along with the falseness scores for each cognate pair. The resource is the largest of the kind that we are aware of, both in terms of languages covered and number of word pairs.
Ana Sabina Uban, Liviu P. Dinu
LREC1
2019 A Computational Approach to Measuring the Semantic Divergence of Cognates
Ana Sabina Uban, Alina Maria Cristea, Liviu P. Dinu
CICLing (2)1
2019 The Myth of Double-Blind Review Revisited: ACL vs. EMNLP
abstract
Cornelia Caragea, Ana Uban, Liviu P. Dinu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Cornelia Caragea, Ana Sabina Uban, Liviu P. Dinu
EMNLP/IJCNLP (1)2
2018 Analyzing Stylistic Variation Across Different Political Regimes
Liviu P. Dinu, Ana Sabina Uban
CICLing (1)2