Liviu P. Dinu

dblp:50/3644 · also Liviu Petrisor Dinu · DBLP profile ↗
← Back
68ranked-venue papers
28as first author
17since 2021 · last 2026
0000-0002-7559-6756ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 18 first-author · 15 since 2021Theory of computation · 12 · 9 first-authorDatabases, data management, data science and information retrieval · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Voices and Echoes in Fictional Dialogue: A Study of Linguistic Coordination in Literary Texts
Ioana-Roxana Boriceanu, Alina Iacob, Liviu P. Dinu
LREC3
2026 PsihoRo: Depression and Anxiety Romanian Text Corpus
Alexandra Ciobotaru, Ana-Maria Bucur, Liviu P. Dinu
LREC3
2026 The Spectrum of Sentiment: Optimistic, Pessimistic, and Neutral Voices in Online Depression Discourse
Stefana Arina Tabusca, Ana-Maria Bucur, Liviu P. Dinu
LREC3
2025 Friend or Foe? A Computational Investigation of Semantic False Friends across Romance Languages
abstract
In this paper we present a comprehensive analysis of lexical semantic divergence between cognate words and borrowings in the Romance languages.We experiment with different algorithms for false friend detection including deceptive cognate and deceptive borrowings and correction and evaluate them systematically on cognate and borrowing pairs in the five Romance languages.We use the most complete and reliable dataset of cognate words and borrowings based on etymological dictionaries for the five main Romance languages (Italian, Spanish, Portuguese, French and Romanian) to extract deceptive cognates and borrowings automatically based on usage, and freely publish the lexicon of obtained true and deceptive cognate and borrowings in every Romance language pair.
Ana Sabina Uban, Liviu P. Dinu, Ioan-Bogdan Iordache, Simona Georgescu, Claudia Vlad
EMNLP2
2025 Benchmarking Large Language Models for Biomedical Relation Extraction
abstract
Extracting SNP-phenotype associations from biomedical literature is vital but challenging. We benchmarked diverse NLP models, including MLMs, hybrid architectures, and state-of-the-art LLMs (Gemini 2.0, OpenAI O-series, Qwen, Mistral), on the SNPPhenA corpus across three tasks: sentence-level, abstract-level, and association strength classification. OpenAI O1 achieved state-of-the-art (SOTA) results using few-shot learning for non-finetuned sentence-level classification (F1 0.89) and established a new SOTA for abstract-level classification (F1 0.82). Association strength classification proved difficult, though fine-tuned Gemini 2.0 Pro performed best (F1 0.60) in the first LLM evaluation of this task. Proprietary LLMs, especially in few-shot (O1) or fine-tuned (Gemini 2.0 Pro) settings, significantly outperformed other models. These findings confrm the power of modern LLMs for genomic knowledge extraction.
Claudiu Creanga, Teodor-George Marchitan, Liviu P. Dinu
KES3
2025 On the State of NLP Approaches to Modeling Depression in Social Media: A Post-COVID-19 Outlook
abstract
Computational approaches to predicting mental health conditions in social media have been substantially explored in the past years. Multiple reviews have been published on this topic, providing the community with comprehensive accounts of the research in this area. Among all mental health conditions, depression is the most widely studied due to its worldwide prevalence. The COVID-19 global pandemic, starting in early 2020, has had a great impact on mental health worldwide. Harsh measures employed by governments to slow the spread of the virus (e.g., lockdowns) and the subsequent economic downturn experienced in many countries have significantly impacted people's lives and mental health. Studies have shown a substantial increase of above 50% in the rate of depression in the population. In this context, we present a review on natural language processing (NLP) approaches to modeling depression in social media, providing the reader with a post-COVID-19 outlook. This review contributes to the understanding of the impacts of the pandemic on modeling depression in social media. We outline how state-of-the-art approaches and new datasets have been used in the context of the COVID-19 pandemic. Finally, we also discuss ethical issues in collecting and processing mental health data, considering fairness, accountability, and ethics.
Ana-Maria Bucur, Andreea-Codrina Moldovan, Krutika Parvatikar, Marcos Zampieri, Ashiqur R. KhudaBukhsh, Liviu P. Dinu
IEEE J. Biomed. Health Informatics6
2024 Pater Incertus? There Is a Solution: Automatic Discrimination between Cognates and Borrowings for Romance Languages
abstract
Identifying the type of relationship between words (cognates, borrowings, inherited) provides a deeper insight into the history of a language and allows for a better characterization of language relatedness. In this paper, we propose a computational approach for discriminating between cognates and borrowings, one of the most difficult tasks in historical linguistics. We compare the discriminative power of graphic and phonetic features and we analyze the underlying linguistic factors that prove relevant in the classification task. We perform experiments for pairs of languages in the Romance language family (French, Italian, Spanish, Portuguese, and Romanian), based on a comprehensive database of Romance cognates and borrowings. To our knowledge, this is one of the first attempts of this kind and the most comprehensive in terms of covered languages.
Liviu P. Dinu, Ana Sabina Uban, Ioan-Bogdan Iordache, Alina Maria Cristea, Simona Georgescu, Laurentiu Zoicas
LREC/COLING1
2024 Verba volant, scripta volant? Don't worry! There are computational solutions for protoword reconstruction
abstract
Liviu P Dinu, Ana Sabina Uban, Alina Maria Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Liviu P. Dinu, Ana Sabina Uban, Alina Maria Cristea, Ioan-Bogdan Iordache, Teodor-George Marchitan, Simona Georgescu, Laurentiu Zoicas
EMNLP1
2024 Fine-Tuning Models for Biomedical Relation Extraction
abstract
Next-Generation Sequencing has revolutionized the study of genetic mutations, enabling large-scale investigations into their roles in disease development. However, extracting meaningful insights from the vast amount of biomedical literature remains a complex challenge that cannot be addressed manually. In this paper, we present pre-trained models (PTMs) for the automatic extraction of relations from biomedical text, specifically targeting the variant-phenotype domain. Our evaluation on the SNPPhenA corpus demonstrates that fine-tuning small BERT-based models, particularly DeBERTa, yields strong performance, approaching the current state-of-the-art (SOTA). Additionally, our results indicate that carefully fine-tuning Google’s Gemini Pro 1.0 outperforms the existing SOTA for both sentence-level tasks (where the model processes only the target sentence) and abstract-level tasks (where the model processes the entire abstract).
Claudiu Creanga, Liviu P. Dinu, Daniela Gîfu
KES2
2023 It's Just a Matter of Time: Detecting Depression with Time-Enriched Multimodal Transformers
Ana-Maria Bucur, Adrian Cosma, Paolo Rosso, Liviu P. Dinu
ECIR (1)4
2023 RoBoCoP: A Comprehensive ROmance BOrrowing COgnate Package and Benchmark for Multilingual Cognate Identification
abstract
The identification of cognates is a fundamental process in historical linguistics, on which any further research is based.Even though there are several cognate databases for Romance languages, they are rather scattered, incomplete, noisy, contain unreliable information, or have uncertain availability.In this paper we introduce a comprehensive database of Romance cognates and borrowings based on the etymological information provided by the dictionaries (the largest known database of this kind, in our best knowledge).We extract pairs of cognates between any two Romance languages by parsing electronic dictionaries of Romanian, Italian, Spanish, Portuguese and French.Based on this resource, we propose a strong benchmark for the automatic detection of cognates, by applying machine learning and deep learning based methods on any two pairs of Romance languages.We find that automatic identification of cognates is possible with accuracy averaging around 94% for the more difficult task formulations.
Liviu P. Dinu, Ana Sabina Uban, Alina Maria Cristea, Anca P. Dinu, Ioan-Bogdan Iordache, Simona Georgescu, Laurentiu Zoicas
EMNLP1
2023 SART & COVIDSentiRo: Datasets for Sentiment Analysis Applied to Analyzing COVID-19 Vaccination Perception in Romanian Tweets
abstract
Vaccination is an important subject of discussion adjacent to the COVID-19 pandemic. Sentiments generated online by this topic are worth analyzing using opinion mining tools, and it is interesting to do so in online content written in an under-researched language, like Romanian. For this reason, we modified and enlarged an existing sentiment analysis dataset comprised of Romanian tweets labeled as negative or positive. The resulting dataset, SART (Sentiment Analysis from Romanian Tweets), comprised of three classes (positive, negative, and neutral) containing 1300 Romanian tweets each, was used to train two different sentiment analysis models: a fastText-based one and a fine-tuned BERT model. We further show the usefulness of the sentiment analysis model by analyzing the sentiment of Romanian tweets regarding vaccination using a corpus created and collected by the authors between January 2021 and February 2022 (COVIDSentiRo).
Alexandra Ciobotaru, Liviu P. Dinu
KES2
2023 Veracity Analysis of Romanian Fake News
abstract
Today, with so much free information online, it becomes increasingly difficult to make sense of what content is based on fact, half-truths or lies. Furthermore, in accordance with the events (e.g., COVID-19) that are increasingly alarming, the speed of false news spread is unprecedented. Of course, fake news impairs social stability and public trust, which calls for increasing demand for their detection. How to spot as faithful as possible fake news? The response that this article gives, by exploring the textual features using artificial intelligence and machine learning. This research addresses the problem of automatic fake news detection for Romanian language. First, we present a new corpus for automatic fake news detection that contains two subsets of 977 and 29 154 news articles in Romanian, separated according to labelling and collection methods. Second, we explore several text-based approaches for automatic fake news detection by using machine learning and artificial intelligence which resulted in an accuracy of 93%.
Liviu P. Dinu, Elena-Casiana Fusu, Daniela Gîfu
KES1
2022 Life is not Always Depressing: Exploring the Happy Moments of People Diagnosed with Depression
abstract
In this work, we explore the relationship between depression and manifestations of happiness in social media. While the majority of works surrounding depression focus on symptoms, psychological research shows that there is a strong link between seeking happiness and being diagnosed with depression. We make use of Positive-Unlabeled learning paradigm to automatically extract happy moments from social media posts of both controls and users diagnosed with depression, and qualitatively analyze them with linguistic tools such as LIWC and keyness information. We show that the life of depressed individuals is not always bleak, with positive events related to friends and family being more noteworthy to their lives compared to the more mundane happy events reported by control users.
Ana-Maria Bucur, Adrian Cosma, Liviu P. Dinu
LREC3
2022 RED v2: Enhancing RED Dataset for Multi-Label Emotion Detection
abstract
RED (Romanian Emotion Dataset) is a machine learning-based resource developed for the automatic detection of emotions in Romanian texts, containing single-label annotated tweets with one of the following emotions: joy, fear, sadness, anger and neutral. In this work, we propose REDv2, an open-source extension of RED by adding two more emotions, trust and surprise, and by widening the annotation schema so that the resulted novel dataset is multi-label. We show the overall reliability of our dataset by computing inter-annotator agreements per tweet using a formula suitable for our annotation setup and we aggregate all annotators’ opinions into two variants of ground truth, one suitable for multi-label classification and the other suitable for text regression. We propose strong baselines with two transformer models, the Romanian BERT and the multilingual XLM-Roberta model, in both categorical and regression settings.
Alexandra Ciobotaru, Mihai Vlad Constantinescu, Liviu P. Dinu, Stefan Dumitrescu
LREC3
2022 Detecting Optimism in Tweets using Knowledge Distillation and Linguistic Analysis of Optimism
abstract
Finding the polarity of feelings in texts is a far-reaching task. Whilst the field of natural language processing has established sentiment analysis as an alluring problem, many feelings are left uncharted. In this study, we analyze the optimism and pessimism concepts from Twitter posts to effectively understand the broader dimension of psychological phenomenon. Towards this, we carried a systematic study by first exploring the linguistic peculiarities of optimism and pessimism in user-generated content. Later, we devised a multi-task knowledge distillation framework to simultaneously learn the target task of optimism detection with the help of the auxiliary task of sentiment analysis and hate speech detection. We evaluated the performance of our proposed approach on the benchmark Optimism/Pessimism Twitter dataset. Our extensive experiments show the superior- ity of our approach in correctly differentiating between optimistic and pessimistic users. Our human and automatic evaluation shows that sentiment analysis and hate speech detection are beneficial for optimism/pessimism detection.
Stefan Cobeli, Ioan-Bogdan Iordache, Shweta Yadav 0001, Cornelia Caragea, Liviu P. Dinu, Dragos Iliescu
LREC5
2022 Investigating the Relationship Between Romanian Financial News and Closing Prices from the Bucharest Stock Exchange
abstract
A new data set is gathered from a Romanian financial news website for the duration of four years. It is further refined to extract only information related to one company by selecting only paragraphs and even sentences that referred to it. The relation between the extracted sentiment scores of the texts and the stock prices from the corresponding dates is investigated using various approaches like the lexicon-based Vader tool, Financial BERT, as well as Transformer-based models. Automated translation is used, since some models could be only applied for texts in English. It is encouraging that all models, be that they are applied to Romanian or English texts, indicate a correlation between the sentiment scores and the increase or decrease of the stock closing prices.
Ioan-Bogdan Iordache, Ana Sabina Uban, Catalin Stoean, Liviu P. Dinu
LREC4
2020 Random Steinhaus Distances for Robust Syntax-Based Classification of Partially Inconsistent Linguistic Data
Laura Franzoi, Andrea Sgarro, Anca P. Dinu, Liviu P. Dinu
IPMU (3)4
2020 Automatic Reconstruction of Missing Romanian Cognates and Unattested Latin Words
abstract
Producing related words is a key concern in historical linguistics. Given an input word, the task is to automatically produce either its proto-word, a cognate pair or a modern word derived from it. In this paper, we apply a method for producing related words based on sequence labeling, aiming to fill in the gaps in incomplete cognate sets in Romance languages with Latin etymology (producing Romanian cognates that are missing) and to reconstruct uncertified Latin words. We further investigate an ensemble-based aggregation for combining and re-ranking the word productions of multiple languages.
Alina Maria Cristea, Liviu P. Dinu, Laurentiu Zoicas
LREC2
2020 Automatically Building a Multilingual Lexicon of False Friends With No Supervision
abstract
Cognate words, defined as words in different languages which derive from a common etymon, can be useful for language learners, who can leverage the orthographical similarity of cognates to more easily understand a text in a foreign language. Deceptive cognates, or false friends, do not share the same meaning anymore; these can be instead deceiving and detrimental for language acquisition or text understanding in a foreign language. We use an automatic method of detecting false friends from a set of cognates, in a fully unsupervised fashion, based on cross-lingual word embeddings. We implement our method for English and five Romance languages, including a low-resource language (Romanian), and evaluate it against two different gold standards. The method can be extended easily to any language pair, requiring only large monolingual corpora for the involved languages and a small bilingual dictionary for the pair. We additionally propose a measure of “falseness” of a false friends pair. We publish freely the database of false friends in the six languages, along with the falseness scores for each cognate pair. The resource is the largest of the kind that we are aware of, both in terms of languages covered and number of word pairs.
Ana Sabina Uban, Liviu P. Dinu
LREC2
2019 MorphoGen: Full Inflection Generation Using Recurrent Neural Networks
Octavia-Maria Sulea, Liviu P. Dinu
CICLing (2)3
2019 A Computational Approach to Measuring the Semantic Divergence of Cognates
Ana Sabina Uban, Alina Maria Cristea, Liviu P. Dinu
CICLing (2)3
2019 The Myth of Double-Blind Review Revisited: ACL vs. EMNLP
abstract
Cornelia Caragea, Ana Uban, Liviu P. Dinu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Cornelia Caragea, Ana Sabina Uban, Liviu P. Dinu
EMNLP/IJCNLP (1)3
2019 Algorithms for Closest and Farthest String Problems via Rank Distance
Liviu P. Dinu, Bogdan Dumitru, Alexandru Popa 0001
TAMC1
2019 Automatic Identification and Production of Related Words for Historical Linguistics
abstract
Language change across space and time is one of the main concerns in historical linguistics. In this article, we develop tools to assist researchers and domain experts in the study of language evolution. First, we introduce a method to automatically determine whether two words are cognates. We propose an algorithm for extracting cognates from electronic dictionaries that contain etymological information. Having built a data set of related words, we further develop machine learning methods based on orthographic alignment for identifying cognates. We use aligned subsequences as features for classification algorithms in order to infer rules for linguistic changes undergone by words when entering new languages and to discriminate between cognates and non-cognates. Second, we extend the method to a finer-grained level, to identify the type of relationship between words. Discriminating between cognates and borrowings provides a deeper insight into the history of a language and allows a better characterization of language relatedness. We show that orthographic features have discriminative power and we analyze the underlying linguistic factors that prove relevant in the classification task. To our knowledge, this is the first attempt of this kind. Third, we develop a machine learning method for automatically producing related words. We focus on reconstructing proto-words, but we also address two related sub-problems, producing modern word forms and producing cognates. The task of reconstructing proto-words consists of recreating the words in an ancient language from its modern daughter languages. Having modern word forms in multiple Romance languages, we infer the form of their common Latin ancestors. Our approach relies on the regularities that occurred when words entered the modern languages. We leverage information from several modern languages, building an ensemble system for reconstructing proto-words. We apply our method to multiple data sets, showing that our approach improves on previous results, also having the advantage of requiring less input data, which is essential in historical linguistics, where resources are generally scarce.
Alina Maria Cristea, Liviu P. Dinu
Comput. Linguistics2
2018 Analyzing Stylistic Variation Across Different Political Regimes
Liviu P. Dinu, Ana Sabina Uban
CICLing (1)1
2018 Full Inflection Learning Using Deep Neural Networks
Octavia-Maria Sulea, Liviu P. Dinu, Bogdan Dumitru
CICLing (1)2
2018 Ab Initio: Automatic Latin Proto-word Reconstruction
abstract
Proto-word reconstruction is central to the study of language evolution. It consists of recreating the words in an ancient language from its modern daughter languages. In this paper we investigate automatic word form reconstruction for Latin proto-words. Having modern word forms in multiple Romance languages (French, Italian, Spanish, Portuguese and Romanian), we infer the form of their common Latin ancestors. Our approach relies on the regularities that occurred when the Latin words entered the modern languages. We leverage information from all modern languages, building an ensemble system for proto-word reconstruction. We use conditional random fields for sequence labeling, but we conduct preliminary experiments with recurrent neural networks as well. We apply our method on multiple datasets, showing that our method improves on previous results, having also the advantage of requiring less input data, which is essential in historical linguistics, where resources are generally scarce.
Alina Maria Cristea, Liviu P. Dinu
COLING2
2018 Exploring Optimism and Pessimism in Twitter Using Deep Learning
abstract
Identifying optimistic and pessimistic viewpoints and users from Twitter is useful for providing better social support to those who need such support, and for minimizing the negative influence among users and maximizing the spread of positive attitudes and ideas.In this paper, we explore a range of deep learning models to predict optimism and pessimism in Twitter at both tweet and user level and show that these models substantially outperform traditional machine learning classifiers used in prior work.In addition, we show evidence that a sentiment classifier would not be sufficient for accurately predicting optimism and pessimism in Twitter.Last, we study the verb tense usage as well as the presence of polarity words in optimistic and pessimistic tweets.
Cornelia Caragea, Liviu P. Dinu, Bogdan Dumitru
EMNLP2
2018 Steinhaus Transforms of Fuzzy String Distances in Computational Linguistics
Anca P. Dinu, Liviu P. Dinu, Laura Franzoi, Andrea Sgarro
IPMU (1)2
2017 Towards a Map of the Syntactic Similarity of Languages
Alina Maria Cristea, Liviu P. Dinu, Andrea Sgarro
CICLing (1)2
2017 Romanian Word Production: An Orthographic Approach Based on Sequence Labeling
Liviu P. Dinu, Alina Maria Cristea
CICLing (1)1
2016 The Minimum Entropy Submodular Set Cover Problem
Gabriel Istrate, Cosmin Bonchis, Liviu P. Dinu
LATA3
2016 A Computational Perspective on the Romanian Dialects
Alina Maria Cristea, Liviu P. Dinu
LREC2
2016 A Corpus of Native, Non-native and Translated Texts
Sergiu Nisioi, Ella Rabinovich, Liviu P. Dinu, Shuly Wintner
LREC3
2016 Using Word Embeddings to Translate Named Entities
Octavia-Maria Sulea, Sergiu Nisioi, Liviu P. Dinu
LREC3
2015 Using NLP Specific Tools for Non-NLP Specific Tasks. A Web Security Application
Octavia-Maria Sulea, Liviu P. Dinu, Alexandra Peste
ICONIP (4)2
2015 Coding Theory: A General Framework and Two Inverse Problems
abstract
We put forward an ample framework for coding based on upper probabilities, or more generally on normalized monotone set-measures, and model accordingly noisy transmission channels and decoding errors. Two inverse problems are considered. In the first case, a decoder is given and one looks for chann els of a specified family over which that decoder would work properly. In the second and more ambitious case, it is codes which are given, and one looks for channels over which those codes would ensure the required error correction capabilities. Upper probabilities allow for a solution of the two inverse problems in the case of usual codes based on checking Hamming distances between codewords: one can equivalently check suitable upper probabilities of the decoding errors. This soon extends to “odd” codeword distances for DNA strings as used in DNA word design, where instead, as we prove, not even the first unassuming inverse problem admits of a solution if one insists on channel models based on “usual” probabilities.
Luca Bortolussi, Liviu P. Dinu, Laura Franzoi, Andrea Sgarro
Fundam. Informaticae2
2014 Predicting Romanian Stress Assignment
abstract
We train and evaluate two models for Romanian stress prediction: a baseline model which employs the consonant-vowel structure of the words and a cascaded model with averaged perceptron training consisting of two sequential models ‐ one for predicting syllable boundaries and another one for predicting stress placement. We show in this paper that Romanian stress is predictable, though not deterministic, by using data-driven machine learning techniques.
Alina Maria Cristea, Anca P. Dinu, Liviu P. Dinu
EACL3
2014 Temporal Text Ranking and Automatic Dating of Texts
abstract
Vlad Niculae, Marcos Zampieri, Liviu Dinu, Alina Maria Ciobanu. Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers. 2014.
Vlad Niculae, Marcos Zampieri, Liviu P. Dinu, Alina Maria Cristea
EACL3
2014 An Etymological Approach to Cross-Language Orthographic Similarity. Application on Romanian
abstract
In this paper we propose a computational method for determining the orthographic similarity between Romanian and related languages.We account for etymons and cognates and we investigate not only the number of related words, but also their forms, quantifying orthographic similarities.The method we propose is adaptable to any language, as far as resources are available.
Alina Maria Cristea, Liviu P. Dinu
EMNLP2
2014 Building a Dataset of Multilingual Cognates for the Romanian Lexicon
Liviu P. Dinu, Alina Maria Cristea
LREC1
2014 On the Romance Languages Mutual Intelligibility
Liviu P. Dinu, Alina Maria Cristea
LREC1
2014 Using a machine learning model to assess the complexity of stress systems
Liviu P. Dinu, Alina Maria Cristea, Ioana Chitoran, Vlad Niculae
LREC1
2014 Aggregation methods for efficient collocation detection
Anca P. Dinu, Liviu P. Dinu, Ionut Sorodoc
LREC2
2014 Syllabic Languages and Go-through Automata
abstract
In this paper we define a new class of contextual grammars and study how the languages generated by such grammars can be accepted by go-through automata. The newly introduced class of grammars is a generalization of the formalism previously used to describe the linguistic process of syllabification. Go-through automata which are used here to recognize, and also parse, the languages generated by this new class of grammars are generalizations of push-down automata in the area of context-sensitivity; they have been proved to be an efficient tool for the recognition of languages generated by contextual grammars. The main results of the paper show how the newly introduced generative model is related with other classes of Marcus contextual languages, and how syllabic languages are recognized and parsed using go-through automata.
Liviu P. Dinu, Radu Gramatovici, Florin Manea
Fundam. Informaticae1
2014 Clustering based on median and closest string via rank distance with applications on DNA
Liviu P. Dinu, Radu Tudor Ionescu
Neural Comput. Appl.1
2012 The Naive Bayes Classifier in Opinion Mining: In Search of the Best Feature Set
Liviu P. Dinu, Iulia Iuga
CICLing (1)1
2012 On the Closest String via Rank Distance
Liviu P. Dinu, Alexandru Popa 0001
CPM1
2012 Learning How to Conjugate the Romanian Verb. Rules for Regular and Partially Irregular Verbs
Liviu P. Dinu, Vlad Niculae, Octavia-Maria Sulea
EACL1
2012 Clustering Based on Rank Distance with Applications on DNA
Liviu P. Dinu, Radu Tudor Ionescu
ICONIP (5)1
2012 Local Patch Dissimilarity for Images
Liviu P. Dinu, Radu Tudor Ionescu, Marius Popescu
ICONIP (1)1
2012 The Romanian Neuter Examined Through A Two-Gender N-Gram Classification System
Liviu P. Dinu, Vlad Niculae, Octavia-Maria Sulea
LREC1
2012 Spearman Permutation Distances and Shannon's Distinguishability
abstract
Spearman distance is a permutation distance which might be used for codes in permutations beside Kendall distance. However, Spearman distance gives rise to a geometry of strings, which is rather unruly from the point of view of error correction and error detection. Special care has to be taken to discriminate between the two notions of codeword distance and codeword distinguishability. This stresses the importance of rejuvenating the latter notion, extending it from Shannon's zero-error information theory to the more general setting of metric string distances.
Luca Bortolussi, Liviu P. Dinu, Andrea Sgarro
Fundam. Informaticae2
2011 Meta-search: aggregating search engine results using rank-distance
abstract
In this demonstration, we present a meta-search engine based on rank-distance, which is a similarity measure between partial rankings. Our meta-search engine fetches results from four main search engines: Google, Yahoo, Bing and Ask.com and combines them into one aggregated ranking, indicating for each result the position that it had in the original rankings. This combination of multiple rankings into one that is as close as possible to them is known as the rank aggregation problem. For our search engine, we use rank-distance to solve this problem. We also give a conjecture that improves the time complexity of the existing rank aggregation algorithm based on rank-distance.
Tiberiu Danet, Liviu P. Dinu
SISAP2
2010 Rank Distance Aggregation as a Fixed Classifier Combining Rule for Text Categorization
Liviu P. Dinu, Andrei A. Rusu
CICLing1
2009 On Insertion Grammars with Maximum Parallel Derivation
abstract
In this paper we investigate insertion grammars and explore their capacity to generate words parallelly by introducing parallel derivation (where more than one rule can be applied to the string in parallel) and maximum parallel derivation (where as many rules as possible are applied to the string in parallel). We compare the generative power of these grammars with context sensitive and context free grammars and with different variants of contextual grammars. We apply these grammars to syllabification in Romanian and provide arguments that they can also be used in a cognitive perspective.
Liviu P. Dinu
Fundam. Informaticae1
2008 Authorship Identification of Romanian Texts with Controversial Paternity
Liviu P. Dinu, Marius Popescu, Anca P. Dinu
LREC1
2008 A Multi-Criteria Decision Method Based on Rank Distance
Liviu P. Dinu, Marius Popescu
Fundam. Informaticae1
2007 On the syllabification of words via go-through automata
Liviu P. Dinu, Radu Gramatovici, Florin Manea
LATA1
2006 On the data base of Romanian syllables and some of its quantitative and cryptographic aspects
Liviu P. Dinu, Anca P. Dinu
LREC1
2006 A Low-complexity Distance for DNA Strings
Liviu P. Dinu, Andrea Sgarro
Fundam. Informaticae1
2006 An efficient approach for the rank aggregation problem
Liviu P. Dinu, Florin Manea
Theor. Comput. Sci.1
2005 A Parallel Approach to Syllabification
Anca P. Dinu, Liviu P. Dinu
CICLing2
2005 On the Syllabic Similarities of Romance Languages
Anca P. Dinu, Liviu P. Dinu
CICLing2
2005 Rank Distance with Applications in Similarity of Natural Languages
Liviu P. Dinu
Fundam. Informaticae1
2003 On the Classification and Aggregation of Hierarchies with Different Constitutive Elements
Liviu P. Dinu
Fundam. Informaticae1
2002 Possibilistic Entropies and the Compression of Possibilistic Data
abstract
We re-take the possibilistic model for information sources recently put forward by the first author, as opposed to the standard probabilistic models of information theory. Based on an interpretation of possibilistic source coding inspired by utility functions, we define a notion of possibilistic entropy for a suitable class of interactive possibilistic sources, and compare it with the possibilistic entropy of stationary non-interactive sources. Both entropies have a coding-theoretic nature, being obtained as limit values for the rates of optimal compression codes. We list properties of the two entropies, which might support their use as measures of "possibilistic ignorance".
Andrea Sgarro, Liviu P. Dinu
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2