VLDB 2026 Research / reviewers in the wild / expert
Valentina Dragos
dblp:09/591 · also Valentina Ceausu
· DBLP profile ↗
31ranked-venue papers
23as first author
6since 2021 · last 2024
0000-0002-1829-7207ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 21 · 14 first-author · 3 since 2021Artificial intelligence and machine learning · 10 · 9 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Exploring the Emotional Dimension of French Online Toxic ContentabstractOne of the biggest hurdles for the effective analysis of data collected on social platforms is the need for deeper insights on the content and meaning of this data. Emotion annotation can bring new perspectives on this issue and can enable the identification of content–specific features. This study aims at investigating the ways in which variation in online content can be explored through emotion annotation and corpus-based analysis. The paper describes the emotion annotation of three data sets in French composed of extremist, sexist and hateful messages respectively. To this end, first a fine-grained, corpus annotation scheme was used to annotate the data sets and then several empirical studies were carried out to characterize the content in the light of emotional categories. Results suggest that emotion annotations can provide new insights for online content analysis and stronger empirical background for automatic content detection. Valentina Dragos, Delphine Battistelli, Fatou Sow, Aline Étienne |
LREC/COLING | 1 |
| 2023 | Comparison of Classification Techniques for Extremism Detection in French Social MediaabstractWith social platforms being used by an increasing number of users, the Internet became a perfect place for extremist ideas and opinions to be created and propagated. Detection of extremist content can be an asset for security agencies but it comes with technical challenges and requires semi-automatic approaches, as reading the sheer amount of data released online is an impossible task for analysts. This paper investigates the use of learning models to detect extremist content in French corpora and focuses on right-wing extremist detection. Several learning models have been developed including unsupervised approaches and neural ones. The models were applied to a data set gleaned online. Experiments show that data representation and parameters of models may affect the overall performance and extremist content can be accurately detected when parameters and thresholds are tuned correctly. These results are novel as they contribute to the analysis of social data conveying extremist ideas in French. Valentina Dragos, Yolène Constable |
FUSION | 1 |
| 2023 | Uncertain about ChatGPT: enabling the uncertainty evaluation of large language modelsabstractChatGPT, OpenAI’s chatbot, has gained consider-able attention since its launch in November 2022, owing to its ability to formulate articulated responses to text queries and comments relating to seemingly any conceivable subject. As impressive as the majority of interactions with ChatGPT are, this large language model has a number of acknowledged shortcomings, which in several cases, may be directly related to how ChatGPT handles uncertainty. The objective of this paper is to pave the way to formal analysis of ChatGPT uncertainty handling. To this end, the ability of the Uncertainty Representation and Reasoning Framework (URREF) ontology is assessed, to support such analysis. Elements of structured experiments for reproducible results are identified. The dataset built varies Information Criteria of Correctness, Non-specificity, Self-confidence, Relevance and Inconsistency, and the Source Criteria of Reliability, Competency and Type. ChatGPT’s answers are analyzed along Information Criteria of Correctness, Non-specificity and Self-confidence. Both generic and singular information are sequentially provided. The outcome of this preliminary study is twofold: Firstly, we validate that the experimental setup is efficient in capturing aspects of ChatGPT uncertainty handling. Secondly, we identify possible modifications to the URREF ontology that will be discussed and eventually implemented in URREF ontology Version 4.0 under development. Anne-Laure Jousselme, Johan Pieter de Villiers, Allan De Freitas, Erik Blasch, Valentina Dragos, Gregor Pavlin, Paulo C. G. Costa, Kathryn B. Laskey, Claire Laudy |
FUSION | 5 |
| 2023 | Qualitative Models of Data Generation Processes: Facilitating Data-Intensive AI SolutionsabstractAI-based decision support solutions require life cycles that adequately address critical steps, such as (i) finding suitable machine learning (ML) methods for the problem at hand, (ii) preparing and executing adequate data acquisition processes and (iii) tractable evaluation of the overall solution. Understanding the data generating processes is key in achieving this. Training and test data can be seen as a result of a causal data generation process, a sampling process in which the data is collected from different sources that are influenced by multiple interdependent phenomena. This is represented by a Qualitative Model of Data Generation Processes (QM-DGP), a causal graphical model. QM-DGP facilitates analysis of the complexity of the underlying data generating processes that can inform the development of trustable ML-based solutions in multiple ways. Firstly, this analysis is the basis for the determination of the required complexity of the ML models. Secondly, it facilitates the determination of the quantities of training data supporting good learning results. Thirdly, it can provide guidance for a systematic simplification of the models, supporting tractable solutions without significantly reduced performance. The construction of QM-DGP and the analysis benefit from sound theoretical concepts, such as d-separation and I-Maps. Experimental results with simulated data indicate that the approach can be effective in predicting the required quantities of training data and the determination of the modelling complexity using different types of models. Gregor Pavlin, Kathryn B. Laskey, Franck Mignet, Filip S. Slijkhuis, Erik Blasch, Valentina Dragos, Johan Pieter de Villiers, Lennard Jansen |
FUSION | 6 |
| 2022 | Ontology Adaptation for Opinion Mining in French CorporaabstractThis paper presents the development of an ontology for opinion mining in French corpora. Since ontology construction from scratch is expensive and time consuming, the model is built by adapting an existing ontology modeling appraisal categories in English. The construction method consists of two main steps: first, the ontology of appraisals is translated in French by using a concept-to-concept approach; then different adaptation strategies are implemented to improve the result of this translation. Adaptation strategies are based on text mining, weakly supervised methods and the use of WordNet to refine the model. The goal of the adaptation phase is to cope with limitations of concept-to-concept translation, to integrate concepts and relations from external corpora and resources, and to build a model that captures various features of opinions. The ontology was designed and build to incorporate linguistic and extra linguistic information on the description of opinions in French. The model was validated by using both expert insights to validate the relevance of concepts and relations and formal criteria to describe the ontology qualities for practical use. Valentina Dragos, Adrien Legros |
KES | 1 |
| 2022 | Angry or Sad ? Emotion Annotation for Extremist Content CharacterisationabstractThis paper examines the role of emotion annotations to characterize extremist content released on social platforms. The analysis of extremist content is important to identify user emotions towards some extremist ideas and to highlight the root cause of where emotions and extremist attitudes merge together. To address these issues our methodology combines knowledge from sociological and linguistic annotations to explore French extremist content collected online. For emotion linguistic analysis, the solution presented in this paper relies on a complex linguistic annotation scheme. The scheme was used to annotate extremist text corpora in French. Data sets were collected online by following semi-automatic procedures for content selection and validation. The paper describes the integrated annotation scheme, the annotation protocol that was set-up for French corpora annotation and the results, e.g. agreement measures and remarks on annotation disagreements. The aim of this work is twofold: first, to provide a characterization of extremist contents; second, to validate the annotation scheme and to test its capacity to capture and describe various aspects of emotions. Valentina Dragos, Delphine Battistelli, Aline Étienne, Yolène Constable |
LREC | 1 |
| 2020 | Trend analysis in online data with unsupervised classification and appraisal categoriesabstractThis paper presents a hybrid approach for identifying trends in social media datasets. This approach uses jointly unsupervised methods for text classification and an ontology of appraisal categories. First, data gleaned on social media are classified with unsupervised methods in order to produce more or less homogeneous clusters. Then, topics are detected within each cluster by using Latent Dirichlet Allocation (LDA). At cluster level, appraisal categories are identified thanks to an ontology build according to the principles of the Appraisal Theory. The joint identification of topics and appraisals offers a basis to analyze trending topics in social data. The paper investigates a novel means to detect trending topics in social data by utilizing unsupervised classification methods and focusing on subjective states such as affect, attitude, denial, disapproval, rejection, endorsement or support, associated with each class. Negative or positive polarity and intensity degrees are also identified, thanks to the appraisal ontology. The approach identifies the most dominant trends in social data as associations of topics and appraisal categories. The paper also discusses experiments carried out to detect trends on Twitter collections and the evaluation of theirs results in the light of manually validated data. Valentina Dragos, Jérôme Besombes, Aurélien Mascaro |
FUSION | 1 |
| 2020 | Is hybrid AI suited for hybrid threats? Insights from social media analysisabstractSocial media create the opportunity for a truly connected world and change the way people communicate, exchange ideas and organize themselves into virtual communities. Both understanding online behavior and processing online content are of strategic importance for security applications. However, high volumes, noisy data and rapid changes of topics impose challenges that hinder the efficacy of classification models and the relevance of semantic models. This paper performs a comparative analysis on supervised, unsupervised and semantic-driven approaches used to analyze social data streams. The goal of the paper is to determine whether empirical findings support the enhancement of decision support and pattern recognition applications. The paper reports on research that has used various approaches to identify hidden patterns in social data collections where text is highly unstructured, comes with a mix of modalities and has potentially incorrect spatial-temporal stamps. The conclusion reports that the disconnected use of machine learning models and semantic-driven approaches in mining social media data has several weaknesses. Valentina Dragos, Bruce Forrester, Kellyn Rein |
FUSION | 1 |
| 2020 | Use cases for social data analysis with URREF criteriaabstractSocial data analysis has gained prominence in a wide range of domains as it provides users with the opportunity to communicate and share posts and topics. Automated analysis and reasoning about such data potentially derive meaningful insights, with tremendous potential for applications. However, the sheer volume, noise, and high dynamics of social data impose challenges that hinder the efficacy of algorithms. Automated approaches and classification models require then significant resources to be developed and prove to be often relevant to only a limited number of tasks. Imperfections of inputs, precision of techniques and accuracy of results need to be accounted and assessed as the process runs. This paper discuses two use cases allowing the investigation of implicit and explicit uncertainty arising when processing data gleaned on social media. The objective of this paper is twofold. The first objective is to set up the ETUR use case on social media analysis by adopting two tasks on opinion mining for cyberspace surveillance and information extraction for crisis analysis, respectively. The second objective is to discuss an overall methodology allowing the identification and assessment of uncertainties underlying each task The paper introduces two illustrations of social data analysis, investigates various sources of uncertainty and describes a methodology to select criteria for uncertainty assessment. Claire Laudy, Valentina Dragos |
FUSION | 2 |
| 2020 | Building a formal model for hate detection in French corporaabstractThis paper investigates the development of a formal model in order to analyse online hate in French corpora. Relevant concepts are identified by exploiting several sources: the cognitive foundations of the appraisal theory, according to which people’s emotional response are based on their own evaluative judgments or appraisals of situations, events or objects; a linguistic model of how different kinds of modalities applied to predicative contents are expressed in textual data; several definitions of a hate speech. Based on those inputs, a formal model was developed to describe online hate speech. The model highlights different categories of hate targets and actions, and emphasizes the importance of context for online hate detection. Delphine Battistelli, Cyril Bruneau, Valentina Dragos |
KES | 3 |
| 2020 | A formal representation of appraisal categories for social data analysisabstractHuman behavior is impacted by subjective states although currently the cyberspace becomes a replacement for real-life spaces and interactions. As social media platforms change people’s lives and impacts the way they communicate and group themselves into virtual networks of like-minded individuals, the analysis of online content offers valuable insights of processes taking place on the Internet. Social data mining revolves around subjective content analysis, which deals with the computational processing of texts conveying people’s evaluations, beliefs, attitudes and emotions. Opinion mining and sentiment analysis are the main paradigm of social media exploration and both concepts are often interchangeable. This paper investigates the use of appraisal categories to explore data gleaned for social media and describes the construction of a formal model describing the way language is used in the cyberspace to evaluate, express mood and affective states, construct personal standpoints and manage interpersonal interactions. The ontology offers a mean to investigate subjective content going beyond the traditional notions of opinion and sentiment. Pitfalls of building a formal model for appraisal categories are examined and limitations of using the model for social data exploration are discussed. Valentina Dragos, Delphine Battistelli, Emmanuelle Kelodjoue |
KES | 1 |
| 2019 | Assessment of Trust in Opportunistic Reporting using Belief Functions
Valentina Dragos, Jean Dezert, Kellyn Rein |
FUSION | 1 |
| 2019 | Entropy-Based Metrics for URREF Criteria to Assess Uncertainty in Bayesian Networks for Cyber Threat Detection
Valentina Dragos, Jürgen Ziegler 0003, Johan Pieter de Villiers, Alta de Waal, Anne-Laure Jousselme, Erik Blasch |
FUSION | 1 |
| 2018 | Beyond Sentiments and Opinions: Exploring Social Media with Appraisal CategoriesabstractThe digital era arrives with a whole set of disruptive technologies that creates both risk and opportunity for open sources analysis. Although the sheer quantity of online conversations makes social media a huge source of information, their analysis is still a challenging task and many of traditional methods and research methodologies for data mining are not fit for purpose. Social data mining revolves around subjective content analysis, which deals with the computational processing of texts conveying people's evaluations, beliefs, attitudes and emotions. Opinion mining and sentiment analysis are the main paradigm of social media exploration and both concepts are often interchangeable. This paper investigates the use of appraisal categories to explore data gleaned for social media, going beyond the limitations of traditional sentiment and opinion-oriented approaches. Categories of appraisal are grounded on cognitive foundations of the appraisal theory, according to which people's emotional response are based on their own evaluative judgments or appraisals of situations, events or objects. A formal model is developed to describe and explain the way language is used in the cyberspace to evaluate, express mood and subjective states, construct personal standpoints and manage interpersonal interactions and relationships. A general processing framework is implemented to illustrate how the model is used to analyze a collection of tweets related to extremist attitudes. Valentina Dragos, Delphine Battistelli, Emmanuelle Kelodjoue |
FUSION | 1 |
| 2018 | Application of URREF Criteria to Assess Knowledge Representation in Cyber Threat ModelsabstractSystems for threat analysis enable users to understand the nature and behavior of threats and to undertake a deeper analysis for detailed exploration of threat profile and risk estimation. Models for threat analysis require significant resources to be developed and are often relevant to limited application tasks. This paper investigated the implicit and explicit uncertainty assessments to be taken into account for threat analysis systems to be effective for providing a relevant threat characterization. The intent of this paper is twofold. The first is to present and discuss an approach to define a model for cyber threats within a simplified expert model and to translate it into a Bayesian network as a tool for the development of practical scenarios for cyber threats analysis. The second is to address the question of assessing the Bayesian network build and its intrinsic knowledge representation model and to show how modeling decisions impact the outcome of the system. The paper describes the construction of an expert model and the corresponding BN to analyze cyber threats, investigates various types of induced uncertainty with the URREF criteria simplicity and expressiveness and implements an assessment procedure to evaluate the overall approach. Valentina Dragos, Jürgen Ziegler 0003, Johan Pieter de Villiers |
FUSION | 1 |
| 2017 | Evaluation metrics for the practical application of URREF ontology: An illustration on data criteriaabstractThe International Society of Information Fusion (ISIF) Evaluation Techniques for Uncertainty Representation Working Group (ETURWG) investigates the quantification and evaluation of all types of uncertainty regarding the inputs, reasoning and outputs of the information fusion process. The ETURWG is developing an Uncertainty Representation and Reasoning Framework (URREF) ontology for this purpose. This paper outlines a start towards the process of defining metrics for the URREF data criteria, which will align the URREF ontology with practical application. A criterion can be evaluated according to several metrics, and a metric can be applied to several criteria. As such, the ontology would have to reflect the nature of a many-to-many mapping between criteria and metrics. The main findings and suggestions of the paper advancing the use of URREF are: 1) The Weight of Information (WoI) is dependent on data criteria, which in turn depend on source criteria. 2) Criteria and metrics that apply to evidence (typically an input of the fusion system), could equally apply to the fusion system outputs or internal information, which in turn could form the inputs of another system. As such the word “Evidence” in the terms “Piece of Evidence” and “Weight of Evidence” should be replaced by the word “Information”. 3) Accuracy and precision and associated metrics are ubiquitous in the URREF ontology and can evaluate many parts of the fusion system. 4) The weight of information also assumes an important position in the ontology, as it depends on several source and data criteria. Johan Pieter de Villiers, Richard W. Focke, Gregor Pavlin, Anne-Laure Jousselme, Valentina Dragos, Kathryn B. Laskey, Paulo C. G. Costa, Erik Blasch |
FUSION | 5 |
| 2017 | Subjects under evaluation with the URREF ontologyabstractThe question addressed in this paper is “what” is to be evaluated by the Uncertainty Representation and Reasoning Evaluation Framework (URREF) ontology. We thus identify the elements composing uncertainty representation and reasoning approaches, which constitute various subjects being assessed. We distinguish between primary evaluation subjects (Uncertainty Representation and Reasoning components of the fusion algorithm), and secondary evaluation subjects (source of information, piece of information, fusion method and mathematical model). This paper proposes a list of source quality criteria to be added to the ontology and establishes formal links between the secondary and primary evaluation subjects. The key contribution of the paper is the update of the definitions of sub-criteria of the Expressiveness criterion together with suggestions for complementary concepts to be included in the ontology (type of scale, type of uncertainty expression). Conclusions are drawn to extend the work in using the expressiveness criterion for information fusion analysis. Johan Pieter de Villiers, Gregor Pavlin, Paulo C. G. Costa, Anne-Laure Jousselme, Kathryn B. Laskey, Valentina Dragos, Erik Blasch |
FUSION | 6 |
| 2017 | Detection of contradictions by relation matching and uncertainty assessmentabstractContradiction detection is a difficult task in the field of natural language processing given the variety of ways contradictions occur across texts. If blunt negations, antonyms and numerical mismatches are obvious features to convey contradictions, they also arise from inconsistent domain knowledge, uncertain co-references or differences in the structures of assertions. In this paper, we investigate the problem of contradictions detection for uncertain statements when the author provides not only factual information but also clues about its plausibility. The problem is of particular interest for application fields relying on reported information, when decision makers receive information from various sources. Along with hints as to the derivation of the content, authors often embed clues as to how strong they support reported facts, in the form on confidence, skepticism, doubt or strong conviction. For such assertions, contradictions highlight not only impossible versions of reported events and actions but also discrepancies in the assessment of their plausibility. After analyzing various types of contradictions in subjective statements, we describe a model to detect contradictions thanks to a joint analysis of functional relations and uncertainty assessments. Valentina Dragos |
KES | 1 |
| 2017 | On-the-fly integration of soft and sensor data for enhanced situation assessmentabstractSituation assessment is at the core of many critical tasks in the civilian and military domains: border monitoring, surveillance of areas and facilities, entity tracking and identification, all require accurate and up-to-day descriptions of the course of events. For all those applications, situations to be built are complex, dynamic and uncertain and their assessment is based on the integration of diverse sources, including sensors and their row values, images, observations, tactical information and knowledge expressed by domain experts or synthesized through discovery techniques. This paper presents a method to combine soft and sensor data to create enhanced situation assessment for a track-and-detect application. First we create a situation of entities and relationships by using only hard data provided by sensors and then we enrich this situation thanks to soft data, in the form of succinct or more complex observation reports. The system relies on semantic mediation to combine observations and sensor data by using ontologies as a common ground creating a bridge between two complementary yet incomplete representations of the world. The result is an augmented situation, having more precise, accurate or complete descriptions of entities and which is easier to analyze. This enhanced assessment allows for the situation to be understood and processed in a meaningful way by decision makers. Valentina Dragos, Sylvain Gatepaille |
KES | 1 |
| 2016 | Refining relation identification by combining soft and sensor data
Valentina Dragos, Sylvain Gatepaille, Xavier Lerouvreur |
FUSION | 1 |
| 2016 | What's in a message? Exploring dimensions of trust in reported information
Valentina Dragos, Kellyn Rein |
FUSION | 1 |
| 2015 | A critical assessment of two methods for heterogeneous information fusion
Valentina Dragos, Xavier Lerouvreur, Sylvain Gatepaille |
FUSION | 1 |
| 2014 | Assessment of uncertainty in soft data: A case study
Valentina Dragos |
FUSION | 1 |
| 2014 | Integration of soft data for information fusion: Pitfalls, challenges and trends
Valentina Dragos, Kellyn Rein |
FUSION | 1 |
| 2014 | Description of a Semantic-based Navigation Model to Explore Document Collections in the Maritime DomainabstractThis paper proposes a novel approach to explore collection of documents in the maritime domain. Documents are reports created by experts in charge of analyzing suspicious behaviors in the maritime field. The goal of this work is twofold: it improves knowledge exploitation and reuse for situation assessment and it provides support to analysts in charge of incident interpretation. Semantic integration is at the core of our navigation model. Semantic integration is the process of interrelating information from diverse sources, by using a commonly adopted description of the application field. For this work, reports are not enriched by semantic annotations, but they are processed in order to represent each document in the form of vectors of numerical values and sets of concepts augmented by corresponding weight values. Weight values are used to take into account the relevance of each concept for a given document. In a similar way, user queries are defined by numerical values and sets of ontological entities. The navigation model implements two information retrieval strategies: finding retrieves specific events occurring in specific areas while explaining highlights clues to explain abnormal vessel behaviors. Search results are provided by a ranking scheme based on both the semantic similarity between document and query and values of weights. Supported with complex domain knowledge, our navigation model offers intelligent means to assist experts while exploring the collection of interpretation reports. The paper also presents remarks on model validation and evaluation of its performances. Valentina Dragos |
KES | 1 |
| 2013 | URREF reliability versus credibility in information fusion (STANAG 2511)
Erik Blasch, Kathryn B. Laskey, Anne-Laure Jousselme, Valentina Dragos, Paulo C. G. Costa, Jean Dezert |
FUSION | 4 |
| 2013 | An ontological analysis of uncertainty in soft data
Valentina Dragos |
FUSION | 1 |
| 2012 | Shallow semantic analysis to estimate HUMINT correlation
Valentina Dragos |
FUSION | 1 |
| 2012 | Ontology modeling for intelligence: the ONTO-CIF modelabstractIn highly dynamic and heterogeneous environments, such as coalition missions or joint operations can be, providing commanders with decision making support requires a through understanding of processes involved and the development of underlying knowledge models upon which reasoning mechanisms can be based. This paper presents the construction of ONTO-CIF, a formal ontology created to support intelligence activities. ONTO-CIF is a core ontology, developed by following a methodology based on textual documents, which allows us to accomplish a satisfactory accuracy level in terms of domain coverage. The paper also illustrates several scenarios using ONTO-CIF to support intelligence analysis, a central task of military domain. Valentina Dragos |
KES | 1 |
| 2011 | Same world, different words: Augmenting sensor output through semantics
Anne-Laure Jousselme, Valentina Dragos, Anne-Claire Boury-Brisset, Patrick Maupin |
FUSION | 2 |
| 2006 | An Ontology Supported Approach to Learn Term to Concept Mapping
Valentina Dragos, Sylvie Desprès |
PKAW | 1 |