Valentina Dragos

dblp:09/591 · also Valentina Ceausu · DBLP profile ↗
← Back
21ranked-venue papers in the field
14as first author
3since 2021 · last 2023
0000-0002-1829-7207ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 21 (14 first)
YearPublicationVenuePosition
2023 Comparison of Classification Techniques for Extremism Detection in French Social Media
abstract
With social platforms being used by an increasing number of users, the Internet became a perfect place for extremist ideas and opinions to be created and propagated. Detection of extremist content can be an asset for security agencies but it comes with technical challenges and requires semi-automatic approaches, as reading the sheer amount of data released online is an impossible task for analysts. This paper investigates the use of learning models to detect extremist content in French corpora and focuses on right-wing extremist detection. Several learning models have been developed including unsupervised approaches and neural ones. The models were applied to a data set gleaned online. Experiments show that data representation and parameters of models may affect the overall performance and extremist content can be accurately detected when parameters and thresholds are tuned correctly. These results are novel as they contribute to the analysis of social data conveying extremist ideas in French.
Valentina Dragos, Yolène Constable
FUSION1
2023 Uncertain about ChatGPT: enabling the uncertainty evaluation of large language models
abstract
ChatGPT, OpenAI’s chatbot, has gained consider-able attention since its launch in November 2022, owing to its ability to formulate articulated responses to text queries and comments relating to seemingly any conceivable subject. As impressive as the majority of interactions with ChatGPT are, this large language model has a number of acknowledged shortcomings, which in several cases, may be directly related to how ChatGPT handles uncertainty. The objective of this paper is to pave the way to formal analysis of ChatGPT uncertainty handling. To this end, the ability of the Uncertainty Representation and Reasoning Framework (URREF) ontology is assessed, to support such analysis. Elements of structured experiments for reproducible results are identified. The dataset built varies Information Criteria of Correctness, Non-specificity, Self-confidence, Relevance and Inconsistency, and the Source Criteria of Reliability, Competency and Type. ChatGPT’s answers are analyzed along Information Criteria of Correctness, Non-specificity and Self-confidence. Both generic and singular information are sequentially provided. The outcome of this preliminary study is twofold: Firstly, we validate that the experimental setup is efficient in capturing aspects of ChatGPT uncertainty handling. Secondly, we identify possible modifications to the URREF ontology that will be discussed and eventually implemented in URREF ontology Version 4.0 under development.
Anne-Laure Jousselme, Johan Pieter de Villiers, Allan De Freitas, Erik Blasch, Valentina Dragos, Gregor Pavlin, Paulo C. G. Costa, Kathryn B. Laskey, Claire Laudy
FUSION5
2023 Qualitative Models of Data Generation Processes: Facilitating Data-Intensive AI Solutions
abstract
AI-based decision support solutions require life cycles that adequately address critical steps, such as (i) finding suitable machine learning (ML) methods for the problem at hand, (ii) preparing and executing adequate data acquisition processes and (iii) tractable evaluation of the overall solution. Understanding the data generating processes is key in achieving this. Training and test data can be seen as a result of a causal data generation process, a sampling process in which the data is collected from different sources that are influenced by multiple interdependent phenomena. This is represented by a Qualitative Model of Data Generation Processes (QM-DGP), a causal graphical model. QM-DGP facilitates analysis of the complexity of the underlying data generating processes that can inform the development of trustable ML-based solutions in multiple ways. Firstly, this analysis is the basis for the determination of the required complexity of the ML models. Secondly, it facilitates the determination of the quantities of training data supporting good learning results. Thirdly, it can provide guidance for a systematic simplification of the models, supporting tractable solutions without significantly reduced performance. The construction of QM-DGP and the analysis benefit from sound theoretical concepts, such as d-separation and I-Maps. Experimental results with simulated data indicate that the approach can be effective in predicting the required quantities of training data and the determination of the modelling complexity using different types of models.
Gregor Pavlin, Kathryn B. Laskey, Franck Mignet, Filip S. Slijkhuis, Erik Blasch, Valentina Dragos, Johan Pieter de Villiers, Lennard Jansen
FUSION6
2020 Trend analysis in online data with unsupervised classification and appraisal categories
abstract
This paper presents a hybrid approach for identifying trends in social media datasets. This approach uses jointly unsupervised methods for text classification and an ontology of appraisal categories. First, data gleaned on social media are classified with unsupervised methods in order to produce more or less homogeneous clusters. Then, topics are detected within each cluster by using Latent Dirichlet Allocation (LDA). At cluster level, appraisal categories are identified thanks to an ontology build according to the principles of the Appraisal Theory. The joint identification of topics and appraisals offers a basis to analyze trending topics in social data. The paper investigates a novel means to detect trending topics in social data by utilizing unsupervised classification methods and focusing on subjective states such as affect, attitude, denial, disapproval, rejection, endorsement or support, associated with each class. Negative or positive polarity and intensity degrees are also identified, thanks to the appraisal ontology. The approach identifies the most dominant trends in social data as associations of topics and appraisal categories. The paper also discusses experiments carried out to detect trends on Twitter collections and the evaluation of theirs results in the light of manually validated data.
Valentina Dragos, Jérôme Besombes, Aurélien Mascaro
FUSION1
2020 Is hybrid AI suited for hybrid threats? Insights from social media analysis
abstract
Social media create the opportunity for a truly connected world and change the way people communicate, exchange ideas and organize themselves into virtual communities. Both understanding online behavior and processing online content are of strategic importance for security applications. However, high volumes, noisy data and rapid changes of topics impose challenges that hinder the efficacy of classification models and the relevance of semantic models. This paper performs a comparative analysis on supervised, unsupervised and semantic-driven approaches used to analyze social data streams. The goal of the paper is to determine whether empirical findings support the enhancement of decision support and pattern recognition applications. The paper reports on research that has used various approaches to identify hidden patterns in social data collections where text is highly unstructured, comes with a mix of modalities and has potentially incorrect spatial-temporal stamps. The conclusion reports that the disconnected use of machine learning models and semantic-driven approaches in mining social media data has several weaknesses.
Valentina Dragos, Bruce Forrester, Kellyn Rein
FUSION1
2020 Use cases for social data analysis with URREF criteria
abstract
Social data analysis has gained prominence in a wide range of domains as it provides users with the opportunity to communicate and share posts and topics. Automated analysis and reasoning about such data potentially derive meaningful insights, with tremendous potential for applications. However, the sheer volume, noise, and high dynamics of social data impose challenges that hinder the efficacy of algorithms. Automated approaches and classification models require then significant resources to be developed and prove to be often relevant to only a limited number of tasks. Imperfections of inputs, precision of techniques and accuracy of results need to be accounted and assessed as the process runs. This paper discuses two use cases allowing the investigation of implicit and explicit uncertainty arising when processing data gleaned on social media. The objective of this paper is twofold. The first objective is to set up the ETUR use case on social media analysis by adopting two tasks on opinion mining for cyberspace surveillance and information extraction for crisis analysis, respectively. The second objective is to discuss an overall methodology allowing the identification and assessment of uncertainties underlying each task The paper introduces two illustrations of social data analysis, investigates various sources of uncertainty and describes a methodology to select criteria for uncertainty assessment.
Claire Laudy, Valentina Dragos
FUSION2
2019 Assessment of Trust in Opportunistic Reporting using Belief Functions
Valentina Dragos, Jean Dezert, Kellyn Rein
FUSION1
2019 Entropy-Based Metrics for URREF Criteria to Assess Uncertainty in Bayesian Networks for Cyber Threat Detection
Valentina Dragos, Jürgen Ziegler 0003, Johan Pieter de Villiers, Alta de Waal, Anne-Laure Jousselme, Erik Blasch
FUSION1
2018 Beyond Sentiments and Opinions: Exploring Social Media with Appraisal Categories
abstract
The digital era arrives with a whole set of disruptive technologies that creates both risk and opportunity for open sources analysis. Although the sheer quantity of online conversations makes social media a huge source of information, their analysis is still a challenging task and many of traditional methods and research methodologies for data mining are not fit for purpose. Social data mining revolves around subjective content analysis, which deals with the computational processing of texts conveying people's evaluations, beliefs, attitudes and emotions. Opinion mining and sentiment analysis are the main paradigm of social media exploration and both concepts are often interchangeable. This paper investigates the use of appraisal categories to explore data gleaned for social media, going beyond the limitations of traditional sentiment and opinion-oriented approaches. Categories of appraisal are grounded on cognitive foundations of the appraisal theory, according to which people's emotional response are based on their own evaluative judgments or appraisals of situations, events or objects. A formal model is developed to describe and explain the way language is used in the cyberspace to evaluate, express mood and subjective states, construct personal standpoints and manage interpersonal interactions and relationships. A general processing framework is implemented to illustrate how the model is used to analyze a collection of tweets related to extremist attitudes.
Valentina Dragos, Delphine Battistelli, Emmanuelle Kelodjoue
FUSION1
2018 Application of URREF Criteria to Assess Knowledge Representation in Cyber Threat Models
abstract
Systems for threat analysis enable users to understand the nature and behavior of threats and to undertake a deeper analysis for detailed exploration of threat profile and risk estimation. Models for threat analysis require significant resources to be developed and are often relevant to limited application tasks. This paper investigated the implicit and explicit uncertainty assessments to be taken into account for threat analysis systems to be effective for providing a relevant threat characterization. The intent of this paper is twofold. The first is to present and discuss an approach to define a model for cyber threats within a simplified expert model and to translate it into a Bayesian network as a tool for the development of practical scenarios for cyber threats analysis. The second is to address the question of assessing the Bayesian network build and its intrinsic knowledge representation model and to show how modeling decisions impact the outcome of the system. The paper describes the construction of an expert model and the corresponding BN to analyze cyber threats, investigates various types of induced uncertainty with the URREF criteria simplicity and expressiveness and implements an assessment procedure to evaluate the overall approach.
Valentina Dragos, Jürgen Ziegler 0003, Johan Pieter de Villiers
FUSION1
2017 Evaluation metrics for the practical application of URREF ontology: An illustration on data criteria
abstract
The International Society of Information Fusion (ISIF) Evaluation Techniques for Uncertainty Representation Working Group (ETURWG) investigates the quantification and evaluation of all types of uncertainty regarding the inputs, reasoning and outputs of the information fusion process. The ETURWG is developing an Uncertainty Representation and Reasoning Framework (URREF) ontology for this purpose. This paper outlines a start towards the process of defining metrics for the URREF data criteria, which will align the URREF ontology with practical application. A criterion can be evaluated according to several metrics, and a metric can be applied to several criteria. As such, the ontology would have to reflect the nature of a many-to-many mapping between criteria and metrics. The main findings and suggestions of the paper advancing the use of URREF are: 1) The Weight of Information (WoI) is dependent on data criteria, which in turn depend on source criteria. 2) Criteria and metrics that apply to evidence (typically an input of the fusion system), could equally apply to the fusion system outputs or internal information, which in turn could form the inputs of another system. As such the word “Evidence” in the terms “Piece of Evidence” and “Weight of Evidence” should be replaced by the word “Information”. 3) Accuracy and precision and associated metrics are ubiquitous in the URREF ontology and can evaluate many parts of the fusion system. 4) The weight of information also assumes an important position in the ontology, as it depends on several source and data criteria.
Johan Pieter de Villiers, Richard W. Focke, Gregor Pavlin, Anne-Laure Jousselme, Valentina Dragos, Kathryn B. Laskey, Paulo C. G. Costa, Erik Blasch
FUSION5
2017 Subjects under evaluation with the URREF ontology
abstract
The question addressed in this paper is “what” is to be evaluated by the Uncertainty Representation and Reasoning Evaluation Framework (URREF) ontology. We thus identify the elements composing uncertainty representation and reasoning approaches, which constitute various subjects being assessed. We distinguish between primary evaluation subjects (Uncertainty Representation and Reasoning components of the fusion algorithm), and secondary evaluation subjects (source of information, piece of information, fusion method and mathematical model). This paper proposes a list of source quality criteria to be added to the ontology and establishes formal links between the secondary and primary evaluation subjects. The key contribution of the paper is the update of the definitions of sub-criteria of the Expressiveness criterion together with suggestions for complementary concepts to be included in the ontology (type of scale, type of uncertainty expression). Conclusions are drawn to extend the work in using the expressiveness criterion for information fusion analysis.
Johan Pieter de Villiers, Gregor Pavlin, Paulo C. G. Costa, Anne-Laure Jousselme, Kathryn B. Laskey, Valentina Dragos, Erik Blasch
FUSION6
2016 Refining relation identification by combining soft and sensor data
Valentina Dragos, Sylvain Gatepaille, Xavier Lerouvreur
FUSION1
2016 What's in a message? Exploring dimensions of trust in reported information
Valentina Dragos, Kellyn Rein
FUSION1
2015 A critical assessment of two methods for heterogeneous information fusion
Valentina Dragos, Xavier Lerouvreur, Sylvain Gatepaille
FUSION1
2014 Assessment of uncertainty in soft data: A case study
Valentina Dragos
FUSION1
2014 Integration of soft data for information fusion: Pitfalls, challenges and trends
Valentina Dragos, Kellyn Rein
FUSION1
2013 URREF reliability versus credibility in information fusion (STANAG 2511)
Erik Blasch, Kathryn B. Laskey, Anne-Laure Jousselme, Valentina Dragos, Paulo C. G. Costa, Jean Dezert
FUSION4
2013 An ontological analysis of uncertainty in soft data
Valentina Dragos
FUSION1
2012 Shallow semantic analysis to estimate HUMINT correlation
Valentina Dragos
FUSION1
2011 Same world, different words: Augmenting sensor output through semantics
Anne-Laure Jousselme, Valentina Dragos, Anne-Claire Boury-Brisset, Patrick Maupin
FUSION2