Irene Díaz

dblp:01/2014 · DBLP profile ↗
← Back
59ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0002-3024-6605ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 26 · 5 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Idempotence and Internality of Aggregations of Random Variables
Juan Baz, Irene Díaz, Susana Montes
EUSFLAT (1)2
2025 Understanding Data Properties in the Mallows Model: Impact of Voter Count Variability
Mario Villar, Noelia Rico, Irene Díaz
EUSFLAT (1)3
2025 Uniform random fuzzy measures
abstract
Random generation of fuzzy measures is an important computational task in applied problems related to fuzzy integrals such as the Choquet, Sugeno or Shilkret integrals. In general, a desirable property is the uniformity over the set of fuzzy measures. However, testing this property is not an easy task. In this paper, properties of uniform random fuzzy measures are derived. Special attention is being paid to the families of balanced fuzzy measures, belief measures and possibilities measures. Then, based of such properties, statistical tests for the uniformity of random fuzzy measures are developed. Finally, the uniformity of the most used algorithms is tested using the proposed methods.
Juan Baz, Gleb Beliakov, Irene Díaz, Susana Montes
Fuzzy Sets Syst.3
2025 Nutriconv: multitask learning framework for digital dietary tracking trained on EFSA's pancake dataset
abstract
Abstract The growing prevalence of nutrition-related health conditions calls for advanced tools to support reliable and efficient dietary monitoring. This paper presents NutriConv, a lightweight multitask convolutional neural network designed to simultaneously perform food classification and weight estimation from single-item food images. Trained on the institutionally validated PANCAKE dataset from the European Food Safety Authority, NutriConv combines classification and regression objectives within a unified architecture, optimized via a hybrid loss function. While its classification accuracy remains lower than that of specialized single-task models, NutriConv achieves competitive regression performance and offers a practical balance between both tasks. Its compact design enables deployment on resource-constrained platforms such as smartglasses and mobile health devices, expanding its usability in real-world dietary tracking scenarios. Extensive experiments confirm its robustness, including external validation on the Nutrition5K dataset, underscoring the model’s generalizability. This work highlights the potential of multitask learning for integrated, scalable, and accessible AI-based nutrition assessment.
Enol Junquera, Noelia Rico, Irene Díaz, Sonia González, Beatriz Remeseiro
Neural Comput. Appl.3
2024 Efficient GPU-algorithms for the combination of evidence in Dempster-Shafer theory
Noelia Rico, Luigi Troiano, Irene Díaz
Future Gener. Comput. Syst.3
2024 Aggregation of random elements over bounded lattices
abstract
Aggregation functions are widely used to fuse information from different sources in a unique value. In many cases, the aggregated information is related to some experimental measure or random sampling of a population. In this direction, it is reasonable to consider aggregation of random elements. In this paper, the concept of aggregation functions of random elements over bounded lattices, which are measurable functions from a probability space to a bounded lattice, is presented. In particular, starting from a partially ordered set, a measurable space is constructed. Random elements are considered to be measurable functions from a probability space to the measurable space. The concept of aggregation of random elements over bounded lattices is defined by generalizing the monotonicity and the boundary conditions in terms of stochastic orders. Several types, such as the induced, random and degenerated aggregations of random elements over bounded lattices are defined and some coherence properties are studied. Particular examples regarding the aggregation of random variables, random graphs and random semi-positive matrices are provided.
Juan Baz, Irene Díaz, Susana Montes
Int. J. Approx. Reason.2
2024 Stochastically ordered aggregation operators
abstract
In aggregation theory, there exists a large number of aggregation functions that are defined in terms of rearrangments in increasing order of the arguments. Prominent examples are the Ordered Weighted Operator and the Choquet and Sugeno integrals. Following a probability approach, ordering random variables by means of stochastic orders can be also a way to define aggregations of random variables. However, stochastic orders are not total orders, thus pairs of incomparable distributions can appear. This paper is focused on the definition of aggregations of random variables that take into account the stochastic ordination of the components of the input random vectors. Three alternatives are presented, the first one by using expected values and admissible permutations, then a modification for multivariate Gaussian random vectors and a third one that involves a transformation of the initial random vectors in new ones whose components are ordered with respect to the usual stochastic order. A deep theoretical study of the properties of all the proposals is made. A practical example regarding temperature prediction is provided
Juan Baz, Franco Pellerey, Irene Díaz, Susana Montes
Int. J. Approx. Reason.3
2024 Computable aggregations of random variables
abstract
Aggregation theory is devoted to the fusing of several values into a unique output that summarizes the given information. Typically, the aggregation process is formalized in terms of an increasing mathematical function that maps the input values to the result, fulfilling some boundary conditions. However, this formalization can be too restrictive for some scenarios. In some cases, the inputs can be seen as observations of random variables, the aggregation result being also a random variable. In others, the aggregation process can be identified as a program that performs the aggregation rather than a mathematical function. In this direction, the concepts of aggregation of random variables and computable aggregation have been defined in the literature. This paper is devoted to the definition of computable aggregation of random variables, which are computer programs, not functions, that aggregate random variables, not numbers. Special attention is given to different possible alternatives to modelize random variables and monotonicity. The implementation of some examples is also provided.
Juan Baz, Irene Díaz, Luis Garmendia, Daniel Gómez 0001, Luis Magdalena, Susana Montes
Inf. Sci.2
2023 Measures of embedding for interval-valued fuzzy sets
abstract
Interval-valued fuzzy sets are a generalization of classical fuzzy sets where the membership values are intervals. The epistemic interpretation of interval-valued fuzzy sets assumes that there is one real-valued membership degree of an element within the membership interval of possible membership degrees. Considering this epistemic interpretation, we propose a new measure, called IV-embedding, to compare the precision of two interval-valued fuzzy sets. An axiomatic definition for this concept as well as a construction method are provided. The construction method is based on aggregation operators and the concept of interval embedding, which is also introduced and deeply studied.
Agustina Bouchet, Mikel Sesma-Sara, Gustavo Ochoa, Humberto Bustince, Susana Montes, Irene Díaz
Fuzzy Sets Syst.6
2023 Ranking the effect of chronodisruption-based biomarkers in reproductive health
abstract
Abstract Chronodisruption alters circadian rhythms, which has negative consequences on different pathologies and mental disorders. This work studies whether factors related to chronodisruption of circadian rhythms motivated by shift works influence on reproductive health or not. In particular, this influence is studied on four particular aspects related to reproductive health: reproductive health disease, first pregnancy attempt, problems during pregnancy and gestation period. Some explainable machine learning models based on trees have been employed. These methods provided information about the importance of each predictor. The most important variables provided by each method were aggregated using a ranking aggregation function in order to reach a consensus ranking of variables that made possible to understand whether the chronodisruption factors had an effect on each of the aspects studied. The data have been obtained from 697 health professionals. Information about classical biomarkers, sleep quality indices and also other new variables related to eating jet lag, sleep hygiene and how the sleep is affected by shift works were considered as input data. Experiments have shown how some of these novel biomarkers are ranked in the top positions of the issues studied in relation to reproductive health. In particular, the light level and the use of electronic devices, which are features related to chronodisruption, are highlighted as biomarkers.
Ana G. Rúa, Noelia Rico, Ana Alonso, Elena Díaz, Irene Díaz
Neural Comput. Appl.5
2023 VCI-LSTM: Vector Choquet Integral-Based Long Short-Term Memory
abstract
Choquet integral is a widely used aggregation operator on 1-D and interval-valued information, since it is able to take into account the possible interaction among data. However, there are many cases where the information taken into account is vectorial, such as long short-term memories (LSTM). LSTM units are a kind of recurrent neural networks that have become one of the most powerful tools to deal with sequential information since they have the power of controlling the information flow. In this article, we first generalize the standard Choquet integral to admit an input composed by$n$-dimensional vectors, which produces an$n$-dimensional vector output. We study several properties and construction methods of vector Choquet integrals (VCIs). Then, we use this integral in the place of the summation operator, introducing in this way the new VCI-LSTM architecture. Finally, we use the proposed VCI-LSTM to deal with two problems: 1) sequential image classification; 2) text classification.
Mikel Ferrero-Jaurrieta, Zdenko Takác, Javier Fernández 0002, Lubomíra Horanská, Graçaliz Pereira Dimuro, Susana Montes, Irene Díaz, Humberto Bustince
IEEE Trans. Fuzzy Syst.7
2023 Kemeny ranking aggregation meets the GPU
abstract
Abstract Ranking aggregation, studied in the field of social choice theory, focuses on the combination of information with the aim of determining a winning ranking among some alternatives when the preferences of the voters are expressed by ordering the possible alternatives from most to least preferred. One of the most famous ranking aggregation methods can be traced back to 1959, when Kemeny introduces a measure of distance between a ranking and the opinion of the voters gathered in a profile of rankings. Using this, he proposed to elect as winning ranking of the election the one that minimizes the distance to the profile. This is factorial on the number of alternatives, posing a handicap in the runtime of the algorithms developed to find the winning ranking, which prevents its use in real problems where the number of alternatives is large. In this work we introduce the first algorithm for the Kemeny problem designed to be executed in a Graphical Processing Unit. The threads identifiers are codified to be associated with rankings by means of the factorial number system, a radix numeral system that is then used to uniquely pair a ranking with the thread using Lehmer’s code. Results guarantee constant execution time up to 14 alternatives.
Noelia Rico, Pedro Alonso 0001, Irene Díaz
J. Supercomput.3
2022 A more informed clustering algorithm through the aggregation of linkage methods
abstract
Agglomerative hierarchical clustering algorithms for finding groups in unlabelled data achieve different clusters of objects depending on how the similarity between two clusters is measured. In this context, the linkage method refers to the process to decide if two clusters are merged. However, it is not possible to establish in advance the linkage method that provides the best grouping of the data. In the classic hierarchical clustering algorithm, the two most similar clusters according to one linkage method are merged together into a single cluster. In this work, we propose a hierarchical clustering algorithm that aggregates in each step the criteria of the single, complete and average linkage methods in order to determine the two most similar clusters to merge. In each step, each linkage method gives a ranking of pairs of clusters, where the pairs are ordered from most to least similar. The rankings given by each linkage method are aggregated using ranking rules to achieve a consensus of the different criteria about which one is the pair to be merged into a single cluster. Results obtained from the validation of the new algorithm with the Rand Index metric in data sets with different characteristics show how the proposed algorithm is useful for reducing the impact of the linkage method chosen in the final clusters obtained.
Noelia Rico, Irene Díaz
FUZZ-IEEE2
2022 Flexible-Dimensional EVR-OWA as Mean Estimator for Symmetric Distributions
Juan Baz, Diego García-Zamora, Irene Díaz, Susana Montes, Luis Martínez-López 0001
IPMU (1)3
2022 A New Similarity Measure for Real Intervals to Solve the Aliasing Problem
Pedro Huidobro, Noelia Rico, Agustina Bouchet, Susana Montes, Irene Díaz
IPMU (1)5
2022 Multidistances and inequality measures on abstract sets: An axiomatic approach
abstract
Starting from the notion of a multidistance, we formalize, through a suitable system of axioms, the concept of an inequality measure defined on a nonempty set with no additional structure implemented a priori. Among inequality measures, apart from multidistances we pay special attention to dispersions, and study their main features. Classical concepts will be generalized to this abstract setting. Multidistances are then revisited, and some new methods to generate them are implemented. A wide spectrum of interdisciplinary applications is outlined in the final section.
María J. Campión, Irene Díaz, Esteban Induráin, Javier Martín, Gaspar Mayor, Susana Montes, Armajac Raventós-Pujol
Fuzzy Sets Syst.2
2022 Similarity measures for interval-valued fuzzy sets based on average embeddings and its application to hierarchical clustering
abstract
Clustering algorithms create groups of objects based on their similarity. As objects are usually defined by data points, this similarity is commonly measured by a distance function. When the objects are defined by variables that are intervals, it is more difficult to determine how to measure the similarity between the objects of the dataset. In this work, we propose some similarity measures between intervals based on average embedding functions. Using these, new similarity measures between interval-based objects are proposed. All the proposed similarities are based on measuring the similarity between the objects variable by variable and then averaging the obtained results to get a single value. By its definition, the objects can be considered as interval-valued fuzzy sets (IVFS), so the similarities introduced are proved to be valid similarities for IVFS. The measures proposed are used in a hierarchical clustering algorithm with the aim of grouping the objects of the dataset into different clusters based on their similarity to interval-valued data. The described process is applied to real data regarding the Spanish weather in order to cluster the provinces of Spain based on the interval temperature of each month in 2021, showing different results that the ones obtained using non-interval-valued data.
Noelia Rico, Pedro Huidobro, Agustina Bouchet, Irene Díaz
Inf. Sci.4
2021 Axiomatization and construction of orness measures for aggregation functions
abstract
The notion of an orness measure for aggregation functions has been a relevant study subject whose history can be traced back to the early works of Dujmović in 1973. Intuitively, an orness measure quantifies the similarity of an aggregation function to the “or” function and results in an essential tool for decision engineering, field in which the choice of aggregation function is sometimes restricted to a desired value of orness (orness-directed aggregation). In 1988, Yager presented a particular example of orness measure for ordered weighted averaging (OWA) functions and initiated a series of contributions aiming at proposing an axiomatic definition of orness measure for OWA functions. In this paper, we go much further and present an axiomatic definition of orness measure for the whole family of aggregation functions. We end by proposing two natural construction methods for an orness measure for aggregation functions. The particular examples of the (discrete) Choquet integral and uninorms are studied in detail.
Raúl Pérez-Fernández, Gustavo Ochoa, Susana Montes, Irene Díaz, Javier Fernández 0002, Daniel Paternain, Humberto Bustince
Int. J. Intell. Syst.4
2020 An Interval-Valued Divergence for Interval-Valued Fuzzy Sets
Susana Díaz, Irene Díaz, Susana Montes
IPMU (2)2
2020 A Genetic Approach to the Job Shop Scheduling Problem with Interval Uncertainty
Hernán Díaz, Inés González Rodríguez, Juan José Palacios 0001, Irene Díaz, Camino R. Vela
IPMU (2)4
2020 On some classes of directionally monotone functions
Humberto Bustince, Radko Mesiar, Anna Kolesárová, Graçaliz Pereira Dimuro, Javier Fernández 0002, Irene Díaz, Susana Montes
Fuzzy Sets Syst.6
2020 Fuzzy sets for decision making in emerging domains
Irene Díaz, Yusuke Nojima
Fuzzy Sets Syst.1
2019 Incorporating ranking rules into k nearest neighbours
abstract
This paper describes how distance-based classification methods could be improved by incorporating ranking rules, which are old acquaintances of aggregation and social choice theorists. Specifically, a novel method incorporating two prominent ranking rules (namely, the plurality and the Borda count ranking rules) into the distance-based classification method of nearest neighbours is presented. Some exploratory experiments have been conducted, obtaining encouraging results. Interestingly, the newly proposed method seems to turn the classic method of nearest neighbours independent of the chosen distance metric while not compromising its performance significantly.
Noelia Rico, Raúl Pérez-Fernández, Irene Díaz
FUZZ-IEEE3
2019 Intelligent decision support to determine the best sensory guardrail locations
Noelia Rico, Irene Díaz, José R. Villar 0001, Enrique A. de la Cal
Neurocomputing2
2019 Minimals Plus: An improved algorithm for the random generation of linear extensions of partially ordered sets
Elías F. Combarro, Julen Hurtado de Saracho, Irene Díaz
Inf. Sci.3
2018 Monotonicity of a Profile of Rankings with Ties
Raúl Pérez-Fernández, Irene Díaz, Susana Montes, Bernard De Baets
IPMU (2)2
2018 On the Problem of Comparing Ordered Ordinary Fuzzy Multisets
Ángel Riesgo, Pedro Alonso 0001, Irene Díaz, Vladimír Janis, Vladimír Kobza, Susana Montes
IPMU (2)3
2018 Basic operations for fuzzy multisets
Ángel Riesgo, Pedro Alonso 0001, Irene Díaz, Susana Montes
Int. J. Approx. Reason.3
2017 Matching media contents with user profiles by means of the Dempster-Shafer theory
abstract
The media industry is increasingly personalizing the offering of contents in attempt to better target the audience. This requires to analyze the relationships that goes established between users and content they enjoy, looking at one side to the content characteristics and on the other to the user profile, in order to find the best match between the two. In this paper we suggest to build that relationship using the Dempster-Shafer's Theory of Evidence, proposing a reference model and illustrating its properties by means of a toy example. Finally we suggest possible applications of the model for tasks that are common in the modern media industry.
Luigi Troiano, Irene Díaz, Ciro Gaglione
FUZZ-IEEE2
2017 Monotonicity-based ranking on the basis of multiple partially specified reciprocal relations
Raúl Pérez-Fernández, Michaël Rademaker, Pedro Alonso 0001, Irene Díaz, Susana Montes, Bernard De Baets
Fuzzy Sets Syst.4
2017 Monotonicity-based consensus states for the monometric rationalisation of ranking rules and how they are affected by ties
Raúl Pérez-Fernández, Pedro Alonso 0001, Irene Díaz, Susana Montes, Bernard De Baets
Int. J. Approx. Reason.3
2017 Identification of Agricultural Management Zones Through Clustering Algorithms with Thermal and Multispectral Satellite Imagery
abstract
Precision Agriculture entails the appropriate management of the inherent variability of soil and crops, resulting in an increase of economic benefits and a reduction of environmental impact. However, site-specific treatments require maps of the soil variability to identify areas of land that share similar properties. In order to produce these maps, we propose a cost-efficient method that combines clustering algorithms with publicly available satellite imagery. The method does not require exploring the parcels with any special equipment or taking samples of the soil for laboratory analysis. The proposed method was tested in a case study for three vineyard parcels with topographical dissimilarities. The study compares different spectral and thermal bands from the Landsat 8 satellite as well as vegetation and moisture indices to determine which one produces the best clustering. The experimental results seem promising for identification of agricultural management zones. The findings suggest that thermal bands produce better clustering than those based on the NDVI index.
R. B. Arango, A. M. Campos, Elías F. Combarro, E. R. Canas, Irene Díaz
Int. J. Uncertain. Fuzziness Knowl. Based Syst.5
2017 New Trends in Information Access
Irene Díaz, Juan M. Fernández-Luna
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2017 Towards an Ontological Model for Technological Systems Structure Representation
abstract
More and more distribution companies need a model to efficiently retrieve physical technological systems. Building ontologies as knowledge representation models with respect to technological systems information is increasingly seen as a key to enable interoperability between software agents and people on the Web or through to network connections. This work presents a first step for the development of an ontological model for technological systems structure representation composed by physical artifacts. The goal of the work is twofold. First, the structure of these systems is described. Besides, a model to adapt the known Bill of Materials (BOM) to the needs of distribution companies focused not only on selling physical artifacts but also to correctly configure a technological system is presented. Finally, our approach proposes a mereotopological relationships set to model the interactions between the parts that make up the structure.
Liudmila Reyes-Alvarez, Jaime Fernández, Luis J. Rodríguez-Muñiz, Irene Díaz
Int. J. Uncertain. Fuzziness Knowl. Based Syst.4
2017 On cardinalities of finite interval-valued hesitant fuzzy sets
Pelayo Quirós, Pedro Alonso 0001, Irene Díaz, Vladimír Janis, Susana Montes
Inf. Sci.3
2016 Representations of votes facilitating monotonicity-based ranking rules: From votrix to votex
Raúl Pérez-Fernández, Michaël Rademaker, Pedro Alonso 0001, Irene Díaz, Susana Montes, Bernard De Baets
Int. J. Approx. Reason.4
2016 On δ-ϵ-Partitions for Finite Interval-Valued Hesitant Fuzzy Sets
abstract
Hesitant fuzzy sets represent a useful tool in many areas such as decision making or image processing. Finite interval-valued hesitant fuzzy sets are a particular kind of hesitant fuzzy sets that generalize fuzzy sets, interval-valued fuzzy sets or Atanassov’s intuitionistic fuzzy sets, among others. Partitioning is a long-standing open problem due to its remarkable importance in many areas such as clustering. Thus, many different partitioning approaches have been developed for crisp and fuzzy sets. This work presents a partitioning method for the so-called finite interval-valued hesitant fuzzy sets. The definition of this partitioning method involves a definition of an ordering relation for finite interval-valued fuzzy sets membership degrees, i.e, finitely generated sets, as well as the definitions of t-norm and t-conorm for these kinds of sets.
Pelayo Quirós, Pedro Alonso 0001, Irene Díaz, Susana Montes
Int. J. Uncertain. Fuzziness Knowl. Based Syst.3
2016 An Analytical Solution to Dujmovic's Iterative OWA
abstract
Iterative OWA (ItOWA) as proposed by Dujmovic, is a two-stage procedure for computing the weighting vector by a double nested iteration: (i) weights at step h are computed as limit to infinity of a matrix power, (ii) the result is used to start the computation at step h + 1, until the OWA operator arity n is reached. Thereafter Dujmovic suggested a computational solution based on the conjecture that the limit exists, and numerical simulations have being supported the hypothesis that the conjecture is correct. In this paper, we prove that the limit actually exists and we provide an analytical solution to the procedure, so the weighting vector can be computed directly instead of an iterative time-consuming procedure. This theoretical result enables a faster computation of the weighting vector and characterization in terms of weights values, attitudinal character and entropy.
Luigi Troiano, Irene Díaz
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2016 Applications of finite interval-valued hesitant fuzzy preference relations in group decision making
Raúl Pérez-Fernández, Pedro Alonso 0001, Humberto Bustince, Irene Díaz, Susana Montes
Inf. Sci.4
2016 Fuzzy mathematical morphology for color images defined by fuzzy preference relations
Agustina Bouchet, Pedro Alonso 0001, Juan Ignacio Pastore, Susana Montes, Irene Díaz
Pattern Recognit.5
2015 Multi-factorial risk assessment: An approach based on fuzzy preference relations
Raúl Pérez-Fernández, Pedro Alonso 0001, Irene Díaz, Susana Montes
Fuzzy Sets Syst.3
2015 Discovering user preferences using Dempster-Shafer theory
Luigi Troiano, Luis J. Rodríguez-Muñiz, Irene Díaz
Fuzzy Sets Syst.3
2015 Ordering finitely generated sets and finite interval-valued hesitant fuzzy sets
Raúl Pérez-Fernández, Pedro Alonso 0001, Humberto Bustince, Irene Díaz, Aranzazu Jurio, Susana Montes
Inf. Sci.4
2015 An entropy measure definition for finite interval-valued hesitant fuzzy sets
Pelayo Quirós, Pedro Alonso 0001, Humberto Bustince, Irene Díaz, Susana Montes
Knowl. Based Syst.4
2014 Measures of Semantic Similarity of Nodes in a Social Network
Ahmad Rawashdeh, Mohammad Rawashdeh, Irene Díaz, Anca L. Ralescu
IPMU (2)3
2014 A Model for Preserving Privacy in Recommendation Systems
Luigi Troiano, Irene Díaz
IPMU (2)2
2014 Statistical analysis of parametric t-norms
Luigi Troiano, Luis J. Rodríguez-Muñiz, Pasquale Marinaro, Irene Díaz
Inf. Sci.4
2013 On random generation of fuzzy measures
Elías F. Combarro, Irene Díaz, Pedro Miranda 0002
Fuzzy Sets Syst.2
2012 Privacy Issues in Social Networks: A Brief Survey
Irene Díaz, Anca L. Ralescu
IPMU (4)1
2012 A Model for Assessing the Risk of Revealing Shared Secrets in Social Networks
Luigi Troiano, Irene Díaz, Luis J. Rodríguez-Muñiz
IPMU (4)2
2010 Identifying the Risk of Attribute Disclosure by Mining Fuzzy Rules
Irene Díaz, José Ranilla, Luis J. Rodríguez-Muñiz, Luigi Troiano
IPMU (1)1
2007 A Hybrid Feature Selection Method for Text Categorization
abstract
Feature Selection is an important task within Text Categorization, where irrelevant or noisy features are usually present, causing a lost in the performance of the classifiers. Feature Selection in Text Categorization has usually been performed using a filtering approach based on selecting the features with highest score according to certain measures. Measures of this kind come from the Information Retrieval, Information Theory and Machine Learning fields. However, wrapper approaches are known to perform better in Feature Selection than filtering approaches, although they are time-consuming and sometimes infeasible, especially in text domains. However a wrapper that explores a reduced number of feature subsets and that uses a fast method as evaluation function could overcome these difficulties. The wrapper presented in this paper satisfies these properties. Since exploring a reduced number of subsets could result in less promising subsets, a hybrid approach, that combines the wrapper method and some scoring measures, allows to explore more promising feature subsets. A comparison among some scoring measures, the wrapper method and the hybrid approach is performed. The results reveal that the hybrid approach outperforms both the wrapper approach and the scoring measures, particularly for corpora whose features are less scattered over the categories.
Elena Montañés, José Ramón Quevedo, Elías F. Combarro, Irene Díaz, José Ranilla
Int. J. Uncertain. Fuzziness Knowl. Based Syst.4
2005 Towards Automatic and Optimal Filtering Levels for Feature Selection in Text Categorization
Elena Montañés, Elías F. Combarro, Irene Díaz, José Ranilla
IDA3
2005 Generating domain representations using a relationship model
Irene Díaz, Juan Llorens Morillo, Gonzalo Génova, José Miguel Fuentes
Inf. Syst.1
2005 Introducing a Family of Linear Measures for Feature Selection in Text Categorization
abstract
Text categorization, which consists of automatically assigning documents to a set of categories, usually involves the management of a huge number of features. Most of them are irrelevant and others introduce noise which could mislead the classifiers. Thus, feature reduction is often performed in order to increase the efficiency and effectiveness of the classification. In this paper, we propose to select relevant features by means of a family of linear filtering measures which are simpler than the usual measures applied for this purpose. We carry out experiments over two different corpora and find that the proposed measures perform better than the existing ones.
Elías F. Combarro, Elena Montañés, Irene Díaz, José Ranilla, Ricardo Mones
IEEE Trans. Knowl. Data Eng.3
2004 Text Categorization by a Machine-Learning-Based Term Selection
Javier Fernández 0002, Elena Montañés, Irene Díaz, José Ranilla, Elías F. Combarro
DEXA3
2004 Improving performance of text categorization by combining filtering and supportvector machines
abstract
Abstract Text Categorization is the process of assigning documents to a set of previously fixed categories. A lot of research is going on with the goal of automating this time‐consuming task. Several different algorithms have been applied, and Support Vector Machines (SVM) have shown very good results. In this report, we try to prove that a previous filtering of the words used by SVM in the classification can improve the overall performance. This hypothesis is systematically tested with three different measures of word relevance, on two different corpus (one of them considered in three different splits), and with both local and global vocabularies. The results show that filtering significantly improves the recall of the method, and that also has the effect of significantly improving the overall performance.
Irene Díaz, José Ranilla, Elena Montañés, Javier Fernández 0002, Elías F. Combarro
J. Assoc. Inf. Sci. Technol.1
2003 Measures of Rule Quality for Feature Selection in Text Categorization
Elena Montañés, Javier Fernández 0002, Irene Díaz, Elías F. Combarro, José Ranilla
IDA3
2002 An algorithm for term conflation based on tree structures
abstract
Abstract This work presents a new stemming algorithm. This algorithm stores the stemming information in tree structures. This storage allows us to enhance the performance of the algorithm due to the reduction of the search space and the overall complexity. The final result of that stemming algorithm is a normalized concept, understanding this process as the automatic extraction of the generic form (or a lexeme) for a selected term.
Irene Díaz, Jorge Morato, Juan Llorens Morillo
J. Assoc. Inf. Sci. Technol.1