Antoon Bronselaer

dblp:65/126 · DBLP profile ↗
← Back
51ranked-venue papers
16as first author
8since 2021 · last 2023
0000-0001-6663-192XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 13 first-author · 3 since 2021Databases, data management, data science and information retrieval · 31 · 9 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 An Orthographic Similarity Measure for Graph-Based Text Representations
Maxime Deforche, Ilse De Vos, Antoon Bronselaer, Guy De Tré
FQAS3
2023 Cost-based analysis of the impact of data completeness and representational consistency
abstract
Data quality is an important topic for businesses and therefore requires appropriate analysis tools. Although several rule-based systems exist today for quality measurement, their results do not always reflect the real impact of quality issues on practical data usability and are therefore not well-suited to base economic decisions on. This work practically implements and evaluates an alternative, cost-based approach for data quality analysis starting from a ‘fitness for use’-perspective. The practical impact of completeness and representational consistency of data stored in an integrated relational database is investigated in an experiment with 218 volunteers. Two alternative versions of this database are then prepared by manually improving their data quality. Participants are randomly assigned to one of three databases and are given a set of questions to resolve by means of SQL. As questions are resolved, we measure several cost-based indicators such as ability to solve, time to solve and number of attempts. Results indicate that the impact of data quality issues can differ significantly from what would be expected when using rule-based measurement. Effects range from almost no impact to a 65% reduction in time needed to solve tasks. Effect sizes up to 0.43 using one-way ANCOVA tests are observed.
Yoram Timmerman, Rihem Nasfi, Guy De Tré, Filip Pattyn, Antoon Bronselaer
Decis. Support Syst.5
2023 A novel approach to assess and improve syntactic interoperability in data integration
Rihem Nasfi, Antoon Bronselaer, Guy De Tré
Inf. Process. Manag.2
2022 Search-based Reinforcement Learning through Bandit Linear Optimization
abstract
The development of AlphaZero was a breakthrough in search-based reinforcement learning, by employing a given world model in a Monte-Carlo tree search (MCTS) algorithm to incrementally learn both an action policy and a value estimation. When extending this paradigm to the setting of simultaneous move games we find that the selection strategy of AlphaZero has theoretical shortcomings, including that convergence to a Nash equilibrium is not guaranteed. By analyzing these shortcomings, we find that the selection strategy corresponds to an approximated version of bandit linear optimization using Tsallis entropy regularization with α parameter set to zero, which is equivalent to log-barrier regularization. This observation allows us to refine the search method used by AlphaZero to obtain an algorithm that has theoretically optimal regret as well as superior empirical performance on our evaluation benchmark.
Milan Peelman, Antoon Bronselaer, Guy De Tré
IJCAI2
2022 Dynamic repair of categorical data with edit rules
abstract
In this paper, a dynamic setting for data quality improvement is studied. In such a setting, there is a repeated search for data quality rules and a fix of their violations until stability is reached. The constraints considered here are simple constant edit rules and searching is done via association analysis. Repair of violations relies on the set cover method. This paper contributes to the field of data quality in three ways. First, it is shown that with appropriate filtering, association analysis is an appealing tool to discover data quality rules with high precision. Second, when edit rules are limited to logical implications such as association rules, then under reasonable circumstances, time complexity of rule implication reduces from exponential to quadratic. This result is formalized as the strong generator theorem. Third, a detailed analysis of data repair in a dynamic setting is provided and the conditions for termination are shown. Empirical results indicate that if the initial precision of rules is high, then repeated search-and-repair offers a boost in recall with a mitigated drop in precision.
Antoon Bronselaer, Toon Boeckling, Filip Pattyn
Expert Syst. Appl.1
2022 Automated monitoring of online news accuracy with change classification models
abstract
In the past decade, news consumption has shifted from printed news media to online alternatives. Although these come with advantages, online news poses challenges as well. Notable here is the increased competition between online newspapers and other online news providers to attract readers. Hereby, speed is often favored over quality. As a consequence, the need for new tools to monitor online news accuracy has grown. In this work, a fundamentally new and automated procedure for the monitoring of online news accuracy is proposed. The approach relies on the fact that online news articles are often updated after initial publication, thereby also correcting errors. Automated observation of the changes being made to online articles and detection of the errors that are corrected may offer useful insights concerning news accuracy. The potential of the presented automated error correction detection model is illustrated by building supervised classification models for the detection of objective, subjective and linguistic errors in online news updates respectively. The models are built using a large news update data set being collected during two consecutive years for six different Flemish online newspapers. A subset of 21,129 changes is then annotated using a combination of automated and human annotation via an online annotation platform. Finally, manually crafted features and text embeddings obtained by four different language models (TF-IDF, word2vec, BERTje and SBERT) are fed to three supervised machine learning algorithms (logistic regression, support vector machines and decision trees) and performance of the obtained models is subsequently evaluated. Results indicate that small differences in performance exist between the different learning algorithms and language models. Using the best-performing models, F2-scores of 0.45, 0.25 and 0.80 are obtained for the classification of objective, subjective and linguistic errors respectively.
Yoram Timmerman, Antoon Bronselaer
Inf. Process. Manag.2
2022 Efficient edit rule implication for nominal and ordinal data
Toon Boeckling, Guy De Tré, Antoon Bronselaer
Inf. Sci.3
2021 Data Quality Management: An Overview of Methods and Challenges
Antoon Bronselaer
FQAS1
2020 Quantifying the Impact of EER Modeling on Relational Database Success: An Experimental Investigation
Yoram Timmerman, Antoon Bronselaer, Guy De Tré
ER2
2019 Measuring data quality in information systems research
Yoram Timmerman, Antoon Bronselaer
Decis. Support Syst.2
2019 Compact representations of temporal databases
Antoon Bronselaer, Christophe Billiet, Robin De Mol, Joachim Nielandt, Guy De Tré
VLDB J.1
2018 Randomness of Data Quality Artifacts
Toon Boeckling, Antoon Bronselaer, Guy De Tré
IPMU (3)2
2018 Operational Measurement of Data Quality
Antoon Bronselaer, Joachim Nielandt, Toon Boeckling, Guy De Tré
IPMU (3)1
2018 An incremental approach for data quality measurement with insufficient information
Antoon Bronselaer, Joachim Nielandt, Guy De Tré
Int. J. Approx. Reason.1
2018 Human Centric Data Management
abstract
With "human centric data management", we denote all kind of practical and theoretical developments that contribute to the improvement of data management for human users, such that it becomes better to understand, easier to handle, and more natural to communicate with.The popularity of digital applications, social media and multimedia created a shift towards "big data" that are characterized by huge data volumes, a large variety of data formats, fast data processing requirements, and veracity problems.The more data we have at our disposal, the more applications arise, but also the more sophisticated these applications become.Along with these technological developments comes the awareness that there is a growing need for human centric data management tools.Indeed, perfect data sets are rare and data imperfections propagate to imperfect data processing solutions.Humans communicate in natural language and cope with imperfect information in their everyday behavior, whereas conventional data management assumes that data are perfect and data manipulation is based on a bivalent Boolean logic.Computational intelligence techniques, more specifically soft computing and fuzzy set theory, offer the tools for bridging the gap between the way humans behave and communicate and the way conventional data management tools work.This is especially the case because they allow to generalize bivalent Boolean logic into multivalued fuzzy logic and offer sound foundations for uncertainty modeling that are less stringent, but broader applicable, than conventional probability theory.This special issue is an initiative of the working group on "Soft Computing in Database Management and Information Retrieval" of the European Society for
Guy De Tré, Janusz Kacprzyk, Gabriella Pasi, Slawomir Zadrozny, Antoon Bronselaer
Int. J. Intell. Syst.5
2018 Handling veracity in multi-criteria decision-making: A multi-dimensional approach
Guy De Tré, Robin De Mol, Antoon Bronselaer
Inf. Sci.3
2018 A Measure-Theoretic Foundation for Data Quality
abstract
In this paper, a novel framework for data quality measurement is proposed by adopting a measure-theoretic treatment of the problem. Instead of considering a specific setting in which quality must be assessed, our approach departs more formally from the concept of measurement. The basic assumption of the framework is that the highest possible quality can be described by means of a set of predicates. Quality of data is then measured by evaluating those predicates and by combining their evaluations. This combination is based on a capacity function (i.e., a fuzzy measure) that models for each combination of predicates the capacity with respect to the quality of the data. It is shown that expression of quality on an ordinal scale entails a high degree of interpretation and a compact representation of the measurement function. Within this purely ordinal framework for measurement, it is shown that reasoning about quality beyond the ordinal level naturally originates from the uncertainty about predicate evaluation. It is discussed how the proposed framework is positioned with respect to other approaches with particular attention to aggregation of measurements. The practical usability of the framework is discussed for several well known dimensions of data quality and demonstrated in a use-case study about clinical trials.
Antoon Bronselaer, Robin De Mol, Guy De Tré
IEEE Trans. Fuzzy Syst.1
2017 On the Need for Explicit Confidence Assessments of Flexible Query Answers
Guy De Tré, Robin De Mol, Antoon Bronselaer
FQAS3
2017 Evaluating flexible criteria on uncertain data
Robin De Mol, Antoon Bronselaer, Guy De Tré
Fuzzy Sets Syst.2
2016 A comparison technique for ill-known time intervals
abstract
Currently, many existing information systems contain large amounts of data. Prior research shows that many of these data represent time (domain) intervals and many of those may be subject to uncertainty. Moreover, great attention has been attracted by defining, finding or using qualitative temporal relationships between time (domain) intervals. In this context, the framework of Allen relationships is one of the most notable proposals. As a consequence, many existing proposals consider qualitative temporal relationships between time (domain) intervals subject to uncertainty on one hand and regular time (domain) intervals on the other. One of the most novel approaches in this context is given by the ill-known constraints framework. However, this framework does not yet support comparisons between time (domain) intervals subject to uncertainty. In this paper, a proposal is presented to expand this ill-known constraints framework in order to allow it to support the aforementioned temporal comparison, based on the Allen relationships.
Christophe Billiet, Antoon Bronselaer, Guy De Tré
FUZZ-IEEE2
2016 Partial absorption aggregators with Fuzzy Integrals
abstract
There are numerous automated strategies in the context of decision making, each with their own strengths and weaknesses. However, the general utility of a decision support systems relates to its expressive capabilities regarding human-like reasoning. In the ongoing effort to bolster decision support, this work aims to enrich fuzzy integrals with the power to model partial absorption operators. Starting from mathematical definitions, it is shown that it is necessary to redefine specific weights of internal sets in fuzzy measures from constants to functions in order to model this behavior. An implementation is proposed for representing the conjunctive partial absorption and the disjunctive partial absorption. Their results are carefully analysed and it is shown the made alterations do not violate the monotonic nature of fuzzy measures.
Robin De Mol, Antoon Bronselaer, Guy De Tré
FUZZ-IEEE2
2016 Indexing possibilistic temporal data in a database of medieval charters
abstract
Querying large databases containing imperfect data requires efficient indexing techniques. Without such techniques query processing would simply take too much time. Considering a possibility based database modelling approach, imperfect data are modelled using a possibility distribution. The possibility distributions used for data modelling in different database records have to be indexed in order to support a faster processing of query conditions that act on the imperfect data. In this paper we study the indexing of imperfect temporal data in the Diplomata Belgica database, which has been co-developed by our research group. Diplomata Belgica is a relational database describing medieval charters written and issued in the southern Low Countries. More specifically, we study how imperfect data on the issuing date of a charter can be modelled and indexed in order to support searches for charters with an issuing date that is compatible with `fuzzy' query preferences provided by the user. A novel, so-called Interval B+-Tree (IBPT) indexing technique is proposed and some illustrative examples of (the handling of) complex, realistic queries are given.
Guy De Tré, Christophe Billiet, Antoon Bronselaer, Carlos D. Barranco
FUZZ-IEEE3
2016 Ordinal Assessment of Data Consistency Based on Regular Expressions
Antoon Bronselaer, Joachim Nielandt, Robin De Mol, Guy De Tré
IPMU (2)1
2016 A Possibilistic Treatment of Data Quality Measurement
Antoon Bronselaer, Guy De Tré
IPMU (2)1
2016 Indexing Possibilistic Numerical Data: The Interval B ^+ -tree Approach
Guy De Tré, Robin De Mol, Antoon Bronselaer
IPMU (2)3
2016 Predicate enrichment of aligned XPaths for wrapper induction
Joachim Nielandt, Antoon Bronselaer, Guy De Tré
Expert Syst. Appl.2
2015 Coreference detection in an XML schema
Marcin Szymczak 0001, Slawomir Zadrozny, Antoon Bronselaer, Guy De Tré
Inf. Sci.3
2015 Using Data Merging Techniques for Generating Multidocument Summarizations
abstract
In this paper, we examine how we can use data merging techniques to summarize a set of coreferent documents that has been clustered while using soft computing techniques. The main focus of this paper lies on the fβ-optimal merge function (a function newly introduced here), which that uses the weighted harmonic mean to find a balance between precision and recall. The global precision and recall measures mentioned are defined by means of a triangular norm receiving local precision and recall values as an input, in order to generate a multiset of key concepts that we can use to generate summarizations. The fβ-optimal merge function is compared with a distance-based merge function and several pointwise merge functions from both a theoretical and an experimental point of view. It will be shown that the fβ-optimal merge function has quite a few advantages over the others, especially if one looks at the practical usage in the context of data merging and summarizing multiple documents concerning the same topic.
Daan Van Britsom, Antoon Bronselaer, Guy De Tré
IEEE Trans. Fuzzy Syst.2
2015 Propagation of Data Fusion
abstract
In a relational database, tuples are called “duplicate” if they describe the same real-world entity. If such duplicate tuples are observed, it is recommended to remove them and to replace them with one tuple that represents the joint information of the duplicate tuples to a maximal extent. This remove-and-replace operation is called a fusion operation. Within the setting of a relational database management system, the removal of the original duplicate tuples can breach referential integrity. In this paper, a strategy is proposed to maintain referential integrity in a semantically correct manner, thereby optimizing the quality of relationships in the database. An algorithm is proposed that is able to propagate a fusion operation through the entire database. The algorithm is based on a framework of first and second order fusion functions on the one hand, and conflict resolution strategies on the other hand. It is shown how classical strategies for maintaining referential integrity, such as DELETE cascading, are highly specialized cases of the proposed framework. Experimental results are reported that (i) show the efficiency of the proposed algorithm and (ii) show the differences in quality between several second order fusion functions. It is shown that some strategies easily outperform DELETE cascading.
Antoon Bronselaer, Daan Van Britsom, Guy De Tré
IEEE Trans. Knowl. Data Eng.1
2014 Bipolar Comparison of 3D Ear Models
Guy De Tré, Dirk Vandermeulen, Jeroen Hermans, Peter Claes, Joachim Nielandt, Antoon Bronselaer
IPMU (3)6
2014 A method based on shape-similarity for detecting similar opinions in group decision-making
Ana Tapia-Rosero, Antoon Bronselaer, Guy De Tré
Inf. Sci.2
2013 Comparing f β -Optimal with Distance Based Merge Functions
Daan Van Britsom, Antoon Bronselaer, Guy De Tré
FQAS2
2013 Enhancing Flexible Querying Using Criterion Trees
Guy De Tré, Jozo J. Dujmovic, Joachim Nielandt, Antoon Bronselaer
FQAS4
2013 Possibilistic Evaluation of Sets
abstract
In the past decades, the theory of possibility has been developed as a theory of uncertainty that is compatible with the theory of probability. Whereas probability theory tries to quantify uncertainty that is caused by variability (or equivalently randomness), possibility theory tries to quantify uncertainty that is caused by incomplete information. A specific case of incomplete information is that of ill-known sets, which is of particular interest in the study of temporal databases. However, the construction of possibility distributions in the case of ill-known sets is known to be overly complex. This paper contributes to the study of ill-known sets by investigating the inference of uncertainty when constraints are specified over ill-known values. More specific, in this paper it is investigated how the knowledge about constraint satisfaction can be inferred if the constraints themselves are defined by means of ill-known values. It is shown how such reasoning can contribute to the study of (fuzzy) temporal databases.
Antoon Bronselaer, Jose Enrique Pons, Guy De Tré, Olga Pons
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2012 Similarity of Membership Functions - A Shaped based Approach
Ana Tapia-Rosero, Antoon Bronselaer, Guy De Tré
IJCCI2
2012 Concept Identification in Constructing Multi-Document Summarizations
Daan Van Britsom, Antoon Bronselaer, Guy De Tré
IPMU (2)2
2012 Robustness of Multiset Merge Functions
Antoon Bronselaer, Daan Van Britsom, Guy De Tré
IPMU (1)1
2012 Coreference Detection of Low Quality Objects
Joachim Nielandt, Antoon Bronselaer, Guy De Tré
IPMU (1)2
2012 On the Applicability of Multi-criteria Decision Making Techniques in Fuzzy Querying
Guy De Tré, Jozo J. Dujmovic, Antoon Bronselaer, Tom Matthé
IPMU (1)3
2012 A framework for multiset merging
Antoon Bronselaer, Daan Van Britsom, Guy De Tré
Fuzzy Sets Syst.1
2012 Concept-relational text clustering
abstract
The ongoing exponential growth of online information sources has led to a need for reliable and efficient algorithms for text clustering. In this paper, we propose a novel text model called the relational text model that represents each sentence as a binary multirelation over a concept space \documentclass{article}\usepackage{amssymb}\pagestyle{empty}\begin{document}${\mathcal{C}}$\end{document}. Through usage of the smart indexing engine (SIE), a patented technology of the Belgian company i.Know, the concept space adopted by the text model can be constructed dynamically. This means that there is no need for an a priori knowledge base such as an ontology, which makes our approach context independent. The concepts resulting from SIE possess the property that frequency of concepts is a measure for relevance. We exploit this property with the development of the CR-algorithm. Our approach relies on the representation of a data set \documentclass{article}\usepackage{amssymb}\pagestyle{empty}\begin{document}${\mathcal{D}}$\end{document} as a multirelation, of which k-cuts can be taken. These cuts can be seen as sets of relevant patterns with respect to the topics that are described by documents. Analysis of dependencies between patterns allows to produce clusters, such that precision is sufficiently high. The best k-cut is the one that best approximates the estimated number of clusters to ensure recall. Experimental results on Dutch news fragments show that our approach outperforms both basic and advanced methods. © 2012 Wiley Periodicals, Inc.
Antoon Bronselaer, Guy De Tré
Int. J. Intell. Syst.1
2011 Automatically generating multi-document summarizations
abstract
This paper describes the News Summarization or NEWSUM algorithm designed to automatically generate multi-document summarizations, hereby focusing on textual documents that concern news items. The NEWSUM algorithm has been implemented and tested in several ways. An overview of both the implementation and the test results are covered in this document.
Daan Van Britsom, Antoon Bronselaer, Guy De Tré
ISDA2
2011 Bipolar database querying using bipolar satisfaction degrees
abstract
When expressing their information needs in a (database) query, users sometimes prefer to state what has to be rejected rather than what has to be accepted. In general, what has to be rejected is not necessarily the complement of what has to be accepted. This phenomenon is commonly known as the heterogeneous bipolar nature of expressing information needs. Satisfaction degrees in regular fuzzy querying approaches are based on the “symmetric'' assumption that the extent to which a database record, respectively, satisfies and does not satisfy a given query are complements of each other and are therefore less suited to adequately handle heterogeneous bipolarity in query specifications and query processing. In this paper, we present a bipolar query satisfaction modeling framework which is based on pairs that consist of an independent degree of satisfaction and degree of dissatisfaction. The use and advantages of the framework are illustrated in the context of fuzzy query evaluation in regular relational databases. More specifically, the evaluation of heterogeneous bipolar queries that contain both positive, negative, and bipolar criteria is studied. © 2011 Wiley Periodicals, Inc.
Tom Matthé, Guy De Tré, Slawomir Zadrozny, Janusz Kacprzyk, Antoon Bronselaer
Int. J. Intell. Syst.5
2010 Consistently Handling Geographical User Data - Context-Dependent Detection of Co-located POIs
Guy De Tré, Antoon Bronselaer, Tom Matthé, Nico Van de Weghe, Philippe De Maeyer
IPMU (2)2
2010 Properties of Possibilistic String Comparison
abstract
The problem of detecting coreferent objects of arbitrary complexity is a challenging topic in current research. A possibilistic solution for this problem is to treat it as an uncertain Boolean problem. This means that two objects are either coreferent or not (i.e., a Boolean matter), but uncertainty about this decision must be dealt with. An operator that determines the uncertainty about the coreference of two objects is called anevaluator. When we deal withstructuredobjects, decomposition into attributes (i.e., atomic subobjects) allows the definition of evaluators on well-known subdomains. This paper proceeds previous research on evaluators for strings, which is a widely used data type for attributes. First of all, the Sugeno integral based on the framework of conditional necessity is shown to be related to the existing technique. More specifically, a special case of this Sugeno integral is equivalent to regular conjunction of transformed possibilistic truth values, which is used by existing evaluators for strings. As a consequence, a subfamily of the existing evaluator is obtained for strings. This subfamily is shown to satisfy several interesting properties, which are used to construct an efficient optimization algorithm for string evaluators. Next, the use of a frequency filter is investigated. Finally, novel and advanced techniques like interlevel-information exchange and the use of multiple quantifiers are defined and investigated. A series of tests on diverse datasets shows the high accuracy and robustness of the approach that is introduced in this paper.
Antoon Bronselaer, Guy De Tré
IEEE Trans. Fuzzy Syst.1
2010 Handling Bipolarity in Elementary Queries to Possibilistic Databases
abstract
Making data-querying and representation easier and more human consistent is an important research topic. In this context, fuzzy logic with its capability to model linguistic expressions provides an interesting framework, which has been adopted by many researchers. However, there are still some aspects that have not been adequately covered. In particular, it becomes widely advocated that while communicating, humans give both positive and negative information to state what they desire and what they reject. Because positive and negative statements do not necessarily mirror each other, this results in so-called heterogeneous bipolar information. Traditional fuzzy approaches do not adequately support the handling of heterogeneous bipolar information in information systems. Therefore, there is a need for more advanced techniques. In this paper, how bipolarity can be dealt with in the formulation and evaluation of selection conditions in fuzzy querying within a possibilistic, relational database framework is presented. Three novel query-evaluation techniques based on interval-valued fuzzy sets, Atanassov fuzzy sets, and twofold fuzzy sets are presented and compared with each other. Possibility theory is used to deal with uncertainty. Special attention is paid to the description of the semantics, use, benefits, and drawbacks of each formalism.
Guy De Tré, Slawomir Zadrozny, Antoon Bronselaer
IEEE Trans. Fuzzy Syst.3
2009 Dealing with Positive and Negative Query Criteria in Fuzzy Database Querying
Guy De Tré, Slawomir Zadrozny, Tom Matthé, Janusz Kacprzyk, Antoon Bronselaer
FQAS5
2009 Extensions of fuzzy measures and Sugeno integral for possibilistic truth values
abstract
The Sugeno integral has been identified and used as an aggregation operator many times in the past. In this paper, an extension of the Sugeno integral for the framework of possibilistic truth values is presented, resulting in a powerful domain-specific aggregation operator. Next, it is shown how the presented integral can be plugged into a reasoning framework for identification of coreferent objects, which are entity descriptions that refer to the same entity in a different way. The concept of hierarchical fuzzy measures, linked to an object structure is introduced, offering a new conditional possibilistic reasoning framework for object matching. © 2008 Wiley Periodicals, Inc.
Antoon Bronselaer, Axel Hallez, Guy De Tré
Int. J. Intell. Syst.1
2009 Performance optimization of object comparison
abstract
Comparing objects can be considered as a hierarchical process. Separate aspects of objects are compared to each other, and the results of these comparisons are combined into a single result in one or more steps by aggregation operators. The set of operators used to compare the objects and the way these operators are related with each other is called the comparison scheme. If a threshold is applied to the final result of the object comparison, the mathematical properties of the operators in the comparison scheme can be used to derive thresholds on the intermediate results. These derived threshold can be used to break of a comparison early, thus offering a reduction of the comparison cost. Using this information, we show that the order in which the operators are evaluated has an influence on the average cost of comparing two objects. Next, we proceed with a study of the properties that allow us to find an optimal order, such that this average cost is minimized. Finally, we provide an algorithm that calculates an optimal order efficiently. Although specifically developed for object comparison, the algorithm can be applied to all kinds of selection processes that involve the combination of several test results. © 2009 Wiley Periodicals, Inc.
Axel Hallez, Guy De Tré, Antoon Bronselaer
Int. J. Intell. Syst.3
2009 Comparison of Sets and Multisets
abstract
The comparison of sets of objects is a research topic with applications in diverse fields such as computer science, biology and psychology. Since the introduction of the Jaccard index, many techniques have been proposed. This paper aims at extending an existing framework of comparison indices for sets. Firstly, the novel indices account for similarities between elements, rather than identity of elements as is the case for existing techniques. As a result, a richer framework of comparison indices is obtained. The use of fuzzy quantifiers in this framework is shown. Secondly, the machinery for sets is extended to the case of multisets, which results in two classes of comparison indices. The first class considers each element instance as a separate element, while the second class considers groups of elements instances as an atomic entity. The number of instances is then a property of this group, that is taken into account when calculating similarity between element groups.
Axel Hallez, Antoon Bronselaer, Guy De Tré
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2009 A Possibilistic Approach to String Comparison
abstract
In this paper, comparison of strings is tackled from a possibilistic point of view. Instead of using the concept of similarity between strings, coreference between strings is adopted. The possibility of coreference is estimated by means of a possibilistic comparison operator. In literature, two important classes of comparison methods for strings have been distinguished: character-based methods and token-based methods. The first class treats a string as a sequence of characters, while the second class treats a string as a vector of substrings. The first contribution of this paper is to propose a new character-based method that is able to detect typographical errors and abbreviations. The main advantage of the proposed technique is the very low complexity in comparison with existing character-based techniques. In a second contribution, two-level systems are investigated and a new approach is described. The novelty of the proposed two-level system is the use of multiset comparison rather than vector comparison. It is shown how an ordered weighted conjunctive operator that uses a parameterized fuzzy quantifier to deliver weights is competitive with frequency-based weights. In addition, the use of a quantifier is significantly faster than the use of existing weight techniques. In a third contribution, a novel class of hybrid techniques is proposed that combines the advantages of several methods. Finally, comparative tests regarding accuracy and execution time are performed and reported.
Antoon Bronselaer, Guy De Tré
IEEE Trans. Fuzzy Syst.1